Memory controller and three-dimensional stacked memory
By setting the read cache area and the prefetch cache area in the memory controller, prefetching and error correction of DRAM data is achieved, solving the problems of DRAM data reliability and system performance improvement, and improving the access speed and bandwidth of the memory.
Patent Information
- Application Number
- CN202510906523.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-01
AI Technical Summary
How to further improve the data reliability and system performance of dynamic random access memory (DRAM).
Set the read cache area and the prefetch cache area in the memory controller. By determining whether the data to be read has been cached in the prefetch cache area in advance. If not, error checks and corrections are performed when reading the data of the hit layer and non-hit layer of the multi-layer memory chip, and the data is cached to the corresponding cache area, and the data is prefetched and corrected before returning to the main chip.
Improves the data reliability and system performance of the memory, reduces access time to multi-layer memory chips, enhances bandwidth, and improves the reliability of stored data without affecting performance.
Smart Images

Figure CN120407468A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of memories, and particularly to a memory controller and a three-dimensional stacked memory. Background Art
[0002] Dynamic Random Access Memory (DRAM) has been widely used in the main memory of various applications, covering multiple fields such as high-performance computing and mobile applications, due to its advantages of high density, simple architecture, low latency, and low power consumption.
[0003] Please refer to Figure 1 , a current DRAM has a memory controller 2 and stacked multi-layer memory chips (i.e., corresponding to the DRAM dies in Figure 1 , which can also be referred to as "DRAM chips") 30, 31, 32..., and the memory controller 2 is coupled to the main chip (logic die, which can also be referred to as "master die" or "compute die") 1, enabling the main chip 1 to directly communicate with each layer of memory chips (i.e., DRAM chips) 30, 31, 32... through Through-Silicon Vias (TSVs) or through a buffer chip (Buffer die, also known as "base die" or "interface die") and TSVs, thereby enabling the main chip 1 to access each layer of memory chips (i.e., DRAM chips) 30, 31, 32... through an AXI interface or the like, thereby expanding the memory capacity managed by the main chip 1. Among them, the memory controller 2 is disposed in the main chip 1 or in the buffer chip, and each layer of memory chips (i.e., DRAM chips) has multiple storage units unit (also known as "DRAM unit") with the same capacity. For the main chip 1, the multiple storage units in each layer of memory chips (i.e., DRAM chips) managed by the memory controller 2 are independent of each other.
[0004] In order to improve the system stability of DRAM, an ECC (Error Checking and Correcting) (also known as "ECC error correction") function is currently introduced to timely detect and correct data storage errors inside the DRAM, thereby extending the normal running time of the DRAM system.
[0005] How to further improve the data reliability and system performance of DRAM remains one of the hot issues that those skilled in the art need to solve. Summary of the Invention
[0006] The object of the present invention is to provide a memory controller and a three-dimensional stacked memory, which can improve the system performance of the memory.
[0007] To achieve the above object, the present invention provides a memory controller, which is coupled between a main chip and stacked multi-layer memory chips. Each layer of the memory chips includes a plurality of memory cells. Among them, the memory controller has a data buffer, and the data buffer includes a read buffer and a prefetch buffer. The memory controller is configured to: Receive and parse a corresponding read instruction, and determine whether the data read by the read instruction has been previously cached in the prefetch buffer. If not, then perform the following steps: Open the word lines sharing the execution row address of the read instruction in the memory chips of different layers to read the data of the corresponding hit memory cells and the data of the non-hit memory cells in the multi-layer memory chips; Perform error checking and correction on the read data; Cache the data after error checking and correction corresponding to the hit memory cells into the read buffer, and cache the data after error checking and correction corresponding to the non-hit memory cells into the prefetch buffer; Take out the data cached in the read buffer and return it to the main chip.
[0008] Optionally, the memory controller is configured to: determine a hit layer in the multi-layer memory chips according to the chip select address of the read instruction, where the hit memory cells belong to the hit layer, and the non-hit memory cells belong to at least one non-hit layer other than the hit layer in the multi-layer memory chips; and / or The memory controller is further configured to: if it is determined that the data read by the read instruction has been previously cached in the prefetch buffer, take out the data corresponding to the read instruction from the prefetch buffer and return it to the main chip.
[0009] Optionally, the memory controller further includes at least one IO interface, a control and management module, and several memory control modules. The data buffer is disposed in the control and management module and is correspondingly disposed and coupled to the IO interface; each of the memory control modules is coupled to the control and management module and the multi-layer memory chips; where: The control and management module is used to parse the read instruction to obtain the hit layer, the hit memory cells in the hit layer, and the execution row address of the hit memory cells, and determine whether the data read by the read instruction has been previously cached in the prefetch buffer; The storage control module is used to turn on the word lines sharing the execution row address in the memory chips of different layers including the hit layer, so as to read the data of the hit storage unit and the data of the non-hit storage unit, perform error checking and correction on the read data, cache the data after error checking and correction corresponding to the hit storage unit into the corresponding read buffer area, and cache the data after error checking and correction corresponding to the non-hit storage unit into the corresponding prefetch buffer area; The IO interface is used to receive the read instruction, and according to the judgment result of the control management module, take out the data corresponding to the read instruction from the corresponding read buffer area or the prefetch buffer area and return it to the main chip.
[0010] Optionally, the control management module further includes: An address decoder, configured to parse the read instruction to obtain the hit layer of the read instruction, the hit storage unit in the hit layer, and the execution row address in the hit storage unit; A discriminator, configured to judge whether the data read by the read instruction has been cached in the prefetch buffer area in advance according to the parsing result of the address decoder.
[0011] Optionally, the storage control module includes an error checking and correction control module, and the error checking and correction control module includes a plurality of error checking and correction control units. Each error checking and correction control unit is correspondingly arranged and coupled to each layer of the memory chips managed by the storage control module. The error checking and correction control unit is used to perform error checking and correction on the data read from the memory chip to which it is coupled.
[0012] Optionally, the data read by the read instruction corresponds to the data continuously stored in a plurality of hit storage units in the hit layer. A plurality of storage control modules managing the plurality of hit storage units cache data into the read buffer area corresponding to the IO interface synchronously, so that the main chip obtains the data read by the read instruction.
[0013] Optionally, each storage control module manages a plurality of storage units on the same through-silicon via hybrid bonding path in the multi-layer memory chips. Each storage unit includes multiple word lines corresponding to multiple row addresses and multiple bit lines corresponding to multiple column addresses; Wherein, the control management module parses the address of the storage control module hit by the read instruction, the layer address of the memory chip in the hit layer, and the execution row address and column address in the hit storage unit from the read instruction.
[0014] Optionally, the memory chip is a DRAM chip, and the IO interface is an AXI interface, a CHI interface, or an AHB interface.
[0015] Optionally, when the storage control module or the memory controller performs error checking and correction on the read data, if a data error is found and corrected, the corrected data is written back to the corresponding storage unit.
[0016] Optionally, the operation of retrieving the data corresponding to the read instruction from the read buffer or the prefetch buffer and returning it to the main chip is defined as the first operation; the operation of writing the corrected data back to the corresponding storage unit is defined as the second operation. The first operation and the second operation are performed in parallel, or the first operation and the second operation are performed serially, and the first operation is executed before the second operation.
[0017] Optionally, the memory controller is disposed in the buffer chip, and the main chip, the buffer chip, and the multi-layer memory chip are stacked together; or the memory controller is disposed in the main chip, and the main chip and the multi-layer memory chip are stacked together.
[0018] Based on the same inventive concept, the present invention also provides a three-dimensional stacked memory, which includes stacked multi-layer memory chips and the memory controller as described in the present invention.
[0019] Compared with the prior art, the technical solution of the present invention has at least one of the following beneficial effects: 1. A read buffer and a prefetch buffer are provided in the memory controller. Each time data is read, it is first determined whether the data to be read has been cached in the prefetch buffer in advance. If not, while reading the corresponding data in the hit layer of the multi-layer memory chip (i.e., the data stored in the storage unit), at least one non-hit layer in the multi-layer memory chip is also read (i.e., prefetch), and error checking and correction (ECC) are further performed on the read and prefetch data. After error checking and correction, the data in the hit layer is cached in the read buffer, the prefetch data is cached in the prefetch buffer, and the data cached in the read buffer this time is retrieved and returned to the main chip, realizing the functions of prefetching data from the multi-layer memory chip, performing ECC error correction on the read and prefetch data first, and then caching them separately, improving the data reliability of the memory and the system performance.
[0020] 2. Each time data is read, if it is determined that the data to be read has been pre-cached in the prefetch buffer, the data can be directly fetched from the prefetch buffer instead of fetching it from the stacked multi-layer memory chips, thereby saving the time for switching the word lines of the stacked multi-layer memory chips, improving the access speed of the stacked multi-layer memory chips, and further increasing the bandwidth of the stacked multi-layer memory chips.
[0021] 3. If a data storage error is detected through the error checking and correction (ECC) function, after correcting the error, the corrected data is written back to the stacked multi-layer memory chips, thereby improving the reliability of the stored data without affecting the performance of the stacked multi-layer memory chips. Description of the Drawings
[0022] Those of ordinary skill in the art will understand that the provided drawings are used to better understand the present invention and do not limit the scope of the present invention in any way. Among them: Figure 1 is a schematic diagram of the system architecture of a DRAM.
[0023] Figure 2 is a schematic diagram of the system architecture of the three-dimensional stacked memory and the memory controller according to a specific embodiment of the present invention.
[0024] Figure 3 is a schematic diagram of the package structure of the three-dimensional stacked memory according to a specific embodiment of the present invention.
[0025] Figure 4 is a schematic diagram of the structure of the memory cells in the three-dimensional stacked memory according to a specific embodiment of the present invention.
[0026] Figure 5 is a schematic diagram of the architecture of the error checking and correction (ECC) control module in the memory controller according to a specific embodiment of the present invention.
[0027] Figure 6 is a schematic diagram of an example system architecture of the three-dimensional stacked memory and the memory controller according to a specific embodiment of the present invention.
[0028] Figure 7 is Figure 6 a schematic diagram of the architecture of the error checking and correction (ECC) control module in the memory controller shown. Detailed Embodiments
[0029] In the following description, numerous specific details are set forth to provide a more thorough understanding of the present invention. However, it will be apparent to one of ordinary skill in the art that the present invention may be practiced without one or more of these specific details. In other instances, well-known features have not been described in order to avoid obscuring the present invention. It should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present invention to those skilled in the art. Like reference numerals refer to like elements throughout. It should be understood that when an element is referred to as being "connected to" or "coupled to" another element, it can be directly connected to the other element or intervening elements may be present. In contrast, when an element is referred to as being "directly connected to" another element, there are no intervening elements. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "comprising" is used to specify the presence of the stated features, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups. As used herein, the term "and / or" includes any and all combinations of the associated listed items.
[0030] Please refer to Figure 2 and Figure 3 , an embodiment of the present invention provides a memory controller 2, which is coupled between a main chip 1 and stacked memory chips of layer j + 1 (where Figure 3 DRAM chips are shown as an example in, and in other examples, it may also be other types of memory chips) 30 to 3j. Each layer of the memory chips 30 to 3j includes a plurality of storage units (units) with the same capacity, such as Figure 2 shown as U0 to Uk in. Wherein, j ≥ 1 and j is an integer.
[0031] The memory controller 2 can perform interface conversion between the main chip 1 and the memory chips 30 to 3j, so as to convert instructions such as read, write, and refresh issued by the main chip 1 into signals that the stacked j + 1-layer memory chips 30 to 3j can recognize, and complete the address decoding, data format conversion (such as data bit width), and operation instruction transfer between the main chip 1 and the stacked j + 1-layer memory chips 30 to 3j, thereby realizing the necessary control for access such as refresh operation and read / write operation of the stacked j + 1-layer memory chips 30 to 3j (including the control of address signals, data signals, and various instruction signals), enabling the main chip 1 to access (or "use", "operate") the storage resources (i.e., corresponding storage units) on the stacked j + 1-layer memory chips 30 to 3j according to the needs of users.
[0032] It should be understood that the present invention does not specifically limit the specific setting position of the memory controller 2.
[0033] For example, the memory controller 2 is integrated with the stacked j + 1-layer memory chips 30 to 3j and is independent of the main chip 1, forming a memory system chip that the main chip 1 can access. For another example, please refer to Figure 3 , the memory controller 2 can also be integrated in a buffer chip (Buffer die, also known as base die or interface die) 20. The buffer chip 20 is arranged between the main chip 1 and the stacked j + 1-layer memory chips 30 to 3j. The main chip 1, the buffer chip 20, and the j + 1-layer memory chips 30 to 3j are stacked together through through-silicon via (TSV) and hybrid bonding (HB) technology and are stacked on the substrate 4. At this time, the main chip 1 can communicate with each layer of memory chips 30 to 3j through the buffer chip 20 and the through-silicon via (TSV) hybrid bonding path.
[0034] For another example, the memory controller 2 can also be integrated inside the main chip 1. At this time, the main chip 1 can directly communicate with each layer of memory chips 30 to 3j through the through-silicon via (TSV) hybrid bonding path (or other non-vertically stacked methods).
[0035] Among them, the main chip 1 can include any type of processor device with computing and processing capabilities such as a central processing unit (CPU), a digital signal processor (DSP), a network processor, an application processor (AP), a field programmable gate array (FPGA), and a dedicated processor. The processor device can be configured to execute instructions or software (including code, operating system, or application program, etc.), firmware, or a combination thereof that can be executed by one or more computers.
[0036] Each of the memory chips 30 to 3j can be a DRAM or any other suitable type of memory die structure. Among them, the DRAM can be any suitable type such as synchronous DRAM (SDRAM), wide I / O DRAM, etc. The stacked multi-layer memory chips 30 to 3j form a memory stack, and this memory stack can be implemented as an unbuffered dual in-line memory module (UDIMM), a registered DIMM (RDIMM), a load-reduced DIMM (LRDIMM), a fully buffered DIMM (FBDIMM), a small outline DIMM (SODIMM), etc.
[0037] Please refer to Figure 4 , each memory cell in each of the memory chips 30 to 3j has a memory array array with a corresponding capacity. This memory array has multiple cells (memory elements) determined by the intersection of multiple word lines WL (each word line can be regarded as a row) and multiple bit lines BL (each bit line can be regarded as a column). Each cell is a memory address. Among them, each cell (corresponding to a "memory address" or "instruction execution address") is determined by a corresponding word line WL and a bit line BL. Correspondingly, the word line WL is addressed by the row address (RA) in the memory address, and the bit line BL is addressed by the column address (CA) in the memory address. The memory controller 2 can manage access such as read operations on the stacked multi-layer memory chips 30 to 3j from the cell level (a cell address is a 1-bit memory address) to the memory unit level. Among them, a memory unit can be any suitable management unit at a level higher than the cell level, such as a memory block (block), a sector (sector), a page (page), etc. of the stacked multi-layer memory chips 30 to 3j. Among them, a page contains multiple bytes (whose address range can be determined by multiple word lines and multiple bit lines), a sector contains multiple pages, and a memory block contains multiple sectors.
[0038] It should be noted that in the memory chips 30 to 3j, the corresponding word lines of each memory cell on the memory chips that are coupled to the same TSV hybrid bonding path and are located on different layers can share the same execution row address, such as Figure 2The storage unit U0 in the even-layer memory chips managed by IP0. When the storage unit (unit) hit by a read instruction is U0 in the memory chip 30 managed by IP0, and the execution row address of this read instruction is WL_800, each U0 in the memory chips of each even layer such as the memory chip 32 managed by IP0 has a corresponding word line WL that shares this execution row address WL_800, that is, each U0 in the memory chips of each layer managed by IP0 has a word line WL hit by this execution row address WL_800.
[0039] In this embodiment, the memory controller 2 has a data buffer 21a, and the data buffer 21a includes a read buffer 210 and a prefetch buffer 211. The memory controller 2 is configured to: Receive and parse the corresponding read instruction (read), and determine whether the data read by this read instruction has been pre-cached in the corresponding prefetch buffer 211. If not, then execute the following steps: Open the word lines WL that share the execution row address of this read instruction in the memory chips 30 to 3j of different layers, that is, open the word line WL coupled to the hit storage unit in the hit layer of this read instruction, and at the same time open the word line WL coupled to the non-hit storage unit in the non-hit layer of this read instruction. Among them, the word line WL coupled to the non-hit storage unit and the hit storage unit shares the execution row address of this read instruction. After opening these word lines WL, read the data of the hit storage unit and the data of the non-hit storage unit (whose coupling is the opened word line). Perform error checking and correction (ECC) on the read data. Cache the data after error checking and correction corresponding to the hit storage unit into the corresponding read buffer 210, and cache the data after error checking and correction corresponding to the non-hit storage unit into the corresponding prefetch buffer 211. Take out the data cached in this read buffer 210 and return it to the main chip 1.
[0040] Optionally, the memory controller 2 determines the hit layer among the memory chips 30 to 3j according to the chip select address (CS, also known as the "chip select signal") parsed from the read instruction. The hit memory cells belong to this hit layer, and the non-hit memory cells belong to other layers outside the hit layer among the memory chips 30 to 3j. That is to say, the non-hit memory cells and the hit memory cells are located in different layers of memory chips among the memory chips 30 to 3j. When the data length read by the read instruction exceeds the data length stored in one memory cell, the hit memory cells can be multiple continuously distributed memory cells in the same layer of memory chips (i.e., the hit layer). The number of non-hit memory cells in each layer that are simultaneously read (i.e., prefetch data) is the same as the number of hit memory cells in the hit layer that are read. When parsing the read instruction, the memory controller 2 can parse out which specific storage control modules IP the read instruction hits, which layer of memory chips (i.e., the hit layer) the read instruction specifically hits or selects among the storage control modules IP managed by these hits, which one or several memory cells in the hit layer the read instruction specifically hits (i.e., the hit memory cells, and this information can also be obtained from the information of which storage control modules IP are hit), which row address (i.e., the execution row address, which is also which word line in the hit memory cells the read instruction specifically hits) and which column addresses (i.e., which one or several bit lines in the hit memory cells the read instruction specifically hits) in the hit memory cells the read instruction specifically hits.
[0041] It should be understood that a memory cell has several word lines WL, such as 16 * 1024 WL. Which specific word line WL is hit is obtained by the memory controller 2 through address decoding of the current read instruction to be executed. That is, if the execution row address (which is also the row address of the hit memory cell) of the read instruction decoded by the memory controller 2 is what (for example, WL_800), then the memory controller 2 will send this execution row address to the hit storage control module IP. The word lines opened by the memory cells managed by the hit storage control module IP are all addressed by this execution address. Thus, while reading the data of the hit memory cells in the hit layer of the read instruction, the chip select signal of the non-hit layer is also opened, and then the data of the non-hit memory cells whose word lines are opened in these non-hit layers are prefetched.
[0042] Optionally, if the memory controller 2 determines that the data read by the read instruction has been cached in the prefetch buffer 211 in advance, the memory controller 2 fetches the data corresponding to the read instruction from the prefetch buffer 211 and returns it to the main chip 1. At this time, the main chip 1 does not need to fetch data from the stacked memory chips 30 to 3j of layer j + 1, but directly fetches data from the prefetch buffer 211, thus saving the time for switching the word lines of the stacked memory chips 30 to 3j of layer j + 1, improving the access speed of the stacked memory chips of layer j + 1, and further increasing the bandwidth of the stacked memory chips of layer j + 1.
[0043] Optionally, the memory controller 2 is further configured to, when performing error checking and correction on the read data, if it is found that the stored data has an error and the error is corrected, write the corrected data back to the corresponding storage unit.
[0044] Wherein, when executing the read instruction, the memory controller 2 has the following two operations: (1) The first operation: fetch the data corresponding to the read instruction from the read buffer 210 or the prefetch buffer 211 and return it to the main chip 1; (2) The operation of writing the data after ECC error correction back to the corresponding storage unit is defined as the second operation; wherein, the first operation and the second operation can be performed in parallel, so there is no additional time consumption; or, the first operation and the second operation are performed serially, and the first operation is executed before the second operation, so there is an additional time for write after read, but this time is not very long and will not have a great impact on the system performance.
[0045] Please continue to refer to Figure 2 , the memory controller 2 of this embodiment includes m + 1 IO interfaces IO0 to IOm, a control and management module 21, and a plurality of memory control modules IP, where m is an integer and m ≥ 0 (that is, the memory controller 2 includes at least 1 IO interface), preferably m ≥ 1. At this time, the m + 1 IO interfaces are all parallel communication protocol interfaces supporting multi-IO.
[0046] Wherein, the data buffer 21a is disposed in the control and management module 21 and is correspondingly arranged and coupled to each of the IO interfaces IO0 to IOm, that is, the number of data buffers 21a in the control and management module 21 is also m + 1. Each data buffer 21a is coupled to n + 1 memory control modules IP0 to IPn. Thus, n + 1 memory control modules IP0 to IPn correspond to one IO interface, that is, when the bit width of the data read by a memory control module IP0 from the hit layer is x, the bit width of the data returned by this IO interface to the main chip 1 is equal to (n + 1) * x, where n is an integer and n ≥ 1.
[0047] Please combine Figure 2 and Figure 5, each storage control module IP0~IPn is coupled to the control management module 21 and the p+1 layer memory chips among the memory chips 30~3j, and is used to manage k+1 storage units U0~Uk in each layer of memory chips it is coupled to. p, k, and j are all integers and k≥1, p≥1, j≥1, that is, each storage control module manages (p+1)*(k+1) storage units. Among them, in the p+1 layer memory chips (such as even-layer memory chips or odd-layer memory chips) managed by each storage control module IP0~IPn, each storage unit U0 is coupled to the same through-silicon via hybrid bonding path, each storage unit U1 is coupled to the same through-silicon via hybrid bonding path, and so on. Each storage unit Uk is coupled to the same through-silicon via hybrid bonding path. Thus, the storage units on the same through-silicon via hybrid bonding path can share the corresponding execution row address.
[0048] Each storage control module IP0~IPn is also used to turn on the word line WL of the execution row address that shares the read instruction in the p+1 layer memory chips, so as to read the data of the hit storage unit in the hit layer of the read instruction, and at the same time read (or "prefetch") the data of the non-hit storage unit corresponding to the execution row address of the read instruction in the p non-hit layers, and perform error checking and correction (ECC) on the read data, and cache the error-checked and corrected data corresponding to the hit storage unit into the corresponding read buffer 210, and cache the error-checked and corrected data corresponding to the non-hit storage unit into the corresponding prefetch buffer 211.
[0049] Optionally, please refer to Figure 2 and Figure 5, each storage control module IP includes an error checking and correcting control module (i.e., ECC control module) 220. The error checking and correcting control module 220 includes p + 1 error checking and correcting control units ECC_ctrl0 to ECC_ctrlp. Each of the error checking and correcting control units ECC_ctrl0 to ECC_ctrlp is set and coupled one-to-one with the p + 1 layers of memory chips managed by the storage control module IP. Each of the error checking and correcting control units ECC_ctrl0 to ECC_ctrlp is used to perform error checking and correction on the data read from the memory chip to which it is coupled. For example, the storage control module IP0 manages the corresponding storage units in the even layers of the memory chips 30 to 3j. The error checking and correcting control unit ECC_ctrl0 is coupled to the memory chip 30 (which can also be referred to as the "first layer memory chip") and is used to perform error checking and correction on the data read from the memory chip 30. The error checking and correcting control unit ECC_ctrl1 is coupled to the memory chip 32 (which can also be referred to as the "third layer memory chip") and is used to perform error checking and correction on the data read from the memory chip 32, and so on. The error checking and correcting control unit ECC_ctrlp is coupled to the memory chip 3(2*p) (which can also be referred to as the "2p + 1 layer memory chip") and is used to perform error checking and correction on the data read from this memory chip. Wherein, when j is even, 2p = j; when j is odd, 2p = j - 1.
[0050] Further optionally, after each of the error checking and correcting control units ECC_ctrl0 to ECC_ctrlp finds a data error and corrects the error data, it writes the corrected data back to the corresponding storage unit.
[0051] The control management module 21 is used to parse the read instruction to obtain the hit layer (i.e., the hit or chip-selected memory chip) of the read instruction, the hit storage unit in the hit layer, and the execution row address of the hit storage unit, and determine whether the data read by the read instruction has been cached in the corresponding prefetch buffer 211 in advance.
[0052] In one example, in addition to including m + 1 data buffers 21a, the control and management module 21 further includes an address decoder 213 and a discriminator 214. The address decoder 213 is used to parse a read instruction to obtain the address of the memory control module IP hit by the read instruction (further, it may include the number x of memory cells hit), the layer address of the memory chip hit (i.e., the hit layer), and the row address (i.e., the execution row address of the hit memory cell) and column address of the memory cell. The discriminator 214 is used to determine, according to the parsing result of the address decoder 213, whether the data read by the read instruction has been cached in advance in the corresponding prefetch buffer 211. It should be noted that, in some embodiments, the address decoding logic of the address decoder 213 of the present invention for the read / write instruction address can be designed such that the logical addresses corresponding to the same row addresses of adjacent memory layers are continuous, that is, the address decoder 213 can be designed to map consecutive logical addresses sent by the main chip 1 to the same WL of adjacent layers (for example, after storing in WL_800 of the first-layer memory chip 30, it is then stored in WL_800 of the second-layer memory chip 31). In this way, the probability that the data prefetched and error-corrected from the non-hit layer (i.e., the data in the prefetch buffer 211) of the present invention is hit by subsequent read instructions is increased.
[0053] Each of the IO interfaces IO0 to IOm is used to communicate and connect with the main chip 1 and the control and management module 21, to implement interface conversion between the main chip 1 and the control and management module 21, and to receive instructions from the main chip 1 and data to be written into the memory chips 30 to 3j, as well as to return the read data to the main chip 1, etc. The IO interfaces IO0 to IOm can be any suitable parallel communication protocol interfaces supporting multiple IOs, such as an AXI (Advanced eXtensible Interface) interface, etc. Among them, the AXI interface is an on-chip bus interface of a master-slave architecture oriented to high performance, high bandwidth, and low latency. Its address, instruction, and data phases are separated, supporting unaligned data transmission. At the same time, in burst transmission, only the first address is required, and the separated read and write data channels support the transmission access and out-of-order access of a relatively large number of outstanding instructions to be executed (such as the number of unfinished transactions such as read and write instructions), and it is easier to perform timing convergence, which is suitable for high-speed memory access. It should be noted that although the AXI protocol is shown in the specification drawings, the present invention is not limited thereto, and the IO interfaces IO0 to IOm can also adopt any other suitable high-bandwidth interface protocols, such as the AHB (Advanced High-performance Bus) protocol or the CHI (Coherent Hub Interface) protocol, etc.
[0054] In this embodiment, each of the IO interfaces IO0 to IOm is used to receive a read instruction sent by the main chip 1, and according to the judgment result of the control and management module 21, take out the data corresponding to the read instruction from the read buffer 210 or the prefetch buffer 211 to which it is coupled, and return it to the main chip 1. For example, when the control and management module 21 determines that the data read by the read instruction has been pre-cached in the corresponding prefetch buffer 211 in advance, the corresponding IO interface takes out the corresponding data from the prefetch buffer 211 and transmits it to the main chip 1; when the control and management module 21 determines that the data read by the read instruction has not been pre-cached in the corresponding prefetch buffer 211 in advance, the corresponding storage control module IP reads out the data of the hit storage unit in the hit layer (i.e., the hit memory chip) of the read instruction, and after performing ECC error correction on the read data and caching it in the corresponding read buffer 210, the corresponding IO interface takes out the corresponding data from the read buffer 210 and transmits it to the main chip 1.
[0055] Among them, when a read instruction needs to read data from multiple storage units in the same layer (i.e., the hit layer), that is, when the data read by a read instruction corresponds to data continuously stored in multiple hit storage units (hitunit) in the same layer (i.e., the hit layer), the multiple storage control modules IP that manage the multiple hit storage units (hit unit) synchronously cache data into the read buffer 210 corresponding to the corresponding IO interface, so that the IO interface can take out the data read by the read instruction from the read buffer 210 at one time, and the main chip 1 can obtain the data read by the read instruction.
[0056] That is to say, the data bit width of each read instruction is equal to the data bit width returned by an IO interface to the main chip 1, and the data bit width of each IO interface is equal to the sum of the data bit widths of the corresponding n + 1 storage control module IPs cached in the read buffer 210 corresponding to the IO interface according to the read instruction. Among them, when a storage control module IP manages p + 1 layers of memory chips and manages k + 1 storage units U0 to Uk in each layer of memory chips, when an IO interface receives a read instruction, the read buffer 210 corresponding to the IO interface can cache the data of (n + 1) * (k + 1) storage units (i.e., hit storage units) hit by the read instruction (the data will be subjected to ECC error correction before being cached), and the prefetch buffer 211 corresponding to the IO interface can cache the data of (n + 1) * (k + 1) * p storage units (i.e., non-hit storage units) not hit by the read instruction (the data will also be subjected to ECC error correction before being cached). For example, when n + 1 = 4, k + 1 = 2, and p + 1 = 4, when an IO interface receives a read instruction, the read buffer 210 corresponding to the IO interface will cache the data of 4 * 2 = 8 units of the read instruction (i.e., cache the data to be read by the read instruction, and cache the data read from 1 word line WL of each of these 8 units and subjected to ECC error correction), and the prefetch buffer 211 corresponding to the IO interface will cache the data of 4 * 2 * 3 = 24 units of the read instruction (i.e., cache the prefetch data according to the read instruction, and cache the data read from 1 word line WL of each of these 24 units and subjected to ECC error correction), and the data returned by the IO interface to the main chip 1 is the data of these 8 units cached in the read buffer 210 (i.e., the data read from the 8 hit storage units of the read instruction and subjected to ECC error correction).
[0057] It should be understood that in addition to the above-mentioned IO interface, control and management module 21, and storage control module IP, the memory controller 2 of this embodiment may also have other control logic modules (not shown) for implementing other functions. For example, other control logic modules in the memory controller 2 may include circuits for clock and frequency control circuits (such as phase-locked loop PLL), circuits for managing power consumption and temperature, first-in first-out queue registers (FIFO), and so on for any required circuits. These circuits are not the focus of the present invention, so they will not be elaborated here.
[0058] In addition, the division and internal composition of each module such as the IO interfaces IO0 to IOm, the control and management module 21, and the storage control module IP in the memory controller 2 of this embodiment are illustrative and mainly a logical function division. There may be other division methods in actual implementation. Each functional module in the embodiments of the present application may be integrated in a processing module, or each module may exist physically alone, or two or more modules may be integrated in a physical module. The above integrated modules may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as an SSD), etc.
[0059] Based on the same inventive concept, please refer to Figures 2 to 5 , an embodiment of the present invention further provides a three-dimensional stacked memory, which includes j + 1 stacked memory chips 30 to 3j and the memory controller 2 as described in this embodiment.
[0060] To better understand the technical solutions of the memory controller 2 and the three-dimensional stacked memory of the present invention, taking j = 7, p = 3, m = 3, k = 2, n = 4 and each layer of memory chips being DRAM as an example, a detailed description will be given.
[0061] Please refer to Figure 6 and Figure 7, the memory controller 2 of this example manages 8-layer memory chips 30 to 37, which are internally provided with 4 parallel communication IO interfaces IO0 to IO3, 4 data buffer areas 21a, and 16 memory control modules. Each data buffer area 21a is coupled to 4 corresponding memory control modules IP0 to IP3 among them. Each memory control module IP manages 4-layer memory chips and manages two memory cells U0 and U1 in each layer of memory chips. One read instruction can access 8 continuously distributed memory cells in the same layer of memory chips through an IO interface and the control management module 21 (that is, one read instruction will hit the data of 8 hit memory cells located in the same layer of memory chips, and these 8 memory cells are managed by the corresponding 4 memory control modules IP0 to IP3). Therefore, the ECC control module 220 inside each memory control module IP includes 4 error checking and correction control units ECC_ctrl0 to ECC_ctrl3. Taking Figure 6 the memory control module IP0 of the first data buffer area in
[0062] as an example, it manages memory chips 30, memory chip 32, memory chip 34, and memory chip 36. The error checking and correction control unit ECC_ctrl0 is used to perform error checking and correction on the data read from memory chip 30, the error checking and correction control unit ECC_ctrl1 is used to perform error checking and correction on the data read from memory chip 32, the error checking and correction control unit ECC_ctrl2 is used to perform error checking and correction on the data read from memory chip 34, and the error checking and correction control unit ECC_ctrl3 is used to perform error checking and correction on the data read from memory chip 36.
[0063] From the perspective of the storage control module IP, a storage control module IP manages four layers of memory chips (DRAM wafer, DRAM die). A read instruction will only hit one layer of memory chips, and specifically hits one word line WL corresponding to one of the two storage units in each layer. The remaining three layers of memory chips are in the non-hit state. For example, if the execution row address of the read instruction is WL_800, the hit storage control module IP includes IP0. The memory chip 30 managed by IP0 is the hit layer for this read instruction and is used to complete the task of outputting the data read by this read instruction. The memory chips 32, 34, and 36 managed by IP0 are the non-hit layers for this read instruction and are used to provide corresponding prefetch data. The memory chips 30, 32, 34, and 36 managed by IP0 simultaneously open the word line WL corresponding to the execution address WL_800. The data of the two storage units U0 and U1 in the memory chip 30 corresponding to WL_800 and hit by this read instruction (that is, the two hit storage units managed by IP0, which means the number x of the storage units managed by IP0 hit by the read instruction is 2) are read out. After being corrected by ECC (that is, through the error checking and correction of ECC_ctrl0 in IP0), they are cached in the corresponding read buffer 210. The corresponding data of the six storage units U0 and U1 in the memory chips 32, 34, and 36 managed by IP0 corresponding to WL_800 are read out (the column address range of these six storage units managed by IP0 read is the same as the column address range of the two hit storage units U0 and U1 in the memory chip 30). After being corrected by ECC of ECC_ctrl1~ECC_ctrl3 in IP0 respectively, they are cached in the corresponding prefetch buffer 211.
[0064] From the perspective of the control management module 21 (or the memory controller 2), eight consecutive storage units on the same layer of memory chips may simultaneously open one word line WL to cooperate with the data output. That is, the maximum reading range of this read instruction can be on eight storage units on the same layer of memory chips of the four storage control modules IP0~IP3 corresponding to one data buffer. Each storage unit only opens one word line WL to output data.
[0065] In this example, when a read instruction is received through the IO interface IO0 and data is read from these eight layers of memory chips, the data of the hit layer is read and the data of the non-hit layer is prefetched. ECC error correction, caching, and error-corrected write-back are performed on both the read data and the prefetched data. The specific steps are as follows: (1) Decode the read instruction to obtain the address of the hit memory control module IP (i.e., the address of IP0), the layer address of the memory chip (assuming it is memory chip 30), the hit unit, and its execution row address WLcounter (which can also be denoted as "WL number", assuming WL number = WL_800). In a further embodiment, the decoding further obtains the number x of hit memory units (assuming Figure 6 8 memory units of the same layer memory chip 30 corresponding to the first data buffer 21a in Figure 6 . The minimum number x of hit memory units can be 1, and the maximum can be 8).
[0066] (2) Determine whether the data to be read has been pre-cached in Figure 6 the prefetch buffer 211 of the first data buffer 21a in Figure 6 . If so, take the data from the prefetch buffer 211 of this data buffer 21a and send it to the IO interface IO0, and then return it to the main chip 1, thereby improving the access speed; if not, execute the following step (3).
[0067] (3) Open 1 word line WL of each of the 8 memory units of memory chips 30, 32, 34, and 36. These opened word lines WL share the execution row address WL_800 (i.e., they are all addressed through this execution row address WL_800). Therefore, from the perspective of the IO interface IO0, 4 * 8 = 32 word lines WL are opened without additional time consumption.
[0068] (4) The 32 word lines WL sharing the execution row address WL_800 output data simultaneously (i.e., 4 * x memory units output data simultaneously) without additional time consumption.
[0069] (5) The data data of memory chip 30 is correspondingly sent to Figure 6 ECC_ctrl0 in the memory control modules IP0 - IP3 coupled to the first data buffer 21a in Figure 6 for error correction; the data data of memory chip 32 is correspondingly sent to Figure 6 ECC_ctrl1 in the memory control modules IP0 - IP3 coupled to the first data buffer 21a in Figure 6 for error correction; the data data of memory chip 34 is correspondingly sent to Figure 6 ECC_ctrl2 in the memory control modules IP0 - IP3 coupled to the first data buffer 21a in Figure 6 for error correction; the data data of memory chip 36 is correspondingly sent to Figure 6The ECC_ctrl3 in the storage control modules IP0-IP3 coupled to the first data buffer 21a performs error correction. The storage control modules IP0-IP3 simultaneously perform ECC error correction on the data read from the four-layer memory chips 30, 32, 34, and 36 without wasting extra time.
[0070] (6) The ECC-corrected data corresponding to the memory chip 30 is cached to the storage control module IP0~IP3. Figure 6 The data after ECC error correction corresponding to the memory chips 32, 34, 36 are all cached to the storage control module IP0~IP3. Figure 6 The first data buffer 21a of the pre-fetch buffer 211 is stored in the memory control module. The four paths IP0 to IP3 send data to the memory control module at the same time. Figure 6 The first data buffer area 21a is read without wasting extra time.
[0071] (7) Figure 6 The data cached in the read cache area 210 of the first data cache area 21a is sent to the IO interface IO0, and then sent to the main chip 1 through the IO interface IO0.
[0072] (8) If data errors are found in the data corresponding to memory chips 32, 34, and 36 during ECC error correction, they are written back to the storage unit of the corresponding memory chip after ECC error correction. Steps (7) and (8) can be executed in parallel. If they are executed in parallel, no additional time is consumed. If steps (7) and (8) are executed serially, an additional read-after-write time is added, but this time is not very long and will not affect performance.
[0073] In summary, the memory controller and three-dimensional stacked memory of the present invention set a read cache area and a pre-fetch cache area in the memory controller. Each time data is read, it is first determined whether the data to be read has been cached in the pre-fetch cache area in advance. If not, while reading the corresponding data in the hit layer in the multi-layer memory chip (that is, the data stored in the hit storage unit), the corresponding data in at least one non-hit layer in the multi-layer memory chip is also read (that is, pre-fetched), and error checking and correction (ECC) is further performed on the read and pre-fetched data. After the error checking and correction, the data in the hit layer is cached in the read cache area, the pre-fetched data is cached in the pre-fetch cache area, and the data cached in the read cache area this time is taken out and returned to the main chip, thereby realizing the function of pre-fetching data of the multi-layer memory chip, performing ECC error correction on the read and pre-fetched data first, and then caching them separately, thereby improving the data reliability and system performance of the memory.
[0074] Optionally, each time data is read, if it is determined that the data to be read has been cached in the prefetch buffer in advance, the data can be fetched directly from the prefetch buffer instead of fetching data from the stacked multi-layer memory chips, thereby saving the time of switching the word lines of the stacked multi-layer memory chips, improving the access speed of the stacked multi-layer memory chips, and further increasing the bandwidth of the stacked multi-layer memory chips.
[0075] In addition, if a data storage error is found through the error checking and correction (ECC) function, after correcting the error, the corrected data is written back to the stacked multi-layer memory chips, thereby improving the reliability of the stored data without affecting the performance of the stacked multi-layer memory chips.
[0076] The above description is only a description of the preferred embodiments of the present invention and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the art of the present invention based on the above disclosure fall within the protection scope of the technical solution of the present invention.
Claims
1. A memory controller is coupled between a main chip and a stacked multi-layer memory chip, and each layer of the memory chip includes a plurality of memory cells, characterized in that, The memory controller has a data buffer, and the data buffer includes a read buffer and a prefetch buffer. The memory controller is configured to: Receive and parse a corresponding read instruction, and determine whether the data read by the read instruction has been pre-cached in the prefetch buffer. If not, then perform the following steps: Open the word lines sharing the execution row address of the read instruction in different layers of the memory chips to read the data of the corresponding hit memory cells and the data of the non-hit memory cells in the multi-layer memory chips; Perform error checking and correction on the read data; Cache the error-checked and corrected data corresponding to the hit memory cells into the read buffer, and cache the error-checked and corrected data corresponding to the non-hit memory cells into the prefetch buffer; Take out the data cached in the read buffer and return it to the main chip.
2. The memory controller according to claim 1, wherein The memory controller is configured to: determine a hit layer in the multi-layer memory chips according to the chip select address of the read instruction, wherein the hit memory cells belong to the hit layer, and the non-hit memory cells belong to at least one non-hit layer other than the hit layer in the multi-layer memory chips; and / or The memory controller is further configured to: if it is determined that the data read by the read instruction has been pre-cached in the prefetch buffer, take out the data corresponding to the read instruction from the prefetch buffer and return it to the main chip.
3. The memory controller according to claim 2, wherein, The memory controller further includes at least one IO interface, a control and management module, and a plurality of memory control modules. The data buffer is disposed in the control and management module, and is correspondingly disposed and coupled to the IO interface; each of the memory control modules is coupled to the control and management module and the multi-layer memory chips; wherein: The control and management module is used to parse the read instruction to obtain the hit layer, the hit memory cells in the hit layer, and the execution row address of the hit memory cells, and determine whether the data read by the read instruction has been pre-cached in the prefetch buffer; The memory control module is used to open the word lines sharing the execution row address in different layers of the memory chips including the hit layer to read the data of the hit memory cells and the data of the non-hit memory cells, perform error checking and correction on the read data, and cache the error-checked and corrected data corresponding to the hit memory cells into the corresponding read buffer, and cache the error-checked and corrected data corresponding to the non-hit memory cells into the corresponding prefetch buffer; The IO interface is used to receive the read instruction, and according to the judgment result of the control and management module, take out the data corresponding to the read instruction from the corresponding read buffer or prefetch buffer and return it to the main chip.
4. The memory controller according to claim 3, wherein The control and management module further includes: An address decoder, which is used to parse the read instruction to obtain the hit layer of the read instruction, the hit memory cells in the hit layer, and the execution row address in the hit memory cells. A discriminator, configured to determine whether the data read by the read instruction has been pre-cached in the prefetch buffer according to the parsing result of the address decoder.
5. The memory controller according to claim 3, characterized in that, The storage control module includes an error checking and correcting control module, which includes a plurality of error checking and correcting control units. Each of the error checking and correcting control units is correspondingly arranged and coupled to each layer of the memory chips managed by the storage control module. The error checking and correcting control unit is configured to perform error checking and correction on the data read from the memory chip to which it is coupled.
6. The memory controller according to claim 3, wherein The data read by the read instruction corresponds to the data continuously stored in a plurality of the hit storage units in the hit layer. A plurality of the storage control modules managing the plurality of the hit storage units synchronously cache the data into the read buffer corresponding to the IO interface, so that the main chip can obtain the data read by the read instruction.
7. The memory controller according to claim 3, wherein Each of the storage control modules manages a plurality of the storage units on the same through-silicon via hybrid bonding path in the multi-layer memory chips. Each of the storage units includes a plurality of word lines corresponding to a plurality of row addresses and a plurality of bit lines corresponding to a plurality of column addresses. Wherein, the control management module parses out the address of the storage control module hit by the read instruction, the layer address of the memory chip in the hit layer, and the execution row address and column address in the hit storage unit from the read instruction.
8. The memory controller according to claim 3, wherein The memory chip is a DRAM chip, and the IO interface is an AXI interface or a CHI interface or an AHB interface.
9. The memory controller according to claim 3, wherein The storage control module or the memory controller is further configured to, when performing error checking and correction on the read data, if a data error is found and error correction is performed, write the corrected data back to the corresponding storage unit.
10. The memory controller according to claim 9, wherein The operation of taking out the data corresponding to the read instruction from the read buffer or the prefetch buffer and returning it to the main chip is defined as the first operation; the operation of writing the corrected data back to the corresponding storage unit is defined as the second operation. The first operation and the second operation are performed in parallel, or the first operation and the second operation are performed serially, and the first operation is executed prior to the second operation.
11. The memory controller according to any one of claims 1-10, characterized in that, The memory controller is disposed in the buffer chip, and the buffer chip and the multi-layer memory chips are stacked together; or the memory controller is disposed in the main chip, and the main chip and the multi-layer memory chips are stacked together.
12. A three-dimensional stacked memory, characterized in that, It includes stacked multi-layer memory chips and the memory controller according to any one of claims 1-11.
Citation Information
Patent Citations
Motion compensating module pixel prefetching device in AVS video hardware decoder
CN101022551A
ECC control circuits, multi-channel memory systems and operation methods thereof
CN101373449A
Method for reading data from flash memory, related memory controller and storage system
CN118642647A
Data reading method of memory, memory controller and related equipment
CN119376632A
Memory controller and three-dimensional stacked memory
CN120183461A