Method and Apparatus for Bad Block Management and Repair in Memory System
Patent Information
- Application Number
- KR1020240051244
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-04-17
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2044-04-17
Smart Images

Figure 112024042091130-PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present disclosure relates to a memory system utilizing a bad block management and repair method and apparatus. Background Technology
[0002] The following description merely provides background information related to the present invention and does not constitute prior art.
[0003] Since its first appearance in the 1970s, Dynamic Random Access Memory (DRAM) has evolved by improving three key performance aspects: capacity, bandwidth, and power. However, by the 2010s, DRAM vendors faced technical limitations in simultaneously improving DRAM performance while maintaining low bit costs. Consequently, methods were attempted to improve the core performance of the memory system based on the overall optimization of the computer system, rather than improving individual DRAM components. A representative method is to configure the memory system into a two-level structure consisting of near and far memory. In this method, "hot data" is stored in near-side memory, which operates at high speeds despite having low capacity, while "cold data," which is accessed relatively infrequently, is stored in far-side memory, which offers larger capacity and lower cost, although its performance is somewhat lower.
[0004] Furthermore, various methods have been proposed in the form of Processing in Memory (PIM) or Near Data Processing (NDP) to intelligently maintain DRAM health and perform simple yet critical computational functions in-memory or near-memory. These approaches are an inevitable result of the trend in which the improvement of DRAM performance metrics—such as capacity, bandwidth, and power—is redefined not merely as a problem for DRAM alone, but as a collective issue for the entire computing system, including the CPU, Memory Controller (MC), memory buffer (e.g., Registering Clock Driver (RCD)), and DRAM. In particular, for High Bandwidth Memory (HBM), which has recently garnered significant attention, methods are being proposed to incorporate computational functions into the base die located at the bottom layer, alongside multiple DRAM dies stacked on top. In addition, from the perspective of JEDEC (Joint Electron Device Engineering Council) standardization, various RAS (Reliability, Availability, Serviceability) functions have already been built into the base die of HBM to facilitate testing of memory BIST (Built-in Self Test) or KGD (Known Good Die) after assembly. Consequently, standard or non-standard RAS functions are being built into DRAM (in-DRAM) or near DRAM, advancing to a stage where the reliability of the entire memory system is improved, costs are lowered, and characteristics are optimized.
[0005] Meanwhile, if a DIMM (Dual In-line Memory Module) possesses such intelligent RAS capabilities, the DIMM can be used in applications such as a CXL (Computer Express Link) memory pool. By resolving consistency / coherency issues during the implementation of a 2-level system, various types of memory, such as HBM and CXL memory, which possess their own unique advantages in addition to existing memory modules, have been commercialized. Methods for distinguishing hot / cold data from a software perspective have been proposed, and memory systems with multi-level hierarchies extending beyond 2 levels are being experimented with.
[0006] For example, the following memory hierarchy can be considered: (Level 1) HBM, a low-capacity, high-performance memory; (Level 2) DIMM, a medium-capacity, medium-performance memory; and (Level 3) CXL memory, a high-capacity, low-cost memory. Depending on the implementation, the memory subsystem can be implemented in various combinations such as (Level 1)-(Level 2)-(Level 3), (Level 1)-(Level 2), and (Level 2)-(Level 3), and the most optimized capacity, bandwidth, and power performance can be achieved depending on the nature of the workload. In this memory hierarchy, CXL memory is fundamentally far-side memory and stores cold data, thus requiring a low bit cost.
[0007] CXL memory is implemented as either 1) a product with a memory card form factor, or ii) a board-type product containing multiple existing DRAM modules. For CXL memory, which occupies the lowest level in terms of memory hierarchy, to be widely adopted, its price must be low. In the case of the first product, the price of CXL memory can be lowered by using DRAM with a lower bit cost for CXL memory. In this case, the DRAM used for CXL memory can be manufactured using a more refined process. This is because CXL memory can operate at a slightly slower speed compared to DRAM used in DIMMs and can utilize technologies such as enhanced ECC (Error Correction Code). In the case of the second product, implementing low-cost CXL memory is not easy. However, if low-cost "used DIMMs" are recycled for CXL memory instead of new DIMMs, low-cost CXL memory can be realized. Used DIMMs may contain bad blocks in which some cells are dead. If bad blocks can be effectively managed, low-cost CXL memory can be implemented by utilizing DIMMs that would otherwise be discarded, and the price of CXL memory can be drastically reduced.
[0008] Generally, Bad Block Management (BBM) technology corresponds to availability in the Reliability Assistance System (RAS) term. In this context, availability refers to the ability to proceed with normal operation at reduced performance levels despite the presence of errors.
[0009] The memory that most widely uses such BBM is NAND flash. Since NAND cells are destroyed based on the number of programming cycles, it utilizes wear-leveling, a technology that manages destroyed cells. In addition to simply disabling bad blocks, wear-leveling prevents excessive access to specific physical cells by adjusting the physical-logical address mapping. Along with wear-leveling technology, NAND flash can employ high-performance ECC, such as Low Density Parity Check (LDPC) code. This is because NAND flash is not a memory system dependent on the computer system, but rather an independent storage system guaranteed autonomy, allowing for sufficient margin in performance aspects such as required latency.
[0010] However, it is rare for DRAM to actively utilize BBM. This is fundamentally because DRAM cells possess superior endurance characteristics compared to NAND Flash cells. For instance, DRAM cells must not have any durability issues throughout their warranty period (usually 10 years). For instance, the lifespan of a server is about 5 years. Servers that have exceeded 5 years are generally discarded because, due to various issues, they spend more time turned off than actually operating. Aging servers are referred to as "comatose servers." Since DIMMs are discarded or donated along with the server before the actual warranty period expires, it is uncommon for DRAM to perform BBM.
[0011] As an example, there exists a technology for performing BBM within DRAM (see Patent Document 1). A CPU or MC performs BBM (i.e., processing some defective blocks in the DRAM cell array so that they cannot be accessed) to determine the reduced capacity in each DRAM. The CPU or MC can determine the total capacity of the module by using the minimum capacity of all DRAMs mounted within the same module. Consider a Registered Dual In-line Memory Module (RDIMM) equipped with 40 16Gb DRAMs (32 excluding the DRAM for ECC), as shown in the example of FIG. 13. If the results of BBM performed on each DRAM are distributed from DRAMs where all cells are alive to DRAMs where up to 256 Mb are processed as bad blocks, the total capacity of the RDIMM is (16G - 256M) × 4 bytes. The technology for performing BBM within individual DRAMs can have a clear advantage in terms of usable capacity. However, in order to process BBM for each DRAM, the BBM function must be pre-loaded within the DRAM; since there is no standardized protocol regarding the loading of the BBM function, the inclusion of the BBM function cannot be expected in most commercial (commodity) DRAM products.
[0012] As another example, there is a technique for performing BBM on an RCD existing on a module containing DRAMs (see Patent Document 2). The technique for performing BBM on an RCD performs BBM on an RCD, which is a memory buffer chip on the module. In this case, the RCD provides a contiguous address space to the MC using address remapping, as shown in the example of FIG. 14. In the example of FIG. 14, it is assumed that the DIMM contains 8 DRAMs, and each DRAM has 16 blocks ranging from hexadecimal (0 = 0000) to (F = 1111). As shown in the example on the right side of FIG. 14, a contiguous address space is generated according to a simple process of converting a logical address to a physical address. After performing BBM by the RCD, the MC can access cells where the upper row address (corresponding to the block address) of the corresponding module is 0000 to 1010, and recognizes the address space from 1011 to 1111 as an inaccessible bad block. The MC provides the upper row address 4 bits from 0000 to 1010. The RCD converts the input address received from the MC into an actual physical address. For example, if "block address = 1010" is input, the RCD outputs "block address = 1111" to the DRAM. As mentioned above, providing a contiguous address space is very important in BBM. If a contiguous address space is not provided, the MC must avoid bad blocks in a complex manner.
[0013] Meanwhile, the case where each DRAM processes BBM and performs address remapping to provide a "contiguous address space" can be represented as shown in the example of FIG. 15. Comparing the example of FIG. 14 with FIG. 15, the technology for processing BBM in the RCD has a relatively small total capacity of DIMM. When processing BBM in the RCD, if at least one of the blocks at a specific address in the eight DRAMs is a bad block, the blocks at the corresponding address in the entire DIMM are considered bad blocks. In the DIMM exemplified in FIG. 14, address remapping is performed after five blocks are set as bad blocks. On the other hand, in the example of FIG. 15, since BBM and address remapping are performed within each DRAM, the address space corresponding to the bad blocks in each DRAM is not random but is aligned contiguously, so only two blocks in the DIMM are set as bad blocks. In conclusion, when processing BBM in a memory buffer (e.g., RCD), if there is only one "BBM controller," the reduction in memory capacity due to BBM can be significant.
[0014] Meanwhile, repairing defective blocks or sub-blocks that occur after BBM may be considered. Repair technology corresponds to serviceability in terms of RAS. In this case, serviceability refers to the ability to eliminate errors occurring in the product and restore the product's health. To apply repair, redundant rows, columns, or blocks must be prepared to replace the defective rows, columns, or blocks. As an example, DRAM repair can be performed in an RCD (Patent Document 3). That is, through cooperation between a DRAM manufacturer and an RCD manufacturer, if the DRAM manufacturer provides additional redundant columns accessible from the outside, the RCD can replace the defective columns using the provided redundant columns. However, since DRAM generally does not provide information about additional redundant cells, the RCD cannot utilize the redundant cells.
[0015] Therefore, when performing BBM in RCD, measures to address the reduction in overall module capacity and methods to effectively repair defects occurring after the application of BBM must be considered. Prior art literature
[0016] Patent Document 1: U.S. Patent No. US 9824755 B2 (SEMICONDUTOR MEMORY DEVICE AND MEMORY SYSTEM INCLUDING THE SAME). Patent Document 2: U.S. Patent No. US 9318168 B2 (MEMORY SYSTEM FOR CONTINUOUSLY MAPPING ADDRESSES OF MEMORY MODULE HAVING DEFECTIVE LOCATIONS). Patent Document 3: U.S. Patent No. US 9087614 B2 (MEMORY MODULES AND MEMORY SYSTEMS). The problem to be solved
[0017] The present disclosure has a primary objective to provide a method and apparatus for performing Bad Block Management (BBM) using a plurality of BBM controllers and supporting a contiguous address space based on address remapping in a memory system (e.g., DIMM, HBM) comprising a plurality of DRAMs and memory buffers (e.g., RCD of DIMM, base die of HBM), wherein part or all of the DRAMs contain defective cells and bad blocks containing defective cells are set according to a memory test.
[0018] Furthermore, the main objective of the present disclosure is to provide a method and apparatus for repairing defective cells in block or sub-block units based on healthy cells detected during the BBM process, using a repair controller within the memory buffer, in a memory system comprising a plurality of DRAMs and a memory buffer, when defective cells progressively occur additionally after BBM execution. means of solving the problem
[0019] According to an embodiment of the present disclosure, a memory system is provided comprising a memory module including a memory controller (C); a plurality of memories; and a buffering device located between the MC and the plurality of memories, wherein the memory module further comprises a plurality of bad block management (BBM) controllers corresponding to a plurality of memory regions divided from the plurality of memories, and converts a continuous logical address provided from the MC into a physical address provided to the memories using the plurality of BBM controllers.
[0020] According to another embodiment of the present disclosure, the memory module further comprises a repair controller, wherein the repair controller identifies a normal block or a normal subblock within a physical address range corresponding to an area processed by the bad blocks, and repairs the bad cells using the normal block or the normal subblock when additional bad cells occur in an area not processed by the bad blocks after the execution of the BBM.
[0021] According to another embodiment of the present disclosure, a method for managing and repairing bad blocks performed by a memory module comprises: performing Bad Block Management (BBM) on bad blocks in a plurality of memories using a plurality of Bad Block Management (BBM) controllers, wherein the plurality of BBM controllers each correspond to a plurality of memory regions partitioned from the plurality of memories; and, when a continuous logical address is provided from an MC, converting the logical address to a physical address provided to the memories using the plurality of BBM controllers, wherein the MC and the memory module constitute a memory system, and the memory module includes the plurality of memories and a buffering device located between the MC and the plurality of memories.
[0022] According to another embodiment of the present disclosure, a method is provided comprising: a step of finding a normal block or a normal subblock within a physical address range corresponding to an area processed by the bad blocks using a repair controller; and a step of repairing the bad cells based on the normal block or the normal subblock using the repair controller when additional bad cells occur in an area not processed by the bad blocks after the execution of the BBM. Effects of the invention
[0023] As described above, according to the present embodiment, in a memory system (e.g., DIMM, HBM) comprising a plurality of DRAMs and memory buffers (e.g., RCD of DIMM, base die of HBM), when some or all of the DRAMs contain defective cells and bad blocks containing defective cells are set according to memory test, by performing BBM using a plurality of BBM controllers and supporting a contiguous address space based on address remapping, it is possible to minimize the capacity of the memory system reduced by BBM and improve the availability of the memory system.
[0024] In addition, according to the present embodiment, in a memory system comprising a plurality of DRAMs and a memory buffer, when defective cells progressively occur after BBM, by providing a method and apparatus for repairing defective cells in block or sub-block units based on healthy cells detected during the BBM process using a repair controller within the memory buffer, it is possible to overcome the occurrence of defective cells after BBM and improve the serviceability of the memory system. Brief explanation of the drawing
[0025] FIG. 1 is a block diagram showing a memory system according to one embodiment of the present disclosure. Figure 2 is an example diagram showing a C / A (Command & Address) connection considering groups and channels in a memory module. FIG. 3 is an exemplary diagram showing the execution of BBM by a group-based BBM (Bad Block Management) controller according to one embodiment of the present disclosure. FIG. 4 is an exemplary diagram showing BBM execution by a group / rank BBM controller according to one embodiment of the present disclosure. FIG. 5 is an exemplary diagram showing the securing of resources for BBM repair according to one embodiment of the present disclosure. FIG. 6 is an exemplary diagram showing a case in which BBM repair is mounted in a DRAM (Synamic Random Access Memory) according to one embodiment of the present disclosure. FIG. 7 is an illustrative diagram showing a remapping technique according to one embodiment of the present disclosure. FIG. 8 is an exemplary diagram showing an offset table according to one embodiment of the present disclosure. FIG. 9 is a repair according to one embodiment of the present disclosure This is an example diagram showing a controller. FIG. 10 is a block diagram illustrating a memory system including a BBM and a repair function after BBM according to another embodiment of the present disclosure. FIG. 11 is an exemplary diagram showing a High Bandwidth Memory (HBM) including a BBM and a repair function after BBM according to one embodiment of the present disclosure. FIG. 12 is a flowchart illustrating a method for managing and repairing bad blocks according to one embodiment of the present disclosure. Figure 13 is an example diagram showing the configuration of a memory module and C / A (Command & Address) connection. Figure 14 is an example diagram showing the execution of BBM by a memory buffer (e.g., RCD). Figure 15 is an example diagram showing the execution of BBM by DRAM. Specific details for implementing the invention
[0026] Hereinafter, embodiments of the present invention will be described in detail with reference to the exemplary drawings. It should be noted that in assigning reference numerals to the components of each drawing, the same components are given the same reference numeral whenever possible, even if they are shown in different drawings. Furthermore, in describing these embodiments, if it is determined that a detailed description of related known components or functions could obscure the essence of these embodiments, such detailed description is omitted.
[0027] In addition, terms such as first, second, A, B, (a), (b), etc. may be used when describing the components of the embodiments. These terms are intended only to distinguish the components from other components, and the essence, order, or sequence of the components is not limited by these terms. Throughout the specification, when a part is described as 'comprising' or 'equipped' with a certain component, unless specifically stated otherwise, this means that it may include additional components rather than excluding other components. Furthermore, terms such as '…part' or 'module' described in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware, software, or a combination of hardware and software.
[0028] The detailed description disclosed below, together with the attached drawings, is intended to describe exemplary embodiments of the present invention and is not intended to represent the only embodiment in which the present invention can be practiced.
[0029] The present embodiment discloses a memory system comprising a memory buffer-based bad block management and repair method and apparatus. More specifically, in a memory system comprising an MC, a plurality of DRAMs, and a memory buffer (e.g., an RCD of a DIMM, a base die of an HBM), when part or all of the DRAMs contain defective cells and bad blocks containing defective cells are set according to a memory test, a method and apparatus are provided to perform Bad Block Management (BBM) using a plurality of BBM controllers and to support a contiguous address space based on address remapping.
[0030] In addition, the present embodiment provides a method and apparatus for repairing defective cells in block or sub-block units based on healthy cells detected during the BBM process using a repair controller in the memory buffer when defective cells progressively occur after BBM.
[0031] Hereinafter, BBM and subsequent repairs according to the present disclosure are described with reference to a memory buffer (e.g., RCD) that is essential to a server-oriented DIMM. The following descriptions are also applicable to HBMs that include a base die as a memory buffer.
[0032] Below, terms and assumptions related to DRAM size are explained from the perspective of RCD based on DDR5 (Double Data Rate 5).
[0033] RA (Row Address) represents the row address of a bank, which is a unit of a collection of memory cells included in DRAM.
[0034] CA(Column Address) represents the column address of the bank.
[0035] BA (Bank Address) represents the address of a bank included in BG (Bank Group).
[0036] BG includes multiple banks.
[0037] A DRAM die includes multiple BGs.
[0038] A Logical Rank includes multiple DRAM dies and is represented by a CID (Chip Identifier) or DID (Die ID). A CID is an identifier for each chip in a package where DRAM chips are stacked in a 3DS (3-Dimensional Stack) form using Through Silicon Via (TSV). A DID is an identifier for each die when two dies are packaged together without using TSV, and up to two dies can be packaged.
[0039] DRAM packages typically include logical ranks.
[0040] Rank is a unit that includes multiple DRAM chips and produces an 8-byte (or 9-byte) output.
[0041] From the perspective of the RCD, a channel typically contains two ranks. Each rank is accessed at different times.
[0042] A DIMM (Dual In-line Memory Module) typically contains two channels. Since the operation of most DRAMs (i.e., the control of the RCD) is performed independently between channels, in the case of DDR5, two identical control structures exist within the RCD.
[0043] Meanwhile, based on DDR4, the DIMM contains one channel from the perspective of the RCD.
[0044] Below, commands related to the operation of DRAM are briefly explained.
[0045] The Activate command is a command to open the corresponding row by passing RA for one bank to the DRAM side.
[0046] The Precharge command is a command to fill (charge) all rows within a bank with a preset value. Since the Precharge command is executed before the Read command for the new row, it also serves the role of closing the previous row.
[0047] The Refresh command is used to read and rewrite data from all cells within the DRAM at regular intervals to prevent charge loss.
[0048] The Read command is a command to read data from a cell by passing CA to an open row within a bank.
[0049] The Write command is a command to write data to a cell by passing CA to an open row within a bank.
[0050] The following DIMM, RDIMM, LRDIMM (Load Reduced DIMM), and memory modules are used compatiblely.
[0051] FIG. 1 is a block diagram showing a memory system according to one embodiment of the present disclosure.
[0052] In the example of FIG. 1, the memory system includes a memory module (100) and a memory controller (MC, 150). The memory module (100) includes DRAMs which are memory (110) and an RCD which is a memory buffer (120), and is connected to the MC (150) or the CPU. The memory module exemplified in FIG. 1 may be an RDIMM or an LRDIMM. The RCD receives a Command & Address (C / A) and a Clock (CLK) from the MC (150), regenerates them into a clean signal with noise removed, and transmits them to the DRAM. That is, the RCD improves the signal integrity (SI) of the overall memory system, thereby enhancing bandwidth characteristics. In the example of FIG. 1, the buffer function of the RCD as described above corresponds to 'Normal Functions'.
[0053] As illustrated in the example of FIG. 1, the RCD includes a bad block management device and a repair device according to the present disclosure. That is, the RCD includes a plurality of BBM controllers (122) and a repair controller (124). The plurality of BBM controllers (122) perform Bad Block Management (BBM) and support a contiguous address space based on address remapping. By addressing the disadvantage that the BBM technology within the RCD (120) can reduce the capacity of the memory system, the plurality of BBM controllers (122) can improve the availability of the memory system. In this case, the bad blocks may be set by memory testing. BBM manages memory capacity considering the bad blocks by restricting access to each bad block. The plurality of BBM controllers (122) perform this BBM and support a contiguous logical address space provided by the MC based on address remapping.
[0054] Meanwhile, if the use of a DRAM containing bad blocks continues until the DRAM approaches or exceeds its product warranty period, there is a high probability that additional defective cells will occur even after the application of BBM. The repair controller (124) monitors the health of each DRAM periodically (Periodic or Patrol) or as needed (Demand or Non-periodic). If additional defective cells occur, a memory test may be performed again, for example. Depending on the additional setting of bad blocks resulting from the additional memory test, the capacity of the memory system may be further reduced. As another example, the capacity of the memory system may be maintained by repairing the defective cells instead of performing the additional memory test. The repair controller (124) can improve the serviceability of the memory system by performing dynamic repair on a block or sub-block basis. Hereinafter, the repair function performed by the repair controller (124) is referred to as "Post-BBM Repair (PBBMR)."
[0055] In the example of FIG. 1, one channel includes two groups (Group A, Group B) or two ranks (Rank 0, Rank 1). In the example of FIG. 1, k, n, m, m1, and m2 each represent the number of bits of the corresponding connection. In DQ, D represents input data and Q represents output data. DCA_A and DCA_B represent the Command & Address (C / A) transmitted from MC (150) to Channel A and Channel B for Channel A and Channel B. As outputs of the RCD, QACA_A and QACA_B represent the C / A representing Group A and Group B of Channel A, and QACA_B and QACA_B represent the C / A representing Group A and Group B of Channel B. The reason for using groups will be described later.
[0056] Below, a BBM controller (122) that resolves the disadvantage that BBM technology within an RCD can reduce the capacity of the memory system is described considering the structure of a DDR5 RDIMM. A repair controller (124) that performs PBBMR when additional defective cells occur in the RDIMM after applying BBM is described. Additionally, a method for the repair controller (124) to provide healthy cells to replace the defective cells when performing PBBMR is described.
[0057] Figure 2 is an example diagram showing a C / A (Command & Address) connection considering groups and channels in a memory module.
[0058] FIG. 2 is an example of the illustration of FIG. 13 re-represented by considering only the connection state, regardless of whether the DRAM chip is located at the front or back of the DIMM. As in the example of FIG. 2, the DDR5 RDIMM includes DRAMs and a memory buffer (120, e.g., RCD) that constitute a memory module (100). In the example of FIG. 2, the memory module (100) includes 40 DRAMs (32 excluding the DRAM for ECC).
[0059] A DDR5 RDIMM includes two channels. Each channel is distinguished by a different command decoder. In the example of FIG. 13, the DRAMs on the left form Channel A, and the DRAMs on the right form Channel B. DCA[6:0]_A and DCA[6:0]_B represent different C / A that are transferred from MC (150) to Channel A and Channel B.
[0060] Each channel includes two groups. In this case, each group has a spatially different connection. For example, as shown in the example of FIG. 13, DRAMs located facing each other with respect to the substrate can form the same group. The reason for using groups will be described later.
[0061] Additionally, each channel includes two ranks. In this case, each rank is accessed at different times. As exemplified in FIG. 2, DRAMs included in a group may have different ranks. A Chip Select (CS) signal may be used for accessing each rank.
[0062] As previously mentioned, the DDR5 RCD includes two channels. Since the channels generally have separate C / A decoders, a BBM controller that provides a contiguous address space with an address remapping function must also exist for each channel. The left 20 DRAMs are included in Channel A, and the right 20 DRAMs are included in Channel B. In this disclosure, a BBM controller is provided for each channel to perform BBM separately for the left 20 DRAMs and the right 20 DRAMs.
[0063] The present disclosure improves the availability of a memory system by providing a plurality of BBM controllers (122) on the same channel. , Improvements in availability are described by comparing the example of FIG. 14, where one BBM controller (122) exists in the RCD for one channel, with the example of FIG. 3, where two BBM controllers (122) exist for two groups in which one channel is divided.
[0064] In the example of Fig. 2, since the DDR5 RCD applies 1:2 registering, the 10 Group A DRAMs and 10 Group B DRAMs on the DRAM side have separate output C / A lines. That is, the RCD splits the input C / A information into Group A and Group B for output. This is intended to improve signal integrity during high-speed operation by reducing line loading. Since different C / A lines are used, when the same logical block address is input, the two C / A lines can output different physical block addresses. QACA[13:0]_A and QBCA[13:0]_A represent the physical addresses of Group A and Group B in Channel A, and QACA[13:0]_B and QBCA[13:0]_B represent the physical addresses of Group A and Group B in Channel B.
[0065] In the present disclosure, a BBM controller (122) is provided separately for each group. By comparing the example of FIG. 3 with the example of FIG. 14, an increase in availability due to the use of the group-specific BBM controller (122) can be confirmed. That is, the final capacity of the memory module (100) is improved from 11 blocks of Block 0 to Block A (0000 to 1010) as in the example of FIG. 14 to 12 blocks of Block 0 to Block B (0000 to 1011) as in the example of FIG. 3.
[0066] In the example of FIG. 2, five of the DRAMs belonging to the same group correspond to Rank 0, and the remaining five correspond to Rank 1. Since Rank 0 and Rank 1 are never accessed simultaneously in time, a BBM controller (122) may be provided separately for Rank 0 and Rank 1. As in the example of FIG. 4, the final capacity of the DIMM depends on the smallest case among the address spaces aligned by the four BBM controllers (122). In the example of FIG. 4, the output of the BBM controller (122) located in the lower right becomes the total capacity of the DIMM. In the example of FIG. 4, the outputs of the BBM controllers (122) are compared two at a time. As another example, four outputs of the BBM controllers (122) may be compared at once.
[0067] As described above, the present disclosure provides a configuration that maximizes the total capacity of a DIMM by positioning a plurality of BBM controllers (122) within a memory buffer (120, e.g., RCD). For example, if the BBM controllers (122) are positioned by group, the DDR5 RCD includes two BBM controllers (122) per channel. Additionally, if positioned by group and by Rank, there are four (=2×2) BBM controllers (122) per channel. Since the DDR5 memory module includes two channels, the DDR5 RCD can be equipped with up to eight (=2×2×2) BBM controllers (122). As another example, the DDR4 memory module includes one channel and two groups. Since the DDR4 memory module includes up to four Ranks, the DDR4 RCD can be equipped with up to eight (=1×2×4) BBM controllers (122). By using the aforementioned multiple BBM controllers (122), the present disclosure can maximize the available capacity of the DIMM.
[0068] According to the example of FIG. 2, the foregoing describes a case where there are two channels, two groups, and two ranks, but is not necessarily limited thereto. Based on the design of the memory module, a plurality of BBM controllers (122) may be used based on a preset number of channels, a preset number of groups, and a preset number of ranks. For example, a plurality of BBM controllers (122) corresponding to "a preset number of channels × a preset number of groups × a preset number of ranks" may be used. Among the capacities of the plurality of BBM controllers (122), the minimum capacity value becomes the total capacity value of the memory module.
[0069] As mentioned above, DRAMs requiring BBM application may be those that are close to or have exceeded their product warranty period. Therefore, even if the DIMM is recycled after reducing its total capacity using BBM (from an availability perspective), the number of failed cells may increase further with continued use. Since the circuits constituting the cell or cell array (e.g., DRAM Cells, Sub-Wordline Driver, Bit-Line Sense Amplifier, Row Decoder, Column Decoder, etc.) may age and cause failures, the occurrence of failed cells is a natural phenomenon. That is, "cell failures" where cells die irregularly, "row failures" where one or more rows die, "column failures" where one or more columns die, or "block failures" where an entire block fails may occur. Additionally, if a block containing defective cell(s) occurs, the defective cell(s) can be resolved, for example, by re-running the memory test and further reducing the DIMM capacity. However, as another example, if repair is possible, it is not necessary to reduce the DIMM capacity, so repair can be the best solution. Repair is a function that fixes the defects that have occurred and is a solution from a serviceability perspective among the RAS functions. It is important to note that a serviceability approach to eliminate defects using repair is not possible before performing BBM, but repair becomes possible after performing BBM.
[0070] For repair, healthy redundant cell(s) are essential to replace defective cell(s). These resources do not exist prior to BBM execution. Generally, DRAM inherently carries redundant row(s) or column(s) to replace defective row(s), column(s), or cell(s), but DIMMs do not have a way to access the mounted redundant DRAM cell(s). However, during the BBM process, 'healthy redundant cell(s) to replace defective cell(s)' can naturally occur as follows.
[0071] In the following, Method 1 utilizes healthy blocks corresponding to the difference between the total capacity of the DIMM and the capacity of the single DRAM as repair resources when one of the DRAMs corresponding to "a preset number of channels × a preset number of groups × a preset number of ranks" has a capacity greater than the total capacity of the DIMM. Method 2 utilizes sub-blocks without defective cells as repair resources for row defects or random cell defects when the bad blocks are not block defects or thermal defects.
[0072] (Method 1) As shown in the example of FIG. 3, a case is described where a BBM controller (122) is provided separately for each group. After performing alignment to secure a continuous address space, in the case of Group A, there are two top bad blocks, and in the case of Group B, there are four top bad blocks. Since the DIMM cannot have different address space sizes for each group, the total capacity of the DIMM consists of 12 blocks, Block 0 to Block B (0000 to 1011), excluding the four bad blocks. That is, the capacity of the DIMM is reduced to 3 / 4. In Group A, there are Block D and Block C, which are healthy but treated as bad blocks. Assume a case where a progressive failure occurs in Block 3 of DRAM2 included in Group A after performing BBM. When Block 3, which is a defective block in Group A, is accessed, the repair controller (124) can perform a repair by remapping Block 3 to Block C. In the example of FIG. 4, BBM is performed by group and by rank as described above. Rank 0 of the A-side (i.e., Group A) secures Block D / E / F as healthy cell(s) for repair, and Rank 1 secures Block D / E. For example, if a defect occurs in Block C at a physical address in one of the DRAMs corresponding to A-side Rank 0, the repair controller (124) can perform repair by remapping Block C to one of Block D / E / F.
[0073] (Method 2) As shown in the example of FIG. 5, row defects or random cell defects may occur in a part of a block processed as a bad block, rather than in the entire block. In this case, healthy cell(s) to be used for repair can be secured in block units such as 1 / 2, 1 / 4, 1 / 8, ... For example, if row defects or random cell defects occur in the upper half region as shown in the example of FIG. 5, addresses where the bit immediately below the block address 4 bit is 0 represent healthy subblocks. In the Block E exemplified in FIG. 5, addresses (1110-1X...X) are bad subblocks and addresses (1110-0X...X) are healthy subblocks. That is, if the repair granularity is set to half the block size, the repair controller (124) can perform repair by examining the block address 4 bits and the lower row address 1 bit. For example, if the upper sub-block (0010-1X...X) in Block 2 of DRAM2 exemplified in FIG. 5 is defective, the repair controller (124) can replace the defective sub-block with the lower sub-block (1110-0X...X) of Block E. The repair granularity can be further reduced by increasing the number of bits to be compared. For example, by comparing the block address 4 bits and the lower row address 2 bits, the repair controller (124) can perform repair in units of 1 / 4 of the block.
[0074] In the above-described (Method 1), the entire healthy block can be used for block-unit repair. Alternatively, the repair subdivision can be reduced so that units such as 1 / 2, 1 / 4, 1 / 8, ... are used for repair, thereby improving repair efficiency. Since the repair subdivision increases when the repair subdivision is reduced, the repair subdivision can be appropriately determined according to the yield improvement effect.
[0075] The reason it is easy to apply a repair scheme after BBM is that the resources of healthy surplus cells required for repair can be naturally secured during the BBM process. Therefore, by equipping a repair controller (124) that performs a Post-BBM Repair (PBBMR) function in addition to the BBM controller (122), the memory system can provide a means of serviceability to continue service against additional degradation phenomena of the DRAM to which BBM has been applied. With the application of BBM and PBBMR according to the present disclosure, reliability can be improved when using low-cost used DIMMs to which BBM has been applied in CXL Memory. In addition, by lowering the price of CXL Memory, investment costs for memory infrastructure in data centers and clouds can be reduced. The present disclosure can contribute significantly to expanding the market for CXL Memory based on low-cost used DIMMs in the future, as well as to expanding the reach of CXL Memory. In addition, by recycling DIMMs that would otherwise be discarded, the present disclosure can also be helpful from an environmental conservation perspective. The present disclosure addresses practical issues regarding availability and service availability that must be resolved in the actual application of BBM.
[0076] Meanwhile, the method of repairing row(s) defects that occur after BBM using healthy row(s) secured during the BBM execution time according to the above-described (Method 1) and (Method 2) can be performed not only in the buffering block between MC and DRAM (memory buffer of DIMM (e.g., RCD), base die of HBM, etc.), but can also be used inside DRAM. That is, the repair controller (124) can be included inside DRAM. Since (Method 2) secures repair resources from a block processed as a bad block, it can naturally be performed inside DRAM. If there are multiple BBM controllers (122) for each bank or bank group within the DRAM, (Method 1) can also be performed within the DRAM. It is not difficult to control the DRAM capacity differently for each bank or bank group, but the decrease in efficiency relative to the increase in complexity is obvious. That is, if there is a method of i) storing a physical address to be used as a repair resource among the blocks processed as bad blocks within the DRAM, and ii) a means of remapping the physical bad address to the physical address of a healthy repair resource within the stored bad block address space when additional failures occur among blocks not processed as bad blocks, then PBBMR can be implemented within the DRAM.
[0077] As an example, PBBMR can be implemented in DRAM as shown in the example of FIG. 6. In the example of FIG. 6, the DRAM includes two banks, and each bank is provided with a BBM controller (122). Since there are two bad blocks in Bank 0 and three bad blocks in Bank 1, the top three blocks are treated as bad blocks for the chip as a whole. Therefore, Block 6 of Bank 0 can be used as a PBBMR resource. Block 7 of Bank 0, Block 5 and Block 1 of Bank 1 can be used as PBBMR resources because, although they are bad blocks at the block level, the internal sub-blocks are healthy.
[0078] FIG. 7 is an exemplary diagram illustrating a remapping technique according to one embodiment of the present disclosure.
[0079] The BBM controller (122) includes a Logical-to-Physical (L / P) address converter (710, hereinafter referred to as 'address converter') that performs remapping to generate a physical address space from a logical address, as shown in the example of FIG. 7. To perform remapping, the L / P address converter (710) includes an offset selector (712). In the example of FIG. 7, the BBM controller (122) is provided by rank. Accordingly, rank information along with BLA_from_MC (block address) is input to the offset selector (712) from the Command & Address (C / A) transmitted from the MC (150). Generally, the block address is the upper row address. The size of the block is determined by the number of cells per bit line, the prefetch size in column access, etc., among the elements constituting the DRAM array, and may vary depending on the DRAM. For example, the total number of row addresses is 17 and within the block 2K(=2 11 If there are ) rows, the block addresses are at most 6 (=17-11). In this case, multiple physical blocks can be grouped together to perform BBM in logical block units. In the example of FIG. 7, there are 4 block addresses. When an upper row address corresponding to a block address is input, the offset selector (712) generates different offset values depending on how many bad blocks there are in an address space smaller than the block containing the row address. In the example of FIG. 7, the offsets for Block 3 to Block 6 are '1'.
[0080] FIG. 8 is an exemplary diagram showing an offset table according to one embodiment of the present disclosure.
[0081] As shown in the example of FIG. 8, the offset selector (712) may use an offset table containing offset values. The offset selector (712) determines the offset value by comparing the reference column of the table with the input address. For example, if the input block address is 2 (=0010), the input block address is greater than the reference value (=1) in the first row and less than or equal to the reference value (=5) in the second row, i.e., "1 < input block address (=2) ≤ 5", so the offset is determined to be '+1' in the second row. If all blocks are healthy, such as Rank 2, the value in the first row of the reference column is set to the maximum value F. Since no input value can be greater than F (=1111), address remapping is not performed. That is, for all input blocks, the offset becomes 0 (=0000).
[0082] FIG. 9 is a repair according to one embodiment of the present disclosure This is an example diagram showing a controller.
[0083] The repair controller (124) performs "post-BBM repair" as in the example illustrated in FIG. 9. The repair controller (124) according to the present disclosure includes an address matcher (912) and a mux (914). The repair controller (124) utilizes a Non-volatile Memory (NVM, 916) as a storage device for storing repair information (e.g., a bad address). The repair controller (124) replaces the output of the L / P address converter (710). In the example of FIG. 9, the address matcher (912) compares the 'q+r bits', which is the sum of the block address 'q bits' and the added lower row address 'r bits', with the bad address 'q+r bits' stored in the NVM (916) or volatile register within the RCD. If the two addresses are identical, the address matcher (912) outputs a 'match' signal as high (=1) and outputs a healthy address 'q+r bit' to be accessed, replacing the faulty address. If the 'match' signal, which is the control signal of the MUX (914), is '1', the upper q bit of the output of the address matcher (912) is output as a physical address, replacing the output of the L / P address converter (710) exemplified in FIG. 7. Additionally, as the lower r bit of the block address according to the repair subdivision, the address r bit provided by the address matcher (912) is output, replacing the address input from the outside. For example, as in the example of FIG. 5, if the repair subdivision is 1 / 2 block, r=1.
[0084] In the example of FIG. 9, the NVM (916) can be implemented as a One-time Programmable (OTP) memory or an E-Fuse. Alternatively, a Code Word (CW) register can be used instead of the NVM (916). In the example of FIG. 9, s represents the number of bits of the address transmitted from the MC (150), and q represents the block address as the upper bits.
[0085] Meanwhile, memory testing to set bad blocks can be performed during an offline testing phase in which defective DIMMs are tested and shipped as normal DIMMs with reduced capacity. For offline testing, dedicated testing equipment may be used. Alternatively, memory testing can be performed online according to the RAS algorithm while mounted on a server.
[0086] In the BBM process, the offset table exemplified in FIG. 8 is stored in the NVM within the memory buffer (120) or in the NVM within the SPD (Serial Presence Detect)-Hub on the DIMM. If the NVM is not present in the memory buffer (120), the offset table may be read from the NVM within the memory module (100) by the MC (150) upon power-up and then written to a register (e.g., CW register) within the memory buffer (120).
[0087] To perform BBM and PBBMR, the following information, along with the offset value, must also be stored in the NVM on the RCD or DIMM in the same manner as above. If the NVM is not present in the memory buffer (120), the information stored in the NVM on the DIMM can be read by the MC (150) upon power-up and then written to a register in the memory buffer (120).
[0088] Information is stored indicating healthy blocks that are treated as bad blocks on a DIMM-wide basis by referencing defective cells, but are free of defects in the corresponding group or rank. For example, the maximum capacity is stored per channel / group / rank.
[0089] The address representing the bad block, that is, the bad address, is stored.
[0090] In addition, a master bit (e.g., 1 bit) indicating whether a healthy subblock exists in a bad block, and the upper r bit immediately below the block address q bit that distinguishes the healthy subblock are stored. For example, if the repair level is 1 / 2 block, additional information of r=1 bit must be prepared for each bad block so that PBBMR can proceed.
[0091] FIG. 10 is a block diagram illustrating a memory system including a BBM and a repair function after BBM according to another embodiment of the present disclosure.
[0092] To carry out the aforementioned processes, control by the MC (150) is required. As shown in the example of FIG. 10, the memory system may be configured to include control by the MC (150). The MC (150) or CPU includes a BB manager (152) that processes BBM in real time and a repair manager (154) that performs PBBMR.
[0093] In the case of CXL memory, the BB manager (152) and the repair manager (154) exist within the CXL memory buffer containing the MC that converts PCIe (Peripheral Component Interconnect Express) signals into DDR signals. By utilizing the BB manager (152) and the repair manager (154), the availability and serviceability levels of the CXL memory can be improved, thereby enhancing the stability of the entire system and reducing memory infrastructure costs.
[0094] Additionally, the MC (150) or CPU may further include an ECC severity checker (158) and a memory BIST (156) for checking ECC severity. The MC (150) or CPU may perform memory tests in an on-line state using the memory BIST (156). The MC (150) or CPU performs ECC using the ECC severity checker (158) and determines the severity of errors detected by ECC, i.e., the ECC severity.
[0095] The criteria for determining whether to perform PBBMR are, in principle, the same as the criteria for determining whether to perform the existing PPR (Post-Package Repair). Depending on how many errors are detected by ECC and whether the health of the memory system can be maintained by correcting errors through scrubbing and rewriting the corrected values to the DRAM cells, it may be determined whether to perform repair on the corresponding block or sub-block. For example, based on the ECC severity, the MC (150) can perform repair using the repair manager (154). At this time, the MC (150) can determine the repair of defective cells based on the ECC severity after interpreting the ECC severity using the ECC severity verifier (158).
[0096] Meanwhile, scrubbing is applied to the DRAM in response to error correction by ECC. When a cell error is detected by ECC, the MC (150) outputs a corrected value. That is, if a single bit error occurs in a DRAM cell, the MC (150) outputs a value corrected by ECC. Scrubbing is a process of rewriting single bit errors remaining in the DRAM cell. Scrubbing, which corrects errors in the cell itself, can be performed periodically or irregularly as needed.
[0097] Meanwhile, the example of FIG. 10 is based on DDR5 RDIMM. That is, the memory module (100) includes DRAMs and a memory buffer (120, e.g., RCD). The memory module (100) includes an SPD-hub including an NVM or an EEPROM (Electrically Erasable Programmable Read-Only Memory) which is an NVM. As previously mentioned, the NVM includes the capacity and repair information of the DIMM. As previously mentioned, the DRAMs may be classified according to channels, groups, or ranks. In the example of FIG. 10, the memory module (100) includes two channels, two groups, and two ranks.
[0098] FIG. 11 is a block diagram showing an HBM including a BBM and a repair function after the BBM, according to another embodiment of the present disclosure.
[0099] The embodiment according to the present disclosure as described above can also be applied to HBM. HBM is in the form of a chip implemented in a single package, but functionally it may correspond to a memory module. Hereinafter, HBM is considered as a type of memory module. As shown in the example of FIG. 11, HBM and RDIMM corresponding to memory modules (100) can be compared by function. The DRAM dies stacked on top of the HBM have the same role as the DRAM chips on the RDIMM, and the base die of the lowest layer of the HBM corresponds to a memory buffer (120) and has a function similar to the RCD of the DIMM. As shown in the example of FIG. 11, the dies stacked on top and the base die are connected by TSV. Not only are internal DRAM functions uploaded to the base die of the HBM, but the functions of the MC (150) are also offloaded. If the embodiment according to the present disclosure is applied to HBM, a BBM controller (122) and a repair controller (124) are positioned on the base die of the HBM, and the base die can perform BBM and PBBMR in the same way as a DIMM. For example, the base die performs BBM to control the stacked DRAMs to have the same capacity and a contiguous address space. Additionally, the base die can repair defects that may occur after BBM by using healthy surplus blocks or subblocks obtained during the BBM process. The HBM has a built-in memory BIST and a built-in function to report ECC severity to the CPU or MC (150). Therefore, the base die of the HBM may additionally include an ECC severity checker (158) and a memory BIST (156). When the CPU or MC (150) issues a simple command to perform BBM and PBBMR, the base die may execute the command itself. For example, the base die can perform a memory test online using memory BIST (156).The base die performs ECC using an ECC severity checker (156) and determines the severity of the error detected by the ECC, i.e., the ECC severity. Additionally, the base die determines the repair of defective cells based on the ECC severity.
[0100] Similar to HBM, in the example of FIG. 11, functions embedded in the MC (150) or CPU can, in principle, be offloaded to a memory buffer (120, e.g., RCD) on a DIMM. However, in the case of, for example, an RCD, due to constraints such as chip size, whether to place some functions within the RCD or in the MC (150) or CPU comes down to optimization.
[0101] Meanwhile, a memory buffer (e.g., RCD) equipped with the intelligent RAS functions (Reliability, Availability, Serviceability) described above is expected to play a key role in improving the performance of future memory systems. Additionally, if computational functions are integrated within the memory buffer, further performance enhancement of the memory system can be expected.
[0102] Accordingly, the present disclosure can improve the reliability of a memory system by incorporating an intelligent RAS function into components such as a memory buffer (e.g., RCD) on a DIMM or a base die of HBM. Additionally, the present disclosure can be applied to the implementation of low-cost CXL memory. The implementation of CXL memory according to the present disclosure can promote the formation of a CXL memory market that is still in its nascent stage.
[0103] The embodiment according to the present disclosure is based on DIMMs equipped with server-oriented DDR4 / DDR5, but as described above, it can also be applied to CXL memory, HBM, etc. In the case of HBM, since various test infrastructures such as memory BIST are embedded in the bottom-most base die, the role of the MC (150) or CPU can be minimized when performing BBM and PBBMR. When the BBM and PBBMR provided in the present disclosure are applied to HBM, a significant effect is expected in terms of RAS functionality. Since HBM is an expensive memory and can be assembled together with a CPU or GPU (Graphic Processing Unit), a failure of HBM can result in the disposal of the entire chip, including the memory, CPU / GPU, etc. Even if bad blocks occur within the HBM, if the HBM operates normally with reduced capacity, that is, reduced performance due to BBM, the availability of the HBM can be improved. Furthermore, even if additional failures occur after BBM, repairs can be performed using PBBMR instead of re-performing BBM to further reduce capacity (i.e., further reduce performance). Maintaining HBM performance based on repairs improves serviceability in the entire computing system and can intelligently enhance the reliability of memory systems using HBM. In the case of CXL memory, BBM and PBBMR allow for the configuration of low-cost products using "used DIMMs." In contrast, in the case of HBM, BBM and PBBMR can continue to function as a reduced-capacity memory in situations where a failure occurs in the high-cost HBM. Additionally, by maintaining the reduced capacity for a longer period based on repairs, PBBMR can increase the lifespan of computing devices equipped with HBM. That is, by employing the BBM and PBBMR technologies according to the present disclosure to improve reliability, "low-cost CXL memory" and "long-lasting high-cost HBM" can be realized.
[0104] Hereinafter, using the example of FIG. 12, a method for managing and repairing bad blocks performed by a memory module is described.
[0105] FIG. 12 is a flowchart illustrating a method for managing and repairing bad blocks according to one embodiment of the present disclosure.
[0106] The memory system includes an MC and a memory module. The memory module includes a plurality of memories and a buffering device located between the MC and the plurality of memories.
[0107] The memory module performs BBM using multiple BBM controllers to restrict access to bad blocks within multiple memories (S1200). Here, the multiple BBM controllers correspond to multiple memory regions partitioned from the multiple memories.
[0108] Multiple BBM controllers determine the capacity of each memory region based on the BBM for the memory region allocated to each BBM controller, and determine the available address range of the memories based on the memory region with the minimum capacity among the memory regions.
[0109] Multiple BBM controllers are included in a buffering device, and each BBM controller is allocated to memory regions that are separated temporally and spatially. Alternatively, multiple BBM controllers are included in each memory, and each BBM controller is allocated to cell regions within each memory according to a channel, bank, or bank group.
[0110] If the memory module is a DRAM module, the buffering device is a memory buffer on the DRAM module. Alternatively, if the memory module is HBM, the buffering device may be a base die within the HBM.
[0111] When a series of logical addresses are provided from the MC, the memory module converts the logical addresses into physical addresses provided to the memories using multiple BBM controllers (S1202).
[0112] The memory module uses a repair controller to find a normal block or normal subblock within a physical address range corresponding to an area processed with bad blocks (S1204).
[0113] The repair controller is included in the buffering device. Alternatively, the repair controller may be included in each memory.
[0114] If additional defective cells occur in the area that was not processed as bad blocks after the execution of BBM, the memory module repairs the defective cells based on a normal block or normal sub-block using a repair controller (S1206).
[0115] The repair controller utilizes blocks as repair resources that are included in the address ranges corresponding to bad blocks in multiple memories but do not contain defective cells in the memory area allocated to each BBM controller.
[0116] Meanwhile, as shown in the example of FIG. 13, the DDR5 RDIMM is implemented in a form in which DRAMs and memory buffers (120, e.g., RCD) are attached to the front and back of a substrate constituting the memory module (100). In the example of FIG. 13, the memory module (100) includes 40 DRAMs (32 excluding DRAMs for ECC). The memory module (100) includes temperature sensors TS00 and TS01. The memory module (100) includes an SPD-hub that controls the chips within the memory module (100), and a PMIC (Power Management IC) that controls the power. For example, the SPD-hub controls chips such as the memory buffer, temperature sensor, and PMIC.
[0117] A DDR5 RDIMM includes two channels. Each channel is distinguished by a different command decoder. In the example of FIG. 13, the DRAMs on the left form Channel A, and the DRAMs on the right form Channel B. DCA[6:0]_A and DCA[6:0]_B represent C / A transferred from MC (150) to Channel A and Channel B.
[0118] Each channel includes two groups. In this case, each group has a spatially different connection. In the example of FIG. 13, DRAMs located opposite each other with respect to the substrate form the same group. As previously mentioned, C / A lines are used to output different physical block addresses for each group to reduce line loading. QACA[13:0]_A and QBCA[13:0]_A represent the physical addresses of Group A and Group B in Channel A, and QACA[13:0]_B and QBCA[13:0]_B represent the physical addresses of Group A and Group B in Channel B.
[0119] Although not shown in the example of FIG. 13, each channel includes two ranks. In this case, each rank is accessed differently in time. For example, as exemplified in FIG. 2, DRAMs included in a single group also have different ranks.
[0120] Although the flowcharts and timing diagrams in this specification describe each process as being executed sequentially, this is merely an illustrative explanation of the technical concept of one embodiment of the present disclosure. In other words, a person skilled in the art to which one embodiment of the present disclosure belongs may modify and adapt the flowcharts and timing diagrams in various ways, such as changing the order described in the flowcharts and timing diagrams or executing one or more of the processes in parallel, without departing from the essential characteristics of one embodiment of the present disclosure; therefore, the flowcharts and timing diagrams are not limited to a chronological order.
[0121] The above description is merely an illustrative explanation of the technical concept of the present embodiment, and a person skilled in the art to which the present embodiment belongs would be able to make various modifications and variations within the scope of the essential characteristics of the present embodiment. Accordingly, the present embodiments are intended to explain, not limit, the technical concept of the present embodiment, and the scope of the technical concept of the present embodiment is not limited by these embodiments. The scope of protection of the present embodiment shall be interpreted by the claims below, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of rights of the present embodiment. Explanation of the symbols
[0122] 100: Memory module 110: Memory 120: Memory buffer 122: BBM Controller 124: Repair Controller 150: MC or CPU 710: L / P Address Converter 712: Offset Selector 912: Address Matcher
Claims
Claim 1 A memory system comprising a memory module including a memory controller (MC); a plurality of memories; and a buffering device located between the MC and the plurality of memories, wherein the memory module further comprises a plurality of Bad Block Management (BBM) controllers that each correspond to a plurality of memory regions divided from the plurality of memories and control bad blocks within the plurality of memories, and using the plurality of BBM controllers to convert a continuous logical address provided from the MC into a physical address provided to the memories, wherein the plurality of memories are Dynamic Random Access Memory (DRAM), and the plurality of BBM controllers determine the capacity of each memory region based on the BBM for a memory region allocated to each BBM controller, and determine the available address range of the memories based on the memory region having the minimum capacity among the memory regions. Claim 2 delete Claim 3 A memory system according to claim 1, wherein the plurality of BBM controllers are included in the buffering device, and each BBM controller is allocated to memory regions separated temporally and spatially. Claim 4 A memory system according to claim 1, wherein the plurality of BBM controllers are included in each memory, and each BBM controller is allocated to a cell area within each memory according to a channel, bank, or bank group. Claim 5 A memory system according to claim 1, wherein the memory module further comprises a repair controller, and the repair controller identifies a normal block or a normal subblock within a physical address range corresponding to an area processed by the bad blocks, and repairs the bad cells using the normal block or the normal subblock when additional bad cells occur in an area not processed by the bad blocks after the execution of the BBM. Claim 6 In paragraph 5, a memory system in which the repair controller is included in the buffering device. Claim 7 In paragraph 5, a memory system in which the repair controller is included in each memory. Claim 8 A memory system according to claim 1, wherein, when the memory module is a DRAM (Dynamic Random Access Memory) module, the buffering device is a memory buffer on the DRAM module. Claim 9 A memory system according to claim 1, wherein, when the memory module is HBM (High Bandwidth Memory), the buffering device is a base die within the HBM. Claim 10 A memory system according to claim 1, wherein the number of the plurality of BBM controllers is determined based on the number of different instruction decoders. Claim 11 A memory system according to claim 1, wherein the number of the plurality of BBM controllers is determined based on the number of cases having spatially different connections. Claim 12 A memory system according to claim 1, wherein the number of the plurality of BBM controllers is determined based on the number of cases of temporally different accesses. Claim 13 A memory system according to claim 1, wherein the number of the plurality of BBM controllers is determined based on a combination of all or part of the number of different instruction decoders, the number of cases with spatially different connections, and the number of cases with temporally different accesses. Claim 14 In paragraph 5, the repair controller utilizes as a repair resource a block that is included in the address range corresponding to the bad blocks in the plurality of memories but is free of the bad cells in the memory area allocated to each BBM controller. Claim 15 In claim 14, the above repair controller utilizes the block created by dividing the block without the defective cells based on a multiple of 1 or 2 as the repair resource, in a memory system. Claim 16 In paragraph 5, the repair controller utilizes all or part of the subblocks as repair resources when the subblocks included in the bad blocks are healthy. Claim 17 In claim 16, the above repair resource is created by dividing a block including the above subblock based on a multiple of 2, and the multiple of 2 is 2 or greater, in a memory system. Claim 18 A memory system according to claim 5, further comprising a Non-volatile Memory (NVM) that stores repair information used by the repair controller. Claim 19 In paragraph 18, a memory system in which the above NVM is included in the above buffering device. Claim 20 A memory system according to claim 18, wherein the NVM exists in the memory module in place of the buffering device, and the MC accesses the NVM and transmits the repair information to the buffering device. Claim 21 In claim 1, the above bad blocks are set by an offline memory test of the memory module, in a memory system. Claim 22 In claim 1, the above bad blocks are set by an online memory test of the memory module, in a memory system. Claim 23 In paragraph 22, the above bad blocks are a memory system configured by the memory BIST (Built-in Self Test) within the MC when the memory module is undergoing the online memory test. Claim 24 In paragraph 22, the above bad blocks are a memory system configured by a memory BIST in the buffering device when the memory module is undergoing the online memory test. Claim 25 In paragraph 5, the repair controller repairs the defective cells when the memory module is online, in a memory system. Claim 26 In paragraph 25, the repair controller repairs the defective cells based on the ECC (Error Correction Code) severity, in a memory system. Claim 27 In claim 26, the above buffering device is a memory system that determines the repair of the above defective cells by interpreting the above ECC severity. Claim 28 In paragraph 26, the above MC is a memory system that determines the repair of the above defective cells by interpreting the above ECC severity. Claim 29 A method for managing and repairing bad blocks performed by a memory module, comprising the step of performing Bad Block Management (BBM) on bad blocks within a plurality of memories using a plurality of Bad Block Management (BBM) controllers, wherein the plurality of BBM controllers each correspond to a plurality of memory regions divided from the plurality of memories; and, when a continuous logical address is provided from an MC, the step of converting the logical address into a physical address provided to the memories using the plurality of BBM controllers, wherein the MC and the memory module constitute a memory system, and the memory module includes the plurality of memories and a buffering device located between the MC and the plurality of memories, and the plurality of memories are Dynamic Random Access Memory (DRAM), and the plurality of BBM controllers determine the capacity of each memory region based on the BBM for a memory region allocated to each BBM controller, and determine the available address range of the memories based on the memory region having the minimum capacity among the memory regions. Claim 30 A method according to claim 29, further comprising: a step of finding a normal block or a normal subblock within a physical address range corresponding to an area processed by the bad blocks using a repair controller; and, if additional defective cells occur in an area not processed by the bad blocks after the execution of the BBM, a step of repairing the defective cells based on the normal block or the normal subblock using the repair controller.
Citation Information
Patent Citations
Memory system and operating method of memory system
KR1020230036680A
Memory device for outputting test results
KR1020230086553A
Non-volatile memory storage system with two-stage controller architecture
US20100017556A1
Reusing partial bad blocks in NAND memory
US20150187442A1