Memory module, memory repair method, controller, medium and system
By configuring a controller in the memory module for error correction and read operation verification, the system identifies and utilizes redundant storage units to repair permanently damaged units in the memory chip, thus solving the problem of frequent memory chip failures during use and improving the reliability and resource utilization of data storage.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING SUPERSTRING ACAD OF MEMORY TECH
- Filing Date
- 2025-12-16
- Publication Date
- 2026-07-23
AI Technical Summary
Failures in memory chips during use, especially in applications like data centers where data storage reliability is critical, are difficult to detect and replace with faulty or unreliable memory cells during integrated circuit manufacturing due to limitations in current technology. This results in a high incidence of failures after the chips leave the factory.
A memory module is provided, including a controller and memory chips. The controller is configured to perform error correction on read data, determine the erroneous address, actively initiate a read operation, verify the data to identify permanently damaged failed memory cells, and repair them using redundant memory cells.
It improves the reliability of memory modules, promptly identifies and repairs permanently damaged and failed storage units, and enhances storage resource utilization and data storage reliability.
Smart Images

Figure CN2025142999_23072026_PF_FP_ABST
Abstract
Description
Memory modules, methods for repairing memory, controllers, media, and systems
[0001] This application claims priority to Chinese Patent Application No. 2025100658739, filed on January 15, 2025, entitled "Memory Module, Method for Repairing Memory, Controller, Medium and System", the contents of which are to be understood as incorporated herein by reference. Technical Field
[0002] This disclosure relates to, but is not limited to, the field of data storage technology, and particularly to a memory module, a method for repairing memory, a controller, a medium, and a system. Background Technology
[0003] Currently, when memory chips (also known as memory chips or dies) are manufactured, some of their memory cells may be faulty or unreliable. Control chips need to be tested before leaving the factory to detect these faulty or unreliable cells and replace them with pre-determined redundant memory cells. This replacement information is then burned into the on-chip eFuse. However, because the stress accumulated inside the chip during integrated circuit manufacturing takes a long time to dissipate, the first six months after a memory chip leaves the factory are a high-risk period for failures, and these potential problems cannot be detected on the production line. Therefore, for applications like data centers that have high reliability requirements for data storage, memory chip failures during use remain a pressing issue that needs to be addressed. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.
[0005] In a first aspect, embodiments of this disclosure provide a memory module, including a controller and memory chips connected to the controller. The controller is configured to: correct read data and determine the error address corresponding to the erroneous data; actively initiate a read operation to the target storage area pointed to by the error address; and verify the data read by the read operation and determine whether there is a permanently damaged failed storage unit in the target storage area corresponding to the error address based on the verification result.
[0006] Secondly, this disclosure also provides a method for repairing memory, applied to the memory module in the above embodiments. The method includes: correcting read data and determining the error address corresponding to the erroneous data; actively initiating a read operation to the target storage area pointed to by the error address; and verifying the data read by the read operation and determining whether there is a permanently damaged failed storage unit in the target storage area corresponding to the error address based on the verification result.
[0007] Thirdly, embodiments of this disclosure also provide a controller applied to a memory module, including a processor configured to perform a method for repairing memory according to any embodiment of this disclosure.
[0008] Fourthly, embodiments of this disclosure also provide a non-transient computer storage medium storing a computer program, which, when executed by a processor, implements the memory repair method in any embodiment of this disclosure.
[0009] Fifthly, embodiments of this disclosure provide a computer system, including a host and the memory module described in the above embodiments.
[0010] After reading and understanding the accompanying diagrams and detailed descriptions, the other aspects can be understood.
[0011] Overview of the attached figures
[0012] The accompanying drawings are used to provide an understanding of the technical solutions of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.
[0013] Figure 1 is a schematic diagram of the structure of a memory module in one embodiment of this disclosure;
[0014] Figure 2 is a flowchart illustrating a method for repairing memory according to an embodiment of this disclosure;
[0015] Figure 3 is a flowchart illustrating a method for repairing memory according to an embodiment of this disclosure;
[0016] Figure 4 is a flowchart illustrating a method for repairing memory according to an embodiment of this disclosure;
[0017] Figure 5 is a schematic diagram of the redirection process in a memory repair method according to an embodiment of the present disclosure;
[0018] Figure 6 is a flowchart illustrating a method for repairing memory according to an embodiment of this disclosure;
[0019] Figure 7 is a flowchart illustrating the process of initiating a read operation in a method for repairing memory according to an embodiment of this disclosure;
[0020] Figure 8 is a schematic diagram of the controller structure in one embodiment of this disclosure.
[0021] Detailed Explanation
[0022] The embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. Unless otherwise specified, the embodiments and features described herein can be combined arbitrarily.
[0023] The embodiments disclosed herein are not necessarily limited to the dimensions shown in the drawings, and the shapes and sizes of the components in the drawings do not reflect actual proportions. Furthermore, the drawings schematically illustrate ideal examples, and the embodiments of this disclosure are not limited to the shapes or values shown in the drawings.
[0024] The ordinal numbers such as "first" and "second" in this disclosure are used to avoid confusion among the constituent elements and do not indicate any order, quantity, or importance.
[0025] This disclosure describes several embodiments, but these descriptions are exemplary and not restrictive, and many more embodiments and implementations are possible within the scope of the embodiments described herein, which will be apparent to those skilled in the art. Although many possible combinations of features are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with or in lieu of any other feature or element in any other embodiment.
[0026] This disclosure includes and contemplates combinations of features and elements known to those skilled in the art. The embodiments, features, and elements disclosed in this disclosure may also be combined with any conventional features or elements to form the scheme defined by the claims. Any feature or element of any embodiment may also be combined with features or elements from other disclosed schemes to form another scheme defined by the claims. Therefore, it should be understood that any feature shown and discussed in this disclosure may be implemented individually or in any suitable combination. Therefore, the embodiments are not limited except by the limitations imposed by the appended claims and their equivalents. Furthermore, various modifications and changes may be made within the scope of the appended claims.
[0027] Furthermore, in describing representative embodiments, the specification may have presented methods and processes as a specific sequence of steps. However, the method or process should not be limited to the specific order of steps described in this disclosure to the extent that it does not depend on such a specific order. As will be understood by those skilled in the art, other sequences of steps are also possible. Therefore, the specific order of steps set forth in the specification should not be construed as a limitation of the claims. Moreover, the claims concerning the method and process should not be limited to performing the steps in the written order, and those skilled in the art will readily understand that these orders can be varied and still remain within the spirit and scope of the embodiments of this disclosure.
[0028] In the description of this disclosure, unless otherwise stated, "multiple" means two or more.
[0029] Currently, the ECC (Error Checking and Correcting) unit in a memory module is mainly used to resolve soft errors during data access. When a hardware error occurs in a storage cell within a memory chip, all access operations targeting that storage cell will be detected by the ECC unit.
[0030] As shown in Figure 1, this disclosure provides a memory module, including: a controller 110 and a memory chip 120 (also referred to as a memory chip or die) connected to the controller. The controller 110 is configured to: perform error correction on data read from the memory chip 120 and determine the error address corresponding to the erroneous data; actively initiate a read operation to the target storage area pointed to by the error address; and verify the data read by the read operation and determine whether there is a permanently damaged failed storage unit in the target storage area corresponding to the error address based on the verification result.
[0031] In some implementations of this embodiment, the memory module can be a CXL memory module. CXL is based on PCIe 5.0 and runs on top of the PCIe physical layer. Based on the characteristics of memory access, CXL is divided into three sub-protocols: 1) CXL.io, 2) CXL.mem, and 3) CXL.cache. Correspondingly, CXL devices (also called CXL memory modules) are divided into three types: Type 1 CXL device, Type 2 CXL device, and Type 3 CXL device. Type 2 and Type 3 devices can have multiple device memories, which can be of types such as DDR (Double Data Rate) or HBM (High Bandwidth Memory). Taking Type 2 devices as an example, a Type 2 device can be an accelerator card (such as a GPU (Graphics Processing Unit)) in a real-world application scenario, with the host providing data and the accelerator card handling the computation. CXL Type 3 devices can provide hosts with large-capacity DRAM (Dynamic Random Access Memory) and can also support persistent media (i.e., non-volatile media, such as flash memory chips).
[0032] The memory module in this embodiment can be a second type of CXL device or a third type of CXL device. The controller 110, also known as a CXL controller, may have a CXL interface based on the CXL protocol, allowing the host CPU to interact with the controller 110 via the CXL interface. For example, based on the CXL.io protocol, the host CPU can send instructions to the controller 110; based on the CXL.mem protocol, the host CPU can read data from the memory chip 120 and write external data to the memory chip 120. The control chip 110 may also have a memory controller (as shown in the figure, a DDR controller) to manage the memory chip 120. The memory chip 120 may be, for example, DRAM.
[0033] In this embodiment, an error correction unit 130 can be configured in the memory module controller 110 to implement the ECC function of the memory module. The error correction unit 130 encodes the data written to the memory chip 120 to generate a data checksum. When data needs to be read from the memory chip 120, the error correction unit 130 can use the previously generated checksum to detect whether there are any error bits in the read data and correct any errors. During this process, the error correction unit 130 can record the error address (i.e., the address corresponding to the erroneous data with the error bit) and the error bit (i.e., the bit) in the error correction information and report the error correction information to the controller 110. Here, the error correction information can be the error address and error bit recorded after the error correction unit 130 successfully corrects the error, or it can be the error address and error bit recorded when the error correction unit 130 fails to correct the error.
[0034] When the controller 110 corrects the read data through the error correction unit 130, it can determine the error address corresponding to the erroneous data through the error correction unit 130 and actively initiate a read operation to the target memory area where the error address is located. For example, the controller 110 can generate a read operation request based on the error address and send it to the memory controller (e.g., the DDR controller). The DDR controller generates a read operation command based on the address information carried in the read operation request and sends it to the corresponding memory chip 120, which then performs the decoding and data reading operations.
[0035] When the error correction unit 130 employs different error correction strategies, the number of error bits corrected in each error correction process varies. For example, when using an 8-bit checksum to correct 64-bit data, one error bit can be corrected at a time; while using a more-bit checksum, more error bits can be corrected at a time. When the number of error bits in an error address exceeds the processing limit of the error correction unit 130, the error correction unit 130 cannot correct the error bits. In this embodiment, the target storage area may include the storage unit corresponding to the error address, and may also include other storage units adjacent to the storage unit corresponding to the error address.
[0036] In one example, the error correction unit 130 can identify the location information of the erroneous data in the memory chip 120 through encoding (for example, it can include the location information of the row and column of the memory cell corresponding to the erroneous data in the memory chip 120), and obtain the media physical address containing the erroneous data. Then, the error correction unit 130 can convert the media physical address containing the erroneous data into a device physical address that can be processed by the processor of the controller 110 according to the preset correspondence between media physical addresses and device physical addresses in the memory module, and report it to the controller 110 as the error address so that the controller 110 can directly initiate a read operation request based on the error address.
[0037] Media physical address is the address used by the controller when accessing storage media (such as DRAM) through the memory interface; device physical address refers to the memory module, and the device physical address is the address carried in the access request sent by the processor (CPU) in the controller to the memory controller (such as the DDR controller in the figure).
[0038] For example, a read operation request generated by the processor of controller 110 based on the device physical address can be sent to the DDR controller, which converts the device physical address into a media physical address, and then the corresponding memory chip 120 accesses the media physical address.
[0039] In another example, the error correction unit 130 can report the physical address of the medium corresponding to the erroneous data to the controller 110 as the error address. The controller 110 then converts the error address into a device physical address based on the mapping relationship between the physical address of the medium and the physical address of the device, and initiates a read operation based on the converted device physical address.
[0040] When a permanently damaged faulty memory cell exists in the target storage area corresponding to the error address, each read operation initiated by the controller 110 will trigger the ECC function of the error correction unit 130. After the controller 110 initiates a read operation, it can use the error correction unit 130 to verify the data read by the read operation to determine whether there is still erroneous data, and thus determine whether there is a permanently damaged faulty memory cell in the target storage area corresponding to the error address. As an example, if the controller 110 initiates a read operation to the target storage area pointed to by the error address and does not receive error correction information reported by the error correction unit 130, or if the error correction information reported by the error correction unit 130 does not contain the error address, then there is no permanently damaged faulty memory cell in the target storage area corresponding to the error address; if the controller 110 initiates a read operation and still receives error correction information reported by the error correction unit 130 and the error correction information contains the error address, it can be determined that there is a permanently damaged faulty memory cell in the target storage area corresponding to the error address.
[0041] In the embodiment shown in Figure 1, when the controller corrects the read data, it can determine the error address corresponding to the erroneous data and actively initiate a read operation to the target storage area where the error address is located. Then, it verifies the data read by the read operation to determine whether there is a permanently damaged or failed storage unit in the target storage area where the error address should be. This can identify permanently damaged or failed storage units in memory in a timely and accurate manner so that they can be replaced and repaired in a targeted manner, which helps to improve the reliability of the memory module.
[0042] In some embodiments, the controller is further configured to: correct the data read from the memory chip, and then write the corrected data back to the target storage area according to the error address.
[0043] As an example, after each error correction by the error correction unit (i.e., correcting erroneous bits in the data according to the encoding), the corrected data can be written back to the target storage area based on the error address. When a permanently damaged or failed storage unit exists in the target storage area, data cannot be written to that unit, and therefore, the erroneous data stored in the permanently damaged or failed storage unit cannot be updated through a write-back operation. Consequently, after performing a write-back operation, the data read from the error address still contains erroneous data. However, when no permanently damaged or failed storage unit exists in the target storage area, writing the corrected data back to the target storage area can correct the erroneous data, ensuring that the erroneous data will no longer appear in subsequent reads.
[0044] In this embodiment, the write-back operation after error correction can correct the erroneous data stored in the non-permanently damaged memory cells in the memory chips. In this way, after the controller initiates a read operation to the target storage area after the write-back, it can identify whether there are permanently damaged failed memory cells in the target storage area corresponding to the error address based on the verification result of the read data, which helps to improve the identification accuracy of permanently damaged failed memory cells.
[0045] When a permanently damaged or failed memory cell is identified in the target storage area corresponding to the error address, the memory module can use the controller and / or memory chips to repair the identified permanently damaged or failed memory cell to improve the reliability of the memory module.
[0046] In some embodiments, the memory module can utilize the repair function of the memory chips themselves to repair permanently damaged or failed memory cells. In these embodiments, the controller is also configured to: upon determining that a permanently damaged or failed memory cell exists in the target memory region corresponding to the error address, send a repair command to the memory chip containing the permanently damaged or failed memory cell, notifying the memory chip to repair the permanently damaged or failed memory cell.
[0047] As an example, the controller can generate a Post Package Repair (PPR) instruction based on the error address and send it to the corresponding memory chip. The memory chip then repairs the permanently damaged failed memory cell corresponding to the error address. For instance, the memory chip can replace the permanently damaged failed memory cell with its own reserved redundant memory cell and record the address mapping between the two in eFuse. When receiving an access instruction from the controller, the controller can map the access operation pointing to the permanently damaged failed memory cell to the corresponding redundant memory cell according to the mapping.
[0048] In this embodiment, the memory module can use the memory chip's own repair function to repair the identified permanently damaged and failed storage units, thereby improving the storage resource utilization rate of the memory chip.
[0049] In other embodiments, the memory module can also utilize the controller to repair identified permanently damaged or failed memory cells. In these embodiments, the controller can use a portion of the memory cells as redundant memory cells. Here, the redundant memory cells can be memory cells in a reserved redundant storage area within the memory cell, or memory cells in the regular storage area (i.e., non-redundant storage area) of the memory cell.
[0050] The controller is also configured to: if it is determined that there is a permanently damaged failed storage unit in the target storage area corresponding to the error address, allocate a redundant storage unit to the target storage area corresponding to the error address, and redirect access to the error address to the redundant storage unit.
[0051] In this embodiment, permanently damaged failed storage units and their corresponding redundant storage units can be located in the same memory chip or in different memory chips. The memory module can use a controller to schedule and allocate the storage units of all or some of the memory chips to repair permanently damaged failed storage units in the memory chips. Dynamic repair of memory chips can be performed during the use of the memory module, which helps improve the reliability of the memory module and the utilization rate of storage resources. This is particularly suitable for certain application scenarios, such as when the memory chips themselves do not have repair capabilities, and can significantly improve the reliability of the memory module.
[0052] In some embodiments of this example, the controller can redirect access in the following way: determine the address correspondence between the permanently damaged failed storage unit and the redundant storage unit; after receiving an access request for the erroneous address, replace the erroneous address in the access request with the address of the redundant storage unit according to the address correspondence; and initiate access to the redundant storage unit.
[0053] In one example, the controller can establish an address mapping table to record the correspondence between the physical addresses of permanently damaged or failed memory cells and redundant memory cells. These physical addresses can be device physical addresses or media physical addresses. For instance, the address mapping table can record the correspondence between the media physical addresses of permanently damaged or failed memory cells and redundant memory cells. This address mapping table can be stored in the memory controller. When the memory controller parses the media physical address of a permanently damaged or failed memory cell from an access request, it can replace the media physical address of the permanently damaged or failed memory cell with the media physical address of the redundant memory cell according to the address mapping table. When the controller receives an access request for an erroneous address, it can replace the erroneous address with the media physical address of the redundant memory cell according to the address mapping table, thereby mapping the access operation for the permanently damaged or failed memory cell to the redundant memory cell.
[0054] For example, an address mapping table can record the correspondence between the physical addresses of permanently damaged or failed memory cells and redundant memory cells. This address mapping table can be stored in the controller. When the controller parses the physical address (i.e., the error address) of a permanently damaged or failed memory cell from an access request, the controller's CPU can replace the physical address of the permanently damaged or failed memory cell with the physical address of the redundant memory cell according to the address mapping table. Then, the memory controller accesses the memory cell using that physical address, thereby mapping the access operation for the permanently damaged or failed memory cell to the redundant memory cell.
[0055] In another example, the error address can be the device physical address; the controller can replace the device physical address of the permanently damaged failed storage unit in the mapping relationship between external addresses and device physical addresses with the device physical address of the redundant storage unit to determine the address correspondence between the permanently damaged failed storage unit and the redundant storage unit.
[0056] Here, the external address, also known as the bus domain address or host physical address, refers to the address carried in the data access request sent by the host. After receiving the data access request from the host, the controller can convert the external address into a device physical address according to the mapping relationship between the external address and the device physical address. Then, it sends the device physical address to the DDR controller, which converts it into a media physical address for access.
[0057] In this example, the controller can obtain the address layout inside the memory chip and determine the mapping relationship between the physical address of the medium and the physical address of the device based on the address layout.
[0058] The controller can convert the physical address of a storage cell to a physical address of a device based on the mapping relationship between the physical address of the medium and the physical address of the device. For example, it can convert the physical address of the storage cell corresponding to the error bit to the corresponding physical address of the device (i.e., the physical address of the device of the permanently damaged storage cell). After determining the target area based on the error address, it can also convert the physical address of the target storage area to the corresponding physical address of the device in order to initiate a read operation.
[0059] As an example, the controller includes a processor and a memory controller. The processor of the controller translates the error address into the corresponding device physical address, generates a read operation request carrying the device physical address, and sends it to the memory controller. The memory controller translates the device physical address in the read operation request into the corresponding media physical address, and sends an access instruction to the memory chip based on the media physical address to access the target storage area.
[0060] After the controller identifies the redundant storage unit corresponding to the permanently damaged failed storage unit, it updates the mapping relationship between the external address and the device physical address. When converting the external address to the device physical address, it can replace the external address of the permanently damaged failed storage unit carried in the access instruction with the device physical address of the redundant storage unit, thereby redirecting the access to the erroneous address to the access to the redundant storage unit.
[0061] Here, the error address is the physical address of the medium. The controller can perform statistical analysis based on the location information in the error address to determine the number and location of the error bit, and then determine whether the memory cell corresponding to the error bit is a failed memory cell that has been permanently damaged.
[0062] In one example of this embodiment, the controller includes a processor and a memory controller. The controller initiates a read operation request to the target memory region where the error address is located, including: the processor converting the error address into the corresponding device physical address, generating a read operation request carrying the device physical address and sending it to the memory controller; the memory controller converting the device physical address in the read operation request into the corresponding media physical address, and sending an access instruction to the memory chip based on the media physical address to access the target memory region.
[0063] For example, the controller can also replace the physical address of a permanently damaged or failed storage unit in the mapping relationship between physical addresses of media and physical addresses of devices with the physical address of a redundant storage unit, so as to determine the address correspondence between permanently damaged or failed storage units and redundant storage units, and map access operations pointing to the erroneous address to the redundant storage unit.
[0064] When the controller receives an access command from the host, it converts the external address carried in the access command into a device physical address. Then, the DDR controller uses the device physical address of the permanently damaged failed memory cell as an index to look up the mapping relationship between the media physical address and the device physical address to determine the media physical address to be accessed. After the controller 110 determines the media physical address of the redundant memory cell corresponding to the media physical address of the permanently damaged failed memory cell based on the mapping relationship between the media physical address and the device physical address, the DDR controller can replace the media physical address to be accessed with the media physical address of the redundant memory cell, thereby mapping the access operation pointing to the erroneous address to the redundant memory cell.
[0065] In some embodiments of this example, before redirecting access to the erroneous address to the redundant storage unit, the controller is further configured to write the corrected data to the redundant storage unit. In this way, after the access to the erroneous address is redirected to the redundant storage unit, no further error correction operation will be triggered if the redundant storage unit is not damaged, thus reducing the performance loss of the memory module.
[0066] For example, before redirecting access to the erroneous address to the redundant storage unit, when the controller receives an access request from the host for the erroneous address, it reads the data from the erroneous address, corrects the data using the error correction unit, and then returns it to the host. This ensures normal data access to the memory module during the repair process.
[0067] In some embodiments, the controller is further configured to: after the memory module restarts, send a repair instruction to the memory chip containing the permanently damaged failed memory cell, instructing the memory chip to repair the permanently damaged failed memory cell.
[0068] As an example, after the memory module restarts, the controller can generate a PPR instruction based on the recorded address information and send it to the memory chip corresponding to the address information. The memory chip itself then repairs the permanently damaged or failed memory cell. For instance, the memory chip can replace the permanently damaged or failed memory cell with its own reserved redundant memory cell and record the address mapping between the two in a replacement table. When it receives an access instruction from the controller, it can map the access operation pointing to the permanently damaged or failed memory cell to the corresponding redundant memory cell according to the replacement table.
[0069] In this embodiment, in addition to using the controller to dynamically repair permanently damaged and failed memory units, the memory chips themselves can also be used to repair permanently damaged and failed memory units. This not only enables memory repair during the use of the memory module, but also improves the reliability of the memory module by combining the hardware repair function of the memory chips themselves.
[0070] In some embodiments of this example, the controller is further configured to: initiate a read operation to the target storage area and verify the data read by the read operation; and, if the data read by the read operation does not contain erroneous data, delete the address correspondence between permanently damaged failed storage units and redundant storage units.
[0071] In this embodiment, to verify the effectiveness of the hardware repair of the memory chip, a read operation can be initiated at the error address recorded before the restart after the memory module is restarted. The verification result of the read data is used to determine whether the permanently damaged failed memory cell has been repaired. For example, if the read data does not contain erroneous data, it is determined that the permanently damaged failed memory cell has been repaired by the memory chip. In this case, the address correspondence between the permanently damaged failed memory cell and the redundant memory cell can be deleted to remove the correspondence between the two.
[0072] Here, deleting the address correspondence between permanently damaged failed storage units and redundant storage units can correspond to the way the address correspondence between the two was determined in the above embodiments. The corresponding entries in the address mapping table, or the mapping relationship between external address and device physical address, or the mapping relationship between media physical address and device physical address can be deleted.
[0073] In some embodiments, the controller actively initiates a read operation to the target storage area pointed to by the error address, including: initiating N read operations to the target storage area according to a preset read operation interval, such that the time interval between two consecutive read operations is not less than the read operation interval; and the controller verifies the data read by the read operation and determines whether there is a permanently damaged failed storage unit in the target storage area corresponding to the error address based on the verification result, including: if M error addresses appear in the verification results corresponding to the N read operations, determining that there is a permanently damaged failed storage unit in the target storage area corresponding to the error address; wherein M and N are preset values, and M is not greater than N.
[0074] In this embodiment, the controller may be equipped with a timer to define the read operation interval, so that the interval between two consecutive read operations is not less than the read operation interval, thereby avoiding frequent refresh of the storage unit.
[0075] In this embodiment, the controller can continuously initiate N read operation requests to the target storage area according to the read operation interval, and determine whether there is a permanently damaged failure storage unit corresponding to the error address based on the number of error addresses included in the verification result of the read data. This can avoid misjudgment caused by accidental errors of storage units and help improve the accuracy of identifying permanently damaged failure storage units.
[0076] In some embodiments of this example, the controller is further configured to: determine that there is a permanently damaged failed storage unit in the target storage area corresponding to the new error address when the same new error address appears M times in the verification results corresponding to N read operations.
[0077] In this embodiment, after the controller actively initiates a read operation, it can identify other storage units in the target storage area based on the verification results, which helps to further improve the reliability of the memory module.
[0078] In some embodiments, the controller is further configured to: when initiating a read operation request to the target storage area, and upon receiving a data access request for the target storage area sent by the host, prioritize responding to the data access request.
[0079] As an example, a mutex can be set in the controller to schedule read operation requests initiated by the controller and data access requests sent by the host, and to prioritize responding to data access requests sent by the host, thereby repairing failed storage units without affecting the normal use of the memory module.
[0080] In some embodiments, the controller actively initiates a read operation to the target storage area pointed to by the error address, including: determining a block containing the storage area corresponding to the error address based on the error address, the block further including storage areas corresponding to multiple addresses adjacent to the error address; and using the block as the target storage area to initiate a read operation to determine whether there are permanently damaged failed storage cells in the target storage area.
[0081] In this embodiment, the block containing the storage area corresponding to the error address may include one or more rows, and each row includes at least one address corresponding to multiple storage units.
[0082] In one example, the target row and target column where the error address is located can be determined based on the error address. Then, a first preset number of adjacent rows adjacent to the target row and a second preset number of adjacent columns adjacent to the target column can be determined. The target row, target column, adjacent rows and adjacent columns are then used as a block containing the error address.
[0083] Here, the first preset quantity can be 0, 1, or other positive integers, and adjacent rows can be located on one side or both sides of the target row; the second preset quantity can be 0 or an integer multiple of the address. For example, when the smallest unit of the address is a byte (Byete, B), each address includes 8 storage units, then the target column includes 8 consecutive columns. The second preset quantity can be an integer multiple of 8, and adjacent columns can be located on one side or both sides of the target column. This example does not limit this.
[0084] If the error address is a device physical address, the controller can determine the device physical addresses of adjacent rows and columns based on the error address, and use the target row, target column, and adjacent rows and columns as the target storage area. Alternatively, if the error address is a media physical address, the controller can first convert the error address to a device physical address, and then determine the device physical addresses of adjacent rows and columns based on the error address; or, if the error address is a media physical address, the controller can determine the media physical addresses of adjacent rows and columns based on the error address, and then convert the media physical address of the target storage area consisting of the target row, target column, and adjacent rows and columns to a device physical address.
[0085] In practice, the probability of a memory cell adjacent to a permanently damaged memory cell being damaged is relatively high. Considering this situation, this embodiment can use other memory cells around the memory cell corresponding to the error address as the target storage area. The controller actively initiates a read operation to identify whether there are other failed memory cells around the memory cell corresponding to the error address, thereby improving the identification efficiency of permanently damaged memory cells and the repair efficiency of the memory module.
[0086] In another example, when the controller corrects errors in the read data, it can determine the error bit corresponding to the erroneous data, and then determine the block containing the error address in the following way: based on the error address, determine a third preset number of storage units adjacent to the storage unit corresponding to the error bit, and then use the storage unit corresponding to the error address and the third preset number of storage units as the block containing the error address.
[0087] Here, the third preset number of storage units can correspond to one or more addresses, and each address includes the preset number of storage units.
[0088] In this example, during error correction processing of the read data, the erroneous bit (also known as the error bit) can be identified in the erroneous data, allowing the controller to pinpoint the individual memory cell that has failed. Considering the high probability of surrounding memory cells failing due to permanent damage, after error correction, the controller can also initiate read requests to adjacent memory cells to determine if other memory cells around the permanently damaged cell have also suffered permanent damage, further improving the reliability of the memory module.
[0089] In the above embodiments, the first preset quantity, the second preset quantity, and the third preset quantity can be preset according to requirements or the performance and usage status of the memory module.
[0090] In some embodiments, memory chips are connected to the controller in parallel.
[0091] As an example, memory chips can be connected to the controller chip via DIMM (Dual Inline Memory Module) strips or similar parallel combinations.
[0092] As shown in Figure 2, this disclosure also provides a method for repairing memory, which is applied to the memory module in any of the above embodiments. The method includes the following steps.
[0093] Step 210: Correct the data read from the memory chip and determine the error address corresponding to the erroneous data.
[0094] In this embodiment, the memory module includes a controller and memory chips connected to the controller. The controller is equipped with an error correction unit. The error correction unit can perform error correction processing on the data in the memory chips. The processor (CPU) of the controller can execute the memory repair method in this embodiment.
[0095] Step 220: Actively initiate a read operation to the target storage area pointed to by the error address.
[0096] Step 230: Verify the data read by the read operation, and determine whether there is a permanently damaged failed storage unit in the target storage area corresponding to the error address based on the verification result.
[0097] In this embodiment, when the controller performs error correction processing on the data read from the memory chip, it can determine the error address corresponding to the erroneous data and actively initiate a read operation to the target storage area where the error address is located. Then, it verifies the data read by the read operation and determines whether there is a permanently damaged failed storage unit in the target storage area corresponding to the error address based on the verification result. This can identify permanently damaged failed storage units in the memory in a timely and accurate manner so as to replace and repair them in a targeted manner, which helps to improve the reliability of the memory module.
[0098] In some embodiments, the method for repairing memory may further include: correcting the read data and then writing the corrected data back to the target storage area according to the error address.
[0099] In this embodiment, the write-back operation after error correction can correct the erroneous data stored in the non-permanently damaged storage cells of the memory chip. In this way, after the controller initiates a read operation to the target storage area after the write-back, it can identify whether there are permanently damaged failed storage cells in the target storage area corresponding to the error address based on the verification result of the read data, which helps to improve the identification accuracy of permanently damaged failed storage cells.
[0100] Figure 3 shows a flowchart of an embodiment of the method for repairing memory according to the present disclosure. As shown in Figure 3, the process may include the following steps.
[0101] Step 310: Correct the data read from the memory chip and determine the error address corresponding to the erroneous data.
[0102] Step 320: Actively initiate a read operation to the target storage area pointed to by the error address.
[0103] Step 330: Verify the data read by the read operation, and determine whether there is a permanently damaged or failed storage unit in the target storage area corresponding to the error address based on the verification result.
[0104] Step 340: If it is determined that there is a permanently damaged failed memory cell in the target storage area corresponding to the error address, a repair command is sent to the memory chip where the permanently damaged failed memory cell is located, notifying the memory chip to repair the permanently damaged failed memory cell.
[0105] In this embodiment, the memory module can use the memory chip's own repair function to repair the identified permanently damaged and failed storage units, thereby improving the storage resource utilization rate of the memory chip.
[0106] Referring now to Figure 4, which shows a flowchart of an embodiment of the method for repairing memory according to the present disclosure, the process may include the following steps.
[0107] Step 410: Correct the data read from the memory chip and determine the error address corresponding to the erroneous data.
[0108] In this embodiment, the controller in the memory module uses a portion of the memory chips' storage units as redundant storage units.
[0109] Step 420: Actively initiate a read operation to the target storage area pointed to by the error address.
[0110] Step 430: Verify the data read by the read operation, and determine whether there is a permanently damaged failed storage unit in the target storage area corresponding to the error address based on the verification result.
[0111] Step 440: If it is determined that there is a permanently damaged failed storage unit in the target storage area corresponding to the error address, allocate a redundant storage unit to the storage unit corresponding to the error address and redirect access to the error address to the redundant storage unit.
[0112] In this embodiment, the memory module can use the controller to schedule and allocate the storage units of all or some of the memory chips in order to repair the permanently damaged and failed storage units in the memory chips. The memory chips can be dynamically repaired during the use of the memory module, which helps to improve the reliability of the memory module and the utilization rate of storage resources. It is especially suitable for some specific application scenarios, such as when the memory chips themselves do not have repair functions, and can significantly improve the reliability of the memory module.
[0113] In some implementations of this embodiment, before step 440, the process shown in FIG4 may also include: step 450, writing the error-corrected data into the redundant storage unit.
[0114] In this way, after access to the erroneous address is redirected to the redundant storage unit, no error correction operation will be triggered if the redundant storage unit is not damaged, which can reduce the performance loss of the memory module.
[0115] In some implementations of this embodiment, before step 440, the process shown in FIG4 may also include: step 460, when receiving an access request from the host for the erroneous address, reading data from the erroneous address, correcting the data, and returning it to the host.
[0116] In this embodiment, step 460 ensures normal data access for the memory module during the repair process.
[0117] In some embodiments of this example, step 440 may also include the process shown in FIG5. As shown in FIG5, the process may include the following steps.
[0118] Step 510: Determine the address correspondence between permanently damaged failed storage units and redundant storage units.
[0119] Step 520: After receiving the access request for the erroneous address, replace the erroneous address in the access request with the address of the redundant storage unit according to the address correspondence.
[0120] Step 530: Initiate access to the redundant storage unit.
[0121] In this embodiment, the controller can replace the erroneous address in the access request with the address of the redundant storage unit according to the address mapping relationship, thereby achieving access redirection.
[0122] Figure 6 shows a flowchart of an embodiment of the method for repairing memory according to the present disclosure. As shown in Figure 6, the process may include the following steps.
[0123] Step 610: Correct the data read from the memory chip and determine the error address corresponding to the erroneous data.
[0124] In this embodiment, the controller in the memory module uses a portion of the memory chips' storage units as redundant storage units.
[0125] Step 620: Actively initiate a read operation to the target storage area pointed to by the error address.
[0126] Step 630: Verify the data read by the read operation, and determine whether there is a permanently damaged failed storage unit in the target storage area corresponding to the error address based on the verification result.
[0127] Step 640: If it is determined that there is a permanently damaged failed storage unit in the target storage area corresponding to the error address, allocate a redundant storage unit to the storage unit corresponding to the error address and redirect access to the error address to the redundant storage unit.
[0128] Step 650: After the memory module restarts, a repair command is sent to the memory chip containing the permanently damaged failed memory cell to instruct the memory chip to repair the permanently damaged failed memory cell.
[0129] In this embodiment, in addition to using the controller to dynamically repair permanently damaged and failed memory units, the memory chips themselves can also be used to repair permanently damaged and failed memory units. This not only enables memory repair during the use of the memory module, but also improves the reliability of the memory module by combining the hardware repair function of the memory chips themselves.
[0130] In some embodiments of this example, the embodiment shown in FIG6 may also include steps 660 and 670.
[0131] Step 660: After the memory module restarts, initiate a read operation to the target storage area and verify the data read out by the read operation.
[0132] Step 670: If the data read in the read operation does not contain erroneous data, delete the address correspondence between the permanently damaged failed storage unit and the redundant storage unit.
[0133] In this embodiment, in order to verify whether the hardware repair of the memory chip is effective, a read operation can be initiated to the error address recorded before the restart after the memory module is restarted, and the permanent damage to the failed storage unit can be determined based on the verification result of the read data.
[0134] In some embodiments, actively initiating a read operation to the target storage region pointed to by the error address includes: initiating N read operations to the target storage region according to a preset read operation interval, such that the time interval between two consecutive read operations is not less than the read operation interval; and verifying the data read by the read operation, and determining whether there is a permanently damaged failed storage unit in the target storage region corresponding to the error address based on the verification result includes: if M error addresses appear in the verification results corresponding to the N read operations, determining that there is a permanently damaged failed storage unit in the target storage region corresponding to the error address; wherein M and N are preset values, and M is not greater than N.
[0135] In this embodiment, by limiting the read operation interval, the time between two consecutive read operations is not less than the read operation interval, thus avoiding frequent refreshes of the storage unit. The controller can continuously initiate N read operation requests to the target storage area according to the read operation interval, and determine whether there is a permanently damaged failed storage unit in the target storage area corresponding to the error address based on the number of error addresses included in the verification result of the read data. This can avoid misjudgment caused by accidental errors of storage units and help improve the accuracy of identifying permanently damaged failed storage units.
[0136] In some embodiments of this example, the method may further include: if the same new error address appears M times in the verification results corresponding to N read operations, determining that there is a permanently damaged failed storage unit in the target storage area corresponding to the new error address.
[0137] In this embodiment, after the controller actively initiates a read operation, it can identify other storage units in the target storage area based on the verification results, which helps to further improve the reliability of the memory module.
[0138] In some embodiments, the method for repairing memory may further include: when initiating a read operation to a target storage area, prioritizing the response to a data access request sent by the host for the target storage area.
[0139] In this embodiment, the data access request sent by the host is responded to first, so as to repair the failed storage unit without affecting the normal use of the memory module.
[0140] Figure 7 shows a flowchart of a read operation in one embodiment of the method for repairing memory according to the present disclosure. As shown in Figure 7, the process may include the following steps.
[0141] Step 710: Based on the error address, determine the block containing the storage area corresponding to the error address.
[0142] The block also includes storage areas corresponding to multiple addresses adjacent to the erroneous address.
[0143] Step 720: Select the block as the target storage area and initiate a read operation to determine whether there are any permanently damaged or failed storage cells in the target storage area.
[0144] In this embodiment, other storage units around the storage unit corresponding to the error address can be used as the target storage area. The controller actively initiates a read operation to identify whether there are other failed storage units around the storage unit corresponding to the error address, thereby improving the identification efficiency of permanently damaged failed storage units and the repair efficiency of memory modules.
[0145] As shown in Figure 8, one embodiment of this disclosure also provides a controller applied to the memory module in any of the above embodiments. The controller may include a processor 810, an interface 820 for connecting to a host, a storage interface 830 for connecting to memory chips, and other hardware 840. The processor 810 is configured to execute the memory repair method in any embodiment of this disclosure.
[0146] In this embodiment, the other hardware 840 may be a memory controller (e.g., a DDR controller), and the processor 810 may be a central processing unit (CPU), a network processor (NP), a microprocessor, or other conventional processors. The processor may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), discrete logic or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above devices. That is, the processor in this embodiment may be any processing device or combination of devices that implements the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. If the embodiments of this disclosure are implemented in part in software, then instructions for software can be stored in a suitable non-volatile computer-readable storage medium, and one or more processors can be used to execute the instructions in hardware to implement the methods of the embodiments of this disclosure.
[0147] An embodiment of this disclosure also provides a non-transient computer storage medium storing a computer program that, when executed by a processor, can implement the memory repair method in any embodiment of this disclosure.
[0148] One embodiment of this disclosure provides a computer system including a host and a memory module as described in any embodiment of this disclosure.
[0149] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
Claims
1. A memory module, comprising a controller and memory chips connected to the controller, wherein, The controller is configured to: correct errors in data read from the memory chip and determine the error address corresponding to the erroneous data; and actively initiate a read operation to the target storage area pointed to by the error address. In addition, the data read by the read operation is verified, and the verification result is used to determine whether there is a permanently damaged or failed storage unit in the target storage area corresponding to the error address.
2. The memory module according to claim 1, wherein, The controller is further configured to: correct the data read from the memory chip, and then write the corrected data back to the target storage area according to the error address.
3. The memory module according to claim 1, wherein, The controller is further configured to: when it is determined that there is a permanently damaged failed memory cell in the target storage area corresponding to the error address, send a repair instruction to the memory chip where the permanently damaged failed memory cell is located, and notify the memory chip to repair the permanently damaged failed memory cell.
4. The memory module according to claim 1, wherein, The controller uses a portion of the memory chip's storage units as redundant storage units. The controller is further configured to: if it is determined that there is a permanently damaged failed storage unit in the target storage area corresponding to the error address, allocate a redundant storage unit to the target storage area corresponding to the error address, and redirect access to the error address to the redundant storage unit.
5. The memory module according to claim 4, wherein, The controller redirects access to the erroneous address to the redundant storage unit, including: Determine the address correspondence between the permanently damaged failed storage unit and the redundant storage unit; Upon receiving an access request for the erroneous address, the erroneous address in the access request is replaced with the address of the redundant storage unit according to the address correspondence. Initiate access to the redundant storage unit.
6. The memory module according to claim 4, wherein, The controller is also configured to: After the memory module restarts, a repair command is sent to the memory chip containing the permanently damaged failed storage unit to instruct the memory chip to repair the permanently damaged failed storage unit.
7. The memory module according to claim 6, wherein, The controller is also configured to: After the memory module restarts, a read operation is initiated to the target storage area, and the data read by the read operation is verified; and, If the data read in the read operation does not contain erroneous data, the address correspondence between the permanently damaged failed storage unit and the redundant storage unit is deleted.
8. The memory module according to claim 4, wherein, Before redirecting access to the erroneous address to the redundant storage unit, the controller is further configured to: The corrected data is written into the redundant storage unit.
9. The memory module according to claim 8, wherein, The controller is also configured to: Before redirecting access to the erroneous address to the redundant storage unit, when a host access request for the erroneous address is received, data is read from the erroneous address, the data is corrected, and then returned to the host.
10. The memory module according to claim 1, wherein, The controller actively initiates a read operation to the target storage region pointed to by the erroneous address, including: initiating N read operations to the target storage region according to a pre-set read operation interval, such that the time interval between two consecutive read operations is not less than the read operation interval; and... The controller verifies the data read by the read operation and determines whether there is a permanently damaged failed storage unit in the target storage area corresponding to the error address based on the verification result, including: if the error address appears M times in the verification results corresponding to N read operations, it is determined that there is a permanently damaged failed storage unit in the target storage area corresponding to the error address. Where M and N are preset values, and M is not greater than N.
11. The memory module according to claim 10, wherein, The controller is further configured to: if the same new error address appears M times in the verification results corresponding to N read operations, determine that there is a permanently damaged failed storage unit in the target storage area corresponding to the new error address.
12. The memory module according to claim 1, wherein, The controller is also configured to: when initiating the read operation to the target storage area, and upon receiving a data access request for the target storage area sent by the host, prioritize responding to the data access request.
13. The memory module according to claim 1, wherein, The controller actively initiates a read operation to the target storage area pointed to by the erroneous address, including: Based on the error address, a block containing the storage area corresponding to the error address is determined, and the block further includes storage areas corresponding to multiple addresses adjacent to the error address; The block is used as the target storage area, and the read operation is initiated to determine whether there are permanently damaged or failed storage units in the target storage area.
14. The memory module according to claim 1, wherein, The memory chips are connected to the controller in parallel.
15. A method for repairing memory, applied to a memory module according to any one of claims 1 to 14, the method comprising: The data read from the memory chip is corrected for errors, and the error address corresponding to the erroneous data is determined. Actively initiate a read operation towards the target storage area pointed to by the error address; as well as, The data read by the read operation is verified, and the verification result is used to determine whether there is a permanently damaged or failed storage unit in the target storage area corresponding to the error address.
16. The method of claim 15, further comprising: After correcting the errors in the data read from the memory chip, the corrected data is written back to the target storage area according to the error address.
17. The method of claim 15, further comprising: If it is determined that there is a permanently damaged failed memory cell in the target storage area corresponding to the error address, a repair instruction is sent to the memory chip where the permanently damaged failed memory cell is located, notifying the memory chip to repair the permanently damaged failed memory cell.
18. The method according to claim 15, wherein, The controller in the memory module uses a portion of the memory chip's storage units as redundant storage units. The method further includes: if it is determined that there is a permanently damaged failed storage unit in the target storage area corresponding to the error address, allocating a redundant storage unit to the target storage area corresponding to the error address, and redirecting access to the error address to the redundant storage unit.
19. The method according to claim 18, wherein, The step of redirecting access to the erroneous address to the redundant storage unit includes: Determine the address correspondence between the permanently damaged failed storage unit and the redundant storage unit; Upon receiving an access request for the erroneous address, the erroneous address in the access request is replaced with the address of the redundant storage unit according to the address correspondence. Initiate access to the redundant storage unit.
20. The method of claim 18, further comprising: After the memory module restarts, a repair command is sent to the memory chip containing the permanently damaged failed storage unit to instruct the memory chip to repair the permanently damaged failed storage unit.
21. The method of claim 20, further comprising: After the memory module restarts, a read operation is initiated to the target storage area, and the data read out by the read operation is verified. as well as, If the data read in the read operation does not contain erroneous data, the address correspondence between the permanently damaged failed storage unit and the redundant storage unit is deleted.
22. The method of claim 18, further comprising, before redirecting access to the erroneous address to the redundant storage unit: The corrected data is written into the redundant storage unit.
23. The method of claim 22, further comprising, before redirecting access to the erroneous address to the redundant storage unit: When an access request is received from the host for the erroneous address, data is read from the erroneous address, the data is corrected, and then returned to the host.
24. The method according to claim 15, wherein actively initiating a read operation to the target storage region pointed to by the error address includes: According to a pre-set read operation interval, N read operations are initiated to the target storage area, such that the time interval between two consecutive read operations is not less than the read operation interval; as well as, Verifying the data read by the read operation and determining whether there is a permanently damaged failed storage unit in the target storage area corresponding to the error address based on the verification result includes: if the error address appears M times in the verification results corresponding to N read operations, determining that there is a permanently damaged failed storage unit in the target storage area corresponding to the error address; Where M and N are preset values, and M is not greater than N.
25. The method of claim 24, further comprising: If the same new error address appears M times in the verification results corresponding to N read operations, it is determined that there is a permanently damaged failed storage unit in the target storage area corresponding to the new error address.
26. The method of claim 15, further comprising: When initiating the read operation to the target storage area, if a data access request for the target storage area is received from the host, the data access request shall be responded to first.
27. The method according to claim 16, wherein, The step of actively initiating a read operation towards the target storage region pointed to by the error address includes: Based on the error address, a block containing the storage area corresponding to the error address is determined, and the block further includes storage areas corresponding to multiple addresses adjacent to the error address; The block is used as the target storage area, and the read operation is initiated to determine whether there are permanently damaged or failed storage units in the target storage area.
28. The method of claim 15, further comprising: When correcting errors in read data, determine the error bits corresponding to the erroneous data; Based on the error address, a third preset number of storage units adjacent to the storage unit corresponding to the error bit are determined; A read operation is initiated on the third preset number of storage units to determine whether any of the third preset number of storage units are permanently damaged or failed storage units.
29. A controller applied to a memory module, wherein, The controller includes a processor; the processor is configured to perform the method of repairing memory as claimed in any one of claims 15 to 28.
30. A non-transient computer storage medium storing a computer program, which, when executed by a processor, implements a method for repairing memory as described in any one of claims 15 to 28.
31. A computer system comprising a host and a memory module as described in any one of claims 1-14.