Method of handling failure of memory, memory controller and memory system
By identifying and repairing memory row addresses that meet bounded fault characteristics, the problem of inaccurate storage system fault detection in the prior art is solved, and the reliability and resource utilization efficiency of the system are improved.
Patent Information
- Application Number
- CN202510869893.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-26
AI Technical Summary
When detecting and repairing memory faults, existing technologies are unable to promptly discover and accurately determine bounded faults, which increases the probability of uncorrectable errors in the storage system and wastes resources.
By recording the row addresses of correctable errors and identifying whether they meet the bounded fault characteristics, a register array is used to store address information and count values, and a post-packaging repair operation is performed to repair the target row addresses that meet the conditions.
It improves the reliability and availability of the storage system, reduces the probability of uncorrectable errors, and optimizes resource utilization.
Smart Images

Figure CN120708681A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a method for processing a memory failure, a memory controller, and a memory system. Background Art
[0002] As memory (e.g., dynamic random access memory (DRAM)) size decreases and storage capacity increases, the probability of memory failure increases. To enhance the reliability, availability, and serviceability (RAS) of storage systems and minimize the probability of errors, memory has introduced repair capabilities. For example, the JEDEC memory specification for DDR5 includes a post-package repair (PPR) feature. When a memory system detects a fault in a logical row address of memory, it can use the DRAM's PPR feature to perform a repair operation on the faulty row address, effectively reducing the error rate, improving system stability, and ensuring the normal operation of the memory system. Summary of the Invention
[0003] At least one embodiment of the present disclosure provides a method for handling a memory fault. The method includes: in response to a correctable error in data read from a row address of the memory, determining whether the correctable error meets the characteristics of a bounded fault; in response to the correctable error meeting the characteristics of a bounded fault, performing a first count on the row address to obtain a first count value; in response to the correctable error not meeting the characteristics of a bounded fault, performing a second count on the row address to obtain a second count value; and in response to the presence of a target row address in the memory, performing a repair operation on the target row address, wherein the first count value of the target row address is greater than a first threshold or the second count value is greater than a second threshold.
[0004] For example, in the method provided by at least one embodiment of the present disclosure, determining whether a correctable error meets the characteristics of a bounded fault includes: in response to the correctable error meeting any of the following conditions, determining that the correctable error meets the characteristics of a bounded fault: the number of error data pins corresponding to the data is 1, and the number of error bursts corresponding to the error data pins is greater than a third threshold; the number of error data pins corresponding to the data is 2, the data read from the 2 error data pins belong to the same half byte, and the total number of error bursts corresponding to the 2 error data pins is greater than the third threshold; and the number of error data pins corresponding to the data is greater than half of the number of data pins of the storage particles, and the error data pins all belong to the same storage particle in the memory.
[0005] For example, the method provided by at least one embodiment of the present disclosure also includes: writing the address information of the row address, the first count value and the second count value into a register array, wherein the register array includes one or more table entries, each table entry is used to store the address information, the first count value and the second count value of the corresponding row address.
[0006] For example, the method provided by at least one embodiment of the present disclosure also includes: setting a funnel decrement mechanism for the first count value and the second count value stored in each table entry in one or more table entries, wherein the funnel decrement frequency of the first count value is slower than the funnel decrement frequency of the second count value.
[0007] For example, the method provided by at least one embodiment of the present disclosure also includes: before writing the address information of the row address into the register array, determining whether there is a first table entry for the row address in the register array; and in response to the existence of the first table entry in the register array, updating the first count value or the second count value in the first table entry.
[0008] For example, in the method provided by at least one embodiment of the present disclosure, in response to the presence of a first table entry in a register array, updating the first count value or the second count value in the first table entry includes: in response to the correctable error meeting the characteristics of a bounded fault, adding 1 to the first count value in the first table entry; or in response to the correctable error not meeting the characteristics of a bounded fault, adding 1 to the second count value in the first table entry.
[0009] For example, the method provided by at least one embodiment of the present disclosure further includes: in response to the absence of the first table entry in the register array and the second table entry in the register array being in an idle state, writing the address information of the row address, the first count value and the second count value into the second table entry.
[0010] For example, the method provided by at least one embodiment of the present disclosure further includes: in response to the absence of the first table entry in the register array and all the table entries in the register array are in an occupied state, determining the table entry to be replaced in the register array, and writing the address information of the row address, the first count value, and the second count value into the table entry to be replaced.
[0011] For example, in the method provided by at least one embodiment of the present disclosure, the first count value and the second count value both include a low-order part and a high-order part, and determining the table entry to be replaced in the register array includes: determining the table entry in the register array whose first count value is 0, and selecting, from the table entry whose first count value is 0, the table entry whose high-order part of the second count value has the smallest value and whose second count value does not reach a second threshold as the table entry to be replaced.
[0012] For example, the method provided by at least one embodiment of the present disclosure further includes: in response to the absence of a table entry to be replaced in the register array, determining a target table entry from the register array, and using a row address corresponding to the target table entry as a target row address to perform a repair operation; and after performing the repair operation, using the target table entry as the table entry to be replaced.
[0013] For example, in the method provided by at least one embodiment of the present disclosure, the first count value and the second count value both include a low-order part and a high-order part, and determining a target table entry from the register array includes: in response to the absence of a table entry in the register array whose first count value is greater than a first threshold or a table entry whose second count value is greater than a second threshold, selecting a third table entry from the register array as the target table entry, wherein the third table entry is the table entry in the register array whose high-order part of the first count value has the largest value.
[0014] For example, in the method provided in at least one embodiment of the present disclosure, in response to the absence of an entry in the register array whose first count value is greater than a first threshold or an entry whose second count value is greater than a second threshold, selecting a third entry from the register array as a target entry includes: in response to the presence of multiple third entries in the register array, selecting an entry with the largest value of the high-order portion of the second count value from the multiple third entries as the target entry.
[0015] For example, in the method provided in at least one embodiment of the present disclosure, the repair operation includes a post-packaging repair operation, wherein the repair operation is performed on the target row address, including: based on the target row address, generating a command sequence of the post-packaging repair operation to repair the target row address.
[0016] At least one embodiment of the present disclosure further provides a memory controller. The memory controller is used to control a memory and includes a bounded fault detection module, a first counter, a second counter, and a repair module. The bounded fault detection module is configured to: in response to the presence of a correctable error in the data read for a row address of the memory, determine whether the correctable error meets the characteristics of a bounded fault. The first counter is configured to: in response to the correctable error meeting the characteristics of a bounded fault, perform a first count on the row address to obtain a first count value. The second counter is configured to: in response to the correctable error not meeting the characteristics of a bounded fault, perform a second count on the row address to obtain a second count value. The repair module is configured to, in response to the presence of a target row address in the memory, perform a repair operation on the target row address, wherein the first count value of the target row address is greater than a first threshold or the second count value is greater than a second threshold.
[0017] For example, in the memory controller provided in at least one embodiment of the present disclosure, the bounded fault detection module is configured to determine that the correctable error meets the characteristics of a bounded fault in response to the correctable error meeting any of the following conditions: the number of error data pins corresponding to the data is 1, and the number of error bursts corresponding to the error data pins is greater than a third threshold; the number of error data pins corresponding to the data is 2, the data read from the 2 error data pins belong to the same half byte, and the total number of error bursts corresponding to the 2 error data pins is greater than the third threshold; and the number of error data pins corresponding to the data is greater than half of the number of data pins of the storage particles, and the error data pins all belong to the same storage particle in the memory.
[0018] For example, at least one embodiment of the present disclosure provides a memory controller further comprising a register array. The register array is configured to store address information of a row address, a first count value, and a second count value in response to a correctable error in data read from a row address of the memory. The register array includes one or more entries, each entry being configured to store address information of a corresponding row address, a first count value, and a second count value.
[0019] For example, the memory controller provided by at least one embodiment of the present disclosure also includes a correctable error detection module and an address queue, wherein the address queue is configured to store address information corresponding to each read operation; the correctable error detection module is configured to: in response to determining that the data read for the first row address of the memory has the correctable error, send a first signal to the address queue so that the address queue sends the address information of the first row address to the register array; the bounded fault detection module is also configured to: in response to determining that the correctable error corresponding to the first row address meets the characteristics of a bounded fault, send a high-level second signal to the address queue; in response to determining that the correctable error corresponding to the first row address does not meet the characteristics of a bounded fault, send a low-level second signal to the address queue.
[0020] For example, in the memory controller provided in at least one embodiment of the present disclosure, the first count value and the second count value stored in one or more table entries are provided with a funnel decrement mechanism, and the funnel decrement frequency of the first count value is slower than the funnel decrement frequency of the second count value.
[0021] For example, the memory controller provided in at least one embodiment of the present disclosure further includes a table entry query module. When the table entry query module determines that a first table entry for a row address exists in the register array, in response to the bounded fault detection module determining that a correctable error in the data meets the characteristics of a bounded fault, the value of a first counter corresponding to the first table entry is incremented by 1; or, in response to the bounded fault detection module determining that the correctable error in the data does not meet the characteristics of a bounded fault, the value of a second counter corresponding to the first table entry is incremented by 1.
[0022] For example, the memory controller provided in at least one embodiment of the present disclosure further includes an address selection module. When the entry query module determines that the first entry for the row address does not exist in the register array and all entries in the register array are occupied, the address selection module is configured to: determine a to-be-replaced entry in the register array, where the to-be-replaced entry is used to store address information of the row address, the first count value, and the second count value.
[0023] For example, in the memory controller provided in at least one embodiment of the present disclosure, the first count value and the second count value both include a low-order part and a high-order part, and the address selection module is further configured to: determine the table entry in the register array whose first count value is 0, and select the table entry whose high-order part of the second count value has the smallest value and whose second count value does not reach the second threshold from the table entry whose first count value is 0 as the table entry to be replaced.
[0024] For example, in the memory controller provided in at least one embodiment of the present disclosure, the address selection module is further configured to: in response to determining that there is no table entry to be replaced in the register array, determine a target table entry from the register array, and use the row address corresponding to the target table entry as the target row address; and after the repair module performs a repair operation on the target row address, use the target table entry as the table entry to be replaced.
[0025] At least one embodiment of the present disclosure further provides a memory system, which includes a memory and a memory controller according to any embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, rather than limiting the present disclosure.
[0027] Figure 1 FIG. 2 shows an exemplary schematic diagram of repairing DRAM based on the PPR mechanism;
[0028] Figure 2A A schematic diagram of an exemplary DRAM memory cell array is shown;
[0029] Figure 2B A schematic diagram showing an exemplary memory controller accessing a memory is shown;
[0030] Figure 3 A flowchart illustrating a method for processing a memory failure according to at least one embodiment of the present disclosure is shown;
[0031] Figure 4A A schematic diagram illustrating a memory controller accessing a memory according to at least one embodiment of the present disclosure is shown;
[0032] Figure 4B Shown Figure 4A A schematic diagram of the structure of the register array in the memory controller shown;
[0033] Figure 5 A flowchart of detecting and repairing a faulty row address according to at least one embodiment of the present disclosure is shown;
[0034] Figure 6 A flowchart of repairing row addresses provided by at least one embodiment of the present disclosure is shown;
[0035] Figure 7 A schematic diagram illustrating a memory controller according to at least one embodiment of the present disclosure is shown;
[0036] Figure 8 A schematic diagram of a memory system provided by at least one embodiment of the present disclosure is shown;
[0037] Figure 9 A schematic diagram illustrating an electronic device according to at least one embodiment of the present disclosure; and
[0038] Figure 10 A schematic diagram of a non-transitory readable storage medium according to at least one embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0039] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0040] Unless otherwise defined, the technical or scientific terms used in this disclosure should have the usual meanings understood by people with ordinary skills in the field to which this disclosure belongs. The words "first", "second" and similar words used in this disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "one", "an" or "the" do not indicate a quantity limitation, but rather indicate the existence of at least one. Words such as "include" or "comprise" mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0041] Flowcharts are used in this disclosure to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0042] First, the abbreviations and related terms involved in this application are defined and explained.
[0043] Dynamic Random Access Memory (DRAM): A type of semiconductor memory that temporarily stores data in computer systems. DRAM consists of multiple memory cells, each of which represents binary data by storing charge on a capacitor. The data is retained by periodically refreshing the capacitor to prevent data loss.
[0044] Double Data Rate (DDR): DDR5 generally refers to the fifth generation of double data rate synchronous dynamic random access memory, a type of high-speed computer memory based on DRAM technology.
[0045] Post-Package Repair (PPR): A technology defined in the DDR5 JEDEC memory specification that repairs memory row faults after DRAM packaging using a fixed instruction flow. PPR technologies include hard Post-Package Repair (hPPR) and soft Post-Package Repair (sPPR).
[0046] Hard Post-Package Repair (hPPR): A technology that permanently repairs faulty memory cells after DRAM packaging is completed through physical means, and the repaired information remains valid even after power is removed.
[0047] Soft Post-Package Repair (sPPR): A technology that repairs faulty memory cells by configuring the DRAM's internal logic after the DRAM is packaged. The repair information becomes invalid after power is removed.
[0048] Error checking and correcting (ECC): A technology used to detect and correct errors during data transmission or storage. ECC adds extra check bits to the original data. When reading or receiving data, it performs an algorithm on the check bits and data bits to detect errors. If an error is detected, the error location is determined based on the check information and automatically corrected.
[0049] UE (Uncorrectable Error): Uncorrectable error.
[0050] CE (Correctable Error): Correctable error.
[0051] Soft error: An error that is not permanent and is caused by, for example, temporary environmental factors or transient disturbances within the system.
[0052] Bounded fault: A data error mode defined in the JEDEC memory specification for DDR5. For example, in a storage system, if a memory cell fails and the resulting DQ error is confined to a finite number of DQs, the fault is called a bounded fault.
[0053] Data Quadrature (DQ): Also known as "data pin" or DQ interface, it is a set of parallel data lines for data transmission between DRAM and external devices.
[0054] Nibble: One nibble contains 4 bits of data, and one byte contains 2 nibbles.
[0055] Data Device: A storage device that stores data, such as DDR memory, also known as a data memory device.
[0056] ECC device: A device that stores parity bits. For example, in DDR memory, it is also called an ECC device.
[0057] Storage granules: Data storage granules and parity storage granules are collectively referred to as storage granules. For example, storage granules in DDR memory are also called memory granules, and can be used as a collective term for data memory granules and parity memory granules.
[0058] It is understood that the terms defined above are merely exemplary definitions in specific application scenarios to facilitate a better understanding of the present application. For example, the exemplary definitions described above for a specific memory can be extended to other types of memory.
[0059] Figure 1 FIG. 1 shows an exemplary schematic diagram of repairing DRAM based on the PPR mechanism. Figure 1 The schematic diagram shown involves the application scenario of DRAM memory, see Figure 1 The described storage unit may correspond to a memory unit, and the memory controller may correspond to a memory controller.
[0060] like Figure 1 As shown, DRAM 100 includes a plurality of memory cells, such as Figure 1 The plurality of storage units are divided into a main storage unit and a redundant storage unit. Figure 1 As shown, the memory cells to the left of the dashed line (such as memory cells 102 and 105) are primary memory cells, while the memory cells to the right of the dashed line (such as memory cell 103) are redundant memory cells. A one-to-one mapping relationship exists between primary memory cells and logical row addresses, allowing precise location of the corresponding primary memory cell through the logical row address, thereby achieving efficient data read and write operations. There is no mapping relationship between redundant memory cells and logical row addresses. Under normal circumstances, redundant memory cells do not directly participate in regular data storage and access processes, but instead serve as backup resources for the primary memory cells.
[0061] For example, when a failure occurs in the row address Address_102 corresponding to the storage unit 102, the repair module 111 in the memory controller 110 performs a repair operation (such as a PPR operation) on the row address Address_102. That is, the repair module 111 remaps the row address Address_102 corresponding to the storage unit 102 to a redundant storage unit (such as the storage unit 103), so that the memory controller 110 can still normally access the row address Address_102 for read and write operations without causing read and write errors.
[0062] For example, PPR operations can include sPPR mode and hPPR mode. In sPPR mode, DRAM 100 retains the data in the memory array (i.e., multiple memory cells). After the sPPR operation is completed, the memory controller 110 can perform normal read and write operations on DRAM 100. The time required for the sPPR operation is the same as the time required for normal row address access. After the storage system is powered off or restarted, the repair information corresponding to the sPPR operation (e.g., the mapping relationship between row address Address_102 and memory cell 103) becomes invalid. In hPPR mode, DRAM 100 cannot retain the data in the memory array. That is, the data in the memory array disappears after the hPPR operation is performed. The hPPR operation takes much longer than the sPPR operation. After the storage system is powered off or restarted, the repair information corresponding to the hPPR operation (e.g., the mapping relationship between row address Address_102 and memory cell 103) remains valid.
[0063] For example, Figure 2A A schematic diagram of an exemplary DRAM memory cell array is shown.
[0064] like Figure 2A As shown, the DRAM includes a plurality of memory cells 220_0 to 220_15, a main wordline driver 200, and a plurality of sub-wordline drivers 210_0 to 210_14. The main wordline driver 200 and the sub-wordline drivers 210_0 to 210_14 are coupled via a global wordline. The sub-wordline drivers 210_0 to 210_14 and the plurality of memory cells 220_0 to 220_15 are coupled via local wordlines and in accordance with the Figure 2A Coupled as shown.
[0065] When the memory controller accesses a row address, the main word line driver 200 corresponding to the row address drives multiple sub-word line drivers 210_0~210_14, and the multiple sub-word line drivers 210_0~210_14 respectively drive the corresponding coupled storage cells, so that the capacitance detection circuit (LSA) 230_0~230_15 can respectively read data from the corresponding coupled storage cells and send the read data to the parallel converter (SERDES) 240_0~240_15. The parallel converter can convert parallel data into serial data, and then transmit the serial data through the DQ interface (or data pin). When the main word line driver 200 fails, all data pins (i.e., DQ0~DQ7) will have read and write errors. When the sub-word line driver fails, some data pins will have read and write errors. For example, when Figure 2A When the sub-word line driver 210_13 in the embodiment fails, read and write errors occur on the data pins DQ6 and DQ7 coupled to the sub-word line driver 210_13.
[0066] Errors caused by failures of the main word line driver or sub-word line driver are all permanent errors. Although the ECC module of the memory controller can correct erroneous data (i.e., data transmitted on the data pins where read / write errors occur) through the ECC algorithm, this weakens the ECC module's ability to correct other occasional errors, thus increasing the probability of UE occurring in the storage system. Figure 2B This is described exemplarily.
[0067] 2B shows a schematic diagram of an exemplary memory controller accessing a memory.
[0068] like Figure 2B As shown, the memory includes four data storage particles 220 to 223 and one verification storage particle 224, wherein the structure of the data storage particle 220 can refer to Figure 2A . "X8" means that the data bus width of the storage particle is 8 bits and includes 8 data pins (or 8 DQ interfaces). "BL=16" means that the burst length of the storage particle is 16. The data storage particle 220 transmits 8 bits of data through data pins DQ0~DQ7 in one read operation, where the data transmitted by data pins DQ0~DQ3 (i.e. DQ[3:0]) belongs to the same half byte (e.g. Nibble A), and the data transmitted by data pins DQ4~DQ7 (i.e. DQ[7:4]) belongs to another half byte (e.g. Nibble B). The data storage particle 220 transmits a total of 16 Burst*8 bit=128 bit data in one read operation.
[0069] The memory controller 210 includes a data check module 211, a data correction unit 212, a data receiver 213, and an ECC data receiver 214. The data receiver 213 is coupled to the data storage particles 220-223 to receive data transmitted by data pins DQ0-DQ31 (i.e., DQ[31:0]). The ECC data receiver 214 is coupled to the parity storage particle 224 to receive parity data transmitted by data pins DQ_ECC0-DQ_ECC7 (i.e., DQ_ECC[7:0]).
[0070] For example, when Figure 2AWhen a sub-wordline driver 210_13 in the memory fails, read and write errors occur on data pins DQ6 and DQ7. The data verification module 211 detects errors in the data transmitted by data pins DQ6 and DQ7, and the data correction module 212 corrects the erroneous data. In this case, if errors (such as sporadic or fixed errors) occur in other storage cells (e.g., data storage cells 223) that exceed the error correction capability of the data correction module 212, the probability of a UE occurring in the storage system increases. Furthermore, the data verification module 211 and the data correction module 212 can be incorporated into an ECC module, which is used to detect data errors and correct any erroneous data.
[0071] The inventors of the present disclosure have noted that due to the limited repair capabilities of the ECC module of a storage system (e.g., a memory system), the probability of a UE occurring in the storage system increases after a faulty row address occurs. An exemplary storage system typically interrupts normal operation of the storage system upon detecting an uncorrectable error (UE), monitors the faulty row address through a memory self-detection mode, and then repairs the faulty row address by performing a PPR operation. However, this method requires initiating the detection and repair of the faulty row address based on the reporting of the UE error, i.e., the detection and repair of the faulty row address is initiated only after a UE occurs. Therefore, after a row address fails, this method is unable to promptly detect and repair the faulty row address.
[0072] Another exemplary method is to record the row address where a correctable error (CE) occurs, and perform a PPR operation on the corresponding row address when the CE count value exceeds a threshold. However, CE in a storage system may also be caused by random errors or column address failures within the DRAM. Therefore, based on the CE threshold, it is impossible to accurately determine whether a row address failure has occurred within the DRAM. For example, when a column address failure occurs in a DRAM, the CE count value may also exceed the threshold, thereby triggering a repair operation (e.g., a PPR operation). However, the PPR operation cannot repair the column address failure and will also perform unnecessary repair operations on non-faulty row addresses, resulting in a waste of PPR resources.
[0073] One or more embodiments of the present disclosure provide a method for handling memory faults, a memory controller, and a memory system. By recording the row address where CE (or fault) occurs and identifying whether the row address meets the characteristics of a bounded fault, the faulty row address can be detected more accurately and the detected faulty row address can be repaired in a timely manner, thereby reducing the probability of UE occurring in the memory system and improving the RAS performance of the memory system.
[0074] The present disclosure is described below using several specific embodiments. To keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and components may be omitted. When any component of an embodiment of the present disclosure appears in more than one drawing, the component is represented by the same or similar reference numeral in each drawing.
[0075] Figure 3 A flowchart of a method for handling a memory failure according to at least one embodiment of the present disclosure is shown. The method may include steps S310 to S340.
[0076] Step S310 : In response to a correctable error in data read from a row address of a memory, determining whether the correctable error meets the characteristics of a bounded fault.
[0077] Step S320 : In response to the correctable error meeting the characteristics of a bounded fault, performing a first count on the row address to obtain a first count value.
[0078] Step S330 : In response to the correctable error not meeting the characteristics of a bounded fault, performing a second count on the row address to obtain a second count value.
[0079] Step S340 : in response to the target row address existing in the memory, performing a repair operation on the target row address, wherein the first count value of the target row address is greater than the first threshold or the second count value is greater than the second threshold.
[0080] In step S310, data is read from the memory based on the row address. If a correctable error is detected in the read data, the row address is referred to as a faulty row address, and further determination is made as to whether the correctable error meets the characteristics of a bounded fault. A correctable error (CE) is an error within the error correction capability of the ECC algorithm, while an uncorrectable error (UE) is an error beyond the error correction capability of the ECC algorithm. Bounded faults are a data error pattern defined in the JEDEC memory specification for DDR5; their specific definition is not detailed here.
[0081] In some embodiments of the present disclosure, when a correctable error meets any of the following conditions, it can be determined that the correctable error meets the characteristics of a bounded fault.
[0082] Condition 1: the number of error data pins corresponding to the data is 1, and the number of error bursts corresponding to the error data pins is greater than a third threshold.
[0083] Condition 2: the number of error data pins corresponding to the data is 2, the data read from the two error data pins belongs to the same half byte, and the total number of error bursts corresponding to the two error data pins is greater than a third threshold.
[0084] Condition 3: The number of erroneous data pins corresponding to the data is greater than half the number of data pins of the memory cell, and the erroneous data pins all belong to the same memory cell in the memory.
[0085] For example, the "erroneous data pin" in the above condition is an "erroneous DQ," and the data transmitted by the erroneous DQ contains errors. If the data transmitted by the data pin within a burst contains errors, the burst is considered an "erroneous burst" in the above condition.
[0086] For example, the third threshold may be 4, or may be set according to actual needs, or may be a suitable value obtained based on model training or experience, and this disclosure does not impose any restrictions on this.
[0087] For example, Figure 2B As shown, when the error DQ corresponding to the data is DQ6 and the error burst number corresponding to DQ6 is greater than the third threshold, the correctable error in the data meets condition 1. Therefore, the correctable error in the data meets the characteristics of a bounded fault.
[0088] For example, Figure 2A As shown, when the sub-word line driver 210_13 fails, the data pins DQ6 and DQ7 have read and write errors. At this time, the number of error DQs corresponding to the data is 2. Figure 2B As shown in Figure 1, the data read from the two erroneous data pins DQ6 and DQ7 belong to the same half-byte. If the total number of error bursts corresponding to DQ6 and DQ7 is greater than the third threshold, the correctable error in the data meets condition 2 and therefore meets the characteristics of a bounded fault.
[0089] For example, Figure 2B As shown, the number of data pins of each storage particle (X8 DRAM) is 8. When the number of erroneous DQs corresponding to the data is greater than 4 (for example, the erroneous DQs are DQ0, DQ1, DQ3, DQ6, and DQ7), since DQ0, DQ1, DQ3, DQ6, and DQ7 all belong to the same storage particle 220 in the memory, the correctable error in the data meets condition 3. Therefore, the correctable error in the data meets the characteristics of a bounded fault.
[0090] In some embodiments of the present disclosure, a register array including one or more registers may be provided to store address information of a faulty row address, a first count value, and a second count value, where the "faulty row address" refers to a row address from which a correctable error is detected in the data read. Figure 4A and Figure 4B This is described exemplarily.
[0091] Figure 4AFIG. 1 shows a schematic diagram of a memory controller accessing a memory according to at least one embodiment of the present disclosure. Figure 4B Shown Figure 4A The schematic diagram of the structure of the register array in the memory controller is shown in FIG.
[0092] like Figure 4B As shown, the register array 415 in the memory controller 420 includes one or more entries (e.g., entries Entry0 to Entry2), each of which is used to store address information of a corresponding row address, a first count value, and a second count value. The address information of the row address may include information such as the memory chip code (DRAM Device ID), the bank group (Bank group), the bank address (Bank Address), and the row address (Row Address).
[0093] For example, the address information of the row address may further include DQ group information (DQ group) of the erroneous data pin. The DQ group information of the erroneous DQ is used to indicate the DQ group where the erroneous DQ is located or the position where the erroneous DQ is located.
[0094] For example, for a x4 memory chip, each memory chip includes four data pins (such as DQ0, DQ1, DQ2, and DQ3). The four data pins can be divided into four groups, each including one DQ. Alternatively, DQ0 and DQ1 can be used as the first group, and DQ2 and DQ3 as the second group. The embodiments of the present disclosure do not limit the division method of DQ groups.
[0095] For example, Figure 4B As shown, the first field of each table entry (also referred to as the "bounded fault field") is used to record the first count value of the corresponding row address, and the first count value indicates the number of times that the data read from the row address contains a correctable error and the correctable error meets the characteristics of a bounded fault. The second field in each table entry (also referred to as the "correctable error field") is used to record the second count value of the corresponding row address, and the second count value indicates the number of times that the data read from the row address contains a correctable error and the correctable error does not meet the characteristics of a bounded fault. Both the first field and the second field include a high-order portion and a low-order portion. The number of bits in the high-order portion may be the same as or different from the number of bits in the low-order portion, and the present disclosure does not impose any restrictions on this.
[0096] In step S340, when the target row address exists in the memory, that is, when there is a table entry in the register array whose first count value is greater than the first threshold or the second count value is greater than the second threshold, the row address corresponding to the table entry is the target row address, and the repair operation of the target row address will be triggered.
[0097] In some embodiments, the first threshold and the second threshold can be set according to actual needs, or can be appropriately set based on model training or empirically obtained values. In addition, the first threshold can be smaller than the second threshold. Of course, this disclosure does not limit this.
[0098] For example, the repair operation in the embodiment of the present disclosure includes a post-packaging repair (PPR) operation. The post-packaging repair (PPR) operation may be an sPPR operation or an hPPR operation, which is not limited in the present disclosure.
[0099] For example, the step of performing a repair operation on the target row address may include: based on the target row address, generating a command sequence for the post-packaging repair operation to repair the target row address. For example, the generated command sequence may establish a new mapping relationship for the target row address. For example, Figure 4A As shown, the target row address Address_402 is remapped from the main storage unit 402 to a new storage unit (eg, the redundant storage unit 403 ), thereby repairing the target row address.
[0100] In some embodiments of the present disclosure, taking into account the possibility that the determined fault row address may be misjudged, for example, the fault row address is caused by an occasional error, a leak self-decrement mechanism can be set for the first count value and the second count value stored in each table entry in the register array, so that the first count value and the second count value are automatically reduced at a certain (for example, fixed or random) frequency, thereby reducing the probability of misjudging the fault row address.
[0101] For example, the first count value and the second count value may be decreased every predetermined period of time.
[0102] For example, considering that the steps for determining whether a correctable error meets the characteristics of a bounded fault are more stringent and the probability of false positives is lower, the funnel decrement frequency of the first count value can be slower than the funnel decrement frequency of the second count value. For example, the first count value stored in each entry can be decremented by 1 every 4 hours, and the second count value can be decremented by 1 every 2 hours.
[0103] It should be noted that Figure 4B The register array 415 structure shown in FIG is merely illustrative, and other register array structures, such as SRAM, may also be used, and the present disclosure does not limit this.
[0104] In some embodiments of the present disclosure, before writing the address information of the row address corresponding to the data with correctable errors into the register array, it can be determined whether there is a first table entry for the row address in the register array; in response to the existence of the first table entry in the register array, the first count value or the second count value in the first table entry is updated.
[0105] For example, in Figure 4B In the case where the first table entry Entry1 for the row address Address_402 exists in the register array 415 shown, for step S320, in response to the presence of a correctable error in the data read from the storage unit 402 corresponding to the row address Address_402, and the correctable error meets the characteristics of a bounded fault, a first count is performed on the row address Address_402 (for example, the first count value currently recorded in the first table entry Entry1 is increased by 1) to obtain the first count value of the row address Address_402.
[0106] For example, in Figure 4B In the case where the first table entry Entry1 for the row address Address_402 exists in the register array 415 shown, for step S330, in response to the presence of a correctable error in the data read from the storage unit 402 corresponding to the row address Address_402, and the correctable error does not meet the characteristics of a bounded fault, a second count is performed on the row address Address_402 (for example, the second count value currently recorded in the first table entry Entry1 is increased by 1) to obtain a second count value of the row address Address_402.
[0107] It should be noted that if the address information stored in the register array also includes the DQ group information of the erroneous DQ, then after determining that there is a first table entry for the row address in the register array, it can also be determined whether the DQ group information of the current erroneous DQ of the row address is the same as the DQ group information recorded in the first table entry.
[0108] If the DQ group information of the current error DQ is the same as the DQ group information recorded in the first table entry, the first count value or the second count value in the first table entry can be directly increased by 1. If the DQ group information of the current error DQ is different from the DQ group information recorded in the first table entry, a new table entry can be used to record the address information of the row address, and the value of the first field or the value of the second field in the new table entry is set to 1 based on whether the correctable error meets the characteristics of a bounded error.
[0109] In some embodiments of the present disclosure, when Figure 4B When the first table entry Entry1 for the row address Address_402 does not exist in the register array 415 shown, if there is an idle second table entry (for example, Entry2) in the register array 415, the address information, the first count value and the second count value of the row address Address_402 can be directly written into the idle second table entry Entry2.
[0110] For example, for step S320, in response to the presence of a correctable error in the data read from the storage unit 402 corresponding to the row address Address_402, and the correctable error meets the characteristics of a bounded fault, a first count is performed on the row address Address_402, for example, the first count value in the second table entry Entry2 is set from an initial value (for example, 0) to 1, and the second count value in the second table entry Entry2 remains the initial value.
[0111] For example, for step S330, in response to the presence of a correctable error in the data read from the storage unit 402 corresponding to the row address Address_402, and the correctable error does not meet the characteristics of a bounded fault, a second count is performed on the row address Address_402, that is, the second count value in the second table entry Entry2 is set from the initial value (for example, 0) to 1, and the first count value in the second table entry Entry2 remains the initial value.
[0112] In some embodiments of the present disclosure, Figure 4B The capacity of the register array 415 shown may be limited, and the register array 415 may contain neither the first entry Entry1 for the row address Address_402 nor the second entry in an idle state. In this case, the entry to be replaced in the register array 415 can be first determined, and then the address information, first count value, and second count value of the row address Address_402 can be written into the entry to be replaced. For example, the address information, first count value, and second count value in the entry to be replaced can be first deleted, and then the address information, first count value, and second count value of the row address Address_402 can be written into the entry to be replaced, or the address information, first count value, and second count value of the row address Address_402 can be directly overwritten into the entry to be replaced.
[0113] For example, the step of determining the entry to be replaced in the register array may include: determining an entry in the register array whose first count value is 0, and selecting, from the entries whose first count value is 0, an entry whose high-order portion of the second count value has the smallest value and whose second count value does not reach a second threshold as the entry to be replaced.
[0114] It should be noted that the “entry to be replaced” here may refer to an entry occupied by a row address with a low probability of triggering the UE, or the entry to be replaced may be determined in other ways, which is not limited in the present disclosure.
[0115] In the method for processing memory failures provided in an embodiment of the present disclosure, by deleting the address information in the table entry to be replaced and writing the address information of the new fault row address, it is possible to record as much address information of the fault row address (for example, the fault row address with a higher probability of triggering the UE) as possible when the capacity of the register array is limited, thereby reducing the probability of triggering the UE.
[0116] In some embodiments of the present disclosure, when there is no table entry to be replaced in the register array, a target table entry can be determined from the register array, and the row address corresponding to the target table entry can be used as the target row address to perform a repair operation; after performing the repair operation, the target table entry can be used as the table entry to be replaced.
[0117] For example, the step of determining a target table entry from the register array may include: in response to the absence of a table entry in the register array having a first count value greater than a first threshold or a table entry in which a second count value is greater than a second threshold, selecting a third table entry from the register array as the target table entry, wherein the third table entry is the table entry in the register array having the largest value of a high-order portion of the first count value.
[0118] For example, when determining a target entry from the register array, in response to the presence of multiple third entries in the register array (e.g., the value of the high-order portion of the entries Entry 0, Entry 1, and Entry 5 is the largest), the entry with the largest value of the high-order portion (e.g., the upper 4 bits) of the second count value is selected as the target entry from the entries Entry 0, Entry 1, and Entry 5. If the values of the high-order portion of the second count value of the entries Entry 0, Entry 1, and Entry 5 are 5 (i.e., 0101), 6 (i.e., 0110), and 6 (i.e., 0110), respectively, then an entry is arbitrarily selected from the entries Entry 1 and Entry 5 as the target entry.
[0119] It should be noted that the “target table entry” here may refer to the table entry occupied by the row address with a high probability of triggering the UE, or the target table entry may be determined from the register array in other ways, which is not limited in the present disclosure.
[0120] It should be noted that the above embodiment only compares the high-order parts of the first count value and the second count value, and ignores the low-order parts of the first count value and the second count value, that is, ignores a few erroneous differences.
[0121] For example, the first count value and the second count value each include 8 bits. In response to the presence of four entries in the register array with a first count value of 0 (e.g., Entry 0, Entry 3, Entry 4, and Entry 6), the upper 6 bits of the second count values of Entry 0, Entry 3, Entry 4, and Entry 6 are compared. If the upper 6 bits of the second count values of Entry 0 and Entry 6 are the smallest (e.g., both are 000001), and the second count values of Entry 0 and Entry 6 do not reach a second threshold, then both Entry 0 and Entry 6 are selected as entries to be replaced. During this comparison, the lower 2 bits of the second count values of Entry 0, Entry 3, Entry 4, and Entry 6 are ignored, thereby obscuring 0 to 3 (0 to 22 - 1) error differences. That is, the difference between the second count values of Entry 0 and Entry 6 ranges from 0 to 3.
[0122] The method of comparing the high-order parts of the first count value and the second count value, and the method of dividing the high-order part and the low-order part in the above embodiment are merely illustrative. All bits of the first count value and the second count value can also be compared to determine the table entry to be replaced or the target table entry. The present disclosure does not impose any restrictions on this.
[0123] In the method for processing memory failures provided in an embodiment of the present disclosure, by determining a target table entry in a register array and performing a repair operation on the row address corresponding to the target table entry in advance, a table entry to be replaced is obtained to write a new row address. This allows for recording as much address information of the faulty row address as possible when the capacity of the register array is limited, and for repairing the row address that is prone to triggering the UE in advance to reduce the probability of triggering the UE.
[0124] For example, Figure 5 FIG. 1 shows a flow chart of detecting and repairing a faulty row address according to at least one embodiment of the present disclosure. It can be understood that Figure 5 This is merely an exemplary description in a scenario where the memory is used as an example of storage, and the present disclosure is not limited thereto. Figure 5 The described flowcharts can be applied to other memories with or without modification.
[0125] like Figure 5 As shown, step S501: turning on the power to start the memory system and perform related initialization operations.
[0126] Step S502: the memory system enters a normal operating mode.
[0127] Step S503: Determine whether the read and written data is correct.
[0128] If the read and written data are correct, the process returns to step S502 and continues to execute the normal working mode.
[0129] If the read or written data is incorrect, step S504 is executed.
[0130] Step S504: determining whether the error in the currently read or written data is a correctable error (CE).
[0131] If the error in the data is not a CE but an uncorrectable error (UE), step S505 is executed (ie, restarting the memory system).
[0132] If the error in the data belongs to CE, steps S506 and S507 are executed.
[0133] Step S506: Correct the current data.
[0134] Step S507: Generate a fault flag and record the address information of the row address corresponding to the current data. The fault flag indicates whether the CE in the current data meets the characteristics of a bounded fault. The steps for determining whether the CE meets the characteristics of a bounded fault can refer to conditions 1 to 3 of the above embodiment and are not repeated here.
[0135] Step S508: Determine whether there is a first table entry for the row address corresponding to the current data in the register array.
[0136] If the first table entry for the row address corresponding to the current data exists in the register array, step S509 is executed; otherwise, step S512 is executed.
[0137] Step S509: Determine the count value that needs to be updated according to the fault flag.
[0138] If it is determined according to the fault flag that the CE meets the characteristics of a bounded fault, step S510 is executed to increase the first count value in the first entry (such as the value of the first field in the first entry) by 1.
[0139] If it is determined according to the fault flag that the CE does not meet the characteristics of a bounded fault, step S511 is executed to increase the second count value in the first entry (such as the value of the second field in the first entry) by 1.
[0140] Step S512: When the register array does not have a first table entry for the row address corresponding to the current data, determine whether the register array has a second table entry. The second table entry is any table entry in the register array that is in an idle state.
[0141] If the register array has a second entry, step S513 is executed to store the row address of the current data into the second entry, and the value of the corresponding field in the second entry (the value of the first field or the value of the second field) is set to 1 according to the fault flag.
[0142] If the second entry does not exist in the register array, step S514 is executed.
[0143] Step S514: Query the register array to see whether there is an entry to be replaced. The specific steps can be referred to the above embodiment and will not be repeated here.
[0144] If there is an entry to be replaced in the register array, step S515 is executed.
[0145] Step S515: deleting the information stored in the entry to be replaced, storing the address information of the row address of the current data into the entry to be replaced, and setting the value of the corresponding field in the entry to be replaced to 1 according to the fault flag.
[0146] After executing step S510 , S511 , S513 or S515 , execute step S516 .
[0147] Step S516: Determine whether there is a target row address in the register array whose first field value is greater than the first threshold or whose second field value is greater than the second threshold.
[0148] If the target row address exists in the register array, step S518 is executed to start performing a repair operation on the target row address.
[0149] If there is no entry to be replaced in the register array, step S517 is executed.
[0150] Step S517: Determine the target entry in the register array and perform a repair operation on the row address occupied by the target entry. The specific steps of determining the target entry can refer to the above embodiment and will not be repeated here.
[0151] After executing step S517 , step S518 is executed to start performing a repair operation on the row address of the occupied target entry.
[0152] Figure 6 A flowchart of repairing a row address provided by at least one embodiment of the present disclosure is shown.
[0153] like Figure 6 As shown, the SPPR operation on the target row address is taken as an example.
[0154] Step S600: Start to perform the repair operation. Figure 5 The process is similar to step 518 in FIG.
[0155] Step S601: A row address that meets the repair condition is selected as the address to be repaired in the subsequent SPPR operation. Here, the row address that meets the repair condition is the target row address in the above-mentioned embodiment (i.e., the row address where the first count value is greater than the first threshold or the second count value is greater than the second threshold), or the row address of the target entry in the register array, and will not be further described here.
[0156] Step S602: Based on the selected row address, generate an SPPR command sequence to complete the SPPR operation on the row address. That is, execute the SPPR standard process in DDR5 JEDEC on the selected row address.
[0157] Step S603: Periodically send a refresh instruction to the memory. Step S603 is used to keep the data in the memory unchanged, thereby avoiding data loss in the memory. If it is not necessary to keep the data in the memory unchanged, step S603 can be skipped.
[0158] Step S604: Determine whether there are other row addresses that meet the repair condition.
[0159] If there are other row addresses that meet the repair condition, the process returns to step S601 and selects one row address from the remaining row addresses that meet the repair condition as the address for subsequent SPPR operation repair. If there are no other row addresses that meet the repair condition, the SPPR operation ends.
[0160] A method for processing a memory failure according to at least one embodiment of the present disclosure (eg Figure 3 Correspondingly, at least one embodiment of the present disclosure further provides a memory controller, which is used to control the memory.
[0161] Figure 7 FIG. 1 is a schematic diagram of a memory controller 700 according to at least one embodiment of the present disclosure.
[0162] See also Figure 7 , the memory controller 700 includes a bounded fault detection module 710 , a first counter 730 , a second counter 740 , and a repair module 720 .
[0163] In some embodiments, the memory controller 700 may refer to Figure 4A The memory controller 420 described herein, wherein the bounded fault detection module 710 and the repair module 720 may correspond to the memory controller 420 of FIG. Figure 4A The bounded fault detection module 413 and the repair module 417 are described.
[0164] In additional and alternative aspects, the memory controller 700 may correspond to Figure 4AThe memory controller 420 described above and the memory controller 700 may optionally include a register array 415 , an address selection module 416 , and a refresh module 418 .
[0165] The bounded fault detection module 413 is configured to, in response to a correctable error in data read from a row address (eg, Address_402 ) of the memory 400 , determine whether the correctable error meets the characteristics of a bounded fault.
[0166] The first counter Counter1 (corresponding to Figure 7 The first counter 730 is configured to perform a first count on the row address in response to the correctable error meeting the characteristics of the bounded fault, and obtain a first count value. The second counter Counter2 (corresponding to Figure 7 The second counter 740 is configured to perform a second count on the row address to obtain a second count value in response to the correctable error not meeting the characteristics of a bounded fault.
[0167] The repair module 417 is configured to perform a repair operation on the target row address in response to the target row address existing in the memory, wherein the first count value of the target row address is greater than the first threshold or the second count value is greater than the second threshold.
[0168] For example, the bounded fault detection module 413 can be configured to determine that a correctable error meets the characteristics of a bounded fault in response to the correctable error meeting any of the following conditions: the number of error data pins corresponding to the data is 1, and the number of error bursts corresponding to the error data pins is greater than a third threshold; the number of error data pins corresponding to the data is 2, the data read from the 2 error data pins belong to the same half byte, and the total number of error bursts corresponding to the 2 error data pins is greater than the third threshold; and the number of error data pins is greater than half of the number of data pins of the storage particle, and the error data pins belong to the same storage particle in the memory.
[0169] For example, the memory controller 420 may further include a register array 415. The register array 415 is configured to store address information of a row address (e.g., Address_402), a first count value, and a second count value in response to a correctable error in data read from a row address (e.g., Address_402) of the memory 400. The register array 415 includes one or more entries (e.g., Entry0, Entry1, Entry2, etc.), each of which is used to store address information of a corresponding row address, a first count value, and a second count value.
[0170] For example, one or more entries may share the first counter Counter1 (corresponding to Figure 7"first counter 730") and the second counter Counter2 (corresponding to Figure 7 A first counter and a second counter may also be set for each entry to perform a first counting operation and a second counting operation on the row address corresponding to each entry, and the present disclosure does not impose any limitation on this.
[0171] For example, the first count value and the second count value stored in each entry are provided with a funnel self-decrement mechanism, and the funnel self-decrement frequency of the first count value is slower than the funnel self-decrement frequency of the second count value.
[0172] For example, Figure 4A As shown, the memory controller 420 may also include an ECC module 411, a correctable error detection module 412 (also referred to as "CE detection module 412"), and an address queue 414. The address queue 414 is configured to store address information for each read operation. The ECC module 411 is configured to detect whether the data read for a row address (e.g., Address_402) of the memory 400 contains errors. If the data contains errors, the ECC module 411 corrects the erroneous data based on an ECC algorithm, and the CE detection module 412 is configured to detect the error type of the erroneous data. If the CE detection module 412 determines that the error is a CE error, it sends a first signal to the address queue 414, causing the address queue 414 to send the address information of the row address (e.g., Address_402) to the register array 415. After the CE detection module 412 determines that the error is a CE, the bounded fault detection module 413 further detects whether the CE meets the characteristics of a bounded fault. If the CE meets the characteristics of a bounded fault, the bounded fault detection module 413 may send a high-level second signal to the address queue 414 so that when the address information of the row address (e.g., Address_402) is written into the table entry (e.g., Entry0) of the register array 415, the first count value of the row address (e.g., Address_402) is incremented by 1 (e.g., the value of the first field of Entry0 is incremented by 1). If the CE does not meet the characteristics of a bounded fault, the bounded fault detection module 413 may send a low-level second signal to the address queue 414 so that when the address information of the row address (e.g., Address_402) is written into the table entry (e.g., Entry0) of the register array 415, the second count value of the row address (e.g., Address_402) is incremented by 1 (e.g., the value of the second field of Entry0 is incremented by 1).
[0173] For example, the memory controller 420 provided by an embodiment of the present disclosure further includes an entry query module 419. When the entry query module 419 determines that a first entry (e.g., Entry 1) for a row address (e.g., Address_402) exists in the register array 415, in response to the bounded fault detection module 413 determining that a correctable error in the data meets the characteristics of a bounded fault (i.e., the second signal is at a high level), the first count value in the first entry (e.g., the value of the first field in Entry 1) is incremented by 1; or in response to the bounded fault detection module 413 determining that the correctable error in the data does not meet the characteristics of a bounded fault (i.e., the second signal is at a low level), the second count value in the first entry (e.g., the value of the second field in Entry 1) is incremented by 1.
[0174] For example, when the entry query module 419 determines that the first entry (e.g., Entry 1) does not exist in the register array 415 and the second entry (e.g., Entry 2) in the register array is in an idle state, the address queue 414 writes the address information of the row address (e.g., Address_402) into the second entry Entry 2. The initial values of the first count value and the second count value of the idle second entry Entry 2 are both 0. If the second signal corresponding to the current read operation is at a high level, the current first count value in the second entry Entry 2 is incremented by 1. If the second signal corresponding to the current read operation is at a low level, the current second count value in the second entry Entry 2 is incremented by 1.
[0175] For example, the memory controller 420 provided in an embodiment of the present disclosure further includes an address selection module 416. When the entry query module 419 determines that the first entry (e.g., Entry 1) does not exist in the register array 415 and all entries in the register array 415 are in an occupied state, the address selection module 416 is configured to determine an entry in the register array 415 to be replaced.
[0176] For example, when the address selection module 416 determines the entry to be replaced in the register array 415, specific steps may include: determining the entry in the register array 415 whose first count value is 0, and selecting, from the entries whose first count value is 0, the entry whose high-order portion of the second count value has the smallest value and whose second count value does not reach the second threshold as the entry to be replaced.
[0177] For example, when address selection module 416 determines that there is no entry to be replaced in register array 415, address selection module 416 is further configured to determine a target entry from register array 415. Repair module 417 uses the row address corresponding to the target entry as the target row address to perform a repair operation. After repair module 417 performs the repair operation, the target entry is selected as the entry to be replaced.
[0178] For example, when the address selection module 416 determines a target entry from the register array 415, specific steps may include: in response to the absence of an entry in the register array 415 having a first count value greater than a first threshold or an entry in the register array 415 having a second count value greater than a second threshold, selecting the third entry in the register array 415 as the target entry, where the third entry is the entry in the register array 415 having the largest value for the high-order portion of the first count value.
[0179] For example, the address selection module 416 is further configured to select, in response to the presence of multiple third entries in the register array 415 , an entry having the largest value of the upper portion of the second count value as the target entry from the multiple third entries.
[0180] For example, when the address selection module 416 determines that the target row address or the target table entry exists in the register array 415, the register array 415 may send a third signal to the address selection module 416 to initiate the repair operation. After receiving the third signal, the address selection module 416 is configured to select a row address from the target row address or the row address occupying the target table entry as the address to be repaired by the repair module 417.
[0181] For example, when the repair operation of the repair module 417 includes a post-packaging repair operation, the repair module 417 is configured to generate a command sequence of the post-packaging repair operation based on the target row address to repair the target row address.
[0182] As described above, according to at least one embodiment of the present disclosure, the memory controller can more accurately detect the faulty row address and promptly repair the detected faulty row address by recording the row address where CE (or fault) occurs and identifying whether the row address meets the characteristics of a bounded fault, thereby reducing the probability of UE occurring in the storage system and improving the RAS performance of the storage system.
[0183] The above-mentioned additional aspects of the memory controller according to at least one embodiment of the present disclosure may correspond to the additional aspects of the method for processing memory failures according to at least one embodiment of the present disclosure. Therefore, the technical effects of the additional aspects of the method for processing memory failures according to at least one embodiment of the present disclosure may also be mapped to the additional aspects of the memory controller according to at least one embodiment of the present disclosure, which will not be repeated here.
[0184] Figure 8 A schematic diagram of a memory system provided by at least one embodiment of the present disclosure is shown.
[0185] At least one embodiment of the present disclosure further provides a memory system 800, which includes a memory 801 (for example, including Figure 1 The DRAM memory 100 shown or Figure 4AThe memory 400 shown in FIG. 4 ) and the memory controller 802 provided by any embodiment of the present disclosure (eg, including Figure 4A The memory controller 420 shown or Figure 7 Memory controller 700 shown).
[0186] For example, the memory 801 may be any memory that the memory controller 802 may control.
[0187] The technical effects of the memory system of the above embodiment of the present disclosure are the same as the technical effects of the above method for processing a memory failure, and therefore are not described in detail.
[0188] Figure 9 FIG2 shows a schematic diagram of an electronic device 900 according to at least one embodiment of the present disclosure.
[0189] like Figure 9 As shown, electronic device 900 includes a processing device 910 and a storage device 920. Storage device 920 includes one or more computer program modules 921. One or more computer program modules 921 are stored in storage device 920 and configured to be executed by processing device 910. These one or more computer program modules 921 include instructions for executing the method for handling memory failures according to at least one embodiment of the present disclosure. When executed by processing device 910, these one or more computer program modules 921 can perform one or more steps of the method for handling memory failures according to at least one embodiment of the present disclosure and its additional aspects. Storage device 920 and processing device 910 can be interconnected via a bus system and / or other form of connection mechanism (not shown). For example, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industrial Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0190] For example, the processing device 910 may be a central processing unit (CPU), a digital signal processor (DSP), or other processors with data processing capabilities and / or program execution capabilities, such as a field programmable gate array (FPGA). For example, the central processing unit (CPU) may be an X86 or ARM architecture, a RISC-V architecture, etc. The processing device 910 may be a general-purpose processor or a dedicated processor, and may control other components in the electronic device 900 to perform desired functions.
[0191] For example, the storage device 920 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, a flash memory, etc. One or more computer program modules 921 may be stored on the computer-readable storage medium, and the processing device 910 may execute one or more computer program modules 921 to implement various functions of the electronic device 900. The computer-readable storage medium may also store various applications and various data, as well as various data used and / or generated by the applications.
[0192] For example, the electronic device 900 may also include input devices such as a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, and gyroscope; output devices such as a liquid crystal display, speaker, and vibrator; storage devices such as a magnetic tape and a hard disk (HDD or SDD); and communication devices such as a network interface card (NIC) such as a LAN card or modem. The communication devices may allow the electronic device 900 to communicate with other devices wirelessly or wired to exchange data, performing communication processing via a network such as the Internet. A drive is connected to the I / O interface as needed. Removable storage media, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, is installed in the drive as needed, so that computer programs read from the media can be installed in the storage device as needed.
[0193] For example, the electronic device 900 may further include a peripheral interface (not shown in the figure), etc. The peripheral interface may be various types of interfaces, such as a USB interface, a lightning interface, etc. The communication device may communicate with a network and other devices via wireless communication, such as the Internet, an intranet, and / or a wireless network such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). Wireless communications may use any of a variety of communication standards, protocols, and technologies, including, but not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), Wi-MAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.
[0194] The electronic device 900 may be, for example, a system on a chip (SOC) or a device including such a SOC, and may be any device such as a mobile phone, tablet computer, laptop computer, e-book, game console, television, digital photo frame, navigation system, home appliance, communication base station, industrial controller, server, or the like. It may also be any combination of operating devices and hardware of a memory controller, and the embodiments of the present disclosure are not limited thereto. The specific functions and technical effects of the electronic device 900 may be referred to the above description of the method for handling a memory failure and its additional aspects according to at least one embodiment of the present disclosure, and will not be further elaborated here.
[0195] At least one embodiment of the present disclosure further provides a computer-readable storage medium having computer-readable instructions stored thereon. When a processor executes the computer-readable instructions, the processor executes the method for handling a memory failure provided in any of the above embodiments.
[0196] Figure 10 A schematic diagram of a non-transitory readable storage medium 1000 according to at least one embodiment of the present disclosure is shown.
[0197] like Figure 10 As shown, a non-transitory readable storage medium 1000 stores computer instructions 1010 , which, when executed by a processor, perform one or more steps of the method for processing memory failure and additional aspects thereof as described above.
[0198] For example, the non-transitory readable storage medium 1000 may be any combination of one or more computer-readable storage media.
[0199] For example, when the program code is read by a computer, the computer may execute the program code stored in the computer storage medium and perform, for example, one or more steps of the method for processing memory failure and additional aspects thereof according to at least one embodiment of the present disclosure.
[0200] For example, the non-transitory readable storage medium may include a memory card of a smartphone, a storage component of a tablet computer, a hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a flash memory, and other non-transitory readable storage media or any combination thereof.
[0201] In addition to the above exemplary explanations, the following points should be explained:
[0202] (1) The drawings of the embodiments of the present disclosure only relate to the structures related to the embodiments of the present disclosure. Other structures may refer to conventional designs.
[0203] (2) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to form new embodiments.
[0204] The foregoing description is merely an exemplary embodiment of the present disclosure and is not intended to limit the scope of protection of the present disclosure. The scope of protection of the present disclosure is determined by the appended claims.
Claims
1. A method for handling a memory failure, comprising: In response to a correctable error in data read for a row address of a memory, determining whether the correctable error meets the characteristics of a bounded fault; In response to the correctable error meeting the characteristics of the bounded fault, performing a first count on the row address to obtain a first count value; In response to the correctable error not meeting the characteristics of the bounded fault, performing a second count on the row address to obtain a second count value; as well as In response to a target row address existing in the memory, a repair operation is performed on the target row address, wherein a first count value of the target row address is greater than a first threshold or a second count value of the target row address is greater than a second threshold.
2. The method according to claim 1, wherein Determining whether the correctable error meets the characteristics of a bounded fault includes: In response to the correctable error meeting any of the following conditions, it is determined that the correctable error meets the characteristics of the bounded fault: The number of error data pins corresponding to the data is 1, and the number of error bursts corresponding to the error data pin is greater than a third threshold; The number of error data pins corresponding to the data is 2, the data read from the two error data pins belongs to the same half byte, and the total number of error bursts corresponding to the two error data pins is greater than a third threshold; and The number of error data pins corresponding to the data is greater than half of the number of data pins of a storage particle, and the error data pins all belong to the same storage particle in the memory.
3. The method of claim 1 , further comprising: The address information of the row address, the first count value and the second count value are written into a register array, wherein the register array includes one or more entries, each entry is used to store the address information of the corresponding row address, the first count value and the second count value.
4. The method of claim 3, further comprising: A funnel decrement mechanism is set for the first count value and the second count value stored in each of the one or more entries, wherein the funnel decrement frequency of the first count value is slower than the funnel decrement frequency of the second count value.
5. The method of claim 3, further comprising: Before writing the address information of the row address into the register array, determining whether there is a first table entry for the row address in the register array; as well as In response to the first entry existing in the register array, a first count value or a second count value in the first entry is updated.
6. The method according to claim 5, wherein: In response to the first entry existing in the register array, updating the first count value or the second count value in the first entry includes: In response to the correctable error meeting the characteristics of the bounded fault, increasing the first count value in the first table entry by 1; or In response to the correctable error not meeting the characteristics of the bounded fault, the second count value in the first entry is increased by 1.
7. The method of claim 5, further comprising: In response to the first entry not existing in the register array and the second entry in the register array being in an idle state, the address information of the row address, the first count value, and the second count value are written into the second entry.
8. The method of claim 5, further comprising: In response to the first entry not existing in the register array and all entries in the register array are in an occupied state, determining an entry to be replaced in the register array; as well as The address information of the row address, the first count value, and the second count value are written into the entry to be replaced.
9. The method of claim 8, wherein: The first count value and the second count value both include a low-order portion and a high-order portion, and determining the entry to be replaced in the register array includes: Determine the table entry with the first count value of 0 in the register array, and select the table entry with the smallest value of the high-order part of the second count value and the second count value not reaching the second threshold from the table entry with the first count value of 0 as the table entry to be replaced.
10. The method of claim 8, further comprising: In response to the absence of the entry to be replaced in the register array, determining a target entry from the register array, and using a row address corresponding to the target entry as the target row address to perform the repair operation; as well as After performing the repair operation, the target entry is used as the entry to be replaced.
11. The method according to claim 10, wherein: The first count value and the second count value both include a low-order portion and a high-order portion, and determining a target entry from the register array includes: In response to the absence of an entry in the register array having the first count value greater than the first threshold or an entry in which the second count value is greater than the second threshold, a third entry is selected from the register array as the target entry, wherein the third entry is an entry in the register array having a maximum value of a high-order portion of the first count value.
12. The method of claim 11, wherein: In response to the absence of an entry in the register array having the first count value greater than the first threshold or an entry in which the second count value is greater than the second threshold, selecting a third entry from the register array as the target entry includes: In response to the presence of a plurality of the third entries in the register array, an entry having a maximum value of a high-order portion of the second count value is selected from the plurality of third entries as the target entry.
13. The method of claim 1, wherein: The repair operation includes a post-packaging repair operation, wherein performing the repair operation on the target row address includes: Based on the target row address, a command sequence of the post-package repair operation is generated to repair the target row address.
14. A memory controller for controlling a memory, the memory controller comprising: A bounded fault detection module is configured to: in response to a correctable error in data read from a row address of the memory, determine whether the correctable error meets the characteristics of a bounded fault; A first counter is configured to: in response to the correctable error meeting the characteristics of the bounded fault, perform a first count on the row address to obtain a first count value; a second counter configured to: in response to the correctable error not meeting the characteristics of the bounded fault, perform a second count on the row address to obtain a second count value; as well as The repair module is configured to perform a repair operation on the target row address in response to the target row address existing in the memory, wherein the first count value of the target row address is greater than a first threshold or the second count value is greater than a second threshold.
15. The memory controller of claim 14, wherein: The bounded fault detection module is configured to determine that the correctable error meets the characteristics of the bounded fault in response to the correctable error meeting any of the following conditions: The number of error data pins corresponding to the data is 1, and the number of error bursts corresponding to the error data pin is greater than a third threshold; The number of error data pins corresponding to the data is 2, the data read from the two error data pins belongs to the same half byte, and the total number of error bursts corresponding to the two error data pins is greater than a third threshold; and The number of error data pins corresponding to the data is greater than half of the number of data pins of a storage particle, and the error data pins all belong to the same storage particle in the memory.
16. The memory controller of claim 14, further comprising: A register array is configured to: in response to the presence of the correctable error in the data read for the row address of the memory, store the address information of the row address, the first count value, and the second count value, wherein the register array includes one or more table entries, each table entry is used to store the address information, the first count value, and the second count value of the corresponding row address.
17. The memory controller of claim 16 , further comprising: Correctable error detection module and address queue, Wherein, the address queue is configured to store address information corresponding to each read operation; The correctable error detection module is configured to: in response to determining that the data read from the first row address of the memory has the correctable error, send a first signal to the address queue to cause the address queue to send address information of the first row address to the register array; The bounded fault detection module is further configured to: in response to determining that the correctable error corresponding to the first row address meets the characteristics of the bounded fault, send a high-level second signal to the address queue; in response to determining that the correctable error corresponding to the first row address does not meet the characteristics of the bounded fault, send a low-level second signal to the address queue.
18. The memory controller of claim 16, wherein: The first count value and the second count value stored in the one or more entries are provided with a funnel self-decrement mechanism, and the funnel self-decrement frequency of the first count value is slower than the funnel self-decrement frequency of the second count value.
19. The memory controller of claim 16, further comprising: Table entry query module, Among them, when the table entry query module determines that there is a first table entry for the row address in the register array, in response to the bounded fault detection module determining that the correctable error of the data meets the characteristics of the bounded fault, the value of the first counter corresponding to the first table entry is increased by 1; or in response to the bounded fault detection module determining that the correctable error of the data does not meet the characteristics of the bounded fault, the value of the second counter corresponding to the first table entry is increased by 1.
20. The memory controller of claim 19, further comprising: Address selection module, Among them, when the table entry query module determines that the first table entry for the row address does not exist in the register array and all the table entries in the register array are in an occupied state, the address selection module is configured to: determine the table entry to be replaced in the register array, wherein the table entry to be replaced is used to store the address information, the first count value and the second count value of the row address.
21. The memory controller of claim 20, wherein: The first count value and the second count value both include a low-order portion and a high-order portion. The address selection module is further configured to: determine a table entry in the register array whose first count value is 0, and select, from the table entry whose first count value is 0, a table entry whose high-order portion of the second count value has the smallest value and whose second count value does not reach the second threshold as the table entry to be replaced.
22. The memory controller of claim 20, wherein: The address selection module is further configured to: In response to determining that the entry to be replaced does not exist in the register array, a target entry is determined from the register array, and a row address corresponding to the target entry is used as the target row address; and after the repair module performs the repair operation on the target row address, the target entry is used as the entry to be replaced.
23. A memory system comprising: Memory; as well as The memory controller according to any one of claims 14 to 22.
Citation Information
Patent Citations
Memory device and memory system including the same
CN114627957A
Method and system for supervising DDR5 memory particle errors, storage medium and equipment
CN115543678A
Memory fault early warning method and device, electronic equipment and readable medium
CN115629905A
Semiconductor memory device
CN117393031A
Verification data processing method and equipment of memory and storage medium
CN119718763A