Memory error correction method, memory module, memory controller, and processor
By combining system-level error correction and particle-level error correction in the memory controller, errors in the memory particles are determined and corrected, and the problem of difficult to correct multiple bit errors in the prior art is solved, thereby improving the reliability of memory data.
Patent Information
- Application Number
- CN202410787725.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-31
- Filing Date
- 2023-06-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-06-25
AI Technical Summary
Existing memory systems are difficult to effectively correct errors in multiple bit errors, resulting in an increase in the risk of data silent errors and affecting the reliability of memory data.
By achieving system-level error correction in the memory controller, combining the particle-level error correction of the on-chip error correction engine, the memory particles that have occurred are determined, and error correction is performed based on the particles to reduce the risk of silent errors.
It effectively reduces the risk of data silent errors and data error correction in memory particles, and improves the error correction ability of the memory system and the reliability of data in memory sticks.
Smart Images

Figure CN118838738B_ABST
Abstract
Description
[0001] This application is a divisional application. The application number of the original application is 202310751941.8, and the original application date is June 25, 2023. The entire content of the original application is incorporated herein by reference. Technical Field
[0002] This application relates to the field of storage technologies, and particularly to an in-memory error correction method, a memory module, a memory controller, and a processor. Background Art
[0003] In storage technologies, a memory system includes a memory controller and a double data rate synchronous dynamic random access memory (DDR SDRAM). Among them, the double data rate synchronous dynamic random access memory is also the memory module, which can be referred to as a memory module or memory. The memory module includes multiple memory dies. A memory die is also called a dynamic random access memory (DRAM). Among them, a part of the memory dies are used to store data, called data devices, and another part of the memory dies are used to store error correcting codes (ECC) of the data, that is, redundant information, called ECC devices. The error correcting code is used to verify whether the data stored in the memory die has an error; the memory controller can correct errors in the codeword based on the error correcting code, that is, the memory controller can achieve system-level error correction. Among them, the internal storage array of the memory die stores multiple groups of codewords, and each group of codewords includes data and a check code; the memory die also includes an on-chip error correction engine, which can detect errors in the data belonging to the same codeword as the check code based on the check code. If a single bit in the data has an error, the on-chip error correction engine can correct the error; if multiple bits in the data have an error, it may be corrected by the on-chip error correction engine as a single-bit error, or the on-chip error correction engine may not be able to detect the error, forming a silent error, that is, the on-chip error correction engine in the memory die can achieve die-level error correction.
[0004] In the related art, when reading data, the on-chip error correction engine first checks the memory die. Once the on-chip error correction engine detects an error, it immediately reports it to the memory controller to assist the memory controller in locating the memory die where the error occurred. Then, based on the redundant information in the ECC device, the memory controller recovers the data of the memory die with the error and rereads the data. However, in the above method, in the case of multiple bits in error, there is a high probability that the on-chip error correction engine cannot detect the error, resulting in the memory controller being unable to correct the error of the memory die, bringing the risk of data silent errors and affecting the reliability of the memory data. Summary of the Invention
[0005] Embodiments of the present application provide a memory error correction method, a memory module, a memory controller, and a processor in a memory system, which can reduce the risk of data silent errors, improve the error correction ability of the memory system, and improve the reliability of the data in the memory module. The technical solution is as follows.
[0006] In a first aspect, a memory error correction method is provided, and the method includes:
[0007] When the memory controller fails to correct the data obtained from the memory module, it determines the memory die in the memory module where the error occurred, and then corrects the data again based on the memory die where the error occurred.
[0008] Among them, the results of the memory controller detecting the data in multiple memory dies and the subsequent steps include three cases. The first case: no error is detected, and the memory controller directly returns the data to the processor. The second case: the number of detected errors is within the error correction ability of the error correction algorithm adopted by the memory controller, that is, the error correction is successful. The memory controller corrects the data through the error correction algorithm and returns the corrected data to the processor. The third case: the number of detected errors exceeds the error correction ability of the error correction algorithm, that is, the error correction fails. The memory controller determines the memory die in the memory module where the error occurred, and then corrects the data based on the memory die where the error occurred.
[0009] Among them, the memory controller determines the memory die in the memory module where the error occurred through the error correction status information of the memory die. The memory controller can determine the memory die with the error correction status of uncorrectable error as the memory die where the error occurred; or it can determine the memory dies with the error correction status of correctable error and uncorrectable error as the memory dies where the error occurred.
[0010] In the above method, the system-level error correction based on the memory controller and the granular-level error correction based on the on-chip error correction engine are coupled with each other. The memory controller corrects the data. Compared with only relying on the on-chip error correction engine for error correction, it can correct the silent errors and mis-corrected errors not detected by the on-chip error correction engine. Therefore, it can reduce the risk of data silent errors and data mis-correction in the memory granules. In addition, when the memory controller fails to correct the error for the first time, it uses the error correction result of the on-chip error correction engine to determine the memory granule where the error occurs, and then corrects the data again based on the memory granule where the error occurs. Since when the memory controller corrects the error for the second time, the known information for error correction includes not only the data obtained from the memory module, but also the information of the memory granule where the error occurs. Compared with only relying on the memory controller for error correction, it can improve the error correction ability of the memory controller, thereby improving the error correction ability of the memory system and the reliability of the data in the memory module.
[0011] Optionally, the memory controller determines the granule where the error occurs in the memory module based on the error correction status information of the memory granule recorded in the first register of the memory granule.
[0012] Wherein, the first register is a reserved register with undefined functions in the original memory module, and the value of the first register can be used to represent the error correction status of the memory granule.
[0013] In the above method, the error correction status information of the memory granule is written into the first register of the memory granule. The memory controller can locate the granule where the error occurs by reading the error correction status information in the first register, so that the granular-level error correction and the system-level error correction can cooperate with each other, realizing the full utilization of redundant resources, which is beneficial to improving the error correction ability of the memory system under a fixed redundancy configuration.
[0014] Optionally, when the memory controller fails to correct the data obtained from the memory module, it backpressures the read and write processes of the memory module, and after the backpressure is successful, it obtains the data in the memory module again. When the data obtained twice is consistent, it reads the error correction status information of multiple memory granules from the first registers of multiple memory granules.
[0015] Among them, backpressure means suppressing the generation of memory access from the source of the central processing unit (CPU) or the memory access path. Among them, the memory controller judges whether the memory address corresponding to the read operation is a direct memory access address. Only when the memory address is a memory address other than the direct memory access (DMA) address, the memory controller can backpressure the read process of the memory module, so as to prevent other read and write processes from rewriting the data to be corrected, and then avoid data inconsistency.
[0016] In the above method, the memory controller backpressures the read process of the memory module to prevent other read and write processes from overwriting the data to be corrected, thereby avoiding data inconsistency. In addition, by obtaining the data from the memory module again to determine whether the data has been overwritten before the backpressure takes effect, and only continuing the subsequent data error correction process when the data has not been overwritten, the effectiveness of the data error correction process can be ensured, and thus the data consistency can be ensured.
[0017] Optionally, when the number of memory granules with errors is less than or equal to the number of error correction code granules in the memory module, the memory controller can correct the data based on the memory granules with errors.
[0018] Among them, when determining the memory granules with errors, the error correction ability of the memory controller can meet the requirement of correcting the target number of memory granules with errors, where the target number is the number of error correction code granules in the memory module. The number of memory granules with errors being less than or equal to the number of error correction code granules in the memory module indicates that the number of redundant memory granules in the memory module is greater than or equal to the number of memory granules with errors, that is, the number of memory granules with errors is within the error correction ability of the memory controller, and the memory controller can correct the data in the memory granules with errors based on the data in the memory granules without errors in the memory module.
[0019] Optionally, after correcting the data, the memory controller writes the corrected data back to the memory granules, rereads the memory address corresponding to the read operation, and verifies the reread data. If the verification passes, the memory controller returns the reread data to the processor; if the verification fails, the memory controller reports an error to the processor.
[0020] Among them, the memory controller writes the corrected data back to the memory address corresponding to the read operation, that is, replaces all the data at the corresponding positions of the memory address in each memory granule.
[0021] In the above method, the memory controller writes the corrected data back to the memory granules, so that when the same memory address is read next time, the correct data can be read, which is beneficial to improving the reliability of the data in the memory module. In addition, the memory controller verifies the reread data and only returns the data when the verification passes, which can ensure the data consistency and improve the reliability of the data in the memory module.
[0022] Optionally, different error correction states of the memory die are represented by different values on the target bits of the first register, and the error correction state of the memory die represented by the value of the target bits of the first register is any one of no error, correctable error, and uncorrectable error. For example, the values of the target bits being 00B, 01B, and 10B represent that the error correction states of the memory die are no error, correctable error, and uncorrectable error, respectively.
[0023] Optionally, the error correction state of the memory die is represented by the occupancy state of the target bits of the register. For example, the occupancy state of the target bit is empty, that is, the target bit has no value, indicating that the error correction state of the memory die is no error or correctable error; the occupancy state of the target bit is not empty, that is, the target bit has a value, indicating that the error correction state of the memory die is uncorrectable error.
[0024] Optionally, the memory die with an error refers to at least one of the memory die with a correctable error state and the memory die with an uncorrectable error state.
[0025] Among them, the memory controller can determine the memory die with an uncorrectable error state as the memory die with an error; it can also determine both the memory die with a correctable error state and the memory die with an uncorrectable error state as the memory die with an error.
[0026] In the above method, the memory controller can determine both the memory die with a correctable error state and the memory die with an uncorrectable error state as the memory die with an error. Since the on-chip error correction engine may mis-correct the data during error correction, that is, correct a multi-bit error as a single-bit error, therefore, the memory die with a correctable error state is also determined as the memory die with an error. Furthermore, the memory controller can not only correct the data that cannot be corrected by the on-chip error correction engine, but also correct the data mis-corrected by the on-chip error correction engine, which is beneficial to further improve the reliability of the data in the memory module.
[0027] Optionally, the memory die includes an on-chip error correction engine, which can correct the data read from the memory die, write the error correction state information of the memory die into the first register of the memory die, and output the corrected data to the memory controller.
[0028] In the above method, the on-chip error correction engine writes the error correction state information of the memory die into the first register of the memory die, enabling the memory controller to locate the memory die with an error by reading the error correction state information in the first register, so that the error correction at the die level and the error correction at the system level can cooperate with each other, realizing the full utilization of redundant resources, which is beneficial to improving the error correction ability of the memory system under a fixed redundancy configuration.
[0029] Second aspect, a memory error correction method is provided, which is executed by a memory module. The memory module includes a plurality of memory dies, and each memory die includes a first register. The method includes:
[0030] Correct the data read from the memory die to obtain the error correction status information of the memory die, and write the error correction status information of the memory die into the first register of the memory die.
[0031] Third aspect, a memory controller is provided. The memory controller includes at least one functional module, and the at least one functional module is configured to execute the memory error correction method provided in the foregoing first aspect or any possible implementation manner of the first aspect.
[0032] Fourth aspect, a memory module is provided. The memory module includes a plurality of memory dies, and each memory die includes a first register. The memory module is configured to execute the memory error correction method provided in the foregoing second aspect.
[0033] Fifth aspect, a memory controller is provided. The memory controller is configured to execute the memory error correction method provided in the foregoing first aspect or any possible implementation manner of the first aspect.
[0034] Sixth aspect, a processor is provided. The processor includes a memory controller and a computing core. The processor is configured to execute the memory error correction method provided in the foregoing first aspect or any possible implementation manner of the first aspect, and the computing core is configured to perform a computing operation on the data in the memory die.
[0035] Seventh aspect, a computing device is provided. The computing device includes a memory controller and a memory module. The memory module is used for temporarily storing data, and the memory controller is configured to execute the memory error correction method provided in the foregoing first aspect or any possible implementation manner of the first aspect.
[0036] Eighth aspect, a computer-readable storage medium is provided. The computer-readable storage medium is used for storing at least one program code, and the at least one program code is used for executing the memory error correction method provided in the foregoing first aspect or any possible implementation manner of the first aspect.
[0037] Ninth aspect, a computer program product including at least one program code is provided. When the at least one program code is run by a computing device, the computing device is caused to execute the memory error correction method provided in the foregoing first aspect or any possible implementation manner of the first aspect. Description of the Drawings
[0038] Figure 1 is a schematic diagram of a memory module provided by an embodiment of the present application;
[0039] Figure 2 It is a schematic diagram of a memory chip including an on-chip error correction engine provided by an embodiment of the present application;
[0040] Figure 3 It is a schematic diagram of a memory system provided by an embodiment of the present application;
[0041] Figure 4 It is a flowchart of a memory error correction method provided by an embodiment of the present application;
[0042] Figure 5 It is a schematic diagram of the process of a memory error correction method provided by an embodiment of the present application;
[0043] Figure 6 It is a block diagram of the structure of a memory error correction device provided by an embodiment of the present application;
[0044] Figure 7 It is a schematic diagram of the structure of a computing device provided by an embodiment of the present application. Detailed implementation manners
[0045] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0046] To facilitate the understanding of the technical solutions of the present application, several technical terms involved in the embodiments of the present application will be introduced first.
[0047] 1. Memory module: The von Neumann computer system includes five major components: arithmetic, control, storage, input, and output. Among them, storage includes main memory (memory) and auxiliary storage (solid-state drive, mechanical hard disk, etc.). Among them, the memory module mainly acts as a working memory, which is used to store the instructions and data required for the operation of the computer and is an indispensable part of the von Neumann computer system. Among them, double data rate synchronous dynamic random access memory (DDR SDRAM) is a type of memory module. The application forms of this memory module include chip surface mounting and dual in-line memory module (DIMM). The memory modules connected to the same chip select signal in the memory module are also called a memory rank. Figure 1 It is a schematic diagram of a memory module provided by an embodiment of the present application, as Figure 1As shown in the figure, the memory module includes a dynamic random access memory (DRAM), a registering clock driver (RCD), a serial presence detect with hub (SPD Hub), a power management integrated circuit (PMIC), a temperature sensor (TS), a printed circuit board (PCB), and other resistors and capacitors. The memory module includes two channels (channel A and channel B). In enterprise-level memory for high-reliability application scenarios, the bit width of each channel is 40 bits (bit), of which the data bit width is 32 bit and the error correcting code (ECC) bit width is 8 bit. The memory module is also known as a memory module. In addition to DRAM, the memory module can also be a random access memory (RAM) and a resistive random access memory (RRAM).
[0048] 2. Memory die: A DRAM in the memory is a memory die, which is the storage medium in the memory module. As Figure 1 shown in the figure, a memory module includes multiple memory dies. According to the content stored in the memory die, the memory die is further divided into a data die and an ECC die. According to the number of data lanes of each memory die, the memory die is divided into x4 memory dies and x8 memory dies. x4 and x8 represent that the number of data lanes of each memory die is 4 and 8 respectively.
[0049] Among them, in an x8 memory module of the 5th generation double data rate synchronous dynamic random access memory (DDR5 SDRAM), one rank in each channel of the memory module includes 4 data devices and 1 error correction code device (ECC device). In response to a memory read operation, each memory device outputs 128-bit data to the memory controller. This 128-bit data is distributed across two codewords, and each memory device contributes 64-bit data to one codeword. Since each memory device has 8 external data lines, the data output by each memory device to the memory controller is also an 8×8 data block.
[0050] Among them, in an x4 memory module of the 5th generation double data rate synchronous dynamic random access memory (DDR5 SDRAM), one rank in each channel of the memory module includes 8 data devices and 2 error correction code devices (ECC device). In response to a memory read operation, each memory device outputs 64-bit data to the memory controller. This 64-bit data is distributed across two codewords, and each memory device contributes 32-bit data to one codeword. Since each memory device has 4 external data lines, the data output by each memory device to the memory controller is also a 4×8 data block.
[0051] 3. Memory array: The array in the memory device for storing data, which is composed of multiple storage units. The memory array includes multiple bank groups, and each bank group includes multiple banks.
[0052] 4. On-die ECC (OD ECC) engine: Also known as the in-DRAM ECC engine, it is an error correction engine implemented inside the memory device, which can correct the data read from the memory device, thereby improving the yield of the memory device. The on-die error correction engine can be present on the input / output pins (IO pins) or global input / output (global io, GIO) in the internal data path of the memory device, or it can be present on the local input / output (local io, LIO) in the internal data path of the memory device. Figure 2 It is a schematic diagram of a memory device including an on-die error correction engine provided by an embodiment of the present application, as Figure 2 shown Figure 2It includes memory dies and a double data rate controller (DDRC), which is also the memory controller. The memory die includes 8 bank groups (bank group0 - 7) and an on - chip error correction engine, and each bank group includes 4 banks (bank0 - 3). Figure 2 As shown in (a), the on - chip error correction engine exists on the global input / output of the memory die. When the memory address corresponding to the read operation of the memory module corresponds to bank3 in bank group7, the on - chip error correction engine reads the data in this memory address, corrects the read data, and returns the corrected data to the memory controller. Figure 2 As shown in (b), the on - chip error correction engine exists on the local input / output of the memory die, that is, the data in each bank is corrected by an on - chip error correction engine.
[0053] Among them, the results of the on - chip error correction engine detecting the read data include three cases: The first case: No error is detected. Correspondingly, the error correction status of the memory die is error - free; The second case: A single - bit error is detected, and the on - chip error correction engine can correct this single - bit error. Correspondingly, the error correction status of the memory die is correctable error; The third case: Multiple - bit errors are detected, and the on - chip error correction engine cannot correct the multi - bit errors. Correspondingly, the error correction status of the memory die is uncorrectable error.
[0054] 5. Memory controller: The memory controller is located in the CPU and is an important part of the computer system that internally controls the memory and is responsible for data exchange between the memory module and the CPU. Taking the memory controller as a double data rate controller (DDRC) as an example, this memory controller includes a DDRC ECC engine, a dynamic memory controller (DMC), a micro - controller unit (MCU), a bit pattern generator (BPG), and a physical layer (PHY).
[0055] 6. Reed-Solomon Codes (RS Codes) Algorithm: Each codeword of this RS code algorithm includes k data elements (symbols) and 2t error correction code elements (ECC symbols), where k and t are positive integers greater than 0. One symbol is the data output through two pins. This RS code algorithm can correct 100% of the errors within t symbols, while errors exceeding t symbols are beyond the error correction capability of this RS code algorithm. For example, a DDR5 SDRAM x8 memory module includes 4 data chips and 1 error correction code chip. Among them, the external data line of each memory chip is 8, so the data input to each memory chip can be divided into 4 symbols. The data output from 5 memory chips to the memory controller is divided into 4×4 = 16 data symbols and 1×4 = 4 error correction code symbols, that is, k = 16, t = 2. This RS code algorithm can correct the errors of 2 symbols in the DDR5 SDRAM x8 memory module, while errors above 2 symbols are beyond the error correction capability of the RS code algorithm.
[0056] The implementation environment of the embodiments of the present application will be introduced below.
[0057] Figure 3 is a schematic diagram of a memory system provided by an embodiment of the present application, as Figure 3 shown. This memory system includes a memory module 301 and a memory controller 302. The memory module 301 includes multiple memory chips. The memory module 301 and the memory controller 302 communicate with each other through a DDR bus.
[0058] Among them, the memory module 301 can be any memory module with an on-chip error correction engine. For example, DDR5, DDR6, or Low Power Double Data Rate SDRAM (LPDDR), etc. The form of the memory module 301 can be a dual in-line memory module or a surface-mounted chip. The memory chips in the memory 301 can be x4 memory chips or x8 memory chips. The embodiments of the present application do not make any limitations in this regard. The memory chips in the memory 301 include data chips and error correction code chips. Each memory chip stores multiple groups of codewords. Each group of codewords includes data and a check code. Exemplarily, the on-chip error correction engine in the memory chip corrects the data belonging to the same codeword as the check code based on the check code, writes the error correction status information of the memory chip into a register, and outputs the corrected data to the memory controller 302.
[0059] Among them, the memory controller 302 can be any memory controller with system-level error correction function. For example, a DDRC that uses RS code for error correction is not limited in this embodiment of the present application. Exemplarily, the memory controller 302 can obtain the data in the memory module 301, correct the obtained data, and then write the corrected data back to the memory module 301.
[0060] In the embodiment of the present application, the reserved registers with undefined functions in the original memory module are defined so that the value of the register (i.e., the first register) can be used to represent the error correction state of the memory die. In some embodiments, the error correction state of the memory die is represented by the value of the target bit of the first register. For example, the values of 00B, 01B, and 10B of the target bit represent that the error correction state of the memory die is error-free, correctable error, and uncorrectable error, respectively. In other embodiments, the error correction state of the memory die is represented by the occupied state of the target bit of the register. For example, the occupied state of the target bit is empty, that is, the target bit has no value, indicating that the error correction state of the memory die is error-free or correctable error, and the occupied state of the target bit is not empty, that is, the target bit has a value, indicating that the error correction state of the memory die is uncorrectable error. It should be noted that the above description of the representation method of the error correction state of the memory die is only exemplary, and the representation method of the error correction state of the memory die can be set according to actual needs, and this embodiment of the present application does not make any limitations.
[0061] Taking the DDR5 SDRAM memory as an example below, an example is given to illustrate the representation method in which the error correction state of the memory die is represented by the value of the target bit of the first register.
[0062] The DDR5 SDRAM memory includes 256 8-bit mode registers (MR): MR0 - MR255. Among them, MR41, MR49, MR70 - 102, MR117, MR119, MR125, MR127, MR135, MR143, MR155, MR159, MR167, MR175, MR183, MR191, MR199, MR207, MR215, MR223, MR231, MR239, MR247, MR255 are completely undefined. There are also many mode registers with two or more bits of register space undefined. For example, the values of Operand[6:2] (OP[6:2]) of MR9 can be used to represent the error correction status. In some embodiments, the first register MRx has two or more bits of space undefined. Take two bits of space from it, for example, OP[u], OP[v]. OP[u], OP[v] are the target bits. Then different error correction statuses are represented by different values of OP[u], OP[v], as shown in Table 1 below. Among them, the on-chip error correction engine writes the error correction status information into the first register by assigning values to OP[u], OP[v] in the first register. As shown in Table 2, if the on-chip error correction engine does not detect an error, then assign 00B to MRx.uv (OP[u], OP[v] in the first register MRx); if the on-chip error correction engine detects a correctable error, then assign 01B to MRx.uv; if the on-chip error correction engine detects an uncorrectable error, then assign 10B to MRx.uv. It should be noted that the representation method shown in Table 1 and the error correction status represented by each value shown in Table 2 are only exemplary. The representation method of the error correction status of the memory die can be set according to actual needs, and the embodiments of the present application do not limit this.
[0063] Table 1
[0064]
[0065] Table 2
[0066]
[0067] Among them, the memory controller can read the first register through an in-band mode register read command (MRR), or can also read the first register through an out-of-band method. The embodiments of the present application do not limit this.
[0068] The definition method of the first register in the embodiments of the present application is introduced above. Next, a memory error correction method provided by the embodiments of the present application is introduced. This method can be applied to DDR5 SDRAM x8 memory modules and DDR5 SDRAM x4 memory modules. The following takes the above two types of memory modules as examples for introduction. It should be noted that the above two types of memory modules are only exemplary, and this method can also be applied to other types of memory modules. The embodiments of the present application do not limit this.
[0069] The following takes the DDR5 SDRAM x8 memory module as an example for introduction. Figure 4 is a flowchart of a memory error correction method provided by the embodiments of the present application. As Figure 4 shown, this method includes the following steps 401 to step 408.
[0070] 401. In response to a read operation on the memory module, the on-chip error correction engine in the memory die obtains the data in the memory die, corrects the obtained data, writes the error correction status information into the first register of the memory die, and outputs the corrected data to the memory controller.
[0071] Among them, the memory die stores data and a check code, and the check code is used to check whether the data has an error. The process of the on-chip error correction engine obtaining the data in the memory die and correcting the obtained data includes: in response to a read operation on the memory module, reading a codeword from the memory address corresponding to the read operation, and the read codeword includes data and the check code of the data; based on the read check code, correcting the read codeword.
[0072] Among them, the results of the on-chip error correction engine detecting the data include three cases: the first case: no error is detected; the second case: a single-bit error is detected in the data, and the on-chip error correction engine corrects the single-bit error; the third case: an error of two bits or more is detected in the data, and the on-chip error correction engine cannot correct it. Corresponding to the error correction result, the error correction status of the memory die includes no error, correctable error (CE), and uncorrectable error (UCE). This error correction status information is used to indicate the error correction status of the data in the memory die.
[0073] Among them, the memory die outputs the error-corrected data to the memory controller through the data bus. In an x8 memory module of DDR5 SDRAM, in response to a memory read operation, each memory die outputs 128-bit data to the memory controller. This 128-bit data is distributed across two codewords, and each memory die contributes 64-bit data to one codeword. Since each memory die has 8 external data lines, the data output by each memory die to the memory controller is also an 8×8 data block.
[0074] In the above method, the on-chip error correction engine writes the error correction status information of the memory die into the first register of the memory die, enabling the memory controller to locate the memory die where the error occurs by reading the error correction status information in the first register. Thus, the error correction at the die level and the error correction at the system level can cooperate with each other, achieving the full utilization of redundant resources and facilitating the improvement of the error correction ability of the memory system under a fixed redundancy configuration.
[0075] It should be noted that in response to a read operation on the memory module, multiple memory dies in the memory module write the error correction status information into their respective first registers and synchronously output the error-corrected data to the memory controller. The data output by the memory die is the superposition of the data in the storage array (i.e., the data read) and the error correction result of the on-chip error correction engine for the data.
[0076] It should be noted that the steps in step 401 where the on-chip error correction engine obtains the data in the memory die, corrects the obtained data, and writes the error correction status information into the first register of the memory die are optional steps. In some embodiments, if an error occurs in the memory die, the memory die directly reports the error to the memory controller. Then, after the memory controller fails to correct the error, it can determine the memory die where the error occurs based on the memory die that reports the error, and the memory die does not need to write the error correction status information into the first register, nor does the first register need to read the error correction status information from the first register to determine the memory die where the error occurs, which can save the time cost of error correction and register resources.
[0077] 402. The memory controller obtains the data output by multiple memory dies and corrects the data in the multiple memory dies.
[0078] Among them, the memory controller detects the data in the multiple memory dies. If the number of errors detected by the memory controller is greater than the target threshold, it indicates that the memory controller fails to correct the error.
[0079] In some embodiments, the memory controller detects data in multiple memory dies through the RS code algorithm. The target threshold is the maximum number of errors that can be corrected by the error correction algorithm adopted by the memory controller. For example, a DDR5 SDRAM x8 memory module includes 4 data dies and 1 error correction code die. Among them, the external data line of each memory die is 8, so the data input by each memory die can be divided into 4 symbols. The data output from 5 memory dies to the memory controller is divided into 4×4 = 16 data symbols and 1×4 = 4 error correction code symbols, that is, k = 16, t = 2. The RS code algorithm can correct 2 symbol errors in the DDR5 SDRAM x8 memory, and errors of more than 2 symbols exceed the error correction ability of the RS code algorithm.
[0080] Among them, the results of the memory controller's detection of data in multiple memory dies and subsequent steps include three cases. The first case: no error is detected, and the memory controller directly returns the data to the CPU. The second case: the number of detected errors is within the error correction ability of the error correction algorithm adopted by the memory controller, that is, the error correction is successful. The memory controller corrects the data through this error correction algorithm and returns the corrected data to the CPU. The third case: the number of detected errors exceeds the error correction ability of this error correction algorithm, that is, the error correction fails. The memory controller determines the memory die in which the error occurs in the memory, and then corrects the data based on the memory die in which the error occurs.
[0081] Next, the process of the memory controller correcting data based on the memory die in which the error occurs in the above third case is introduced. This process includes the following steps 403 to 406.
[0082] 403. If the error correction fails, the memory controller backpresses the read-write process of the memory module.
[0083] Among them, backpressure means suppressing the generation of memory access from the CPU source or the memory access path. Among them, the memory controller judges whether the memory address corresponding to the read operation is a direct memory access address. Only when the memory address is a memory address other than the direct memory access (DMA) address, the memory controller can backpress the read process of the memory module, so as to prevent other read-write processes from overwriting the data to be corrected, and thus avoid data inconsistency.
[0084] In some embodiments, the processor includes multiple cores. The process of the memory controller applying backpressure to the read and write processes of the memory module includes: If the number of errors in the data accessed by the memory read operation of any processor core exceeds the target threshold, then that core initiates a core stop broadcast by triggering a software generated interrupt (SGI); Other cores in the processor set a core stop flag after receiving the core stop broadcast to stop memory access and wait to be awakened; The core that initiated the core stop broadcast detects whether other cores have completed core stop after a preset time. If all cores have completed core stop, then the memory controller enters the subsequent data recovery process; If any core has not completed core stop, then the memory controller records a core stop failure, wakes up other cores, and returns an error status. In the above method, by stopping the cores of each processor to apply backpressure to the read process of the memory module, it is possible to prevent other read and write processes from overwriting the data to be corrected from the access source, thereby effectively ensuring data consistency.
[0085] In other embodiments, the memory controller includes a dynamic memory controller and a physical layer (PHY). The dynamic memory controller includes a scheduler for scheduling the read and write commands of the memory module. The memory controller cuts off the path between the scheduler and the PHY by applying backpressure to the command scheduling queue in the memory controller, thereby preventing other read and write processes from overwriting the data to be corrected from the scheduling queue. The backpressure is more refined, the processor cores do not stop working, and the impact on the upper-layer services is smaller.
[0086] It should be noted that step 403 is described by taking the memory address corresponding to the read operation as a memory address other than the DMA address as an example. In some embodiments, if the memory address corresponding to the read operation is the DMA address, then the memory controller abandons the current error correction and returns an error status.
[0087] It should be noted that the process of applying backpressure to the read and write processes of the memory module in step 403 is an optional step. In some embodiments, the process of applying backpressure to the read and write processes of the memory module in step 403 is not executed, and the embodiments of the present application do not make any limitations in this regard.
[0088] 404. The memory controller obtains the data in the memory module again. If the data obtained twice is the same, then the error correction status information of the data in the multiple memory dies is read from the first registers of the multiple memory dies.
[0089] Among them, the process of the memory controller obtaining data from the memory module is the same as that of step 402 above, and will not be elaborated here. If the data obtained by the memory controller twice is the same, it indicates that the data in the memory address corresponding to the read operation has not been rewritten before the backpressure takes effect, and the memory controller continues the subsequent data error correction process; if the data obtained by the memory controller twice is different, it indicates that the data in the memory address corresponding to the read operation has been rewritten before the backpressure takes effect, and the memory controller ends the data error correction process and returns an error status. In the above method, by obtaining data from the memory module again to determine whether the data has been rewritten before the backpressure takes effect, and only continuing the subsequent data error correction process when it has not been rewritten, the effectiveness of the data error correction process can be guaranteed, and thus the data consistency can be guaranteed.
[0090] Among them, the memory controller can read the first register through an in-band mode register read command, or can read the first register through an out-of-band method. The embodiments of the present application do not limit this.
[0091] It should be noted that this step 404 is an optional step. In some embodiments, if the process of backpressuring the read / write process of the memory module in step 403 above is not executed, then this step 404 is not executed. The embodiments of the present application do not limit this.
[0092] 405. If the number of memory grains with errors is less than or equal to the number of error correction code grains in the memory module, the memory controller corrects the data based on the memory grains with errors.
[0093] Among them, the memory controller determines the memory grains with errors based on the error correction status information of multiple memory grains read from the first register. If the error correction status information of any memory grain indicates that the memory grain has an error, the memory controller determines that the memory grain is a memory grain with an error. In the above method, the memory controller determines the memory grains with errors based on the error correction status information in the first register, and then corrects the data based on the memory grains with errors. The error correction at the grain level and the error correction at the system level can cooperate with each other, realizing the full utilization of redundant resources, which is beneficial to improving the error correction ability of the memory system under a fixed redundancy configuration.
[0094] In some embodiments, the memory controller determines memory grains with an uncorrectable error correction status as the memory grains with errors; in other embodiments, the memory controller determines both memory grains with a correctable error and an uncorrectable error correction status as the memory grains with errors. Since the on-chip error correction engine may mis-correct data during error correction, that is, correct a multi-bit error as a single-bit error, therefore, memory grains with a correctable error correction status are also determined as memory grains with errors. Furthermore, the memory controller can not only correct data that cannot be corrected by the on-chip error correction engine, but also correct data mis-corrected by the on-chip error correction engine, which is beneficial to further improve the reliability of data in the memory module.
[0095] Among them, the error correction ability of the memory controller can meet the requirement of error correction for a target number of memory grains with errors, where the target number is the number of error correction code grains in the memory module. The number of memory grains with errors is less than or equal to the number of error correction code grains in the memory module, indicating that the number of redundant memory grains in the memory is greater than or equal to the number of memory grains with errors. That is, the number of memory grains with errors is within the error correction ability of the memory controller, and the memory controller can correct the data in the memory grains with errors based on the data in the memory grains without errors in the memory. For example, a DDR5 SDRAM x8 memory module includes 4 data grains and 1 error correction code grain, that is, the number of redundant memory grains in the memory is 1. When and only when the number of memory grains with errors is less than or equal to 1, the memory controller can correct the data in the memory grains with errors. When the number of memory grains with errors is greater than 1, the memory controller ends the data error correction process and returns an uncorrectable error to the CPU.
[0096] In some embodiments, the memory controller uses an erasure code (EC) algorithm to correct the data in the memory grains with errors. Taking a DDR5 SDRAM x8 memory module as an example, the data error correction process includes the following steps 405A to 405C.
[0097] 405A. Denote the data in the 4 data grains as D1, D2, D3, and D4 respectively, and denote the data in 1 error correction code grain (redundant grain) as C1. Among them, D1, D2, D3, D4, and C1 are all 8×8 data blocks. According to the RS code algorithm, that is, the property of the EC code, there exists a matrix H for generating codewords, where the visible elements of H are all 8×8 block matrices. This process can be represented by the following formula (1).
[0098]
[0099] 405B. For the case where an error occurs in any memory particle, the matrix row corresponding to this memory particle is removed from matrix H, and the matrix is still a full-rank matrix H'. This full-rank matrix has an inverse matrix H'. -1 . For example, if an error occurs in memory particle D1, then the matrix row corresponding to D1 is removed from matrix H. This process can be represented by the following formula (2).
[0100]
[0101] 405C. Multiply the data after removing D1 on the left by the inverse matrix H'. -1 , and the corrected data is obtained. This process can be represented by the following formula (3).
[0102]
[0103] It should be noted that the above steps 403 to 405 are an implementation manner of determining the memory particle in which an error occurs in the memory module and correcting the data based on the memory particle in which the error occurs in case of failed error correction. In the following embodiments, this process is implemented in other ways, and the embodiments of the present application do not make any limitations in this regard.
[0104] 406. The memory controller writes the corrected data back to the memory particle.
[0105] Among them, the memory controller writes the corrected data back to the memory address corresponding to the read operation, that is, replaces all the data at the position corresponding to this memory address in each memory particle. In some embodiments, the memory controller writes the data in the corrected data corresponding to the memory particle in which an error occurs back to the position corresponding to this memory address in this memory particle, that is, only replaces the data at the position corresponding to this memory address in the memory particle in which an error occurs. The embodiments of the present application do not make any limitations in this regard.
[0106] In the above method, the memory controller writes the corrected data back to the memory particle, so that when the same memory address is read next time, the correct data can be read, which is beneficial to improving the reliability of the data in the memory module.
[0107] It should be noted that this step 406 is an optional step. In some embodiments, this step 406 is not executed, and the memory controller directly returns the corrected data to the CPU to shorten the response time of the read operation on the memory module and improve the response efficiency. The embodiments of the present application do not make any limitations in this regard.
[0108] 407. The memory controller releases the backpressure on the read and write processes of the memory module.
[0109] Among them, step 407 corresponds to the above-mentioned step 403. If the memory controller backpresses the read / write process of the memory module by means of powering down the processor core, the memory controller wakes up the processor core to relieve the backpressure; if the memory controller backpresses the read / write process of the memory module by means of the backpressure command scheduling queue, the memory controller relieves the backpressure on the command scheduling queue, so that the path between the scheduler and the physical interface protocol is connected, thereby relieving the backpressure on the read / write process of the memory module.
[0110] It should be noted that step 407 is an optional step. In some embodiments, if the above-mentioned step 403 is not executed, this step 407 is not executed, and the embodiments of the present application do not make any limitations in this regard.
[0111] 408. The memory controller rereads the memory address corresponding to the read operation, checks the reread data. If the check passes, the memory controller returns the reread data to the processor; if the check fails, the memory controller reports an uncorrectable error to the processor.
[0112] Among them, the process of the memory controller rereading the memory address corresponding to the read operation is the same as the process of the memory controller obtaining the data in the memory module in the above-mentioned step 402, and the process of the memory controller checking the reread data is the same as the process of the memory controller detecting the data in the above-mentioned step 402, which will not be elaborated here.
[0113] In the above-mentioned step 408, the memory controller checks the reread data and returns the data only when the check passes, which can ensure data consistency and improve the reliability of the data in the memory.
[0114] It should be noted that step 408 is an optional step. In some embodiments, if the above-mentioned step 406 is not executed, this step 408 is not executed, and the embodiments of the present application do not make any limitations in this regard.
[0115] Next, Figure 5 an example is given to illustrate the process shown in the above steps 401 to 408. Figure 5 is a schematic flowchart of a memory error correction method provided by an embodiment of the present application. As Figure 5 shown, Figure 5 it includes a memory controller and a memory module. Among them, the memory module is a DDR5 SDRAM x8 memory module, and the memory controller is a DDRC. One memory rank in one channel of the memory module includes 4 data grains and 1 error correction code grain. The number of external data lines of each data grain is 8. The memory grain includes an on-chip error correction engine, and each memory grain corresponds to a first register. The memory controller includes an RS code DDRCECC engine, a DMC, an MCU, a BPG, and a PHY.
[0116] Figure 5 Step 1 in the method (corresponding to step 401 above): In response to a read operation on the memory module, the on-chip error correction engine reads data from the storage array of the memory die and performs a check. Based on the check result, the error correction status information is written into the MRx register. Figure 5 Step 2 in the method (corresponding to step 402 above): The data in each memory die is returned to the DDRC ECC engine in the memory controller via the data bus. The DDRC ECC engine in the memory controller uses the RS code algorithm to detect the data in the memory die. If the RS code DDRC ECC engine does not detect an error, the memory controller directly returns the data to the CPU and the process ends prematurely. If the number of symbols with errors detected by the RS code DDRC ECC engine is within t, the memory controller corrects the data via the RS code DDRC ECC engine and returns it to the CPU, and the process ends prematurely. If the number of symbols with errors detected by the RS code DDRC ECC engine is greater than t, it exceeds the error correction capacity limit of the RS code DDRC ECC engine. Figure 5 Step 3 in the method (corresponding to step 403 above): If the error in the memory data accessed by any CPU core exceeds the error correction capacity limit of the RS code DDRC ECC engine, the RS code DDRC ECC engine triggers a synchronous external abort (SEA), and the MCU enters the basic input output system (BIOS) processing flow. The MCU first determines whether the address belongs to a DMA address and performs corresponding operations: if it belongs to a DMA address, the current error correction is abandoned, an error status is returned, and the original processing flow is executed; if it does not belong to a DMA address, the subsequent error correction flow is continued. Figure 5 Step 4 in the method (corresponding to step 403 above): The DMC backpressures the read and write processes of the memory module to prevent data from being rewritten by other read and write processes, causing data inconsistency. Figure 5 Step 5 in the method (corresponding to step 404 above): The MCU initiates a read of the memory die (DRAM) via the BPG, and compares the read data with the error data recorded by the RS code DDRC ECC engine. If they are the same, the processing flow continues; otherwise, it indicates that the data was rewritten before the backpressure took effect, the process ends prematurely, and an error status is returned. The MCU directly reads the first register (MRx.uv) of each memory die via the MRR command, and records the error correction status information obtained by the on-chip error correction engine of each memory die. Figure 5Step 6 (corresponding to steps 405 and 406 above): When and only when the error correction status of a certain particle is 10B and the rest are 00B or 01B, the MCU uses the EC algorithm to perform data error correction; the process of the rest of the scenarios ends in advance and returns non-correctable. After the EC algorithm completes error correction, the MCU writes the corrected data back to the memory particle through BPG, and modifies the interrupt return vector to trigger the system to reread the memory address corresponding to the read operation, that is, to reread the cache line. Figure 5 Step 7 (corresponding to steps 407 and 408 above): The MCU releases the backpressure. After the backpressure is released, the memory controller initiates a reread of the data in the memory particle where the error occurred, and uses the RS code DDRC ECC engine to check the reread data. If the check passes, the data is returned; otherwise, a non-correctable error is reported.
[0117] It should be noted that the above steps 401 to 408 are described by taking the DDR5 SDRAM x8 memory module as an example. The memory error correction method provided in the embodiments of the present application can also be applied to the DDR5 SDRAM x4 memory module. The memory error correction method in the DDR5 SDRAM x4 memory module is the same as the process shown in the above steps 401 to 407. The difference is that in the DDR5 SDRAM x4 memory module, in response to a memory read operation, each memory particle outputs 64-bit data to the memory controller. This 64-bit data is distributed in two codewords, and each memory particle contributes 32-bit data to one codeword. Since each memory particle has 4 external data lines, the data output by each memory particle to the memory controller is also a 4×8 data block; in addition, the DDR5 SDRAM x4 memory module includes 8 data particles and 2 error correction code particles, that is, the number of redundant memory particles in the memory module is 2. When and only when the number of memory particles with errors is less than or equal to 2, the memory controller can correct the data in the memory particle with errors. When the number of memory particles with errors is greater than 2, the memory controller ends the data correction process and returns a non-correctable error to the CPU. In the DDR5 SDRAM x4 memory module, the process of the memory controller correcting data includes the following steps A to C.
[0118] Step A: Denote the data in the 8 data particles as D1, D2…, D8 respectively, and denote the data in the 2 error correction code particles (redundant particles) as C1 and C2. Among them, D1-D8, C1, and C2 are all 4×8 data blocks. According to the properties of the RS code algorithm, that is, the EC code, there is a matrix H for generating codewords, where the visible elements of H are all 4×4 block matrices. This process can be represented by the following (4).
[0119]
[0120] Step B: For the case where any two memory chips have errors, remove the matrix rows corresponding to these two memory chips from matrix H. The matrix is still a full-rank matrix H′, and this full-rank matrix H′ has an inverse matrix H′ -1 . For example, if memory chips D1 and C1 have errors, then remove the matrix rows corresponding to D1 and C1 from matrix H. This process can be represented by the following formula (5).
[0121]
[0122] Step C: Multiply the read data on the left by the inverse matrix H′ -1 , and the corrected data will be obtained. This process can be represented by the following formula (6).
[0123]
[0124] The other steps of the memory error correction method in the DDR5 SDRAM x4 memory module are the same as those in the memory error correction method in the DDR5 SDRAM x8 memory module, and the same parts will not be elaborated here.
[0125] In the above method, the on-chip error correction engine performs the first error correction on the data and writes the error correction status information of the memory chips into the first register, enabling the memory controller to locate the faulty chips by reading the error correction status information in this first register. The error correction at the chip level and the system level can cooperate with each other, achieving the full utilization of redundant resources and being conducive to improving the error correction ability of the memory system under a fixed redundancy configuration; after the memory controller obtains the data, it performs the second error correction on the data. Compared with only relying on the on-chip error correction engine for error correction, it can correct the silent errors and mis-corrected errors that the on-chip error correction engine fails to detect, thereby reducing the risk of data silent errors and data mis-correction errors in the memory chips; if the memory controller fails to correct the errors, the memory controller reads the error correction status information of the memory chips from the register in the memory module, determines the faulty memory chips based on the error correction status information, and then performs the third error correction on the data. Since the result of the chip-level error correction is utilized, compared with only relying on the memory controller for error correction, it can improve the error correction ability of the memory controller, thereby improving the error correction ability of the memory system and the reliability of the data in the memory module. For x4 memory, this method can achieve error correction for two memory chips (dual-chipkill), and for x8 memory modules, this method can achieve error correction for a single memory chip (chipkill).
[0126] It should be noted that the above embodiments are described by taking the module form in the application form of the memory module as an example. In some embodiments, the memory error correction method provided by the embodiments of the present application can be applied to the memory module in the form of chip surface mounting. In other embodiments, the memory error correction method provided by the embodiments of the present application is also applicable to the scenario where the memory particles include on-chip error correction engines and the memory controller uses RS codes for error correction. For example, the scenario of DDR6 or LPDDR4 with sideband ECC. The embodiments of the present application are not limited to the specific scenarios shown above.
[0127] Figure 6 A memory controller provided by an embodiment of the present application, the memory controller includes a data acquisition module 601 and a data error correction module 602.
[0128] The data acquisition module 601 is configured to obtain data from the memory module in response to a read operation on the memory module.
[0129] The data error correction module 602 is configured to correct the data. If the error correction fails, determine the memory particles in the memory module that have errors, and based on the memory particles that have errors, correct the data.
[0130] Optionally, the data error correction module 602 includes:
[0131] A reading unit, configured to correct the data. If the error correction fails, read the error correction status information of multiple memory particles from the first registers of the multiple memory particles.
[0132] A determining unit, configured to determine the memory particle that has an error as the memory particle that has an error if the error correction status information of any memory particle indicates that the memory particle has an error.
[0133] Optionally, the reading unit is configured to:
[0134] Correct the data. If the error correction fails, the memory controller backpresses the read / write process of the memory module.
[0135] Obtain the data in the memory again. If the data obtained twice is consistent, read the error correction status information of multiple memory particles from the first registers of the multiple memory particles.
[0136] Optionally, the data error correction module includes:
[0137] An error correction unit, configured to correct the data based on the memory particles that have errors if the number of memory particles that have errors is less than or equal to the number of error correction code particles in the memory module.
[0138] Optionally, the memory controller further includes:
[0139] A write-back module, configured to write the error-corrected data back to the memory die;
[0140] A re-reading module, configured to re-read the memory address corresponding to the read operation;
[0141] A verification module, configured to verify the re-read data. If the verification passes, the re-read data is returned to the processor. If the verification fails, the memory controller reports an error to the processor.
[0142] Optionally, the error correction status of the memory die is represented by the value of the target bit of the first register, and the error correction status of the memory die represented by the value of the target bit of the first register is any one of no error, correctable error, and uncorrectable error.
[0143] Optionally, the error correction status of the memory die is represented by the occupied status of the target bit of the first register. The occupied status of the target bit of the first register being empty indicates that the error correction status of the memory die is no error or correctable error, and the occupied status of the target bit of the first register not being empty indicates that the error correction status of the memory die is uncorrectable error.
[0144] Optionally, the memory die with an error refers to at least one of the memory die with a correctable error status and the memory die with an uncorrectable error status.
[0145] Among them, both the data acquisition module 601 and the data error correction module 602 can be implemented by software or by hardware. Exemplarily, next, taking the data acquisition module 601 as an example, the implementation manner of the data acquisition module 601 is introduced. Similarly, the implementation manner of the data error correction module 602 can refer to the implementation manner of the data acquisition module 601.
[0146] As an example of a software functional unit, the data acquisition module 601 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the data acquisition module 601 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region, or may be distributed in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ), or may be distributed in different AZs, and each AZ includes one data center or multiple geographically close data centers. Among them, generally, one region may include multiple AZs.
[0147] Similarly, multiple hosts / virtual machines / containers for running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Usually, one VPC is set up within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set up within each VPC, and the interconnection between VPCs is achieved through the communication gateway.
[0148] As an example of a hardware functional unit, the data acquisition module 601 may include at least one computing device, such as a server. Alternatively, the data acquisition module 601 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0149] The multiple computing devices included in the data acquisition module 601 can be distributed in the same region or in different regions. The multiple computing devices included in the data acquisition module 601 can be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the data error correction module 602 can be distributed within the same VPC or across multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0150] It should be noted that in other embodiments, the steps to be implemented by the above modules can be specified as needed. The above modules respectively implement different steps in the above memory error correction method to achieve all the functions of the above device. That is, the memory error correction device provided in the above embodiments is only illustrated by the division of the above functional modules when implementing the memory error correction method. In practical applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the corresponding method embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.
[0151] The present application further provides a computing device, which includes a memory controller and a memory module. The memory module is used for temporarily storing data, and the memory controller is used for executing the memory error correction method provided in the above method embodiments.
[0152] Figure 7 is a schematic structural diagram of a computing device provided by an embodiment of the present application. As Figure 7 shown, the computing device 700 includes: a bus 701, a processor 702, a memory 703, and a communication interface 704. The processor 702, the memory 703, and the communication interface 704 communicate with each other through the bus 701. The computing device 700 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 700.
[0153] The bus 701 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 7 only one line is shown in [the figure], but it does not mean that there is only one bus or one type of bus. The bus 701 can include a path for transmitting information between various components (such as the memory 703, the processor 702, and the communication interface 704) of the computing device 700.
[0154] The processor 702 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0155] The memory 703 can include a volatile memory (VM), such as a random access memory (RAM). The memory 703 can also include a non-volatile memory (NVM), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0156] The executable program code is stored in the memory 703, and the processor 702 executes the executable program code to respectively implement the functions of the foregoing data acquisition module 601 and data error correction module 602, thereby implementing the memory error correction method. That is, the memory 703 stores instructions for executing the memory error correction method.
[0157] The communication interface 704 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 700 and other devices or communication networks.
[0158] An embodiment of the present application provides a processor, which includes a memory controller and a computing core. The processor is used to execute the memory error correction method provided in the foregoing embodiment, and the computing core is used to perform computing operations on the data in the memory particles.
[0159] An embodiment of the present application provides a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, the at least one computing device executes the memory error correction method provided in the foregoing embodiment.
[0160] An embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions. When the instructions are executed by a computing device cluster, the computing device cluster executes the memory error correction method provided in the foregoing embodiment.
[0161] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the data involved in the present application is obtained under full authorization.
[0162] Those of ordinary skill in the art will appreciate that the method steps and units described in the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the steps and components of the embodiments have been generally described in terms of function in the above description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0163] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0164] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling, direct coupling, or communication connection can be an indirect coupling or communication connection through some interfaces, devices, or units, and can also be in the form of electrical, mechanical, or other connections.
[0165] The unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this application.
[0166] In addition, the units in each embodiment of this application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software unit.
[0167] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computing device (which may be a personal computer, a server, or a computing device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0168] In this application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. It should be understood that there is no logical or chronological dependency between "first", "second", and "nth", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first memory particle can be called the second memory particle, and similarly, the second memory particle can be called the first memory particle. Both the first memory particle and the second memory particle can be memory particles, and in some cases, they can be separate and different memory particles.
[0169] In this application, the meaning of the term "at least one" refers to one or more, and the meaning of the term "a plurality of" refers to two or more. For example, a plurality of first memory particles refers to two or more first memory particles. In this article, the terms "system" and "network" are often used interchangeably.
[0170] It should also be understood that the term "if" can be interpreted to mean "when" ("when" or "upon") or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined..." or "if [the stated condition or event] is detected" can be interpreted to mean "when it is determined..." or "in response to determining..." or "when [the stated condition or event] is detected" or "in response to detecting [the stated condition or event]".
[0171] The above description is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present application, and these modifications or substitutions should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
[0172] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer program instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0173] The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid-state drive).
[0174] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0175] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A memory error correction method, characterized in that, Executed by a memory controller, the method includes: In response to a read operation on a memory module, reading data from the memory module, where the data is the data corrected by an on-chip error correction engine of memory dies in the memory module; Performing error correction on the data. If the error correction fails, stop rewriting the data; obtain the data in the memory module again. If the data obtained twice is the same, read the error correction status information of the multiple memory dies from a first register of the multiple memory dies in the memory module; if the error correction status information of any one of the memory dies indicates that the memory die has an error, determine that the memory die is the memory die with an error, and the error correction status information of the memory die is written into the first register by the on-chip error correction engine of the memory die after correcting the data in the memory die; Based on the memory die with an error, perform error correction on the data.
2. The method according to claim 1, wherein The performing error correction on the data based on the memory die with an error includes: If the number of memory dies with an error is less than or equal to the number of error correction code dies in the memory module, perform error correction on the data based on the memory dies with an error.
3. The method according to claim 1 or 2, characterized in that, After performing error correction on the data based on the memory die with an error, the method further includes: Writing the error-corrected data back to the memory die; Re-reading the memory address corresponding to the read operation, and verifying the re-read data; If the verification passes, return the re-read data to the processor; If the verification fails, report an error to the processor.
4. The method according to claim 1 or 2, characterized in that, The error correction status of the memory die is represented by the value of a target bit of the first register, and the error correction status of the memory die represented by the value of the target bit of the first register is any one of no error, correctable error, and uncorrectable error.
5. The method according to claim 1 or 2, characterized in that, The error correction status of the memory die is represented by the occupancy status of a target bit of the first register. The occupancy status of the target bit of the first register being empty indicates that the error correction status of the memory die is no error or correctable error, and the occupancy status of the target bit of the first register not being empty indicates that the error correction status of the memory die is uncorrectable error.
6. The method according to claim 1 or 2, characterized in that The memory die with an error refers to at least one of the memory dies with a correctable error status and the memory dies with an uncorrectable error status.
7. The method according to claim 1 or 2, characterized in that, The memory die includes an on-chip error correction engine, and the method further includes: In response to a read operation on the memory module, the on-chip error correction engine obtains the data in the memory die; The on-chip error correction engine corrects the obtained data, writes the error correction status information of the memory die into the first register of the memory die, and outputs the error-corrected data to the memory controller.
8. A memory error correction method, characterized in that, Executed by a memory module, the memory module includes multiple memory dies, and each memory die includes a first register. The method includes: Performing error correction on the data in the memory die to obtain the error correction status information of the memory die; Write the error correction status information into the first register of the memory die. The error correction status information is used to indicate the memory die in the memory module where an error occurs, and is also used to be read by the memory controller to determine the memory die in the memory module where an error occurs. The memory controller is used to execute the memory error correction method according to any one of claims 1 to 7.
9. A memory controller, characterized in that, The memory controller includes: A data acquisition module, configured to read data from the memory module in response to a read operation on the memory module. The data is the data after being error-corrected by the on-die error correction engine of the memory die in the memory module. A data error correction module, including: A read unit, configured to perform error correction on the data. If the error correction fails, stop rewriting the data; acquire the data in the memory module again. If the data acquired twice is the same, read the error correction status information of multiple memory dies from the first registers of multiple memory dies in the memory module. A determination unit, configured to, if the error correction status information of any one of the memory dies indicates that the memory die has an error, determine that the memory die is the memory die with an error, and perform error correction on the data based on the memory die with the error. The error correction status information of the memory die is written into the first register after the on-die error correction engine of the memory die performs error correction on the data in the memory die.
10. The memory controller according to claim 9, wherein The data error correction module includes: An error correction unit, configured to, if the number of memory dies with errors is less than or equal to the number of error correction code dies in the memory module, perform error correction on the data based on the memory dies with errors.
11. The memory controller according to claim 9 or 10, characterized in that, The memory controller further includes: A write-back module, configured to write the error-corrected data back into the memory die. A re-read module, configured to re-read the memory address corresponding to the read operation. A verification module, configured to verify the re-read data. If the verification passes, return the re-read data to the processor. If the verification fails, report an error to the processor.
12. The memory controller according to claim 9 or 10, characterized in that, The error correction status information of the memory die is the value of the target bit of the first register. The error correction status of the memory die indicated by the value of the target bit of the first register is any one of no error, correctable error, and uncorrectable error.
13. The memory controller according to claim 9 or 10, characterized in that, The error correction status information of the memory die is the occupancy status of the target bit of the first register. The occupancy status of the target bit of the first register being empty indicates that the error correction status of the memory die is no error or correctable error, and the occupancy status of the target bit of the first register not being empty indicates that the error correction status of the memory die is uncorrectable error.
14. The memory controller according to claim 9 or 10, characterized in that, The memory die with an error refers to at least one of the memory dies with a correctable error status and the memory dies with an uncorrectable error status.
15. A memory module, characterized in that, The memory module includes multiple memory dies, and each memory die includes a first register. The memory module is used to execute the memory error correction method according to claim 8.
16. A memory controller, characterized in that, The memory controller is used to execute the memory error correction method according to any one of claims 1 to 7.
17. A processor, characterized in that, The processor includes a memory controller and a computing core. The processor is configured to execute the memory error correction method according to any one of claims 1 to 7, and the computing core is configured to perform computing operations on data in memory chips.
18. A computing device, characterized in that, It includes a memory controller and a memory module. The memory module is used for temporarily storing data, and the memory controller is configured to execute the memory error correction method according to any one of claims 1 to 7.
19. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used for storing at least one segment of program code, and the at least one segment of program code is used for executing the memory error correction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for fault tolerance of grading instruction memory structure capable of actively writing back
CN107885611A
Memory error correction method, memory controller and electronic equipment
CN112579342A