Memory device assisted error correction

By incorporating error history buffers in memory chips to record internal error corrections, the memory controller can better identify and correct faults in memory modules with limited parity information, improving overall error correction efficacy.

WO2025096310A1PCT designated stage expired Publication Date: 2025-05-08RAMBUS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/053112
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-02
Filing Date
2024-10-25
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing memory modules with a limited number of memory chips struggle to detect and correct large faults, as the error correction logic lacks sufficient information to determine which chip has failed.

Method used

Implementing error history buffers on each memory chip to record instances of internal error correction, allowing the memory controller to query these buffers and determine which chip may have failed, thereby facilitating error correction.

Benefits of technology

Enhances the error correction capability of memory modules by providing the memory controller with the necessary information to identify and correct faults, even in configurations with insufficient parity information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024053112_08052025_PF_FP_ABST
    Figure US2024053112_08052025_PF_FP_ABST
Patent Text Reader

Abstract

A memory controller includes a memory interface configured to receive data read from a memory module during a first transaction. The memory module includes a plurality of memory devices, wherein the plurality of memory devices maintain respective error history buffers. Responsive to a triggering event, the memory interface is configured to query the respective error history buffers and receive error history information associated with the first transaction. The memory controller further includes error correction logic, which is executed by a processing device and coupled to the memory interface. The error correction logic is configured to perform an error correction operation on the data using the error history information.
Need to check novelty before this filing date? Find Prior Art

Description

MEMORY DEVICE ASSISTED ERROR CORRECTIONBACKGROUND

[0001] Modem computer systems generally include a data storage device, such as a memory component. The memory component may be, for example a random access memory (RAM) or a dynamic random access memory (DRAM). The memory component includes memory banks made up of storage cells which are accessed by a memory controller or memory client through a command interface and a data interface within the memory component.BRIEF DESCRIPTION OF THE DRAWINGS

[0002] The present disclosure is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings.

[0003] Figure l is a block diagram illustrating a computing environment with a memory module configured for memory device assisted error correction, according to an embodiment.

[0004] Figure 2 is a flow diagram illustrating a method for memory device assisted error correction, according to an embodiment.

[0005] Figure 3 is a block diagram illustrating on-chip error history buffers for memory device assisted error correction, according to an embodiment.

[0006] Figure 4 is a block diagram of an example computer system in which embodiments of the present disclosure may operate.DETAILED DESCRIPTION

[0007] The following description sets forth numerous specific details such as examples of specific systems, components, methods, and so forth, in order to provide a good understanding of several embodiments of the present disclosure. It will be apparent to one skilled in the art, however, that at least some embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known components or methods are not described in detail or are presented in simple block diagram format in order to avoid unnecessarily obscuring the present disclosure. Thus, the specific details set forth are merely exemplary. Particular implementations may vary from these exemplary details and still be contemplated to be within the scope of the present disclosure.

[0008] A memory module, such as a dual in-line memory module (DIMM), may include a series of memory devices (i.e., memory chips), such as dynamic random accessmemory (DRAM) integrated circuits. The memory module may be accessible by a host system implementing a host memory controller to perform memory access operations, such as read or write operations, on the memory module. The host memory controller may include error correction logic that may detect and correct certain errors occurring on the memory module. In some implementations, a certain number of the memory chips in the module may be used to store data, while a remainder of the memory chips are used to store parity information that may be used to correct certain faults and errors in the data. For example, a given memory module might include ten memory chips, where eight memory chips are used to store data and two memory chips are used to store parity information. Such a configuration allows the error correction logic to detect and identify the location of a major fault (e.g., partial or complete chip failure) of any of the memory chips in the module, as well as correct the fault to restore any lost data. Other memory modules, however, in an effort to reduce cost, size, power consumption, etc., may have a lesser number of memory chips, such as nine memory chips, where eight memory chips are used to store data and only one memory chip is used to store parity information. In this configuration, if any memory chip suffers a large-enough fault, the error correction logic of the memory controller will not typically have enough information to detect and correct the error. The size of such a fault will depend on the error correction capability of the system, which may vary, for example, from a single bit error up to half a chip (e.g., 4 bytes on DDR5 memory). More precisely, when a large fault occurs, the error correction logic may not have enough information to determine which chip failed, but would have enough information to correct the error if the location was known.

[0009] In various embodiments, certain memory devices feature internal errorcorrection capability (e.g., implementing a single-error correction code). The worst faults in these memory devices usually cause multiple errors, which might be misinterpreted as singleerrors by the internal error correction, and be correspondingly miscorrected, thereby potentially introducing additional errors. This behavior is largely indiscernible by the associated memory controller, which may obtain long-term statistics about the on-device error correction, but not fine-grained information. The behavior of the internal error correction, however, may be useful to the error correction logic executing on the memory controller.

[0010] If the error correction logic of the memory controller could determine which memory chip or chips attempted to correct an error during a given transaction using the internal on-chip error correction, this may be indicative of which memory chip failed (e.g. if exactly one chip reported a single error correction during the transaction). Thus, if thelocation of the fault were known, the error correction logic of the memory controller would be able to correct the fault using the available parity information in the memory module.

[0011] To improve the error correction capability of the memory module, in certain embodiments, the memory chips may be modified to maintain a record of when the internal error correction capability is activated for a certain number of past transactions (i.e., to maintain an error history buffer), and possibly other conditions that are likely indicative of the occurrence of an error or memory chip fault. For example, the error history buffer could be maintained in a register on each memory chip. The error correction logic on the memory controller could query each memory chip for the information in the error history buffer as needed, such as by using a memory register read command or other communication protocol. In one embodiment, the error history buffer implementation could be a shift register which, on every read transaction, shifts over by one position and records whether an error was observed during that transaction. In other embodiments, the error history buffer implementation could record other indications of an error, whether the error was deemed recoverable on-die, which bit(s) were corrected, etc. This could be stored in the shift register or in some other structure. Depending on the structure, it is possible that only the most recent erroneous transaction or few transactions would be stored in the error history buffer, so that on a query indicating a transaction too far in the past, an “uncertain” or similar result may be returned.

[0012] While in some implementations, it may be possible for the memory chips to alert the memory controller of the use of the internal error correction capability, the controller will typically already be aware of the occurrence of the error by means of its own error correction logic. Thus, the controller may choose to query the error history buffer only on errors that are otherwise unrecoverable, such as apparent chip failures on modules which don’t have enough parity for a chipkill operation (e.g., a 9-chip DDR5 module). Similarly, when using beyond-bound chipkill (i.e., 10 chips, but with per-line metadata putting chipkill errors beyond the Singleton bound) there are rare ambiguous cases where the error correction logic cannot determine which memory chip in the module has failed. Consulting the on- device error history buffers would increase the chance that the controller can determine which chip has failed.

[0013] Figure 1 depicts an environment 100 showing a memory module 120. As an option, one or more instances of environment 100 or any aspect thereof may be implemented in the context of the architecture and functionality of the embodiments described herein.

[0014] As shown in Figure 1, environment 100 comprises a host 102 coupled to a memory module 120 through a system bus 110. In one embodiment, memory module 120 is a dual in-line memory module (DIMM). Such memory modules may be referred to as DRAM DIMMs, registered DIMMs (RDIMMs), or load-reduced DIMMs (LRDIMMs), and may share a memory channel with other DRAM DIMMs.

[0015] In one embodiment, the host 102 further comprises a CPU core 103, a cache memory 104, and a host memory controller 105. Host 102 may comprise multiple instances each of CPU core 103, cache memory 104, and host memory controller 105. The host 102 of environment 100 may further be based on various architectures (e.g., Intel x86, ARM, MIPS, IBM Power, etc.). Cache memory 104 may be dedicated to the CPU core 103 or shared with other cores. The host memory controller 105 of the host 102 communicates with the memory module 120 through the system bus 110 using a physical interface 112 (e.g., compliant with the JEDEC DDR3, DDR4 or DDR5 SDRAM standard, etc.). Specifically, the host memory controller 105 may write data to and / or read data from multiple sets of DRAM devices 124i - 1244 using a data bus 1141 and a data bus 1142, respectively. For example, the data bus 114i and the data bus 1142 may transmit the data as electronic signals such as a data signal, a chip select signal, or a data strobe signal.

[0016] In one embodiment, the host memory controller 105 includes error correction logic 107 that may detect and correct certain errors occurring on the DRAM devices 124i - 1244 of memory module 120. For example, upon receiving data from DRAM devices 124i - 1244 via data bus 1141 and data bus 1142, the error correction logic 107 may perform error correction operations to detect and correct errors in the data. The error correction operations may vary depending on the embodiment, but could include, for example, use of Reed- Solomon codes, basic parity information, Hamming codes, Bose-Chaudhuri-Hocquenghem (BCH) codes, etc. As noted above, depending on the number of memory chips that make up each of DRAM devices 124i - 1244, the error correction logic 107 may not have the capability to identify the location of a major fault (i.e., identify the specific memory chip within one of DRAM devices 124i - 1244 that has failed), but would have the capability to recover from such a fault if the location was known.

[0017] The DRAM devices 124i - 1244 may each comprise an array of some number (e.g., eight, nine, or ten) memory devices (e.g., SDRAM) arranged in various topologies (e.g., A / B sides, single-rank, dual-rank, quad-rank, etc.). In some cases, as shown, the data to and / or from the DRAM devices 124i - 1244 may optionally be buffered by a set of data buffers 122i and data buffers 1222, respectively. Such data buffers may serve to boost thedrive of the signals (e.g., data or DQ signals, etc.) on the system bus 110 to help mitigate high electrical loads of large computing and / or memory systems. In other embodiments, data buffers 122i and data buffers 1222 are not present in memory module 120.

[0018] Further, commands from the host memory controller 105 may be received by a command buffer 126, such as a register clock driver (RCD), at the memory module 120 using a command and address (CA) bus 116. For example, the command buffer 126 might be an RCD such as included in registered DIMMs (e.g., RDIMMs, LRDIMMs, etc.). Command buffers such as command buffer 126 may comprise a logical register and a phase-lock loop (PLL) to receive and re-drive command and address input signals from the host memory controller 105 to the DRAM devices on a DIMM (e.g., DRAM devices 124i, DRAM devices 1242, etc.), reducing clock, control, command, and address signal loading by isolating the DRAM devices from the host memory controller 105 and the system bus 110. In other embodiments, in addition or in the alternative, memory module 120 may include other volatile memory devices, such as synchronous DRAM (SDRAM), Rambus DRAM (RDRAM), static random access memory (SRAM), etc.

[0019] In one embodiment, each of the memory chips within DRAM devices 124i - 1244 may maintain a record of when the internal error correction capability on the chips is activated. For example, for a certain number of past transactions, each memory chip may track whether an error was detected and / or corrected by the internal error correction. These records may be represented by error history buffers 1251-1254. Depending on the embodiment, there may be a single buffer for each of DRAM devices 124i - 1244, or there may be respective buffers for each individual memory chip. When an uncorrectable error is detected by error correction logic 107, the information in the error history buffers 1251-1254 may be queried, such as by using a memory register read (MRR) command sent via CA bus 116 to command buffer 126. Command buffer 126 may then forward the command on to the appropriate one of DRAM devices 124i - 1244. In one embodiment, data buffers 1221-1222 also store data associated with some number of previous transactions (e.g., data read from DRAM devices 124i - 1244). Thus, in addition to querying the error history buffers 1251- 1254, error correction logic 107 may query data buffers 1221-1222 to obtain the associated data, which may be useful for certain error detection and correction operations. Additional details regarding operations of error correction logic 107 are described below.

[0020] The memory module 120 shown in environment 100 presents merely one partitioning. The specific example shown where the command buffer 126 and the DRAM devices 124i - 1244 are separate components is purely exemplary, and other partitioning ispossible. For example, any or all of the components comprising the memory module 120 and / or other components may comprise one device (e.g., system-on-chip or SoC), multiple devices in a single package or printed circuit board, multiple separate devices, and may have other variations, modifications, and alternatives.

[0021] Figure 2 is a flow diagram illustrating a method for memory device assisted error correction, according to an embodiment. The method 200 may be performed by processing logic that may comprise hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions run on a processing device to perform hardware simulation), or a combination thereof. In one embodiment, the method 200 is performed by the memory interface 306 and error correction logic 107 of the host memory controller 105, as shown in Figure 1 and Figure 3.

[0022] At operation 210, the processing logic sends a read command to a memory module to read data. In one embodiment, as shown in Figure 3, host memory controller 105 includes a memory interface component 306 in addition to error correction logic 107, and potentially other components. The memory interface component 306 is responsible for handling communication between host memory controller 105 and memory module 120, and may thus send and receive commands, data, or other messages via physical interface 112 and system bus 110. In one embodiment, memory interface 306 sends the read command to one or more memory chips 324I-324N (i.e., memory devices) of the memory module 120. The read command may be in response to a host request (e.g., initiated by an application or operating system executed by CPU core 103) or may be an internal read command initiated by host memory controller 105 itself. In one embodiment, the memory chips 324I-324N represent the individual memory chips that form one of DRAM devices 124i - 1244, as shown in Figure 1. In one embodiment, the memory chips 324I-324N represent nine memory chips, where eight memory chips are used to store data and one memory chip is used to store parity information pertaining to the data on the other eight memory chips. In other embodiments, there may be some other number of memory chips.

[0023] Referring again to Figure 2, at operation 220, the processing logic receives the data read from memory module 120 during a first transaction. In one embodiment, the data includes a portion of the data from each of the memory chips 324I-324N. In one embodiment, host memory controller 105 executes a sequence of memory transactions to perform various memory access operations on memory module 120. The sequence of transactions may include for example, read operations, program operations, erase operations, or a combination of these and / or other operations. A transaction may represent an individual one of theseoperations, a portion of an operation, or a combination of multiple operations that are performed together. In one embodiment, a transaction includes a request (e.g., sent by memory interface 306 of host memory controller 105) and a response (e.g., sent from memory module 120). For example, a given transaction might include a read request sent to memory module 120 and a response sent to host memory controller 105. In some embodiments, there is a latency period between the different parts of the transaction (i.e., between the request and the response), such that multiple transactions might overlap. For example, if a request for a first transaction is sent in a first time window, a request for a second transaction could be sent in a subsequent second time window, and the response for the first transaction could be received in a subsequent third time window.

[0024] At operation 230, the processing logic performs an initial error correction operation on the data. In one embodiment, upon receiving the data from memory chips 324i- 324N, the error correction logic 107 may perform error correction operations to detect and correct errors in the data. The error correction operations may vary depending on the embodiment, but could include, for example, use of Reed-Solomon codes, basic parity information, Hamming codes, Bose-Chaudhuri-Hocquenghem (BCH) codes, etc. As noted above, depending on the number of memory chips 324I-324N, the error correction logic 107 may not have the capability to identify the location of a major fault (i.e., identify the specific memory chip that has failed), but would have the capability to recover from such a fault if the location was known.

[0025] At operation 240, the processing logic determines whether there has been an occurrence of a triggering event. In one embodiment, the triggering event comprises a determination, based on the initial error correction operation, that one or more uncorrectable errors are present in the data. For example, if error correction logic 107 is unable to identify which of memory chips 324I-324N has suffered a failure, this may be one instance of a triggering event. In another embodiment, the triggering event comprises receipt of a signal from the memory module 120 indicating that internal error correction was activated on at least one of the memory chips 324I-324N during the first transaction. In another embodiment, the triggering event comprises a determination, based on the initial error correction operation, that one or more correctable errors are present in the data. In such an example, additional error correction operations may be performed to confirm the determination, as will be described in more detail below.

[0026] Responsive to the occurrence of a triggering event, at operation 250, the processing logic queries respective error history buffers associated with the memory chipsand receives error history information associated with the first transaction. In one embodiment, as illustrated in Figure 3, the memory chips 324I-324N maintain respective error history buffers 325I-325N. In one embodiment, the respective error history buffers 325I-325N comprise indicators associated with a plurality of transactions, the indicators to indicate whether internal error correction was activated on a corresponding memory chips 324I-324N during any of the plurality of transactions. For example, in the embodiment illustrated in Figure 3, each of error history buffers 325I-325N includes eight bits representing eight previous transactions. In other embodiments, some other number of previous transactions may be represented. In one embodiment, the error history buffers 325 I-325N are implemented as circular (e.g., ring or shift) buffers, where each time a newer transaction occurs, the value representing the oldest transaction is evicted. In Figure 3, the oldest transaction in each of error history buffers 325I-325N is represented by the value on the left and the newest transaction is represented by the value on the right. Thus, in error history buffer 325i, the most recent transaction has a value of “0” (i.e., indicating that internal error correction was not activated on memory chip 324i during that transaction), but the next value is set to “1” (i.e., indicating that internal error correction was activated on memory chip 324i during the second most recent transaction). In error history buffer 3252, the fifth most recent transaction has a value of “1” (i.e., indicating that internal error correction was activated on memory chip 3242 during that transaction).

[0027] In one embodiment, in response to the occurrence of the triggering event, memory interface 306 may query each of the respective error history buffers 325I-325N to determine the error history information stored therein. In one embodiment, the memory module 120 may return the full contents of each of error history buffers 325I-325N in response to this query. In another embodiment, the query may specify a particular transaction, and the memory module 120 may return only the values representing that transaction from each of the error history buffers 325I-325N. In one embodiment, based on an indication in the error history buffers 325I-325N that the internal error correction was activated on a corresponding one of memory chips 324I-324N during a given transaction, the error correction logic 107 may determine that the corresponding memory chip suffered multiple uncorrectable errors. Since the error correction logic 107 determined at operation 240 that there was an uncorrectable error associated with the transaction (e.g., a major fault of failure of one of memory chips 324I-324N), the fact that a certain memory chip invoked internal error correction operations during that same transaction is a strong suggestion thatthe same memory chip suffered multiple uncorrectable errors, which were not properly identified or corrected by the internal error correction of the memory chip.

[0028] Responsive to the occurrence of the triggering event, at operation 260, the processing logic optionally queries respective data buffers associated with the plurality of memory devices and re-receives the data associated with the first transaction. As illustrated in Figure 3, each of memory chips 324I-324N may have a corresponding one of data buffers 322I-322N. The data buffers 322I-322N may store the data associated with some number of previous memory transactions. For example, each of data buffers 322I-322N may store the data read from the corresponding memory chips 324I-324N during one or more previous memory transactions. In one embodiment, in addition to querying error history buffers 3251- 325N, error correction logic 107 may further query the corresponding data buffers 322I-322N to obtain the data that was read from memory chips 324I-324N during the relevant transactions.

[0029] At operation 270, the processing logic performs an error correction operation on the data using the error history information and, optionally, the data re-received from the respective data buffers. In one embodiment, the error correction logic 107 may re-read the data from the memory module 120 to check for errors in transmission of the data instead of in the storage of the data. Transmission errors would not be reflected in the error history buffers 325I-325N because such errors occur downstream of those buffers. If a transient transmission error caused the fault, then re-reading the data might solve the problem, or may show that enough errors have occurred to indicate that the problem is uncorrectable.

[0030] In another embodiment, the error correction logic 107 may use the information in the error history buffers 325I-325N for correcting errors (i.e., corruptions in unknown locations) and / or erasures (i.e., corruptions in known locations). A typical code may correct twice as many erasures as errors. For example, with one extra chip and a Reed-Solomon code, a chipkill can be corrected if it is known which chip died, but only a half-chip error (i.e., a bounded fault) if the location is not known.

[0031] In another embodiment, the error correction logic 107 may use more information about the error to help prevent miscorrections of the data, depending on a model of how the memory chips 324I-324N are likely to corrupt bits. For example, if one of memory chips 324I-324N reports that it corrected an error in bit N, and the memory controller 105 proposes a correction that flips the state of bit N, as well as a few other bits, then this would give confidence that the revised correction is right (and the chip initially miscorrected that bit). Similarly, if the memory chip reports that it detected an error but failed to correctthe error, there is more confidence that that memory chip gave bad data, and may have mathematical information about which bits were bad.

[0032] Figure 4 illustrates an example machine of a computer system 400 within which a set of instructions, for causing the machine to perform any one or more of the methodologies discussed herein, can be executed. In some embodiments, the computer system 400 may correspond to a host system (e.g., the host system 102 of Figure 1) that includes, is coupled to, or utilizes a memory module (e.g., the memory module 120 of Figure 1) or may be used to perform the operations of a controller (e.g., the host memory controller 105 of Figure 1). In alternative embodiments, the machine may be connected (e.g., networked) to other machines in a LAN, an intranet, an extranet, and / or the Internet. The machine may operate in the capacity of a server or a client machine in client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or a client machine in a cloud computing infrastructure or environment.

[0033] The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a Personal Digital Assistant (PDA), a cellular telephone, a web appliance, a server, a network router, a switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. Further, while a single machine is illustrated, the term “machine” shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.

[0034] The example computer system 400 includes a processing device 402, a main memory 404 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc.), a static memory 406 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage system 418, which communicate with each other via a bus 430.

[0035] Processing device 402 represents one or more general-purpose processing devices such as a microprocessor, a central processing unit, or the like. More particularly, the processing device may be a complex instruction set computing (CISC) microprocessor, reduced instruction set computing (RISC) microprocessor, very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets, or processors implementing a combination of instruction sets. Processing device 402 may also be one or more special-purpose processing devices such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), network processor, or the like. The processing device 402 is configured to execute instructions 426 forperforming the operations and steps discussed herein. The computer system 400 may further include a network interface device 408 to communicate over the network 420.

[0036] The data storage system 418 may include a machine-readable storage medium 424 (also known as a computer-readable medium) on which is stored one or more sets of instructions 426 or software embodying any one or more of the methodologies or functions described herein. The instructions 426 may also reside, completely or at least partially, within the main memory 404 and / or within the processing device 402 during execution thereof by the computer system 400, the main memory 404 and the processing device 402 also constituting machine-readable storage media. The machine-readable storage medium 424, data storage system 418, and / or main memory 404 may correspond to the memory module 120 of Figure 1.

[0037] In one embodiment, the instructions 426 include instructions to implement functionality corresponding to the error correction logic 107 of Figure 1. While the machine- readable storage medium 424 is shown in an example embodiment to be a single medium, the term “machine-readable storage medium” should be taken to include a single medium or multiple media that store the one or more sets of instructions. The term “machine-readable storage medium” shall also be taken to include any medium that is capable of storing or encoding a set of instructions for execution by the machine and that cause the machine to perform any one or more of the methodologies of the present disclosure. The term “machine- readable storage medium” shall accordingly be taken to include, but not be limited to, solid- state memories, optical media, and magnetic media.

[0038] Although the operations of the methods herein are shown and described in a particular order, the order of the operations of each method may be altered so that certain operations may be performed in an inverse order or so that certain operation may be performed, at least in part, concurrently with other operations. In certain implementations, instructions or sub-operations of distinct operations may be in an intermittent and / or alternating manner.

[0039] It is to be understood that the above description is intended to be illustrative, and not restrictive. Many other implementations will be apparent to those of skill in the art upon reading and understanding the above description. The scope of the disclosure should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.

[0040] In the above description, numerous details are set forth. It will be apparent, however, to one skilled in the art, that the aspects of the present disclosure may be practicedwithout these specific details. In some instances, well-known structures and devices are shown in block diagram form, rather than in detail, in order to avoid obscuring the present disclosure.

[0041] Some portions of the detailed descriptions above are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of steps leading to a desired result. The steps are those requiring physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0042] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise, as apparent from the following discussion, it is appreciated that throughout the description, discussions utilizing terms such as “receiving,” “determining,” “selecting,” “storing,” “setting,” or the like, refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices.

[0043] The present disclosure also relates to an apparatus for performing the operations herein. This apparatus may be specially constructed for the required purposes, or it may comprise a general purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer readable storage medium, such as, but not limited to, any type of disk including floppy disks, optical disks, CD-ROMs, and magnetic-optical disks, read-only memories (ROMs), random access memories (RAMs), EPROMs, EEPROMs, magnetic or optical cards, or any type of media suitable for storing electronic instructions, each coupled to a computer system bus.

[0044] The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. Various general purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems will appear as set forth in the description. In addition, aspects of the present disclosure are not described with reference to any particular programming language. It will be appreciated that a variety of programming languages may be used to implement the teachings of the present disclosure as described herein.

[0045] Aspects of the present disclosure may be provided as a computer program product, or software, that may include a machine-readable medium having stored thereon instructions, which may be used to program a computer system (or other electronic devices) to perform a process according to the present disclosure. A machine-readable medium includes any procedure for storing or transmitting information in a form readable by a machine (e.g., a computer). For example, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., a computer) readable storage medium (e.g., read only memory (“ROM”), random access memory (“RAM”), magnetic disk storage media, optical storage media, flash memory devices, etc.).

Claims

CLAIMSWhat is claimed is:

1. A memory controller comprising: a memory interface configured to: receive data read from a memory module during a first transaction, the memory module comprising a plurality of memory devices, wherein the plurality of memory devices maintain respective error history buffers; and responsive to a triggering event, query the respective error history buffers and receive error history information associated with the first transaction; and error correction logic, executed by a processing device, coupled to the memory interface, wherein the error correction logic is configured to perform an error correction operation on the data using the error history information.

2. The memory controller of claim 1, wherein the memory interface is further configured to: send a read command to the memory module to read the data.

3. The memory controller of claim 1, wherein the error correction logic is further configured to: perform an initial error correction operation on the data.

4. The memory controller of claim 3, wherein triggering event comprises a determination, based on the initial error correction operation, that one or more uncorrectable errors are present in the data.

5. The memory controller of claim 1, wherein triggering event comprises receipt of a signal from the memory module indicating that internal error correction was activated on at least one of the plurality memory devices during the first transaction.

6. The memory controller of claim 1, wherein the memory interface is further configured to: responsive to the triggering event, query respective data buffers associated with the plurality of memory devices and re-receive the data associated with the first transaction.

7. The memory controller of claim 6, wherein the error correction logic is further configured to: perform the error correction operation on the data using the error history information and the data re-received from the respective data buffers.

8. The memory controller of claim 1, wherein the respective error history buffers comprise indicators associated with a plurality of transactions, the indicators to indicate whether internal error correction was activated on a corresponding memory device during any of the plurality of transactions.

9. The memory controller of claim 8, wherein based on an indication in the respective error history buffers that the internal error correction was activated on a corresponding memory device during the first transaction, the error correction logic is to determine that the corresponding memory device suffered multiple uncorrectable errors.

10. A method comprising: receiving, by a memory interface of a memory controller, data read from a memory module during a first transaction, the memory module comprising a plurality of memory devices, wherein the plurality of memory devices maintain respective error history buffers; responsive to a triggering event, querying the respective error history buffers and receiving error history information associated with the first transaction; and performing, by error correction logic of the memory controller, an error correction operation on the data using the error history information.

11. The method of claim 10, further comprising: sending, by the memory interface, a read command to the memory module to read the data.

12. The method of claim 10, further comprising: performing, by the error correction logic, an initial error correction operation on the data.

13. The method of claim 12, wherein triggering event comprises a determination, based on the initial error correction operation, that one or more uncorrectable errors are present in the data.

14. The method of claim 10, wherein triggering event comprises receipt of a signal from the memory module indicating that internal error correction was activated on at least one of the plurality memory devices during the first transaction.

15. The method of claim 10, further comprising: responsive to the triggering event, querying, by the memory interface, respective data buffers associated with the plurality of memory devices and re-receiving, by the memory interface, the data associated with the first transaction.

16. The method of claim 15, further comprising: performing, by the error correction logic, the error correction operation on the data using the error history information and the data re-received from the respective data buffers.

17. The method of claim 10, wherein the respective error history buffers comprise indicators associated with a plurality of transactions, the indicators to indicate whether internal error correction was activated on a corresponding memory device during any of the plurality of transactions.

18. The method of claim 17, further comprising: determining, based on an indication in the respective error history buffers that the internal error correction was activated on a corresponding memory device during the first transaction, that the corresponding memory device suffered multiple uncorrectable errors.

19. An integrated circuit comprising: a memory; and a processing device operatively coupled to the memory, the processing device to perform operations comprising: receiving data read from a memory module during a first transaction, the memory module comprising a plurality of memory devices, wherein the plurality of memory devices maintain respective error history buffers;responsive to a triggering event, querying the respective error history buffers and receiving error history information associated with the first transaction; and performing an error correction operation on the data using the error history information.

20. The integrated circuit of claim 19, wherein the respective error history buffers comprise indicators associated with a plurality of transactions, the indicators to indicate whether internal error correction was activated on a corresponding memory device during any of the plurality of transactions.

Citation Information

Patent Citations

  • Error detection / correction code which detects and corrects memory module / transmitter circuit failure

    US20040003336A1

  • System for identifying and correcting data errors

    US20190042369A1

  • Runtime sparing for uncorrectable errors based on fault-aware analysis

    US20220350715A1