Memory Device, System, and Method with Error Check and Clear Mode
By implementing the error checksum clearance (ECS) mode in the memory device, the problem of DRAM error accumulation is solved, and the reliability and accessibility of the memory system are improved.
Patent Information
- Application Number
- CN202111091173.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-12-26
- Filing Date
- 2016-08-04
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2036-08-04
AI Technical Summary
In the prior art, DRAM errors accumulate in the memory device and are not detected, resulting in system-level ECC being unable to effectively correct memory device-level errors, affecting the reliability and accessibility of the memory system.
The error checksum clearance (ECS) mode is used to read the memory location through the ECC logic in the memory device, perform error checksum correction, and count the error information, and store it in the register for use by the host system.
Improve the reliability and accessibility of the memory system. Through transparent error information transmission, the host system can effectively improve system-level error correction and reduce the impact of error accumulation in the memory device.
Smart Images

Figure CN113808658B_ABST
Abstract
Description
[0001] Related Applications
[0002] This patent application is a non - provisional application based on U.S. Provisional Application No. 62 / 211,448, filed on August 28, 2015. This application claims the benefit of the priority of that provisional application. That provisional application is hereby incorporated by reference.
[0003] This patent application relates to the following two patent applications, which also claim priority to the same U.S. Provisional Application identified above: Patent Application No. TBD [P88609] entitled "MEMORY DEVICE CHECKBIT READ MODE"; and Patent Application No. TBD [P93260] entitled "MEMORY DEVICE ON - DIE ECC (ERROR CHECKING AND CORRECTING) CODE"; these two applications are filed simultaneously herewith. TECHNICAL FIELD
[0004] The description generally relates to memory management, and more particularly, to error checking and correction in a memory subsystem having a memory device that performs internal error checking and correction.
[0005] COPYRIGHT NOTICE / LICENSE
[0006] The publicly disclosed portion of this patent document may contain copyrighted material. The copyright owner does not oppose the reproduction by anyone of the patent document or the patent disclosure when it appears in the patent and trademark office patent file or records, but otherwise reserves all copyright rights whatsoever. The copyright notice applies to all data described below, and to any software described herein and in the figures hereof: 2015, Intel Corporation. All rights reserved. BACKGROUND OF THE INVENTION
[0007] Volatile memory resources find wide use in current computing platforms, whether for servers, desktop or laptop computers, mobile devices, and consumer and business electronic devices. DRAM (Dynamic Random Access Memory) devices are the most common type of memory devices in use. As the manufacturing processes for producing DRAM continue to scale down to smaller geometries, DRAM errors are expected to increase. One technique for addressing the increased DRAM errors is to employ on-die ECC (Error Checking and Correction). On-die ECC refers to error detection and correction logic that resides on the memory device itself. With on-die ECC logic, DRAM can correct single-bit faults (such as through Single Error Correction (SEC)). On-die ECC can be used in addition to system-level ECC, but system-level ECC has no insight into what error corrections have been performed at the memory device level. Thus, while on-die ECC can handle errors within the memory device, errors can accumulate without being detected by the host system. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The following description includes a discussion of figures having diagrams that are presented as examples of implementations of the present invention. The drawings should be understood as examples and not as limitations. As used herein, reference to one or more "embodiments" is to be understood as describing specific features, structures, and / or characteristics included in at least one implementation of the present invention. Thus, phrases such as "in one embodiment" or "in an alternative embodiment" that appear herein describe various embodiments and implementations of the present invention and do not necessarily all refer to the same embodiment. However, they are not necessarily mutually exclusive either.
[0009] Figure 1 is a block diagram of an embodiment of a system in which a memory device monitors errors in an error checking mode.
[0010] Figure 2 is a block diagram of an embodiment of a system that performs internal error correction and stores error information.
[0011] Figure 3 is a block diagram of an embodiment of a system in which a register stores multiple rows with errors and the maximum error for any row.
[0012] Figure 4A is a block diagram of an embodiment of a command encoding capable of implementing an Error Check and Scrub (ECS) mode.
[0013] Figure 4B is a block diagram of an embodiment of a mode register capable of implementing an Error Check and Scrub (ECS) mode.
[0014] Figure 4C is a block diagram of an embodiment of a multi-purpose register for storing a row address with a maximum error count.
[0015] Figure 4D A block diagram of an embodiment of a general - purpose register for storing a count of multiple lines containing errors.
[0016] Figure 5 A block diagram of an embodiment of logic at a memory device that generates error - correction information and supports an error checksum clear mode.
[0017] Figure 6 A flowchart of an embodiment of a process for monitoring one or more error counts via error - correction operations in an error checksum clear (ECS) mode.
[0018] Figure 7 A block diagram of an embodiment of a computing system in which an error checksum clear mode with error tracking can be implemented.
[0019] Figure 8 A block diagram of an embodiment of a mobile device in which an error checksum clear mode with error tracking can be implemented.
[0020] The following is a description of certain details and implementations, which includes a description of figures that may depict some or all of the embodiments described below, as well as discussion of other potential embodiments or implementations of the inventive concepts presented herein. Detailed Description
[0021] As described herein, a memory - device mode can implement error monitoring. The error - monitoring mode can be referred to by any label. The error - monitoring mode described herein enables the performance of error checksum and correction (ECC) and counting the total number of memory segments with errors (and the segments with the highest number of errors). Generally, a segment or portion will be a row of a memory device, where each data block (e.g., data of a prefetch size) within that row can be tested for errors. ECC, as used herein, refers to any process and / or operation or set of operations for validating the data stored based on error checksum and correcting the data, and performing some form of error - correction routine based on that operation or process.
[0022] In one embodiment, the error monitoring mode is an Error Check and Clear (ECS) mode that enables a DRAM (Dynamic Random Access Memory device) to perform one or more ECC operations and count errors. Additionally, such error monitoring modes can be referred to by any name, but for simplicity, the ECS mode is used as an example herein and is not restrictive. A memory controller associated with the memory device triggers the ECS mode with a trigger sent to the memory device. Thus, the host can control the mode and receive error count information. The memory device includes a plurality of addressable memory locations that can be organized into segments (such as word lines or rows or other portions). The memory locations store data and have associated ECC information. In the ECS mode, the memory device reads one or more memory locations and performs ECC on the one or more memory locations based on the ECC information stored within the memory device. Thus, the memory device performs internal ECC in the ECS mode. The memory device counts error information, which includes a segment count indicating the number of segments with at least a threshold number of errors and a maximum count indicating the maximum number of errors in any segment.
[0023] References to memory devices can apply to different memory types. A memory device generally refers to volatile memory technology. Volatile memory is memory whose state (and thus the data stored thereon) is indeterminate if power is interrupted to the device. Non-volatile memory refers to memory whose state is determined even if power is interrupted to the device. Dynamic volatile memory requires refreshing the data stored in the device to maintain the state. An example of dynamic volatile memory includes DRAM (Dynamic Random Access Memory), or some variant such as synchronous DRAM (SDRAM). The memory subsystem described herein can be compatible with multiple memory technologies such as DDR3 (Double Data Rate version 3, first edition made by JEDEC (Joint Electron Device Engineering Council) on June 27, 2007, currently at edition 21), DDR4 (DDR version 4, initial specification published by JEDEC in September 2012), DDR4E (DDR version 4 extension, currently under discussion by JEDEC), LPDDR3 (Low Power DDR version 3, JESD209-3B, by JEDEC in August 2013), LPDDR4 (Low Power Double Data Rate (LPDDR) version 4, JESD209-4, initially published by JEDEC in August 2014), WIO2 (Wide I / O 2 (WideIO2), JESD229-2, initially published by JEDEC in August 2014), HBM (High Bandwidth Memory DRAM, JESD235, initially published by JEDEC in October 2013), DDR5 (DDR version 5, currently under discussion by JEDEC), LPDDR5 (currently under discussion by JEDEC), HBM2 (HBM version 2, currently under discussion by JEDEC) and / or other memory technologies, and technologies derived or extended based on such specifications.
[0024] In addition to or as an alternative to volatile memory, in one embodiment, a reference to a memory device can refer to a non-volatile memory device whose state is determined even if power is interrupted to the device. In one embodiment, the non-volatile memory device is a block-addressable memory device such as NAND or NOR technology. Thus, the memory device can also include future generations of non-volatile devices such as three-dimensional cross-point memory devices, other byte-addressable non-volatile memory devices, or memory devices using chalcogenide phase change materials (e.g., chalcogenide glass). In one embodiment, the memory device can be or include multi-threshold level NAND flash memory, NOR flash memory, single-level or multi-level phase change memory (PCM), resistive memory, nanowire memory, ferroelectric transistor random access memory (FeTRAM), magnetoresistive random access memory (MRAM) memory combined with memristor technology or spin transfer torque (STT)-MRAM, or any combination of the above or other memories.
[0025] The present description that refers to "DRAM" can apply to any memory device that allows random access, whether volatile or non-volatile. The memory device or DRAM can refer to the die itself and / or the packaged memory product.
[0026] However, in a traditional memory subsystem, due to on-die or internal ECC, errors can accumulate undetected within the DRAM. As described herein, the memory device can count the errors and expose the error information to the host system. On-die or internal ECC within the memory device generally refers to error checking and correction of single-bit errors (SBEs) in the memory array. In one embodiment, a DRAM compatible with DDR4E or a variant or extension can apply ECS combined with error counting or monitoring. Thus, an embodiment of a DDR4E DRAM can clear errors and track the error accumulation of the device. In one embodiment, in the case of using ECS, the memory device maintains an error count indicating the number of rows with at least one error and tracks the row address with the highest number of errors.
[0027] In one embodiment, a DDR4E device incorporates on-die ECC and includes a row or word line having a plurality of addressable locations with 128-bit data blocks. The data blocks can have associated ECC bits (e.g., 8-bit internal error correction data). The ECS mode can provide transparent support for such DDR4E devices. In one embodiment, the ECS mode includes an error checking and clearing operation that combines an error counting mechanism as part of the ECC operation. The ECS mode enables the DRAM to internally read, correct the SBE, and write the corrected data back to the array. Such read, correction, and write-back can be referred to as "clearing" the error.
[0028] In one embodiment, a memory device supporting the ECS mode includes one or more registers for storing error count information. For example, the register can be a DRAM mode register. The register location can include a multi-purpose register (MPR), where the memory device can store error count information including the number of segments having errors and the maximum error count and / or the address of the segment having the maximum error count. In one embodiment, the memory device uses two registers in the ECS mode to track the codeword and parity bit errors detected during ECS mode operation. In one embodiment, the memory device stores the value from the row error counter into one register and stores the value from each row error counter into another register. In one embodiment, the row error counter tracks the number of rows having at least a threshold number of detected codeword and parity bit errors. The threshold number can be one error. The threshold number can be two errors or some other number of errors. In one embodiment, each row error counter tracks the address of the row having the maximum number of codeword and parity bit errors and can include the codeword and parity bit error counts for that row.
[0029] In one embodiment, the memory controller reads the error count information stored in the register by the memory device. The host system (e.g., the host operating system and / or the host CPU (central processing unit)) can utilize the error count information to improve the system-level RAS (reliability, availability, and serviceability) of the memory subsystem. Thus, by providing transparency to the error count information on the memory device through the memory subsystem, the memory devices themselves do not need to expose specific error information but can provide sufficient information to improve error correction at the system level. In one embodiment, the memory controller can extract multi-bit error information from the error information and apply the multi-bit information to determine how to apply ECC to the system (e.g., by knowing where the error occurs in the memory). In one embodiment, the memory controller uses the error information from the memory device as metadata for improving SDDC (single device data correction) ECC operations for multi-bit errors.
[0030] Figure 1 FIG. 7 is a block diagram of an embodiment of a system in which a memory device monitors errors in an error checking mode. System 100 includes elements of a memory subsystem in a computing device. Processor 110 represents the processing unit of a host computing platform that executes an operating system (OS) and applications, which can be collectively referred to as the "host" of the memory. The OS and applications perform operations that result in memory access. Processor 110 can include one or more individual processors. Each individual processor can include a single and / or multi-core processing unit. The processing unit can be a main processor such as a CPU (central processing unit) and / or a peripheral processor such as a GPU (graphics processing unit). System 100 can be implemented as a SOC or with discrete components.
[0031] Memory controller 120 represents one or more memory controller circuits or devices of system 100. Memory controller 120 represents control logic that generates memory access commands in response to the execution of operations by processor 110. Memory controller 120 accesses one or more memory devices 140. Memory device 140 can be DRAM according to any of the memory devices mentioned above. In one embodiment, memory device 140 is organized and managed as different channels, where each channel is coupled to a bus and signal lines, and the bus and signal lines are coupled in parallel to multiple memory devices. Each channel is independently operable. Thus, each channel is independently accessed and controlled, and for each channel, timing, data transfer, command and address exchange, and other operations are separate. In one embodiment, the settings for each channel are controlled by individual mode registers or other register settings. In one embodiment, each memory controller 120 manages a separate memory channel, although system 100 can be configured such that multiple channels are managed by a single controller, or there are multiple controllers on a single channel. In one embodiment, memory controller 120 is part of host processor 110, such as logic implemented on the same die or in the same package space as the processor.
[0032] Memory controller 120 includes I / O interface logic 122 for coupling to the system bus. I / O interface logic 122 (and I / O 142 of memory device 140) can include pins, connectors, signal lines, and / or other hardware for connecting to the device. I / O interface logic 122 can include a hardware interface. As shown, I / O interface logic 122 includes at least drivers / transceivers for signal lines. Generally, the wires within an integrated circuit interface with bond pads or connectors to interface to signal lines or traces between devices. I / O interface logic 122 can include drivers, receivers, transceivers, terminations, and / or other circuitry for transmitting and / or receiving signals on the signal lines between devices. The system bus can be implemented as multiple signal lines that couple memory controller 120 to memory device 140. The system bus includes at least a clock (CLK) 132, command / address (CMD) 134, data (DQ) 136, and other signal lines 138. The signal lines for CMD 134 can be referred to as the "C / A bus" (or ADD / CMD bus, or some other naming indicating the transfer of address information and commands), and the signal lines for DQ 136 can be referred to as the "data bus". In one embodiment, independent channels have different clock signals, C / A buses, data buses, and other signal lines. Thus, system 100 can be considered to have multiple "system buses" in the sense that independent interface paths can be considered separate system buses. It will be understood that in addition to the lines explicitly shown, the system bus can include strobe signaling lines, warning lines, auxiliary lines, and other signal lines.
[0033] It will be understood that the system bus includes a data bus (DQ 136) configured to operate at a certain bandwidth. Based on the design and / or implementation of system 100, DQ 136 can have more or less bandwidth per memory device 140. For example, DQ 136 can support memory devices with x32 interfaces, x16 interfaces, x8 interfaces, or other interfaces. The convention “xN” (where N is a binary integer) refers to the interface size of memory device 140, which represents the number of data lines DQ 136 that exchange data with memory controller 120. The interface size of a memory device is a control factor regarding how many memory devices can be used simultaneously per channel in system 100 or be coupled in parallel to the same data lines.
[0034] Memory device 140 represents a memory resource for system 100. In one embodiment, each memory device 140 is a separate memory die, which can include multiple (e.g., 2) channels per die. Each memory device 140 includes I / O interface logic 142, which has a bandwidth determined by the device implementation (e.g., x16 or x8 or some other interface bandwidth) and enables the memory device to interface with memory controller 120. I / O interface logic 142 can include a hardware interface and can be in accordance with the I / O 122 of the memory controller, but at the end of the memory device. In one embodiment, multiple memory devices 140 are connected in parallel to the same data bus. For example, system 100 can be configured with multiple memory devices 140 coupled in parallel, where each memory device responds to commands and accesses its respective internal memory resource 160. For a write operation, a single memory device 140 can write a portion of an overall data word, and for a read operation, a single memory device 140 can fetch a portion of an overall data word.
[0035] In one embodiment, memory device 140 is directly disposed on the motherboard of a computing device or a host system platform (e.g., a PCB (printed circuit board) on which processor 110 is disposed). In one embodiment, memory device 140 can be organized into memory module 130. In one embodiment, memory module 130 represents a dual in-line memory module (DIMM). In one embodiment, memory module 130 represents other organizations of multiple memory devices that share access or control circuitry for at least a portion thereof, which can be a separate circuit, a separate device, or a separate board from the host system platform. Memory module 130 can include multiple memory devices 140, and the memory module can include support for multiple separate channels to the included memory devices disposed thereon.
[0036] Each memory device 140 includes memory resources 160. The memory resources 160 represent respective arrays of memory locations or storage locations for data. Typically, the memory resources 160 are managed as data rows and are accessed via cache lines (rows) and bit lines (individual bits within a row). The memory resources 160 can be organized into separate channels, ranks, and banks of the memory. A channel is an independent control path to storage locations within the memory device 140. A rank refers to common locations across multiple memory devices (e.g., the same row address within different devices). A bank refers to an array of memory locations within the memory device 140. In one embodiment, the banks of the memory are divided into sub - banks, with at least a portion of shared circuitry for the sub - banks.
[0037] In one embodiment, the memory device 140 includes one or more registers 144. The registers 144 represent storage devices or storage locations that provide settings or configurations for the operation of the memory device. In one embodiment, as part of a control or management operation, the registers 144 can provide storage locations for the memory device 140 to store data for access by the memory controller 120. In one embodiment, the registers 144 include mode registers. In one embodiment, the registers 144 include multi - purpose registers. The location configurations within the registers 144 can configure the memory device 140 to operate in different "modes", where command and / or address information or signal lines can trigger different operations within the memory device 140 depending on the mode. The settings of the registers 144 can indicate configurations for I / O settings (e.g., timing, termination, or ODT (on - die termination), driver configuration, and / or other I / O settings).
[0038] In one embodiment, the memory device 140 includes an ODT 146 as part of the interface hardware associated with the I / O 142. The ODT 146 can be configured as described above and provides settings for the impedance to be applied to specific signal lines. The ODT settings can be changed based on whether the memory device is the selected target or a non - target device of an access operation. The ODT 146 settings can affect the timing and reflection of signaling on the termination lines. Careful control of the ODT 146 enables higher - speed operation with improved matching of the applied impedance and loading.
[0039] The memory device 140 includes a controller 150, which represents control logic within the memory device for controlling internal operations within the memory. For example, the controller 150 decodes commands sent by the memory controller 120 and generates internal operations to execute or satisfy the commands. The controller 150 can be referred to as an internal controller. The controller 150 is capable of determining what mode to select (based on the register 144) and configuring the access and / or execution of operations for the memory resources 160 based on the selected mode. The controller 150 generates control signals for controlling the routing of bits within the memory device 140 to provide an appropriate interface for the selected mode and directs commands to the appropriate memory location or address.
[0040] Referring again to the memory controller 120, the memory controller 120 includes command (CMD) logic 124, which represents the logic or circuitry for generating commands sent to the memory device 140. Generally, signaling in the memory subsystem includes address information within or accompanying the command to indicate or select one or more memory locations where the memory device should execute the command. In one embodiment, the controller 150 of the memory device 140 includes command logic 152 for receiving and decoding the command and address information received from the memory controller 120 via the I / O 142. Based on the received command and address information, the controller 150 can control the operation timing of the logic and circuitry within the memory device 140 for executing the command. The controller 150 is responsible for complying with standards or specifications.
[0041] In one embodiment, the memory controller 120 includes refresh (REF) logic 126. The refresh logic 126 can be used in cases where the memory device 140 is volatile and needs to be refreshed to maintain a determined state. In one embodiment, the refresh logic 126 indicates the location of the refresh and the type of refresh to be performed. The refresh logic 126 can trigger self-refresh within the memory device 140 and / or perform external refresh by sending a refresh command. For example, in one embodiment, the system 100 supports all bank refreshes and per-bank refreshes, or other all-bank commands and per-bank commands. All-bank commands cause the operation of the selected bank in all memory devices 140 coupled in parallel. Per-bank commands cause the operation of a specific bank in a specific memory device 140. In one embodiment, the controller 150 within the memory device 140 includes refresh logic 154 for applying refresh within the memory device 140. In one embodiment, the refresh logic 154 generates internal operations for performing refresh in accordance with an external refresh received from the memory controller 120. The refresh logic 154 can determine whether the refresh is directed to the memory device 140 and what memory resources 160 are to be refreshed in response to the command.
[0042] In one embodiment, memory controller 120 includes error correction and control logic for performing system-level ECC of system 100. System-level ECC refers to applying error correction in memory controller 120 and being able to apply error correction to data bits from multiple different memory devices 140. In one embodiment, memory controller 120 includes ECS 170, which represents circuitry or logic for implementing the ECS mode in one or more memory devices 140. In one embodiment, ECS 170 includes logic for setting mode register 144 of memory device 140 to trigger the ECS mode. In one embodiment, ECS 170 includes logic for encoding and sending commands to trigger the ECS mode in one or more memory devices 140. In one embodiment, ECS 170 includes logic for reading error information from memory device 140.
[0043] In one embodiment, controller 150 includes ECS logic 156, which represents logic in memory device 140 for entering the ECS mode and performing one or more ECS operations. ECS logic 156 can be regarded as including on-die or on-chip ECC logic for performing ECC operations. In one embodiment, ECS 156 reads the settings of register 144 to determine whether the ECS mode can be implemented. In one embodiment, ECS 156 determines that an ECS mode trigger is received via command logic decoding of logic 152. In one embodiment, ECS 170 of memory controller 120 sets one or more bits of register 144 to reset the error information count.
[0044] Typically, DRAM has an on-die or internal oscillator (not shown explicitly) for controlling internal operation timing. For a small number of operations, the difference in timing between host-controlled system timing and the internal oscillator can be synchronized quite easily. For a longer series of instructions, the timing offset between memory device 140 and the host can create synchronization problems. In one embodiment, memory controller 120 issues a series of ECS operation commands for memory device 140 to execute in the ECS mode. In one embodiment, memory controller 120 simply places memory device 140 in the ECS mode and allows controller 150 to control operations internally. In the case of a series of external operations from memory controller 120, controller 150 can generate internal commands from the external commands to perform ECC in sequence through the memory locations of memory resource 160. In the case where memory controller places memory device 140 in the ECS mode, controller 150 can generate internal commands to pass through the memory resource in sequence. In one embodiment, in either case, controller 150 controls at least a part of the generation of the memory location addresses for ECS operations.
[0045] Figure 2 It is a block diagram of an embodiment of a system that performs internal error correction and stores error information. System 200 represents components of a memory subsystem. System 200 provides an example of a memory subsystem according to an embodiment of System 100. System 200 can be included in any type of computing device or electronic circuit that uses memory with internal ECC, where the memory device counts error information. Processor 210 represents any type of processing logic or component that performs operations based on data stored in or to be stored in Memory 230. Processor 210 can be or include a host processor, a central processing unit (CPU), a microcontroller or microprocessor, a graphics processor, a peripheral processor, an application-specific processor, or other processors. Processor 210 can be or include a single-core or multi-core circuit. Figure 1 Memory controller 220 represents the logic for interfacing with Memory 230 and managing access to data in Memory 230. Similar to the memory controller above, Memory controller 220 can be separate from or part of Processor 210. From the perspective of Memory 230, Processor 210 and Memory controller 220 together can be regarded as the "host", and Memory 230 stores data for the host. In one embodiment, Memory 230 includes DDR4E DRAM with internal ECC (which may be referred to as DDR4E in the industry). In one embodiment, System 200 includes multiple memory resources 230. Memory 230 can be implemented in System 200 using any type of architecture that supports access via internal ECC in the memory through Memory controller 220. Memory controller 220 includes I / O (Input / Output) 222, which includes hardware resources for interconnecting with the corresponding I / O 232 of Memory 230.
[0046] Memory 230 includes command execution 234, which represents the control logic within the memory device for receiving and executing commands from Memory controller 220. The commands can include a series of ECC operations for the memory device to perform in the ECS mode to record error count information. In one embodiment, mode register 238 includes one or more multi-purpose registers for storing error count information. In one embodiment, mode register 238 includes one or more fields that can be set by Memory controller 220 to enable resetting of error count information.
[0047]
[0048] Memory 230 includes an array 240, which represents an array of memory locations where data is stored in the memory device. In one embodiment, each address location 244 of array 240 contains associated data and ECC bits. In one embodiment, address location 244 represents an addressable data block, such as a 128-bit block, a 64-bit block, or a 256-bit block. In one embodiment, address locations 244 are organized as segments or groups of memory locations. For example, as shown, memory 230 includes multiple rows 242. In one embodiment, each row 242 is a part or segment of the memory that is checked for errors. In one embodiment, rows 242 correspond to memory pages or word lines. Array 240 includes N rows 242, and rows 242 include M memory locations.
[0049] In one embodiment, address location 244 corresponds to a memory word, and row 242 corresponds to a memory page. A page of memory refers to a certain granularity measure of memory space allocated for a memory access operation. In one embodiment, array 240 has a larger page size to accommodate ECC bits in addition to data bits. Thus, the normal page size will include enough space allocated for data bits, and array 240 allocates enough space for data bits plus ECC bits.
[0050] In one embodiment, the memory controller includes an ECS controller 226 for managing the ECS mode for ECC operations and the error count in memory 230. In one embodiment, memory 230 includes internal ECC managed by ECS control 250. Memory controller 220 manages system-wide ECC and can detect and correct errors across multiple different parallel memory resources (e.g., multiple memory resources 230). Many techniques for system-wide ECC are known and can include managing memory resources in a way that distributes errors across multiple parallel resources. By distributing errors across multiple resources, memory controller 220 can recover data even if there is one or more failures in memory 230. Memory failures are generally classified as software errors or software faults (which are usually transient bit errors resulting from random environmental conditions) or hardware errors or hardware faults (which are non-transient bit errors that occur as a result of hardware failures).
[0051] In one embodiment, the ECS control 250 includes a count 252 of row 242 including the SBE. In one embodiment, the ECS control 250 includes one or more counters. For example, the SBE count 252 can be one of the counters. The ECS control 250 includes ECC logic (not specifically shown) for performing error checking and correction. When an error is detected in row 242, the ECS control 250 increments the SBE count 252. In one embodiment, the ECS control 250 includes a maximum count and / or a maximum address 254. In one embodiment, the maximum count can be maintained in another counter. In one embodiment, the maximum address refers to the address of the segment or row 242 determined to have the highest number of errors.
[0052] In one embodiment, the memory 230 is part of a bank of memory resources. In one embodiment, the memory includes multiple banks of memory resources, and each bank can be accessed individually. In one embodiment, all banks must be precharged and be in an idle state before the ECS mode can be enabled. In one embodiment, the command execution 234 identifies a specific command sequence for the ECS mode. In one embodiment, only a specific command sequence is permitted in the ECS mode. Although different implementations can vary, an example of an allowed command sequence can be as follows: ECS-->DES-->ACT-->DES-->WR-->DES-->PRE-->DES. It will be understood that NOP can be used in place of the deselect command (DES), but the NOP command may require toggling the chip select (CS) bit and requires the memory to decode the command, while the DES command allows the memory to simply remain idle for the cycle.
[0053] The command sequence can be described as an ECS command trigger (which may not be necessary if the ECS mode is triggered via a mode register set), followed by deselect (DES), activate (ACT), another DES, write (WR), another DES, precharge (PRE), and another DES. In one embodiment, the ECS command places the memory 230 in the ECS mode. In one embodiment, the ECS command is followed by an ACT command tMOD; the ACT command is followed by a WR command tRCD; and the WR command is followed by a PRE command WL + tWR + 10CK (assuming tECSc is satisfied). In one embodiment, the minimum time for an ECS mode cycle is tECSc (which can be the larger of 45 CK or 110 ns). It will be understood that tMOD can refer to the timing delay for a mode register set command update, tRCD can refer to the timing delay from the ACT command to an internal read or write, WL can refer to the write wait time, tWR can refer to the write recovery time, and CK can refer to a clock cycle.
[0054] In one embodiment, when in the ECS mode, the memory 230 ignores the data I / O and address inputs of the I / O 232 for a period of tMOD after the ECS command is issued. Ignoring the address input can include ignoring the bank address (BA) value and the bank group (BG) value. In one embodiment, the I / O control for the I / O 232 sets the data I / O and address inputs to a tri-state operation.
[0055] In one embodiment, for the ECS-->ACT-->WR sequence (ignoring the intervening DES commands), the ECS command will be able to implement the ECS mode, and the ACT command will perform an internal row activation. Row activation can occur for the row determined by the internal ECS address counter (not specifically shown) within the ECS control 250. In one embodiment, the WR command will perform an internal read-modify-write cycle for the activated row, where the column address is determined by the internal ECS address counter within the ECS control 250. It will be understood that the ECS control 250 can be part of an internal controller (such as the controller 150 of the system 100). Thus, the counters and controls within the ECS control 250 can be part of the logic within the internal controller. Alternatively, separate ECS control logic, such as a separate logic circuit or microcontroller (used to implement ECS operations).
[0056] In one embodiment, the internal read and write cycles or read-modify-write cycles read the entire codeword and parity bits (e.g., 128 data bits and 8 parity bits) from the array 240, correct the SBE in the codeword or parity bits, and write the resulting codeword and parity bits back to the appropriate row 242 in the array 240. In one embodiment, the PRE command exits the ECS mode and returns the memory 230 to the normal mode.
[0057] As mentioned above, in one embodiment, the ECS control 250 includes logic for internally generating address information for performing ECS operations. In one embodiment, the ECS mode sequentially passes through an entire memory region, such as a bank group. In one embodiment, the ECS control 250 keeps track of the addresses for sequentially passing through the memory 230. In one embodiment, the memory controller 220 keeps track of certain address information, and the memory 230 keeps track of other address information. For example, consider the following implementation where the memory controller 220 identifies the bank group and the memory 230 keeps track of the row and column addresses within the bank group. Alternatively, consider that the memory controller 220 identifies the bank group and columns, and the memory 230 identifies the rows. Many other implementations are possible. The implementation will depend on system design, such as whether the address locations are coordinated on the memory 230 and the memory controller cannot efficiently sequentially traverse the check space. In one embodiment, the memory controller is responsible for providing sufficient ECC operation commands, and the memory device performs all addressing internally. In one embodiment, an ECS command causes an ECC operation on a single memory location, and the memory controller sends a sequence of ECS commands to perform ECC operations on many locations.
[0058] Regardless of whether the memory 230 or the memory controller 220 identifies the addressing, and regardless of whether the ECS command triggers an ECC operation on one or more memory locations, the memory 230 counts the number of errors during ECS mode operation. As described, the error count information can include the total number of rows or other segments with errors, and the highest error count for any segment.
[0059] Figure 3 is a block diagram of an embodiment of a system in which registers store the number of rows with errors and the maximum error for any row. System 300 illustrates the components of an ECC system that maintains error information counts and implements an ECS example according to Figure 1 system 100 and / or Figure 2 system 200. System 300 provides a representation of the components for performing ECS mode operations. In one embodiment, the register storing the final count is a host-accessible register, and the other components are within the controller on the memory device.
[0060] Address control 320 represents control logic used to enable a memory device to manage internal addressing for ECS operations. As previously described, the memory device may be responsible for certain address information and the memory controller for other address information. Thus, system 300 can have different implementations depending on the address information responsibilities of the memory device. In one embodiment, address control 320 includes one or more ECS address counters for incrementing the address used for ECS commands. Address control 320 can include counters for column address, row address, bank address, bank group address, and / or combinations thereof. In one embodiment, address control 320 includes a counter for incrementing the column address for a row or other segment used for ECS operations.
[0061] In one embodiment, address control 320 increments the column address on each ECS WR command, which can be triggered or indicated by performing a logical AND operation on the internal ECS command signal (ECS CMD) and the internal write command indicated by the internal column address strobe for write (CAS WR), as shown by logic 314. Internal signals refer to signals internally generated by the internal controller within the memory device, rather than commands specified to be received from an external memory controller. Internal commands are generated in response to external commands from the host. In one embodiment, once the column counter wraps or rolls over, the row counter will increment sequentially and each codeword and associated parity bit in the next row will be read until all rows within the bank have been accessed. In one embodiment, once the row counter wraps, the bank counter will increment sequentially and the next sequential bank within the bank group will repeat the process of accessing each codeword and associated parity bit until all banks and bank groups within the memory have been accessed.
[0062] In one embodiment, the column, row, bank, bank group, and / or other counters used to generate the internal address in address control 320 can be reset in response to a reset condition for the memory subsystem and / or in response to a bit in the mode register (MRx Ay bit), as illustrated by logic 312. The reset can occur during a power cycle or other system restart. In one embodiment, the address information generated by address control 320 can be used to control decoding logic (e.g., row and column decoder circuits within the memory device) to trigger ECS operations. However, in view that for other operations, the memory device operates on external address information received from the host, in one embodiment, system 300 includes a multiplexer 370 for selecting between an external address 372 and internal address information (from address control 320). Although not specifically shown, it will be understood that the control for multiplexer 370 can be the ECS mode state, where in the ECS mode, internal address generation is selected to access memory locations in the memory array, and when not in the ECS mode, the external address 372 is selected to access memory locations in the memory array.
[0063] When the memory controller provides some address information, the selection logic of system 300 can be selectively more complex. Based on internal address information and / or external address information, system 300 can verify all codewords and verify bits within the memory. More specifically, system 300 can perform ECS operations that include ECC operations to read, correct, and write memory locations for all selected or specified addresses. After operating on system 300 once, system 300 can be configured (e.g., via the size of a counter) to wrap the counter and restart the address count. For example, once all memory locations of the memory have been verified, the bank group counter can wrap and the process can start again with the next ECS command. The total number of ECS commands required to complete one cycle of the ECS mode is density-dependent. Table 1 provides an example of a list of the number of commands required based on the number of codewords for various memory configurations. In one embodiment, the DRAM controller keeps track of the number of ECS commands issued to the DRAM to perform error checking on all codewords and parity bits in the device.
[0064]
[0065] Table 1: Number of Codewords per Bank Group (e.g., for 128-bit CW)
[0066] For an 8Gb device with an x4, x8, or x16 interface, the number of 128-bit codewords is 67,108,864, which will require an equal number of command cycles to pass through all memory locations. For a 12Gb device with an x4, x8, or x16 interface, the number of 128-bit codewords is 100,663,296. For a 16Gb device with an x4, x8, or x16 interface, the number of 128-bit codewords is 134,217,728.
[0067] System 300 includes ECC correction logic 330 for performing error checking and correction operations on selected or addressed memory locations. In one embodiment, when an error is detected at a memory location, logic 330 generates an output flag. Logic 330 can generate a flag in response to detecting and correcting a single-bit error (SBE) with an ECC operation. In one embodiment, an error detection signal from logic 330 triggers the ERC (Error Row Counter) 332 and the per-row error counter 344 to increment. In one embodiment, the ERC 332 has a threshold number of memory locations (identified in some places as threshold N) in a specific row or segment that must be triggered as having an error before counting the row as having an error. For example, the threshold can be 1 error, and all rows with at least one error are counted. As another example, the threshold can be 2 errors, and rows with only one error are not counted, but all rows with at least 2 errors are counted. It will be understood that the threshold can be set based on system configuration considerations. Thus, the definition of what is a "bad row" can be set by the threshold for the ERC 332. If the number of errors is 1, the ERC 332 may be unnecessary, and the output of the ECC logic 330 can be directly fed to the row error counter 342.
[0068] As shown, when the threshold number of errors per row has been reached as determined by the ERC 332, the ERC 332 outputs a signal to cause the row error counter 342 to increment its count. Each error detected by logic 330 increments the per-row error counter 344. It will be understood that the ERC 332 and the per-row error counter 344 generate information on a per-row or per-segment basis; thus, when the row address rolls over, the address control 320 can trigger the ERC 332 and the per-row error counter 344 to reset their counts.
[0069] Thus, in one embodiment, the per-row error counter 344 increments whenever a codeword or parity bit error is detected and is reset whenever the row counter wraps or rolls over. In one embodiment, if one or more codeword or parity bit errors are detected during an ECS WR command on a given row, the row error counter 342 increments by 1. In one embodiment, the row error counter 342 increments by 1 for every two errors (or other configured number of errors) detected in the row.
[0070] The per-line error counter 344 keeps track of the total number of codeword or parity bit errors on a given line. In one embodiment, the counter 344 provides its count for comparison with a previous high or maximum error count. The high error count 352 represents a register or other storage device that holds the high error count (which indicates the maximum number of errors detected in any line). The comparison logic 350 determines whether the count value in the per-line error counter 344 is greater than the high error count 352. In one embodiment, if the current error count in counter 344 is greater than the previous maximum of the high error count 352, the logic 350 generates an output that triggers the high error address 354 to be set to the current line address. Additionally, the high error count 352 is set to the count from the per-line error counter 344. In one embodiment, the per-line codeword or parity bit error count is compared with the previous per-line codeword or parity bit error count to determine the line address within the DRAM that has the highest error count.
[0071] In one embodiment, after performing the ECS operation for all lines in the memory space identified for verification, the address control 320 triggers a signal indicating that the last address has been reached or rolled over. In one embodiment, the signal triggers the row error register 362 to load the count from the row error counter 342, and the row error counter 342 is reset. In one embodiment, the system 300 includes a delay in propagating the signal to ensure that the register 362 latches the count from the counter 342 before the counter is reset. In one embodiment, the signal indicating that the last address has been reached also triggers the per-line error register 364 to load the count from the high error count 352. In one embodiment, the signal indicating that the last address has been reached also triggers the per-line error register 364 to load the address from the high error address 354. In one embodiment, the signal for the register 364 to load the address and / or count also triggers the high error count 352 to be reset. In one embodiment, the row error register 362 represents the MPR page 5 of the DRAM's mode register. In one embodiment, the per-line error register 364 represents the MPR page 4 of the DRAM's mode register.
[0072] Figure 4AFIG. is a block diagram of an embodiment of a command encoding that enables an Error Check and Scrub (ECS) mode. Command encoding 410 represents an example of an ECS command encoding for any embodiment of an ECS that triggers the ECS mode with the command encoding. ECS command 412 illustrates an example of the encoding where the clock enable bit and ACT_n, CAS_n, and WE_n are set high. Chip select CS_n together with RAS_n is set low. In one embodiment, the bank address information (BA and BG) is set low, and the row address bits A11:A0 remain high. In one embodiment, A12 is set low. In such an encoding, the DRAM will internally generate an address for the ECS operation. Alternatively, command 412 can contain some address information.
[0073] Figure 4B FIG. is a block diagram of an embodiment of a mode register that enables an Error Check and Scrub (ECS) mode. Mode register 420 represents an example of a mode register in a DRAM that supports the ECS mode. In one embodiment, the register illustrated as mode register 420 can be two separate mode registers. As shown, address 424 (Az) represents the bit that triggers the ECS mode for any embodiment of an ECS that triggers the ECS mode with the mode register. Address 422 (Ay) represents the bit that triggers the reset of the ECS counter. In an implementation where the command encoding triggers the ECS mode, address 424 may not be used, while address 422 can still provide the ability to reset the ECS counter. In one embodiment, a "1" written to address 424 can trigger the ECS mode to cause the memory device to perform an ECC operation and count errors, while a "0" can exit the ECS mode.
[0074] In one embodiment, the ECS mode requires two additional mode register bits. The first bit is address 422, which enables clearing the counter and the error result register. The encoding can be inverted, but as shown, a "1" can clear the counter and the result register, and a "0" can initialize the register and the counter. In one embodiment, Ay must be written to "0", after which a subsequent 1 can be applied to clear the counter and the register. The second bit is not specifically shown and is a mode register bit that enables the result register. For example, in one embodiment, the mode register bit can enable MPR pages 4 to 7, where the results are stored in MPR 4 and 5 (pages 4 and 5).
[0075] Figure 4CA block diagram of an embodiment of a multi-purpose register for storing the row address with the maximum error count. The MPR page 430 represents the register storing the row address with the maximum number of errors. For example, the MPR page 430 can be the MPR page 4 in a DDR4E implementation. In one embodiment, the addresses BA1:BA0 provide the MPR location. For example, "00" can correspond to MPR0, "01" can correspond to MPR1, "10" can correspond to MPR2, and "11" can correspond to MPR3. In one embodiment, the bits of the MPR page 430 are assigned as follows: A[17:0] is the row address, BA[1:0] is the bank address, BG[2:0] is the bank group address, and EC[5:0] is the number of codeword and parity bit errors, which is limited to a maximum of 64 errors (2 5 ). In one embodiment, only the first row with the maximum number of bits in error will be recorded in the MPR page 430, and if any other rows also have the maximum number of errors, they will not be specifically identified by the address.
[0076] In one embodiment, the MPR page 430 stores information for host access. In one embodiment, the MPR page 430 is not automatically cleared after being read by the host, and it should be read by the host whenever a complete sequence ECS command is executed. In one embodiment, the host reset register is reset to 0 by the host and then subsequently passed through the DRAM in order.
[0077] Figure 4D A block diagram of an embodiment of a multi-purpose register for storing the count of multiple rows containing errors. The MPR page 440 represents the register storing the row address with the maximum number of errors. For example, the MPR page 440 can be the MPR page 5 in a DDR4E implementation. In one embodiment, the addresses BA1:BA0 provide the MPR location. For example, "00" can correspond to MPR0, "01" can correspond to MPR1, "10" can correspond to MPR2, and "11" can correspond to MPR3. In one embodiment, the bits of the MPR page 450 are assigned as follows: EC[N - 1:0] is the number of rows with at least a threshold (e.g., at least 1 or at least 2) codeword or parity bit errors, up to a maximum of 2 N-1 rows. In one embodiment, for a maximum of 65,536 rows, N = 16. As shown, N = 20. The positions marked as RFU can be reserved for future use.
[0078] In one embodiment, the MPR page 440 stores information for access by the host. In one embodiment, the MPR page 440 is not automatically cleared after being read by the host and should be read by the host whenever a full sequence ECS command is executed. In one embodiment, the host reset register is reset to 0 by the host and then subsequently proceeds in order through the DRAM.
[0079] Figure 5 FIG. 4 is a block diagram of an embodiment of logic at a memory device that generates error correction information and supports an error check and scrubbing mode. System 500 is an example of an ECC component operation of a memory subsystem with a memory device having internal ECC that supports counting error information in the ECS mode, in accordance with the embodiments described herein. System 500 provides an example of internal ECC in DRAM that generates and stores internal parity bits. Host 510 includes a memory controller or equivalent or alternative circuitry or components that manages access to memory 520. Host 510 performs external ECC on data read from memory 520.
[0080] System 500 illustrates a write path 532 in memory 520, which represents the path by which data is written from host 510 to memory 520. Host 510 provides data 542 to memory 520 for writing to the memory array. In one embodiment, memory 520 generates parity bits 544 with parity bit generator 522 for storage with the data in the memory, which can be an example of internal ECC bits for codeword checking / correction. Parity bits 544 enable memory 520 to correct errors that may occur during writing to and reading from the memory array. Data 542 and parity bits 544 can be included as a codeword into 546 (which is written to the memory resource). It will be understood that parity bits 544 represent internal parity bits within the memory device. In one embodiment, there is no write path to parity bits 544.
[0081] Read path 534 represents the path by which data is read from memory 520 to host 510. In one embodiment, at least some of the hardware components of write path 532 and read path 534 are the same hardware. In one embodiment, memory 520 fetches codeword out 552 in response to a read command from host 510. The codeword can include data 554 and parity bits 556. Data 554 and parity bits 556 can correspond respectively to data 542 and parity bits 544 written in write path 532 if the address locations of the write command and the read command are the same. It will be understood that error correction in read path 534 can include applying an XOR (exclusive OR) tree to a corresponding H matrix to detect errors and selectively correct errors (in the case of single-bit errors).
[0082] As understood in the art, the H matrix refers to the Hamming code parity check matrix, which shows how linear combinations of the digits of a codeword equal zero. Thus, the H matrix row identifies the coefficients of the parity check equations that the components or digits to be part of the codeword must satisfy. In one embodiment, the memory 520 includes syndrome decoding 524, which enables the memory to apply parity bits 556 to data 554 to detect errors in the read data. The syndrome decoding 524 can generate a syndrome 558 for use in generating appropriate error information for the read data. The data 554 can also be forwarded to error correction 528 for correcting detected errors.
[0083] In one embodiment, the syndrome decoding 524 passes the syndrome 558 to a syndrome generator 526 to generate an error vector. In one embodiment, the parity bit generator 522 and the syndrome generator 526 are completely specified by the corresponding H matrix for the memory device. In one embodiment, if there are no errors in the read data (e.g., zero syndrome 558), the syndrome generator 526 can generate a signal that causes the data 554 to be written back to the memory. In one embodiment, if there is a single-bit error (e.g., a non-zero syndrome 558 that matches one of the columns of the corresponding H matrix), the syndrome generator 526 can generate a CE (corrected error) signal with an error location 564, which is a corrected error indication to the error correction logic 528. The error correction 528 can apply the corrected error to a specified location in the data 554 to generate corrected data 566 for writing back to the memory. In the ECS mode, ECC can use ECC operations to check for and clear errors in memory locations.
[0084] Figure 6 It is a flowchart of an embodiment of a process for monitoring one or more error counts via error correction operations in an error check and clear (ECS) mode. The process for performing ECS operations can be executed in an embodiment of a memory subsystem that supports the ECS mode described herein. The memory subsystem includes a memory controller 602 and a memory 604. In accordance with any of the embodiments described herein, the memory 604 can be a memory device. The memory responds to a trigger to enter the ECS mode. In the ECS mode, the memory device checks for and corrects errors while counting error information that the memory controller can read.
[0085] In one embodiment, memory controller 602 writes data to memory device 612 with a write command. In response to the write command, in one embodiment, memory 604 can generate internal error correction, which includes generating internal parity bits. Memory 604 records or stores the data along with the corresponding ECC bits, 614. Memory 604 can then perform an ECC operation using the ECC bits, such as an ECC operation in the ECS mode according to any of the embodiments described herein.
[0086] In one embodiment, the memory controller determines to perform Error Check and Scrub (ECS), 620. Memory controller 602 generates and sends an ECS mode trigger, 622. In one embodiment, the ECS mode trigger is an ECS command, which is a command sent with a specific encoding to cause memory 604 to perform an ECS operation. In one embodiment, the ECS mode trigger is one or more bits in a mode register. For example, memory controller 602 can execute a mode register set command that arranges the memory ECS mode. In response to the ECS mode trigger, the memory enters the ECS mode, 624.
[0087] In one embodiment, the memory generates internal address information for the ECS operation, 626. For example, memory 604 can include an internal controller that manages the ECS operation, including generating the address information. Such a controller can include one or more counters to track the addresses in order to sequentially pass through the memory for the ECS operation. The memory generates internal signals to perform the ECS operation, 628. The internal signals can include specific operations performed at the address locations for the ECS mode and can include generating and / or decoding the address information for subsequent sequential passage through the memory.
[0088] In one embodiment, the ECS operation includes reading data, correcting the data, and writing the data back to a specified memory location. Thus, in one embodiment, the memory reads one or more memory locations of a portion of the memory, such as a row, and performs error correction on the memory locations, 630. The error correction can be an ECC operation based on the stored ECC bits that were generated and stored with the data. The memory counts the number of portions with errors, 632. In one embodiment, the memory counts the error for each error found within the portion (e.g., within any addressable memory location of the memory portion). In one embodiment, the memory only counts the errors for portions that have at least a threshold number of errors (e.g., 2). Thus, the memory can count the error for every Nth detected error, or increment the count for each portion or segment that has at least N errors. In one embodiment, the memory counts both the number of errors per portion and the number of portions with errors, 634. The number of errors per portion indicates the maximum number of errors in any segment.
[0089] In one embodiment, the memory determines whether the count of the number of errors in each portion exceeds a previous maximum number of errors, 636. If the number of errors does not exceed the maximum value, 638 "No" branch, then the memory can continue to pass through the memory in order at 644. If the number of errors exceeds the previous maximum value, 638 "Yes" branch, then in one embodiment, the memory stores the number of errors as the maximum value, 640, and stores the address information of that portion together with the maximum error count, 642.
[0090] The memory can determine whether there are more memory portions to be verified and cleared, 644. If there are more portions to be verified, 646 "Yes" branch, then the memory can generate subsequent addresses for ECS operations, 626. If there are no more portions to be verified, 646 "No" branch, then in one embodiment, the memory stores the error count and the maximum error count, 648. For example, the memory can store the error information in one or more registers to be available for access by the memory controller 602. After all verifications are complete, the memory can exit the ECS mode, 650. In one embodiment, the memory controller accesses the error information 652.
[0091] Figure 7 FIG. is a block diagram of an embodiment of a computing system in which an error check and scrubbing mode with error tracking can be implemented. System 700 represents a computing device according to any embodiment described herein, and can be a laptop computer, desktop computer, server, gaming or entertainment control system, scanner, copier, printer, routing or switching device, or other electronic device. System 700 includes a processor 720 that provides processing, operation management, and instruction execution for System 700. Processor 720 can include any type of microprocessor, central processing unit (CPU), processing core, or other processing hardware to provide processing for System 700. Processor 720 controls the overall operation of System 700, and can be or include one or more programmable general or special microprocessors, digital signal processors (DSPs), programmable controllers, application specific integrated circuits (ASICs), programmable logic devices (PLDs), etc., or a combination of such devices.
[0092] Memory subsystem 730 represents the main memory of system 700 and provides temporary storage for code to be executed by processor 720 or data values to be used in execution routines. Memory subsystem 730 can include one or more memory devices, such as read-only memory (ROM), flash memory, one or more types of random access memory (RAM), or other memory devices or combinations of such devices. Memory subsystem 730 stores and hosts an operating system (OS) 736, among other things, to provide a software platform for executing instructions in system 700. In addition, other instructions 738 are stored and executed from memory subsystem 730 to provide the logic and processing of system 700. OS 736 and instructions 738 are executed by processor 720. Memory subsystem 730 includes memory device 732, in which it stores data, instructions, programs, or other items. In one embodiment, memory subsystem includes memory controller 734, which is a memory controller used to generate and issue commands to memory device 732. It will be understood that memory controller 734 can be a physical part of processor 720.
[0093] Processor 720 and memory subsystem 730 are coupled to bus / bus system 710. Bus 710 is an abstraction representing any one or more individual physical buses, communication lines / interfaces, and / or point-to-point connections connected through appropriate bridges, adapters, and / or controllers. Thus, bus 710 can include, for example, one or more of the following: a system bus, a Peripheral Component Interconnect (PCI) bus, a HyperTransport or Industry Standard Architecture (ISA) bus, a Small Computer System Interface (SCSI) bus, a Universal Serial Bus (USB), or an Institute of Electrical and Electronics Engineers (IEEE) standard 1394 bus (commonly referred to as "FireWire"). The bus of bus 710 can also correspond to the interface in network interface 750.
[0094] System 700 also includes one or more input / output (I / O) interfaces 740, a network interface 750, one or more internal mass storage devices 760, and a peripheral interface 770 coupled to bus 710. I / O interface 740 can include one or more interface components through which a user interacts with system 700 (such as video, audio, and / or alphanumeric docking). Network interface 750 provides system 700 with the ability to communicate with remote devices (such as servers, other computing devices) on one or more networks. Network interface 750 can include an Ethernet adapter, a wireless interconnect component, a USB (Universal Serial Bus), or other wired or wireless standard-based interface or a proprietary interface.
[0095] The storage device 760 can be or include any conventional medium for storing large amounts of data in a non-volatile manner, such as one or more magnetic, solid-state, or optical disks or combinations thereof. The storage device 760 preserves the code or instructions and data 762 in a persistent state (i.e., the value is retained despite an interruption to the power of the system 700). The storage device 760 can generally be regarded as a "memory", although the memory 730 is used to provide an execution or operating memory for instructions to the processor 720. However, the storage device 760 is non-volatile, and the memory 730 can include volatile memory (i.e., if power is interrupted to the system 700, the value or state of the data is indeterminate).
[0096] The peripheral interface 770 can include any hardware interface not specifically mentioned above. A peripheral generally refers to a device that is dependently connected to the system 700. A dependent connection is a connection in which the system 700 provides a software and / or hardware platform on which operations are performed and with which a user interacts.
[0097] In one embodiment, the memory 732 is DRAM. In one embodiment, the processor 720 represents one or more processors that execute data stored in one or more DRAM memories 732. In one embodiment, the network interface 750 exchanges data with another device at another network location, and the data is the data stored in the memory 732. In one embodiment, the system 700 includes an ECS controller 780 for managing ECS mode operation for the system. The ECS controller 780 represents ECS logic for performing ECS operations to internally checksum and correct errors and count error information within the memory 732 according to any of the embodiments described herein.
[0098] Figure 8 is a block diagram of an embodiment of a mobile device in which an error checksum and scrubbing mode with error tracking can be implemented. The device 800 represents a mobile computing device, such as a computing tablet, mobile phone or smartphone, wireless-enabled e-reader, wearable computing device, or other mobile device. It will be understood that certain components are generally shown, and not all components of such devices are shown in the device 800.
[0099] Device 800 includes a processor 810 that performs the main processing operations of device 800. The processor 810 can include one or more physical devices, such as a microprocessor, application processor, microcontroller, programmable logic device, or other processing components. The processing operations performed by the processor 810 include the execution of an operating platform or operating system on which applications and / or device functions are executed. The processing operations include operations related to I / O (input / output) with a human user or with other devices, operations related to power management, and / or operations related to connecting device 800 to another device. The processing operations can also include operations related to audio I / O and / or display I / O.
[0100] In one embodiment, device 800 includes an audio subsystem 820, which represents the hardware (e.g., audio hardware and audio circuits) and software (e.g., drivers, codecs) components associated with providing audio functionality to a computing device. The audio functionality can include speaker and / or headphone output, and microphone input. Devices for such functionality can be integrated into device 800 or connected to device 800. In one embodiment, the user interacts with device 800 by providing audio commands that are received and processed by the processor 810.
[0101] The display subsystem 830 represents the hardware (e.g., display device) and software (e.g., drivers) components that provide a visual and / or tactile display for user interaction with the computing device. The display subsystem 830 includes a display interface 832, which includes a specific screen or hardware device for providing a display to the user. In one embodiment, the display interface 832 includes logic separate from the processor 810 for performing at least some of the processing related to the display. In one embodiment, the display subsystem 830 includes a touch screen device that provides both output and input to the user. In one embodiment, the display subsystem 830 includes a high definition (HD) display for providing output to the user. High definition can refer to a display having a pixel density of approximately 100 PPI (pixels per inch) or greater, and can include formats such as full HD (e.g., 1080p), retina display, 4K (ultra high definition or UHD), or other formats.
[0102] The I / O controller 840 represents the hardware devices and software components related to interaction with the user. The I / O controller 840 can operate to manage the hardware that is part of the audio subsystem 820 and / or the display subsystem 830. In addition, the I / O controller 840 shows connection points for attaching additional devices to device 800 through which the user can interact with the system. For example, devices that can be attached to device 800 can include a microphone device, a speaker or stereo system, a video system or other display device, a keyboard or keypad device, or other I / O devices for use with specific applications (such as a card reader or other device).
[0103] As mentioned above, the I / O controller 840 can interact with the audio subsystem 820 and / or the display subsystem 830. For example, inputs through a microphone or other audio device can provide input or commands for one or more applications or functions of the device 800. Additionally, an audio output can be provided instead of or in addition to the display output. In another example, if the display subsystem includes a touch screen, the display device also acts as an input device, which can be at least partially managed by the I / O controller 840. There can also be additional buttons or switches on the device 800 to provide I / O functions managed by the I / O controller 840.
[0104] In one embodiment, the I / O controller 840 manages devices such as accelerometers, cameras, light sensors or other environmental sensors, gyroscopes, Global Positioning System (GPS), or other hardware that can be included in the device 800. The inputs can be part of direct user interaction and provide environmental input to the system to affect its operation (such as for noise filtering, adjusting the display for brightness detection, applying a flash for a camera application, or other features). In one embodiment, the device 800 includes a power management 850 that manages battery power usage, battery charging, and features related to power saving operations.
[0105] The memory subsystem 860 includes memory devices 862 for storing information in the device 800. The memory subsystem 860 can include non-volatile (states do not change if power to the memory device is interrupted) and / or volatile (states are uncertain if power to the memory device is interrupted) memory devices. The memory 860 can store application data, user data, music, photos, documents, or other data, as well as system data (whether long-term or temporary) related to the execution of the applications and functions of the system 800. In one embodiment, the memory subsystem 860 includes a memory controller 864 (which can also be considered part of the control of the system 800 and potentially part of the processor 810). The memory controller 864 includes a scheduler for generating and issuing commands to the memory devices 862.
[0106] Connectivity 870 includes hardware devices (such as wireless and / or wired connectors and communication hardware) and software components (such as drivers, protocol stacks) for enabling the device 800 to communicate with external devices. The external devices can be standalone devices such as other computing devices, wireless access points or base stations, and peripherals such as headsets, printers, or other devices.
[0107] Connectivity 870 can include multiple different types of connectivity. Generally speaking, device 800 is shown with cellular connectivity 872 and wireless connectivity 874. Cellular connectivity 872 generally refers to cellular network connectivity provided by a wireless carrier, such as cellular network connectivity provided via GSM (Global System for Mobile Communications) or variations or derivatives thereof, CDMA (Code Division Multiple Access) or variations or derivatives thereof, TDM (Time Division Multiplexing) or variations or derivatives thereof, LTE (Long Term Evolution - also known as "4G"), or other cellular service standards. Wireless connectivity 874 refers to wireless connectivity that is not cellular and can include personal area networks (such as Bluetooth), local area networks (such as WiFi), and / or wide area networks (such as WiMax) or other wireless communications. Wireless communication refers to the transfer of data through the use of modulated electromagnetic radiation through a non-solid medium. Wired communication occurs through a solid communication medium.
[0108] Peripheral connectivity 880 includes the hardware interfaces and connectors as well as software components (such as drivers, protocol stacks) used to make peripheral connections. It will be understood that device 800 can be both a peripheral device to other computing devices ("to" 882) and can have peripheral devices connected to it ("from" 884). Device 800 typically has a "docking" connector for connecting to other computing devices for purposes such as managing (e.g., downloading and / or uploading, changing, synchronizing) the content on device 800. Additionally, the docking connector can allow device 800 to connect to certain peripherals that allow device 800 to control, for example, the content output to an audiovisual system or other systems.
[0109] In addition to proprietary docking connectors or other proprietary connection hardware, device 800 can make peripheral connections 880 via common connectors or standards-based connectors. Common types can include Universal Serial Bus (USB) connectors (which can include any number of different hardware interfaces), display ports including Mini DisplayPort (MDP), High-Definition Multimedia Interface (HDMI), FireWire, or other types.
[0110] In one embodiment, memory 862 is DRAM. In one embodiment, processor 810 represents one or more processors that execute data stored in one or more DRAM memories 862. In one embodiment, system 800 includes an ECS controller 890 for managing the ECS mode operation of the system. ECS controller 890 represents ECS logic that performs ECS operations to internally check for and correct errors in memory 862 and count error information in accordance with any of the embodiments described herein.
[0111] In one aspect, a dynamic random access memory device (DRAM) includes: a memory array including a plurality of memory segments, the memory segments including a plurality of memory locations for storing data and error check and correction (ECC) information associated with the data; I / O (input / output) circuitry for coupling to an associated memory controller, the I / O circuitry receiving a trigger for an error check and scrubbing (ECS) mode when coupled to the associated memory controller; and an internal controller on the DRAM, responsive to the trigger for the ECS mode, for reading one or more memory locations, performing ECC on the one or more memory locations based on the ECC information, and counting error information, the error information including a segment count indicating the number of segments having N or more errors and a maximum count indicating the maximum number of errors in any segment.
[0112] In one embodiment, the DRAM includes a synchronous dynamic random access memory device (SDRAM) compliant with double data rate version 4 extended (DDR4E). In one embodiment, the memory segments include DRAM rows. In one embodiment, the trigger includes an ECS command generated by the memory controller. In one embodiment, the trigger includes a mode register setting set by the memory controller. In one embodiment, the internal controller further generates address information for the memory locations in response to the trigger for the ECS mode. In one embodiment, the memory controller is used to identify a bank group associated with the trigger for the ECS mode, and wherein the internal controller further generates address information for a specific row within the bank group. In one embodiment, the internal controller is used to perform single bit error (SBE) ECC on the one or more memory locations in response to the trigger for the ECS mode. In one embodiment, N equals 1. In one embodiment, N equals 2. In one embodiment, the internal controller further includes a multi-purpose register (MPR) for storing the segment count in a mode register of the DRAM. In one embodiment, the internal controller further includes storing an address of the segment having the maximum count, wherein the internal controller is used to store the address in the multi-purpose register (MPR) of the mode register of the DRAM.
[0113] In one aspect, a method for error correction management in a memory subsystem includes: receiving, at a memory device having a storage array including a plurality of memory segments, a trigger for an error check and scrub (ECS) mode, the memory segments including a plurality of memory locations for storing data and error check and correction (ECC) information associated with the data; reading one or more of the memory locations in response to receiving the trigger for the ECS mode; performing ECC on the one or more memory locations based on the ECC information; and counting error information, the error information including a segment count indicating the number of segments having at least a threshold number of errors and a maximum count indicating the maximum number of errors in any segment.
[0114] In one embodiment, the memory device includes a synchronous dynamic random access memory device (SDRAM) compliant with double data rate version 4 extended (DDR4E). In one embodiment, the memory segments include DRAM rows. In one embodiment, receiving the trigger includes receiving an ECS command generated by the memory controller. In one embodiment, receiving the trigger includes receiving a mode register setting set by the memory controller. In one embodiment, further includes generating address information for the memory locations in response to the trigger for the ECS mode. In one embodiment, further includes identifying a bank group associated with the trigger for the ECS mode and generating address information for a particular row within the bank group. In one embodiment, performing ECC includes performing single-bit error (SBE) ECC on the one or more memory locations in response to the trigger for the ECS mode. In one embodiment, the threshold number is equal to 1. In one embodiment, the threshold number is equal to 2. In one embodiment, further includes storing the segment count in a multi-purpose register (MPR) of a mode register of the memory device. In one embodiment, further includes storing the address of the segment having the maximum count. In one embodiment, further includes: storing the segment count in a multi-purpose register (MPR) of a mode register of the memory device; and storing the address of the segment having the maximum count.
[0115] In one aspect, a system having a memory subsystem includes: a memory controller; and a plurality of double data rate version 4 extended (DDR4E) synchronous dynamic random access memory devices (SDRAMs) that include a memory array having a plurality of memory segments, the memory segments including a plurality of memory locations for storing data and error check and correction (ECC) information associated with the data; an I / O (input / output) circuit coupled to the memory controller, the I / O circuit for receiving from the memory controller a trigger for an error check and scrubbing (ECS) mode; and an internal controller that, in response to the trigger for the ECS mode, reads one or more memory locations, performs ECC on the one or more memory locations based on the ECC information, and counts error information that includes a segment count indicating the number of segments having N or more errors and a maximum count indicating the maximum number of errors in any segment.
[0116] In one embodiment, the DRAM includes a synchronous dynamic random access memory device (SDRAM) compliant with double data rate version 4 extended (DDR4E). In one embodiment, the memory segments include DRAM rows. In one embodiment, the trigger includes an ECS command generated by the memory controller. In one embodiment, the trigger includes a mode register setting set by the memory controller. In one embodiment, the internal controller further generates address information for the memory locations in response to the trigger for the ECS mode. In one embodiment, the memory controller is used to identify a bank group associated with the trigger for the ECS mode, and wherein the internal controller further generates address information for a particular row within the bank group. In one embodiment, the internal controller is used to perform single bit error (SBE) ECC on the one or more memory locations in response to the trigger for the ECS mode. In one embodiment, N is equal to 1. In one embodiment, N is equal to 2. In one embodiment, it further includes an internal controller for storing the segment count in a multi-purpose register (MPR) of the DRAM's mode register. In one embodiment, it further includes an internal controller for storing the segment count in a multi-purpose register (MPR) of the SDRAM's mode register and for storing the address of the segment having the maximum count. In one embodiment, it further includes one or more of the following: at least one processor communicatively coupled to the memory controller; a display communicatively coupled to at least one processor; or a network interface communicatively coupled to at least one processor.
[0117] The flowcharts shown herein provide examples of sequences of various process actions. A flowchart can indicate operations (as well as physical actions) to be performed by a software or firmware routine. In one embodiment, a flowchart can show the states of a finite state machine (FSM), which can be implemented in hardware and / or software. Although shown in a specific order or sequence, the order of actions can be modified unless otherwise specified. Thus, the illustrated embodiments should be understood only as examples, and the processes can be performed in different orders and some actions can be performed in parallel. Additionally, one or more actions can be omitted in various embodiments; thus, not all actions are required in every embodiment. Other process flows are possible.
[0118] To some extent, various operations or functions are described herein, which can be described or defined as software code, instructions, configurations, and / or data. The content can be in directly executable (“object” or “executable” form), source code, or delta code (“delta” or “patch” code). The software content of the embodiments described herein can be provided via an article of manufacture having the content stored thereon or via a method of operating a communication interface to send data via the communication interface. A machine-readable storage medium can cause a machine to perform the described functions or operations and includes any mechanism that stores information in a form accessible by a machine (such as a computing device, an electronic system, etc.), such as a recordable / non-recordable medium (e.g., read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, etc.). A communication interface includes any mechanism that interfaces to any hardwired, wireless, optical, etc. medium to convey to another device, such as a memory bus interface, a processor bus interface, an Internet connection, a disk controller, etc. A communication interface can be configured by providing configuration parameters and / or sending signals to prepare the communication interface to provide a data signal representing the software content. A communication interface can be accessed via one or more commands or signals sent to the communication interface.
[0119] The various components described herein can be parts for performing the described operations or functions. Each component described herein includes software, hardware, or a combination of these. A component can be implemented as a software module, a hardware module, special-purpose hardware (such as application-specific hardware, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), etc.), an embedded controller, a hardwired circuit, etc.
[0120] In addition to the content described herein, various modifications can be made to the disclosed embodiments and implementations of the present invention without departing from their scope. Thus, the descriptions and examples herein should be construed in an illustrative sense rather than a restrictive sense. The scope of the present invention should be measured only with reference to the appended claims.
Claims
1. A dynamic random access memory (DRAM) device, comprising: A memory array including a plurality of memory rows, the memory rows including a plurality of memory locations for storing data and error check and correction (ECC) information associated with the data; Input / output (I / O) circuitry coupled to an associated memory controller, the input / output (I / O) circuitry configured to receive a trigger for an error check and scrub (ECS) mode when coupled to the associated memory controller; And An internal controller on the dynamic random access memory (DRAM) device, the internal controller configured to, in response to the trigger for the error check and scrub (ECS) mode, read one or more memory locations, perform error check and correction (ECC) for the one or more memory locations based on the error check and correction (ECC) information, and count error information, the error information including a row count and a maximum count, the row count indicating the number of rows having N or more errors, the maximum count indicating the address of the row in the rows of the memory array having the largest number of codeword errors detected in the error check and scrub (ECS) mode.
2. The dynamic random access memory (DRAM) device according to claim 1, wherein the trigger includes a mode register setting, and the mode register setting is set by the memory controller.
3. The dynamic random access memory (DRAM) device according to claim 1 or 2, wherein the internal controller further generates address information for the memory locations in response to the trigger for the error check and scrub (ECS) mode.
4. The dynamic random access memory (DRAM) device according to any one of claims 1-3, wherein the internal controller is to perform single-bit error (SBE) error check and correction (ECC) for the one or more memory locations in response to the trigger for the error check and scrub (ECS) mode.
5. The dynamic random access memory (DRAM) device according to any one of claims 1-4, wherein N is equal to 1.
6. The dynamic random access memory (DRAM) device according to any one of claims 1-4, wherein N is equal to 2.
7. A method for error correction management in a memory subsystem, comprising: Receiving, at a dynamic random access memory (DRAM) device, a trigger for an error check and scrub (ECS) mode, the dynamic random access memory (DRAM) device having a memory array including a plurality of memory rows, the memory rows including a plurality of memory locations for storing data and error check and correction error check and correction (ECC) information associated with the data; In response to receiving the trigger for the error check and scrub (ECS) mode, reading one or more memory locations; Perform error checking and correction (ECC) for the one or more memory locations based on the error checking and correction (ECC) information; and Count error information, the error information including a row count and a maximum count, the row count indicating the number of rows having at least a threshold number of errors, and the maximum count indicating the address of the row in the rows of the memory array having the maximum number of codeword errors detected in the error checking and scrubbing (ECS) mode.
8. The method according to claim 7, further comprising: Generate address information for the memory location in response to a trigger for the error checking and scrubbing (ECS) mode.
9. A dynamic random access memory (DRAM) device, comprising: A memory array including a plurality of rows; and Error checking and scrubbing (ECS) logic including error checking and correction (ECC) logic to perform an error checking and scrubbing (ECS) mode in which the dynamic random access memory (DRAM) device sequentially reads data from the memory array, corrects single-bit errors in the rows of the memory array using the error checking and correction (ECC) logic, and writes the corrected data back to the memory array; wherein the error checking and scrubbing (ECS) logic includes address generation logic to internally generate and manage address information for the error checking and scrubbing (ECS) mode to sequentially read data from the memory array.
10. The dynamic random access memory (DRAM) device according to claim 9, wherein the address generation logic counts the address until rollover of the address information.
11. The dynamic random access memory (DRAM) device according to claim 10, wherein the rollover of the address information includes a bank address rollover.
12. The dynamic random access memory (DRAM) device according to claim 10, wherein the rollover of the address information includes a memory array address rollover.
13. The dynamic random access memory (DRAM) device according to claim 12, further comprising: A register to indicate the number of rows having at least one codeword error detected in the error checking and scrubbing (ECS) mode.
14. The dynamic random access memory (DRAM) device according to claim 13, wherein in response to the memory array address rollover, the error checking and scrubbing (ECS) logic resets the register.
15. The dynamic random access memory (DRAM) device according to claim 9, further comprising: An input / output (I / O) interface to receive a command to set bits of a mode register to enter the error checking and scrubbing (ECS) mode.
16. The dynamic random access memory (DRAM) device according to claim 9, wherein the dynamic random access memory (DRAM) device comprises a synchronous dynamic random access memory (SDRAM) device compatible with the double data rate (DDR) standard.
17. A system, comprising: A memory controller; and A dynamic random access memory (DRAM) device coupled to the memory controller, the dynamic random access memory (DRAM) device comprising: A memory array comprising a plurality of rows; and Error check and scrub (ECS) logic, which includes error check and correction (ECC) logic for performing an error check and scrub (ECS) mode, in which the dynamic random access memory (DRAM) device is to sequentially read data from the memory array, correct single-bit errors in the rows of the memory array using the error check and correction (ECC) logic, and write the corrected data back to the memory array; Wherein the error check and scrub (ECS) logic includes address generation logic for internally generating and managing address information for the error check and scrub (ECS) mode in order to sequentially read data from the memory array.
18. The system of claim 17, wherein the address generation logic is to count the addresses until rollover of the address information.
19. The system of claim 18, wherein the rollover of the address information includes a bank address rollover.
20. The system of claim 18, wherein the rollover of the address information includes a memory array address rollover.
21. The system of claim 20, the dynamic random access memory (DRAM) device further comprising: A register for indicating the number of rows having at least one codeword error detected in the error check and scrub (ECS) mode.
22. The system of claim 21, wherein in response to the memory array address rollover, the error check and scrub (ECS) logic is to reset the register.
23. The system of claim 17, further comprising: An input / output (I / O) interface for receiving a command to set bits of a mode register to enter the error check and scrub (ECS) mode.
24. The system of claim 17, wherein the dynamic random access memory (DRAM) device comprises a synchronous dynamic random access memory (SDRAM) device compatible with the double data rate (DDR) standard.
25. A method for memory error management, comprising: Entering an error check and scrub (ECS) mode; Detecting and correcting bit errors in multiple rows of a memory using error check and correction (ECC) logic internal to the memory; Writing the corrected data back to the memory; Storing, in a first register serving as a row error counter, the total number of codeword errors detected in the error check and scrub (ECS) mode or the number of rows having at least one error; And Storing, in a second register serving as a per-row error counter, the address of the row having the maximum number of codeword errors in the error check and scrub (ECS) mode, In the error check and scrub (ECS) mode, a command sequence is received, the command sequence including an error check and scrub (ECS) entry command, an activate command (ACT), a write command (WR), and a precharge command (PRE).
26. The method of claim 25, wherein the first register and the second register include mode registers.
27. The method of claim 25, further comprising: counting the number of errors in the row being checked; comparing the count with a stored maximum value; and if the count is higher than the stored maximum value, writing the address of the row being checked to the second register.
28. The method of claim 25, wherein the memory is a dynamic random access memory (DRAM) device, the dynamic random access memory (DRAM) device including a synchronous dynamic random access memory (SDRAM) device compatible with the double data rate (DDR) standard.
29. The method according to claim 25, wherein, Storing the total number of the codeword errors only for the detected number of errors exceeding a threshold.
Citation Information
Patent Citations
Apparatus, method and system to determine memory access command timing based on error detection
US20140211579A1