Memory scrubbing based on detected correctable errors
By monitoring the accumulation of correctable errors in the memory array, and determining whether to perform a brushing operation based on the relationship between the number of errors and the threshold, the problem of inaccurate triggering of brushing operation in the prior art is solved, and system performance and reliability are improved.
Patent Information
- Application Number
- CN202411758657.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-04
- Filing Date
- 2024-12-03
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively determine when a memory brushing operation is performed, especially in non-ECC memory, which may reduce system performance and increase power consumption.
By monitoring correctable error accumulation in the memory array, and determining whether to perform a data brushing operation based on the relationship between the number of detected errors and the specified threshold.
More precise brushing operation triggering is achieved, unnecessary brushing is avoided, thereby improving system performance, reducing power consumption, and improving the reliability of memory devices.
Smart Images

Figure CN120108475A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to memory scrubbing based on detected correctable errors. Background Art
[0002] Memory devices used in computers or other electronic devices can be classified as volatile and non-volatile memory. Volatile memory requires power to maintain its data and includes random access memory (RAM), dynamic random access memory (DRAM), or synchronous dynamic random access memory (SDRAM), etc. Non-volatile memory can retain stored data when not powered and includes flash memory, read-only memory (ROM), electrically erasable programmable ROM (EEPROM), static RAM (SRAM), erasable programmable ROM (EPROM), resistance variable memory, phase change memory, storage class memory, resistive random access memory (RRAM) and magnetoresistive random access memory (MRAM), etc. Persistent memory is an architectural property of the system, where the data stored in the media is available after a system reset or power cycle. In some examples, non-volatile memory media can be used to build a system with a persistent memory model.
[0003] The memory device may be coupled to a host (e.g., a host computing device) to store data, commands, and / or instructions for use by the host during operation of the computer or electronic system. For example, during operation of the computing or other electronic system, data, commands, and / or instructions may be transferred between the host and the memory device.
[0004] Various protocols or standards may be applied to facilitate communication between a host and one or more other devices, such as a memory buffer, accelerator, or other input / output device. In an example, an out-of-order protocol, such as Compute Express Link (CXL), may be used to provide high bandwidth and low latency connectivity. Summary of the invention
[0005] In one aspect, the present disclosure provides a method comprising: receiving first data from a first portion of an array of memory devices; determining a number of correctable errors in the first data; and determining a relationship between the determined number of correctable errors and a specified threshold number of correctable errors and based on the relationship, performing one of the following: using the first data and the first portion of the memory array to perform a data scrubbing operation or using the first data and the first portion of the memory array to prohibit a data scrubbing operation.
[0006] In another aspect, the present disclosure provides a system comprising: a host device; and a memory device coupled to the host device, wherein the memory device comprises a memory device controller configured to: monitor for accumulation of correctable errors in each of a plurality of regions of a memory array; and conditionally trigger a scrubbing operation based on a number of observed correctable errors.
[0007] On the other hand, the present disclosure provides a non-transitory processor-readable storage medium, which includes instructions that, when executed by a processor circuit, cause the processor circuit to: read first data from a first portion of an array of a memory device; determine, using an error correction code decoder, a number of correctable errors present in the first data; in response to the number of correctable errors being less than a specified threshold number of correctable errors, perform a scrubbing operation using the first data, the scrubbing operation including determining corrected data based on the first data and writing the corrected data to the first portion of the array of the memory device; and in response to the number of correctable errors being greater than or equal to the specified threshold number of correctable errors, not allow the scrubbing operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In the drawings, which are not necessarily drawn to scale, like numerals may describe similar components in different views. Like numerals with different letter suffixes may represent different instances of similar components. The drawings generally illustrate various embodiments discussed in this document by way of example and not limitation.
[0009] Figure 1 A block diagram of an example computing system including a host and a memory system is generally illustrated.
[0010] Figure 2 An example of a Compute Express Link (CXL) system is generally described.
[0011] Figure 3 An example of a CXL system implementing a virtual hierarchy for managing transactions is generally described.
[0012] Figure 4A and Figure 4B An example of a CXL memory device is generally described.
[0013] Figure 5 A portion of a controller for a memory device is described.
[0014] Figure 6 Examples of different examples of memory arrays are generally described.
[0015] Figure 7 A portion of a controller for a memory device having a correctable error address list is described.
[0016] Figure 8 An example of a method for identifying soft and hard errors in a memory device is described.
[0017] Fig. 9 A block diagram illustrating an example machine with which, in which, or by which any one or more of the techniques discussed herein may be implemented. DETAILED DESCRIPTION
[0018] Compute Express Link (CXL) is an open standard interconnect configured for high-bandwidth, low-latency connectivity between host devices and other devices, such as accelerators, memory devices, memory buffers, and other I / O devices. CXL is designed to facilitate high-performance computing workloads by supporting heterogeneous processing and memory systems. CXL implements coherency and memory semantics on top of PCI Express (PCIe)-based I / O semantics for optimized performance.
[0019] In some examples, CXL is used in applications such as artificial intelligence, machine learning, analytics, cloud infrastructure, edge computing devices, communication systems, and others. Data processing in such applications can use various scalar, vector, matrix, and spatial architectures that can be deployed in CPUs, GPUs, FPGAs, smart NICs, or other accelerators that can be coupled using CXL links.
[0020] CXL supports dynamic multiplexing using a set of protocols that include input / output (CXL.io, based on PCIe), cache (CXL.cache), and memory (CXL.memory or CXL.mem) semantics. In an example, CXL can be used to maintain a unified, consistent memory space between a CPU (e.g., a host device or host processor) and any memory on an attached CXL device. This configuration allows the CPU and CXL devices to share resources and operate on the same memory region to achieve higher performance, reduce data movement, and reduce software stack complexity. In an example, the CPU is primarily responsible for maintaining or managing consistency in a CXL environment. Therefore, CXL can be utilized to help reduce device cost and complexity as well as the overhead typically associated with consistency across I / O links.
[0021] CXL runs on top of the PCIe PHY and provides full interoperability with PCIe. In an example, if its link partner supports CXL, the CXL device begins link training at the PCIe Gen 1 data rate and negotiates CXL as its operating protocol (e.g., using the alternative protocol negotiation mechanism defined in the PCIe 5.0 specification). As a result, devices and platforms can more easily adopt CXL by leveraging the PCIe infrastructure and without having to design and validate PHYs, channels, channel extension devices, or other upper layers of PCIe.
[0022] In an example, CXL supports single-stage switching to achieve fan-out to multiple devices. This enables multiple devices in a platform to be migrated to CXL while maintaining CXL's backward compatibility and low latency characteristics. In an example, CXL can provide a standardized computing fabric that supports pooling of multiple logical devices (MLDs) and single logical devices, such as using CXL switches connected to several host devices or nodes (such as root ports). This feature enables servers to pool resources such as accelerators and / or memory that can be assigned based on workload. For example, CXL can help facilitate resource allocation or commitment and release. In an example, CXL can help allocate and deallocate memory to various host devices as needed. This flexibility helps designers avoid over-provisioning while ensuring the best performance.
[0023] Some of the compute-intensive applications and operations mentioned herein may require or use large data sets. Memory devices storing such data sets may be configured for low latency and high bandwidth and persistence. One problem with load-store interconnect architectures includes ensuring persistence. CXL can help solve the problem using software architectural flows and standard memory management interfaces, such as being able to move persistent memory from a controller-based approach to direct memory management.
[0024] The inventors have recognized that a problem to be solved includes improving the reliability of a CXL memory device. The problem may include determining when to perform a memory scrub. An example of a memory scrub may include (1) reading data from a specific memory location, (2) correcting one or more errors in the read data (e.g., using an error correction code (ECC) algorithm), and then (3) writing the corrected data to the same specific memory location or a different memory location. In an example, the problem may include a memory device that does not include or use internal error correction code (ECC) memory to detect and correct data corruption, such as a DDR4 memory device. Non-ECC memory is generally unable to detect errors, however, some non-ECC memory includes parity information that can be used to detect errors. The inventors have further recognized that a scrub operation of a non-ECC memory may be initiated by system reliability, availability, and serviceability (RAS) components, which may degrade performance (e.g., in terms of processing time or bandwidth utilization) and increase power consumption.
[0025] The inventors have recognized that solutions to these and other problems may include performing a scrubbing operation when a particular type of correctable error (CE) occurs or when a pattern of multiple correctable errors occurs in a block of data. In an example, the solution may include or use an error detection algorithm (e.g., configured to use an error detection and / or error correction code, such as a Reed-Solomon (RS) code) to quantify any bit errors detected. That is, the error detection algorithm may determine the number of bit errors in the block of data. If more than a specified threshold number of CE errors are detected, then the probability that the error is caused by a hard failure at the peripheral level of the device is high. In another example, if more than a specified threshold number of CE errors are detected, then the portion of the die corresponding to the erroneous data is more likely to have uncorrectable errors in the near future. Thus, for example, when more than a specified threshold number of CE errors are detected for a particular block, memory scrubbing is avoided because scrubbing is less likely to be effective.
[0026] If the number of bit errors in a particular block is less than a specified threshold number of bit errors, the particular block may be deemed to have correctable errors and a scrubbing operation may be initiated to write the corrected data back to the memory device (e.g., to the same memory or die location from which the data with errors was previously read). In an example, the scrubbing operation may be performed immediately after a CE is detected to avoid error accumulation. In other examples, the scrubbing operation may be deferred or scheduled to be performed later by the system controller.
[0027] Figure 1 A block diagram generally illustrates an example of a computing system 100 including a host device 102 and a memory system 104. The host device 102 includes a central processing unit (CPU) or processor 110 and a host memory 108. In an example, the host device 102 may include a host system such as a personal computer, a desktop computer, a digital camera, a smart phone, a memory card reader and / or an Internet of Things enabled device, as well as various other types of hosts, and may include a memory access device such as the processor 110. The processor 110 may include one or more processor cores, a system of parallel processors, or other CPU arrangements.
[0028] The memory system 104 includes a controller 112, a buffer 114, a cache 116, and a first memory device 118. The first memory device 118 may include, for example, one or more memory modules (e.g., a single in-line memory module, a dual in-line memory module, etc.). The first memory device 118 may include volatile memory and / or non-volatile memory, and may include a multi-chip device including one or more different memory types or modules. In an example, the computing system 100 includes a second memory device 120 that interfaces with the memory system 104 and the host device 102.
[0029] Host device 102 may include a system backplane and may include a number of processing resources (e.g., one or more processors, microprocessors, or some other type of control circuitry). Computing system 100 may optionally include separate integrated circuits for host device 102, memory system 104, controller 112, buffer 114, cache 116, first memory device 118, second memory device 120, any one or more of which may include respective chiplets that may be connected together and used together. In an example, computing system 100 includes a server system and / or a high performance computing (HPC) system and / or a portion thereof. Although Figure 1 The example shown in illustrates a system having a von Neumann architecture, but embodiments of the present disclosure may be implemented in a non-von Neumann architecture, which may not include one or more components typically associated with a von Neumann architecture (e.g., a CPU, an ALU, etc.).
[0030] In an example, the first memory device 118 may provide main memory for the computing system 100, or the first memory device 118 may include auxiliary memory or storage devices for use by the computing system 100. In an example, the first memory device 118 or the second memory device 120 includes one or more memory cell arrays, such as volatile and / or non-volatile memory cells. For example, the array may be a flash array with a NAND architecture. The embodiments are not limited to a particular type of memory device. For example, the memory device may include RAM, ROM, DRAM, SDRAM, PCRAM, RRAM, and flash memory, etc.
[0031] In embodiments where the first memory device 118 includes persistent or non-volatile memory, the first memory device 118 may include a flash memory device, such as a NAND or NOR flash memory device. The first memory device 118 may include other non-volatile memory devices, such as non-volatile random access memory devices (e.g., NVRAM, ReRAM, FeRAM, MRAM, PCM), memory devices including ferroelectric capacitors that may exhibit hysteresis characteristics (e.g., ferroelectric RAM devices), 3D cross point (3D XP) memory devices, etc., or combinations thereof.
[0032] In an example, the controller 112 includes a media controller, such as a non-volatile memory express (NVMe) controller. The controller 112 may be configured to perform operations such as copying, writing, reading, error correction, etc. on the first memory device 118. In an example, the controller 112 may include specially constructed circuit systems and / or instructions for performing various operations. That is, in some embodiments, the controller 112 may include circuit systems and / or may be configured to execute instructions to control the movement of data and / or addresses associated with the data, such as between the buffer 114, the cache 116, and / or the first memory device 118 or the second memory device 120.
[0033] In an example, at least one of the processor 110 and the controller 112 includes a command manager (CM) for the memory system 104. The CM may receive, for example, a read command for a particular logical row address in the first memory device 118 or the second memory device 120 from the host device 102. In some examples, the CM may determine that the logical row address is associated with the first row based at least in part on a pointer stored in a register of the controller 112. In an example, the CM may receive a write command for the logical row address from the host device 102, and the write command may be associated with the second data. In some examples, the CM may be configured to issue an access command associated with the first memory device 118 or the second memory device 120 to the non-volatile memory and between issuing a read command and a write command. In some examples, the CM may issue an access command associated with the first memory device 118 or the second memory device 120 to the non-volatile memory and between issuing a read command and a write command.
[0034] In an example, buffer 114 includes data buffering circuitry that includes an area of physical memory used to temporarily store data, such as while the data is being moved from one location to another. Buffer 114 may include a first-in, first-out (FIFO) buffer, where the oldest (e.g., first-in) data is processed first. In some embodiments, buffer 114 includes a hardware shift register, a circular buffer, or a list.
[0035] In an example, the cache 116 includes an area of physical memory for temporarily storing specific data that may be used again. The cache 116 may include a pool of data entries. In some examples, the cache 116 may be configured to operate according to a write-back policy, where data is written to the cache and not simultaneously written to the first memory device 118. Therefore, in some embodiments, data written to the cache 116 will not have a corresponding data entry in the first memory device 118.
[0036] In an example, the controller 112 may receive write requests (e.g., from the host device 102) involving the cache 116 and cause data associated with each of the write requests to be written to the cache 116. In some examples, the controller 112 may receive write requests at a rate of thirty-two (32) gigatransfers (GT) / second, e.g., according to or using the CXL protocol. The controller 112 may similarly receive read requests and cause data stored in, for example, the first memory device 118 or the second memory device 120 to be retrieved and written to, for example, the host device 102 via the interface 106.
[0037] In an example, the interface 106 may include any type of communication path, bus, or the like that allows information to be transferred between the host device 102 and the memory system 104. Non-limiting examples of the interface may include a peripheral component interconnect (PCI) interface, a peripheral component interconnect express (PCIe) interface, a serial advanced technology attachment (SATA) interface, and / or a micro serial advanced technology attachment (mSATA) interface, etc. In an example, the interface 106 includes a PCIe 5.0 interface that complies with the Compute Express Link (CXL) protocol standard. Therefore, in some embodiments, the interface 106 supports a transfer speed of at least 32 GT / s.
[0038] As similarly described elsewhere herein, CXL is a high-speed central processing unit (CPU) to device or CPU to memory interconnect designed to enhance computing performance. CXL technology maintains memory coherency between the CPU memory space (e.g., host memory 108) and the memory on the attached device or accelerator (e.g., first memory device 118 or second memory device 120), which allows resource sharing to achieve higher performance, reduced software stack complexity, and lower overall system cost. CXL is designed as an industry open standard interface for high-speed communications as accelerators are increasingly used to complement CPUs to support emerging data-rich and compute-intensive applications, such as artificial intelligence and machine learning.
[0039] Figure 2 An example of a CXL system 200 is generally described that uses a bus system (including a CXL link bus 206 and a system management bus 208) to connect a host device 202 and a CXL device 204. In the example, the host device 202 includes or corresponds to the host device 102 and the CXL device 204 includes or corresponds to a host device 102. Figure 1 1. The memory system command manager (CM) may include a portion of the host device 202 or the CXL device 204.
[0040] In an example, the system management bus 208 (eg, corresponding to the Figure 1The system management bus 208 is a portion of the interface 106 of an example of an example of a host device 202 configured to support main-band or side-band communications between the host device 202 and the CXL device 204. The system management bus 208 may use the PCIe and CXL protocols to carry miscellaneous commands or events such as link speed changes, reset commands issued by the host, and other reliability, availability, and serviceability features.
[0041] In an example, CXL link bus 206 (eg, corresponding to Figure 1 104) may support communication using a multiplexed protocol for caching (e.g., CXL.cache), memory access (e.g., CXL.mem or CXL.memory), and data input / output transactions (e.g., CXL.io). CXL.io may include a PCIe-based protocol for functions such as device discovery, configuration, initialization, I / O virtualization, and direct memory access (DMA) using incoherent load-store, producer-consumer semantics. CXL.cache may enable a device to cache data from a host memory (e.g., from host memory 214) using a request and response protocol. CXL.memory may enable a host device 202 to use memory attached to a CXL device 204, such as in or using a virtualized memory space. CXL-based memory devices may include or use, for example, volatile or non-volatile memory that may be characterized by different speeds or latencies. In an example, a CXL-based memory device may include a CXL-based memory controller configured to manage transactions with volatile or non-volatile memory.
[0042] In an example, CXL.memory transactions may be memory load and store operations that are run downstream or external to the host device 202. CXL memory devices may have varying levels of complexity. For example, a simple CXL memory system may include a CXL device that includes or is coupled to a single media controller, such as a memory controller (MEMC). A medium CXL memory system may include a CXL device that includes or is coupled to multiple media drivers. A complex CXL memory system may include a CXL device that includes or is coupled to a cache controller (and its accompanying cache) and one or more media or memory controllers.
[0043] exist Figure 2In an example of , host device 202 includes host processor 216 (e.g., including one or more CPUs or cores) and IO device 228. Host device 202 may include or may be coupled to host memory 214. Host device 202 may include various circuitry or logic configured to facilitate CXL-based communications and transactions with CXL device 204. For example, host device 202 may include coherence and memory logic 220 configured to implement transactions according to CXL.cache and CXL.memory semantics, and host device 202 may include PCIe logic 222 configured to implement transactions according to CXL.io semantics. In an example, host device 202 may be configured to manage the coherence of data cached at CXL device 204 using, for example, its coherence and memory logic 220.
[0044] Host device 202 may further include a host multiplexer 218 configured to modulate communications over CXL link bus 206 (e.g., using a PCIe PHY layer). Multiplexing of protocols ensures that latency-sensitive protocols (e.g., CXL.cache and CXL.memory) have the same or similar latency as native processor-to-processor links. In an example, CXL defines an upper bound on response time for latency-sensitive protocols to help ensure that device performance is not negatively impacted by latency variations between different devices implementing coherency and memory semantics.
[0045] In an example, a symmetric cache coherence protocol may be difficult to implement between host processors because different architectures use different solutions, which in turn compromises backward compatibility. CXL can address this problem by strengthening the coherence function at the host device 202 (eg, using the coherence and memory logic 220).
[0046] CXL devices may include devices with a variety of different architectures and capabilities. For example, a Type 1 CXL device may be a device configured to implement a fully coherent cache without host management. Transaction types used with Type 1 devices may include device-to-host (D2H) coherent transactions and host-to-device (H2D) snoop transactions, among others. Type 2 CXL devices (e.g., which may include or use attached high-bandwidth memory) may be configured to optionally implement a coherent cache and may be managed by the host. CXL.cache and CXL.mem transactions are typically supported by Type 2 devices. Type 3 CXL devices (e.g., memory expanders for hosts) may be configured to include or use host-managed memory. Type 3 devices support CXL.mem transactions.
[0047] The CXL device 204 may include various components or logic blocks, including a CXL host interface 232 and a device management system 234. In an example, the CXL host interface 232 may be configured to receive and manage various requests and transactions. For example, the CXL host interface 232 may be configured to receive and communicate PCIe resets, such as using PERST (PCI Express Reset), Warm Reset, FLR (Function Level Reset), and CXL Reset. In an example, the CXL host interface 232 may be configured to receive and communicate DOE transaction layer packets. In an example, the CXL host interface 232 may be configured to handle sideband requests or other miscellaneous events from PCIe and CXL devices, such as using the CXL link bus 206 or the system management bus 208.
[0048] The CXL host interface 232 may include or use a plurality of CXL interface physical layers 212. The device management system 234 may include a device logic and memory controller 224, among other things. In an example, the CXL device 204 may include a device memory 230 or may be coupled to another memory device. The CXL device 204 may include various circuitry or logic configured to facilitate CXL-based communications and transactions with the host device 202 using the CXL link bus 206. For example, the device logic and memory controller 224 may be configured to implement transactions received using the CXL host interface 232 according to CXL.cache, CXL.memory, and CXL.io semantics. The CXL device 204 may include a CXL device multiplexer 226 configured to control communications over the CXL link bus 206.
[0049] In an example, one or more of the coherency and memory logic 220, the device management system 234, and the device logic and memory controller 224 include a unified assist engine (UAE) or compute fabric with various functional units such as a command manager (CM), a thread engine (TE), a stream engine (SE), a data manager or data mover (DM), or other units. The compute fabric is reconfigurable and can include separate synchronous and asynchronous streams.
[0050] Device management system 234 or device logic and memory controller 224 or portions thereof may be configured to operate in the application space of CXL system 200 and, in some examples, may launch its own threads or sub-threads that may operate in parallel and may optionally use resources or units on other CXL devices 204. Queue and transaction control through the system may be coordinated by the CM, TE, SE, or DM components of the UAE. In an example, each queue or thread may be mapped to a different loop iteration to thereby support multi-dimensional loops. With the ability to launch such nested loops, among other capabilities, the system may achieve significant time savings and latency improvements for compute-intensive operations.
[0051] In an example, command fences may be used to help maintain order throughout such operations, which may be performed locally or throughout the computational space of device logic and memory controller 224. In some examples, CM may be used to route commands to a specific command execution unit (e.g., device logic and memory controller 224 comprising a specific instance of CXL device 204) using an out-of-order interconnect that provides corresponding transaction identifiers (TIDs) to command and response message pairs.
[0052] In an example, the CM may coordinate synchronization flows, such as using an asynchronous fabric of a reconfigurable computing fabric to communicate with other synchronization flows and / or other components of the reconfigurable computing fabric using asynchronous messages. For example, the CM may receive an asynchronous message from a dispatch interface and / or from another flow controller to indicate a new thread at or using a synchronization flow. The dispatch interface may interface between the reconfigurable computing fabric and other system components. In some examples, a synchronization flow may send an asynchronous message to the dispatch interface to indicate thread completion.
[0053] Asynchronous messages may be used by synchronous streams, for example, to access memory. For example, a reconfigurable computing fabric may include one or more memory interfaces. A memory interface is a hardware component that may be used by a synchronous stream or its components to access external memory that is not part of the synchronous stream but is accessible by the host device 202 or CXL device 204. Threads executed using synchronous streams may include sending read and / or write requests to the memory interface. Because reads and writes are asynchronous, a thread that initiates a read or write request to the memory interface may not receive the result of the request. Instead, the result of the read or write request may be provided to a different thread executing at a different synchronous stream. Delay and output registers in one or more of the CXL devices 204 may help coordinate and maximize the efficiency of the first stream, such as by accurately timing the engagement of a particular computing resource of one device with the arrival of data related to the first stream. The registers may help enable a particular computing resource of the same resource to be repurposed for a stream other than the first stream, such as when the first stream is lingering or waiting for other data or operations to complete. Such other data or operations may depend on one or more other resources of the fabric.
[0054] Figure 3 An example of a portion of a CXL system that may include or use a virtual hierarchy for managing transactions, such as memory transactions with CXL memory devices, is generally described. The example may include or use real-time telemetry to help facilitate allocation of new or ongoing queues. Figure 3 Examples of include first virtual layer 304 and second virtual layer 306. First virtual layer 304, second virtual layer 306, or one or more modules or components thereof may be implemented using host device 202, CXL device 204, or multiple instances of host device 202 or CXL device 204.
[0055] exist Figure 3 In the example of FIG. 3 , the first virtual hierarchy 304 includes a first host device 308 and the second virtual hierarchy 306 includes a second host device 310. A CXL switch 302 may be provided to expose a plurality of CXL resources to different hosts in the system. In other words, the CXL switch 302 may be configured to couple each of the first host device 308 and the second host device 310 to the same or different resources, such as using respective virtual CXL switches (VCSs), such as the first VCS 320 and the second VCS 322, respectively. The CXL switch 302 may be statically configured to couple each host device to respective different resources, or the CXL switch 402 may be dynamically configured to different resources, such as depending on the needs of a particular one of the host devices to execute its respective queue or thread. Thus, the CXL switch 302 enables virtual hierarchies and resource sharing between different hosts.
[0056] In an example, a fabric manager (FM) may be provided to assign or coordinate connectivity of CXL switches 302 and may be configured to activate, deactivate, or reconfigure a virtual hierarchy of a CXL system. The FM may include a baseboard management controller (BMC), an external controller, a centralized controller, or other controller.
[0057] exist Figure 3 In the example of FIG. 1 , the CXL switch 302 or the first VCS 320 or the second VCS 322 may coordinate communications between the host device and various accelerators or other CXL devices. For example, the CXL switch 302 may be coupled to various CXL devices (e.g., the first CXL device 318 or the second CXL device 324) via multiple logical devices (MLDs, such as the MLD 312) or to various logical devices, such as a single logical device (LD, such as the first LD 314, the second LD 316, the third LD 326, or the fourth LD 328). Each CXL device and logical device may represent a corresponding accelerator or CXL device with its own corresponding CXL.io configuration space, CXL.mem memory space, and CXL.cache cache space.
[0058] Figure 4A and Figure 4BAn example of a CXL device 402, which may include a memory device, is generally described. In the example, the CXL device 402 includes a CXL controller that manages transactions with a host, and the CXL device 402 includes a memory controller that manages transactions with the memory. The memory may include or use volatile memory, such as DRAM, SDRAM, PCRAM, RRAM, and other kinds of memory. The memory may additionally or alternatively include or use non-volatile memory, such as NAND or NOR flash memory. Although hosts and other CXL devices are discussed in various examples herein as "CXL" host devices and "CXL" accelerators or "CXL" devices, other types of hosts and accelerators that do not include or use the CXL protocol may similarly be used.
[0059] In an example, CXL device 402 is a type of accelerator device that is configured to communicate with one or more hosts via a CXL interface, such as using transactions defined by CXL.io, CXL.mem, and CXL.cache protocols. CXL device 402 may include a Type 3 CXL device, such as a memory device having one or more memories, such as memories of the same type or different types (e.g., memories exhibiting respective different latency characteristics).
[0060] For ease of illustration and discussion, the example of CXL device 402 includes a hypothetical front end portion 404, a mid-end portion 406, and a back end portion 408. Portions of CXL device 402 and its components may be configured or combined differently according to different implementations of CXL device 402.
[0061] exist Figure 4A In an example of the present invention, the front end portion 404 may include a CXL link 412 configured to interface with a host device using a physical layer (CXL PCIe PHY layer 410). The front end portion 404 may further include a CXL data link layer 414 and a CXL transport layer 416 configured to manage transactions between the CXL device 402 and the host. In an example, the CXL transport layer 416 includes registers and operators configured to manage a CXL request queue (e.g., including one or more memory transaction requests) and a CXL response queue (e.g., including one or more memory transaction responses) of the CXL device 402.
[0062] In an example, the CXL device 402 may include a memory device that includes a cache (eg, including SRAM) and includes longer term volatile or non-volatile memory accessible via a memory controller. Figure 4A and Figure 4BIn the example of , CXL device 402 includes cache memory 420 in a mid-end portion 406 of the device. Mid-end portion 406 may include a cache controller 418 configured to monitor requests from CXL transport layer 416 and identify requests that can be satisfied using cache memory 420.
[0063] Various complexities may arise in a CXL system. For example, CXL transactions may be based on a relatively large transaction size (e.g., 64 bytes), while some processes may use a larger granularity or a smaller data size. Thus, in some examples, a cache controller 418 may be included or used in a CXL device 402 to store excess data fetched from a backend media controller or memory (e.g., from one or more memories in a backend portion 408 of a CXL device 402).
[0064] In certain examples (e.g., including or using CXL devices 402 with DDR4 or DDR5 attached memory), sideband ECC may be supported or used to help protect data integrity. When the transaction size is 64 bytes, a relatively large amount of ECC data may be retrieved at one time, while only a portion of the ECC data may be available for a particular transaction. Cache memory 420 may be used to store excess ECC data for more efficient access, thereby helping to reduce latency for future transactions.
[0065] In an example, the cache controller 418 is coupled to a crossbar interface or XBAR interface 422. The XBAR interface 422 can be configured to allow multiple requestors to access multiple memory controllers in parallel, such as including multiple memory controllers in the back end portion 408 of the CXL device 402. In an example, the XBAR interface 422 provides essentially point-to-point access between the requestors and the memory controllers and generally provides higher performance than using a conventional bus architecture. The XBAR interface 422 can be configured to receive responses from the back end portion 408 or receive cache hits from the cache memory 420 and deliver the responses to the front end portion 404 using a cache response queue.
[0066] exist Figure 4B 4. The backend portion 408 of the CXL device 402 includes a plurality of memory controllers, including a first memory controller 424 to an Nth memory controller 428. Each of the memory controllers may have or use a respective memory request and response queue. Each of the memory controllers may be coupled to a respective medium or memory, which may include volatile or non-volatile memory, for example. In the illustrated example, the first memory controller 424 is coupled to a first memory 426 and the Nth memory controller 428 is coupled to an Nth memory 430.
[0067] In an example, each of the multiple memory controllers in the system can manage its own respective queue. In some examples, different memory controllers can be configured to use or interface with memories having respective different latency characteristics. Thus, performance optimization can include coordination of the respective queues of each memory controller. Informed coordination can be based on, for example, request and response path information of each memory controller.
[0068] In an example, a memory device of the CXL device 402 may include a memory array that does not include or use on-die ECC. In this case, random errors in the memory array may be corrected at or using a memory controller (e.g., the first memory controller 424, the Nth memory controller 428, etc.). For example, the controller may be configured to use a Reed-Solomon (RS) code to identify or correct errors in data retrieved from the memory array. In an example, the controller accesses data provided by multiple dies (e.g., 18 dies) accessed in parallel (e.g., using a 72-bit channel). For example, each die may use multiple data input or data output pins (DQ pins), such as 4 pins per die. In an example that may include a CXL memory device, the minimum transaction size may be 64 bytes. Thus, in an 18-die device in which 2 dies include parity information (e.g., Reed-Solomon code data), each die provides 4 bytes of data to thereby provide a 64-byte transaction or data block.
[0069] In an example, a scrubbing operation may be performed. The scrubbing operation may include one or more of reading data, correcting data, and writing data. In an example, the scrubbing operation includes reading data from a specific location or address in the memory array and identifying correctable errors (CEs). CEs may be identified using an error detection algorithm, such as using Reed-Solomon codes, BCH codes, or other codes.
[0070] In an example, the CE may be identified by a decoding engine configured to apply an error correction code. In an example, the decoding engine may be configured to provide corrected data. For example, the scrubbing operation may include writing the corrected data back to the same specific location or address in the memory array from which the data (with the error) was originally or previously read.
[0071] Even after a scrub operation, some cells within the memory array may continue to contain incorrect data. The inventors have recognized that when errors persist or accumulate (including after a scrub operation is performed), future device failure may be indicated.
[0072] The inventors have further recognized that scrubbing operations may be conditionally performed based on or in response to error patterns identified by a decoding engine and controller. In other words, when some scrubbing operations are unlikely to be effective in correcting particular errors or error patterns, such operations may be avoided (and other ECC techniques applied).
[0073] In some examples, a scrubbing operation may include rewriting data immediately after a read operation to take advantage of an already open row and thus avoid a specific activate command. This approach may help optimize the system and avoid error accumulation in memory (eg, DRAM) components.
[0074] Figure 5 An example of a portion of a CXL device 402 including a memory controller 502 and a media subsystem 504 is generally illustrated. The media subsystem 504 may include a memory array 510, for example including a plurality of dies (e.g., 18 dies). The memory controller 502 may include a scrub manager 506, an error encoder-decoder 508, a repair 512 module configured to perform other data repairs (e.g., for errors outside the scope of the error encoder-decoder 508), etc. In an example, the error encoder-decoder 508 includes an enhanced RAS module configured to perform other reliability, availability, and serviceability mechanisms of the memory device.
[0075] exist Figure 5 In the example of , 16 dies include data and 2 dies include parity information used with the data. In the example, the data burst from the media subsystem 504 includes 72 symbols, of which 8 symbols represent parity information, which may include, for example, a Reed-Solomon (RS) code. In this example, the RS code can correct up to 4 erroneous symbols, and each die provides 4 symbols. Therefore, even if a particular die is unavailable or faulty, the RS code can be used to correct the data provided by the chip as long as there are no additional errors.
[0076] The media subsystem 504 may provide a burst of data to the memory controller 502 and the data may be received by the error encoder-decoder 508. The error encoder-decoder 508 may be configured to process the received data to determine whether there are one or more correctable errors in the received data, etc. The output from the error encoder-decoder 508 may include, for example, one or more of an indication of whether an error was detected, corrected data (e.g., whether one or more correctable errors were found), and information about error patterns. In an example, the output from the error encoder-decoder 508 includes information about the results of applying the RS code check and may further include information about uncorrectable errors (if encountered).
[0077] The scrub manager 506 may be configured to use information about the error pattern to determine whether to allow data to be rewritten to the same memory location from which the data was originally or previously read. That is, the scrub manager 506 may receive information about the CE error pattern and may receive corrected data from the error encoder-decoder 508, and in response, the scrub manager 506 may be configured to determine whether and where to write the data back to the media subsystem 504 or take different actions.
[0078] In an example, the scrub manager 506 may be configured to use information from the error encoder-decoder 508 to distinguish between random errors or "soft" errors corresponding to a discrete number of bits or cells (e.g., 1 or 2 cells) and peripheral failures corresponding to a relatively large number of bits or cells (e.g., 3 or more cells). In an example, peripheral failures corresponding to a large number of cells may be considered "hard" errors and uncorrectable or indicative of a device failure. The scrub manager 506 may be configured to allow scrubbing operations for detected soft or random errors. The scrub manager 506 may be configured to not allow scrubbing operations for peripheral failures or hard errors. In an example, the scrub manager 506 may be configured to not allow scrubbing operations for random or soft errors that meet or exceed a threshold count or weight of errors in a particular region or die of the memory array.
[0079] The inventors have recognized that information about correctable error patterns can be used to distinguish random errors from peripheral faults. In an example, the correctable error pattern includes a count of the number of symbols corrected by the error encoder-decoder 508. If the count of corrected symbols is greater than a specified threshold count or weight, then the error is unlikely to be a random error and a corrective action other than a scrubbing operation can be initiated. Functions other than counting can similarly be used to determine the weight of the error observed in the output of the error encoder-decoder 508. In an example, a Hamming weight can be determined and used to determine whether to perform a scrubbing operation. The number of 1s can be counted in the data stream of each die. This count can be considered as the Hamming weight of the pattern in the data stream of each corresponding die. The weight can be calculated by adding the weights of the symbols affected by the error. The weight can be compared to a specified threshold and used to determine whether to continue the scrubbing operation (e.g., when the weight does not exceed the threshold) or to not allow the scrubbing operation (e.g., when the weight exceeds the threshold).
[0080] In other words, if less than a threshold number of errors (e.g., correctable errors) are identified by the error encoder-decoder 508 for a particular die, then a scrubbing operation may be initiated and performed. Since the error count does not exceed the threshold, the system may allow and configure appropriate resources (e.g., power, bandwidth) to write data back to the array or media subsystem 504. If more than a threshold number of correctable errors are identified by the error encoder-decoder 508 for a particular die, then a scrubbing operation is not initiated because the scrubbing operation is unlikely to correct future errors. By avoiding scrubbing operations that are unlikely to succeed and unlikely to correct future errors, time and power may be saved and bandwidth utilization may be preserved for more productive operations.
[0081] Figure 6 A graphical representation of examples of different CE patterns observed by the error encoder-decoder 508 in data from the memory array 510 is generally illustrated. In the first example of the array 602, multiple correctable errors are observed or detected in die 4. The graphical representation indicates that, of the total 32 bits read from die 4, there are clusters with a single isolated bit error in a portion and additional bit errors in another portion. For example, a cluster may represent 4 or more bit errors. A cluster of bit errors may represent bit errors corresponding to physically adjacent or nearby memory cells in the array. If the correctable error threshold for triggering a scrubbing operation is two bit errors, then a scrubbing operation is not triggered for the data read from die 4 of the first example of the array 602 because the number of errors exceeds the threshold. In this case, the errors associated with die 4 are likely to be uncorrectable errors, and scrubbing is unlikely to help avoid future errors.
[0082] In a second example of array 604, multiple correctable errors are observed or detected in multiple dies. For example, a single bit error may be observed in dies 0, 11, and 13 and a pair of bit errors may be observed in die 4. In this example, if the correctable error threshold for triggering a scrubbing operation is two bit errors, then a scrubbing operation will be triggered for data read from each of dies 0, 4, 11, and 13 because each die is responsible for two or fewer correctable errors (i.e., soft errors), which can be handled by error encoder-decoder 508 and may not be uncorrectable errors.
[0083] Figure 7 The general description may include examples of deferring one or more scrubbing operations using the scrub manager 506. For example, the memory controller 502 may include or use the address list 702 to store information about one or more memory locations (e.g., locations in the memory array 510 of the media subsystem 504) to be scrubbed when resources are available or at a specific scheduled time.
[0084] In an example, address list 702 can be populated with information from error encoder-decoder 508 regarding one or more correctable errors identified in data received from memory array 510. In an example, address list 702 can be populated with error address information while errors are identified by error encoder-decoder 508.
[0085] However, deferred scrubbing operations consume additional system resources because additional read commands are used to retrieve the data to be scrubbed from the memory array 510. That is, the scrub manager 506 may send a read command (RD) using one or more addresses from the address list 702, and may then perform a scrubbing operation, including reading the received data, correcting the data using the error encoder-decoder 508, and writing the corrected data back to the memory array 510. However, consuming additional system resources to perform deferred scrubbing operations at scheduled times or during other times that reduce system utilization may be preferable to completing the scrubbing operations while errors are identified.
[0086] Figure 8 A method 800 is generally described for a memory device maintenance process that may include one or more scrubbing operations or cycles of scrubbing operations. In an example, a memory device may enforce a variety of different maintenance and scrubbing strategies. For example, a periodic scrubbing strategy or patrol scrubbing may be performed according to a specified cadence (e.g., daily). For example, after each 24-hour operation cycle, each row in the memory array 510 may be read, corrected, and rewritten.
[0087] In an example, the active CE-triggered scrubbing strategy may include an on-demand strategy. In an on-demand strategy, for each detected CE, corrected data (e.g., provided by the error encoder-decoder 508) may be written back to the memory array 510 at the same location from which the data was read. In an example, the on-demand strategy may be enhanced or improved by including a re-read operation to capture hard failures. For example, reading the scrubbed data entries multiple times may help classify errors as soft errors or hard errors, which in turn may lead to more efficient and effective corrections. In addition, host notifications about error events may be retained for persistent errors, thereby reducing interconnect (e.g., CXL interconnect) traffic. The method 800 includes an example of an on-demand strategy with data re-reading.
[0088] Method 800 may start at operation 802 and perform a data read (RD) at operation 804. The data read from the memory at operation 804 may be processed by a decoder (e.g., an RS decoder) at decision operation 806. The result of the decoder processing may be one of a zero error (ZE), an uncorrectable error (UE), or a correctable error (CE). If an error is identified by the decoder, the decoder may be configured to determine the number of correctable errors and uncorrectable errors in the data. In an example, the decoder may be configured to determine the relationship between the determined number of correctable errors and the specified threshold number of correctable errors and based on the relationship, use the memory location from which the data is read at operation 804 to selectively perform (e.g., allow) or prohibit (e.g., not allow) a data scrubbing operation. In an example, the decoder may be configured to determine the distribution of errors with various granularities. For example, the decoder may be configured to determine the distribution of errors in one or more memory devices, memory arrays, or rows, columns, or cells of an array. The decoder may be further configured to determine whether the error distribution includes a cluster of errors, such as multiple errors that may represent the same row or column or multiple adjacent or nearby rows or columns of a particular die.
[0089] At decision operation 806, if the result is ZE, then there are no errors in the data and other operations of the memory device may proceed. If at decision operation 806, the result is CE, then at operation 808, the scrub manager may perform a scrub operation, such as using corrected data from the decoder and the memory location corresponding to the read at operation 804. In an example, the scrub manager may be configured to perform a scrub operation (e.g., at operation 808) only when the number of correctable errors for a particular die is less than a specified threshold number of correctable errors. At operation 810, the same memory location may be read again and at decision operation 812, may be processed again by a decoder (e.g., the same or a different decoder). Returning to decision operation 806, if the result is UE, then method 800 may proceed to decision operation 812.
[0090] The results from decision operation 812 may be ZE, UE, and CE. If the result of decision operation 812 is ZE, then the CE identified at decision operation 806 was successfully corrected by scrubbing operation 808 and no additional errors were detected. If the result of decision operation 812 is UE, then the controller may notify the host, such as by sending a poison message to the host at operation 822, and the controller may log the event at operation 824. In response, the host may take other corrective actions, such as taking one or more pages of the array offline or employing other RAS mechanisms to handle the detected UE. If the result of decision operation 812 is CE, then the error may be considered a hard error. That is, because the error was encountered twice (once after decision operation 806 and once after decision operation 812), the error is unlikely to be corrected using additional scrubbing operations and therefore the error may be classified as a hard error.
[0091] At decision operation 814, the CE or hard error identified at decision operation 812 may be analyzed to determine whether repair criteria are met. If the repair criteria are not met, the decision may be recorded at operation 820 and the host may optionally determine to perform other mitigation operations. If at decision operation 814, the repair criteria are met, then repair may be performed at operation 816 and data may be copied to a new location in the memory array at operation 818. In an example, operation 816 may include post-packaging repair to remap accesses from a failed row associated with the CE to a different row. In an example, information about the CE identified at decision operation 812 may be stored or added to a CE list (e.g., corresponding to address list 702) at operation 828, for example, after the repair operation is completed.
[0092] Fig. 9A block diagram of an example machine 900 is illustrated in which, in which, or by which any one or more of the techniques discussed herein (e.g., methods) may be implemented. As described herein, an example may include logic or several components or mechanisms in the machine 900 or may be operated by it. A circuit system (e.g., a processing circuit system) is a collection of circuits implemented in a tangible entity of the machine 900 including hardware (e.g., simple circuits, gates, logic, etc.). The members of the circuit system (e.g., belonging to a host-side device or process or belonging to an accelerator-side device or process) may change flexibly over time. The circuit system includes members that can perform specific operations individually or in combination when operating. In an example, the hardware of the circuit system may be designed unchanged to implement specific operations (e.g., hard-wired), such as using device logic and memory controller 224 or host interface circuitry or using its specific command execution unit, such as monitoring or tracking correctable errors in a memory device and based on the correctable error pattern, selectively allowing or not allowing scrubbing operations to help mitigate or avoid future uncorrectable errors in problem die areas. In an example, the hardware of the circuit system may include variably connected physical components (such as command execution units, transistors, simple circuits, etc.) that include machine-readable (such as processor-readable) media that are physically modified (such as magnetically, electrically, removably placing constant mass particles, etc.) to encode instructions for specific operations. For example, when the physical components are connected, the basic electrical properties of the hardware composition change from insulators to conductors, and vice versa. The instructions enable embedded hardware (such as execution units or loading mechanisms) to create members of the circuit system in hardware via variable connections to implement parts of specific operations when operating. Therefore, in an example, the machine-readable media element is part of the circuit system or is communicatively coupled to other components of the circuit system when the device is operating. In an example, any of the physical components can be used in more than one member of more than one circuit system. For example, under operation, the execution unit can be used in a first circuit of a first circuit system at one point in time and reused by a second circuit in the first circuit system or a third circuit in the second circuit system at a different time.
[0093] In alternative embodiments, the machine 900 may operate as a standalone device or may be connected (e.g., networked) to other machines. In a networked deployment, the machine 900 may operate as a server machine, a client machine, or both in a server-client network environment. In an example, the machine 900 may act as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. The machine 900 may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a network appliance, a network router, a switch or a bridge, or any machine capable of executing instructions (sequentially or otherwise) specifying actions to be taken by the machine. In addition, although only a single machine is described, the term "machine" shall also be construed to include any collection of machines that individually or collectively execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein, such as cloud computing, software as a service (SaaS), and other computer cluster configurations.
[0094] Any one or more of the components of the machine 900 may include or use one or more instances of the host device 202 or the CXL device 204 or other components in or attached to the computing system 100. The machine 900 (e.g., a computer system) may include a hardware processor 902 (e.g., a host processor 216, a device logic and memory controller 224, a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), a main memory 904, a static memory 906 (e.g., a memory or storage device for firmware, microcode, basic input output (BIOS), a unified extensible firmware interface (UEFI), etc.), and a mass storage device 908 or a stack of memory dies, a hard drive, a tape drive, a flash storage device, or other block devices, some or all of which may communicate with each other via an interconnect 930 (e.g., a bus). The machine 900 may further include a display device 910, an alphanumeric input device 912 (e.g., a keyboard), and a user interface (UI) navigation device 914 (e.g., a mouse). In an example, the display device 910, the input device 912, and the UI navigation device 914 may be a touch screen display. The machine 900 may additionally include a mass storage device 908 (e.g., a drive unit), a signal generating device 918 (e.g., a speaker), a network interface device 920, and one or more sensors 916, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or other sensors. The machine 900 may include an output controller 928, such as a serial (e.g., universal serial bus (USB)), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection to communicate with or control one or more peripheral devices (e.g., printers, card readers, etc.).
[0095] The hardware processor 902, main memory 904, static memory 906, or a register of a mass storage device 908 may be or include a machine-readable medium 922 on which is stored one or more sets of data structures or instructions 924 (e.g., software) embodying or used by any one or more of the techniques or functions described herein. The instructions 924 may also reside, completely or at least partially, within any of the registers of the hardware processor 902, main memory 904, static memory 906, or a mass storage device 908 during execution thereof by the machine 900. In an example, one or any combination of the hardware processor 902, main memory 904, static memory 906, or a mass storage device 908 may constitute the machine-readable medium 922. Although the machine-readable medium 922 is illustrated as a single medium, the term "machine-readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database or associated caches and servers) configured to store one or more instructions 924.
[0096] The term "machine-readable medium" (or equivalently, "processor-readable medium") may include any medium capable of storing, encoding, or carrying instructions executed by the machine 900 and causing the machine 900 to perform any one or more of the techniques of the present disclosure, or capable of storing, encoding, or carrying data structures used by or associated with such instructions. Non-limiting machine-readable medium examples may include solid-state memory, optical media, magnetic media, and signals (e.g., radio frequency signals, other photon-based signals, sound signals, etc.). In an example, a non-transitory machine-readable medium includes a machine-readable medium with a plurality of particles having a constant (e.g., stationary) mass and is therefore a composition of matter. Thus, a non-transitory machine-readable medium is a machine-readable medium that does not include a transitory propagating signal. Specific examples of non-transitory machine-readable media may include: non-volatile memory, such as semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0097] In an example, information stored or otherwise provided on machine-readable medium 922 may represent instructions 924, such as instructions 924 themselves or a format from which instructions 924 may be derived. Such a format from which instructions 924 may be derived may include source code, encoded instructions (e.g., in compressed or encrypted form), packaged instructions (e.g., divided into multiple packages), or the like. Information in machine-readable medium 922 representing instructions 924 may be processed by processing circuitry into instructions to implement any of the operations discussed herein. For example, deriving instructions 924 from information (e.g., processed by processing circuitry) may include: compiling (e.g., from source code, object code, etc.), interpreting, loading, organizing (e.g., dynamically or statically linking), encoding, decoding, encrypting, decrypting, packaging, unpacking, or otherwise manipulating information into instructions 924.
[0098] In an example, the export of instructions 924 may include assembling, compiling, or interpreting information (e.g., by processing circuitry) to create instructions 924 from some intermediate or pre-processed format provided by machine-readable medium 922. When provided in multiple parts, the information may be combined, unpacked, and modified to create instructions 924. For example, the information may be in multiple compressed source code packages (or object code or binary executables, etc.) on one or several remote servers. The source code packages may be encrypted when transmitted over the network and decrypted at the local machine, decompressed, assembled (e.g., linked) as needed, and compiled or interpreted (e.g., compiled or interpreted into a library, a stand-alone executable, etc.), and executed by the local machine.
[0099] The instructions 924 may further be transmitted or received over a communication network 926 using a transmission medium via the network interface device 920 using any of a number of transmission protocols such as frame relay, Internet Protocol (IP), transmission control protocol (TCP), user datagram protocol (UDP), hypertext transfer protocol (HTTP), etc. Example communication networks may include a local area network (LAN), a wide area network (WAN), a packet data network such as the Internet, a mobile telephone network such as a cellular network, a plain old telephone (POTS) network, and a wireless data network such as the Institute of Electrical and Electronics Engineers (IEEE) 802.11 series of standards (known as ), IEEE 802.16 series of standards (called )), IEEE 802.15.4 family of standards, peer-to-peer (P2P) networks, etc. In an example, the network interface device 920 may include one or more physical jacks (such as Ethernet, coaxial or telephone jacks) or one or more antennas to connect to the network 926. In an example, the network interface device 920 may include multiple antennas to communicate wirelessly using at least one of single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple-input single-output (MISO) technology. The term "transmission media" should be taken to include any intangible media capable of storing, encoding or carrying instructions executed by the machine 900, and including digital or analog communication signals or other intangible media to facilitate the communication of such software. Transmission media is machine-readable media.
[0100] To illustrate the methods and apparatus discussed herein, a non-limiting group of example embodiments are set forth below as a digital recognition example.
[0101] Example 1 is a method comprising: receiving first data from a first portion of an array of memory devices; determining a number of correctable errors in the first data (e.g., performing a count of the number of correctable errors detected); and determining a relationship between the determined number of correctable errors and a specified threshold number of correctable errors. In Example 1, based on the relationship, the method includes one of: (a) performing a data scrubbing operation using the first data and the first portion of the memory array; or (b) inhibiting a data scrubbing operation using the first data and the first portion of the memory array. Determining the number of correctable errors may be performed using a RAS scheme (e.g., including using Reed-Solomon ECC).
[0102] In Example 2, the subject matter of Example 1 includes determining the number of correctable errors in the first data using a Reed-Solomon error correction code.
[0103] In Example 3, the subject matter of Examples 1-2 includes or uses the first portion of the memory device comprising a first die of a plurality of dies in an array of the memory device, and determining the number of correctable errors in the first data includes determining the number of correctable errors in the first die.
[0104] In Example 4, the subject matter of Examples 1 to 3 includes performing or disabling the data scrubbing operation includes performing the data scrubbing operation when the determined number of correctable errors is less than the specified threshold number of correctable errors and disabling the data scrubbing operation when the determined number of correctable errors meets or exceeds the specified threshold number of correctable errors.
[0105] In Example 5, the subject matter of Examples 1-4 includes determining that the first data includes a cluster of bit errors in the first data, wherein the cluster of bit errors corresponds to information from two or more adjacent cells, rows, or columns in a particular die of the memory device array.
[0106] In Example 6, the subject matter of Example 5 includes performing or disabling the data scrubbing operation includes disabling the data scrubbing operation when the first data includes the cluster of bit errors.
[0107] In Example 7, the subject matter of Examples 1-6 includes the first portion of the memory device array comprising a plurality of dies in the memory device array, and determining the number of correctable errors in the first data includes determining a respective number of correctable errors in each of the plurality of dies.
[0108] In Example 8, the subject matter of Example 7 includes determining a distribution of correctable errors in the plurality of dice and performing or inhibiting the data scrubbing operation using the first data based on the determined distribution.
[0109] In Example 9, the subject matter of Examples 1-8 includes: performing the data scrubbing operation using the first data from the first portion of the memory array; receiving second data from the same first portion of the memory device array; determining a second number of correctable errors in the second data; and in response to the second number of correctable errors in the second data satisfying the specified threshold number of correctable errors, inhibiting a second data scrubbing operation.
[0110] In Example 10, the subject matter of Example 9 includes notifying a host device that the first portion of the memory array contains an uncorrectable error.
[0111] In Example 11, the subject matter of Examples 1-10 includes: performing the data scrubbing operation using the first data from the first portion of the memory array; receiving second data from the same first portion of the memory array; determining a second number of correctable errors in the second data; and responsive to the second number of correctable errors in the second data being less than the specified threshold number of correctable errors, determining whether a repair criterion for the first portion of the memory array is satisfied. Example 11 may further include performing a post-package repair operation on the first portion of the memory array.
[0112] Example 12 is a system comprising a host device and a memory device coupled to the host device, wherein the memory device comprises a memory device controller configured to monitor correctable error accumulation in each of a plurality of regions of a memory array and conditionally trigger a scrubbing operation based on a number of correctable errors observed.
[0113] In Example 13, the subject matter of Example 12 includes or uses the memory device controller configured to perform the scrubbing operation on particular data read from a first region of the memory array when the error accumulation indication associated with the first region of the memory array is less than a threshold number of correctable errors.
[0114] In Example 14, the subject matter of Example 13 includes or uses the memory device controller configured to not perform the scrubbing operation on the particular data when the error accumulation associated with the first region of the memory array indicates at least the threshold number of correctable errors.
[0115] In Example 15, the subject matter of Examples 13-14 includes the first region of the memory array comprising a first die of a plurality of dies in the memory array.
[0116] In Example 16, the subject matter of Examples 13-15 includes the first region of the memory array comprising a first row or a first column of memory cells in the memory array.
[0117] In Example 17, the subject matter of Examples 13-16 includes the memory device controller being configured to monitor the correctable error accumulation of the same region of the memory array over multiple scrubbing operation cycles, and in response to identifying the same correctable error in the same region of the memory array over multiple scrubbing operation cycles, the memory device controller being configured to notify the host device that the memory array includes one or more regions for repair or remapping.
[0118] In Example 18, the subject matter of Examples 13-17 includes the memory device being coupled to the host device using a compute express link (CXL) interconnect.
[0119] Example 19 is a non-transitory processor-readable storage medium that includes instructions that, when executed by a processor circuit, cause the processor circuit to: read first data from a first portion of an array of a memory device; determine, using an error correction code decoder, a number of correctable errors present in the first data; in response to the number of correctable errors being less than a specified threshold number of correctable errors, perform a scrubbing operation using the first data, the scrubbing operation including determining corrected data based on the first data and writing the corrected data to the first portion of the array of the memory device; and in response to the number of correctable errors being greater than or equal to the specified threshold number of correctable errors, not allow the scrubbing operation.
[0120] In Example 20, the subject matter of Example 19 includes the processor-readable storage medium including further instructions, which, when executed by the processor circuit, cause the processor circuit to: read second data from the first portion of the array of the memory device; determine whether the second data includes the same number of correctable errors present in the first data; and in response to determining that the second data includes the same number of correctable errors, perform a repair operation on the first portion of the array of the memory device.
[0121] Example 21 is an apparatus comprising means for implementing any of Examples 1-20.
[0122] Example 22 is a system for implementing any of Examples 1-20.
[0123] Each of these non-limiting examples may stand alone or may be combined with one or more of the other examples discussed herein in various permutations or combinations.
[0124] The above detailed description includes reference to the accompanying drawings, which form a part of the detailed description. The drawings show by way of illustration specific embodiments in which the present invention can be practiced. These embodiments are also referred to herein as "examples". Such examples may include elements other than those shown or described. However, the inventors may also consider examples in which only those elements shown or described are provided. In addition, the inventors may also consider examples using any combination or arrangement of those elements (or one or more aspects thereof) shown or described with respect to a specific example (or one or more aspects thereof) or with respect to other examples (or one or more aspects thereof) shown or described herein.
[0125] In this document, the term "a / an", which is common in patent documents, is used to include one or more than one, independent of any other examples or uses of "at least one" or "one or more". In this document, the term "or" is used to refer to a non-exclusive or, so that "A or B" may include "A but not B", "B but not A", and "A and B", unless otherwise indicated. In the appended claims, the terms "including" and "in which" are used as the plain English equivalents of the corresponding terms "including" and "wherein". In addition, in the appended claims, the terms "including" and "comprising" are open-ended, that is, systems, devices, articles or processes that include elements other than the elements listed after this term in the claim are still considered to be within the scope of the claim. In addition, in the appended claims, the terms "first", "second", and "third", etc. are used only as labels and are not intended to impose numerical requirements on their objects.
[0126] The above description is intended to be illustrative rather than limiting. For example, the above examples (or one or more aspects thereof) may be used in combination with each other. For example, a person of ordinary skill in the art may use other embodiments after reviewing the above description. It should be understood that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the "Detailed Description", various features may be grouped together to simplify the present invention. This should not be interpreted as a desire not to claim that the disclosed features are an indispensable part of any claim. Specifically, the subject matter of the invention may have less than all the features of a particular disclosed embodiment. Therefore, the attached claims are hereby incorporated into the "Detailed Description", wherein each claim is independent as a separate embodiment, and it is expected that such embodiments can be combined with each other in various combinations or arrangements. The scope of the present invention should be determined with reference to the attached claims together with the full range of equivalents authorized by this claim.
Claims
1. A method comprising: receiving first data from a first portion of an array of memory devices; determining a number of correctable errors in the first data; and A relationship between the determined number of correctable errors and a specified threshold number of correctable errors is determined and based on the relationship, one of: performing a data scrubbing operation using the first data and the first portion of the memory array or inhibiting a data scrubbing operation using the first data and the first portion of the memory array. 2 . The method of claim 1 , wherein determining the number of correctable errors in the first data comprises using a Reed-Solomon error correction code.
3. The method of claim 1 , wherein the first portion of the memory device array comprises a first die of a plurality of dice in the memory device array, and wherein determining the number of correctable errors in the first data comprises determining the number of correctable errors in the first die.
4. The method of claim 1 , wherein performing or disabling the data scrubbing operation comprises performing the data scrubbing operation when the determined number of correctable errors is less than the specified threshold number of correctable errors and disabling the data scrubbing operation when the determined number of correctable errors meets or exceeds the specified threshold number of correctable errors.
5. The method of claim 1, further comprising determining that the first data includes a cluster of bit errors in the first data, wherein the cluster of bit errors corresponds to information from two or more adjacent cells, rows, or columns in a particular die of the memory device array. 6 . The method of claim 5 , wherein executing or disabling the data scrubbing operation comprises disabling the data scrubbing operation when the first data includes the cluster of bit errors.
7. The method of claim 1, wherein the first portion of the memory device array comprises a plurality of dice in the memory device array, and wherein determining the number of correctable errors in the first data comprises determining a respective number of correctable errors in each of the plurality of dice.
8. The method of claim 7, further comprising determining a distribution of correctable errors in the plurality of dice and performing or inhibiting the data scrubbing operation using the first data based on the determined distribution.
9. The method according to claim 1, further comprising: performing the data scrubbing operation using the first data from the first portion of the memory array; receiving second data from the same first portion of the array of memory devices; determining a second number of correctable errors in the second data; and In response to the second number of correctable errors in the second data satisfying the specified threshold number of correctable errors, a second data scrubbing operation is disabled.
10. The method of claim 9, further comprising notifying a host device that the first portion of the memory array contains uncorrectable errors.
11. The method according to claim 1, further comprising: performing the data scrubbing operation using the first data from the first portion of the memory array; receiving second data from the same first portion of the array of memory devices; determining a second number of correctable errors in the second data; and In response to the second number of correctable errors in the second data being less than the specified threshold number of correctable errors, determining whether repair criteria for the first portion of the memory array are satisfied and selectively performing a post-package repair operation on the first portion of the memory array.
12. A system comprising: Host device; and a memory device coupled to the host device, wherein the memory device includes a memory device controller configured to: monitoring correctable error accumulation in each of a plurality of regions of a memory array; and The scrubbing operation is conditionally triggered based on the number of correctable errors observed.
13. The system of claim 12, wherein the memory device controller is configured to perform the scrubbing operation on specific data read from a first region of the memory array when the error accumulation indication associated with the first region of the memory array is less than a threshold number of correctable errors.
14. The system of claim 13, wherein the memory device controller is configured to not perform the scrubbing operation on the particular data when the error accumulation associated with the first region of the memory array indicates at least the threshold number of correctable errors.
15. The system of claim 13, wherein the first region of the memory array comprises a first die of a plurality of dies in the memory array.
16. The system of claim 13, wherein the first region of the memory array comprises a first row or a first column of memory cells in the memory array.
17. The system of claim 13, wherein the memory device controller is configured to monitor the accumulation of correctable errors in the same region of the memory array over multiple scrubbing operation cycles, and in response to identifying the same correctable errors in the same region of the memory array over multiple scrubbing operation cycles, the memory device controller is configured to notify the host device that the memory array includes one or more regions for repair or remapping.
18. The system of claim 13, wherein the memory device is coupled to the host device using a compute express link (CXL) interconnect.
19. A non-transitory processor-readable storage medium comprising instructions that, when executed by a processor circuit, cause the processor circuit to: reading first data from a first portion of an array of a memory device; determining, using an error correction code decoder, a number of correctable errors present in the first data; in response to the number of correctable errors being less than a specified threshold number of correctable errors, performing a scrubbing operation using the first data, the scrubbing operation comprising determining corrected data based on the first data and writing the corrected data to the first portion of the array of the memory device; and In response to the number of correctable errors being greater than or equal to a specified threshold number of correctable errors, the scrubbing operation is not permitted.
20. The non-transitory processor-readable storage medium of claim 19, wherein the processor-readable storage medium comprises further instructions that, when executed by the processor circuit, cause the processor circuit to: reading second data from the first portion of the array of the memory device; determining whether the second data includes the same number of correctable errors present in the first data; and In response to determining that the second data includes the same number of correctable errors, a repair operation is performed on the first portion of the array of the memory device.