Data corruption tracking for memory reliability

The memory controller forces uncorrectable errors and uses a demand scrub circuit to enhance memory reliability in low-power systems like LPDDR5, addressing the challenge of error tracking and correction in non-server contexts.

JP2025100958AActive Publication Date: 2025-07-04APPLE INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025036593
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-01
Filing Date
2025-03-07
Publication Date
2025-07-04
Estimated Expiration
2043-01-12

AI Technical Summary

Technical Problem

Existing memory systems, particularly in non-server contexts, face challenges in tracking and recording memory errors efficiently, making it difficult to maintain data reliability due to power consumption and circuit area considerations, and existing techniques are not suitable for low-power memory technologies like LPDDR5.

Method used

A memory controller circuit forces uncorrectable errors by writing a specific data and parity combination, propagates a contamination indicator, and uses a demand scrub circuit to correct errors, thereby maintaining error tracking without additional hardware and minimizing power consumption.

Benefits of technology

This approach enhances memory reliability by accurately tracking and correcting errors in low-power memory systems like LPDDR5, reducing the likelihood of further errors and enabling efficient error handling with minimal power and area overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100958000001_ABST
    Figure 2025100958000001_ABST
Patent Text Reader

Abstract

To provide a device, a method, and a non-transitory computer-readable medium for improving memory reliability in a memory circuit.SOLUTION: A memory controller circuit communicates with a memory circuit via an interface supporting link error detection, and transmits a combination of data and parity for a first data block that causes the memory circuit to detect an uncorrectable write interface error based on a damaged indicator. Subsequent reading of the location may cause an uncorrectable error indication, but the memory controller circuit does not need additional tracking of the indicator by the memory circuit or a memory controller, and propagates the damaged indicator as the uncorrectable error in the memory circuit.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the reliability of computer memory, and more particularly to tracking corrupted data and recording data errors.

Background Art

[0002] The reliability of data stored in memory is important in various computing contexts. In the context of a data server, for example, various memory reliability features can be implemented using redundant storage / interfaces, extended ECC fields, etc. These techniques may not be appropriate in non-server contexts, for example, due to power consumption and circuit area considerations. However, improving memory reliability may still be desirable in these contexts. Data can pass through various circuits within a system-on-chip, and it can be difficult to track the state of corrupted data as the data moves through the system. Additionally, it can be difficult to efficiently track and record memory errors and their sources.

Brief Description of the Drawings

[0003]

Figure 1

[0004]

Figure 2

[0005]

Figure 3

[0006]

Figure 4

[0007]

Figure 5

[0008]

Figure 6

[0009]

Figure 7

[0010]

Figure 8A

[0011]

Figure 8B

[0012]

Figure 9

[0013]

Figure 10

[0014]

Figure 11

[0015]

Figure 12

[0016]

Figure 13

[0017]

Figure 14

[0018]

Figure 15

DETAILED DESCRIPTION OF THE INVENTION

[0019] In the disclosed embodiments discussed in detail below, the computing device is configured to track the corruption state of data passing through various circuits (this state may be referred to herein as a corruption indicator or “poison” indicator, which may indicate a detected error that is not correctable) and record the memory errors encountered.

[0020] In some embodiments, the memory controller forces uncorrectable errors when writing contaminated data to a memory circuit (e.g., DRAM) and maintains the contaminated status when later reading the same location from the memory circuit. This can enable the tracking of data stored in and then retrieved from the memory as contaminated without requiring a dedicated memory cell field or a separate memory controller tracking structure. Generally, the disclosed techniques can improve memory reliability with limited or negligible increases in area and power consumption.

[0021] In some embodiments, the memory is an LPDDR5 memory configured to detect both link errors and on-chip errors. The memory controller may, in this context, force write-link error correction code (ECC) errors and maintain a contamination status.

[0022] In some embodiments, the memory circuit is configured to correct correctable on-chip errors and indicate to the memory controller that a correctable error has been corrected, e.g., via a decoded status flag (DSF) interface. Note that a correctable error is a detected error for which error correction information (e.g., the use of an ECC mechanism) provides sufficient information to correct the error (as opposed to an uncorrectable error that is detected but for which there is not enough information to perform a correction). However, the memory circuit may correct data in-flight and leave incorrect data in the memory cell. Thus, in some embodiments, the demand scrub circuit is configured to initiate operations that cause internal read / correct / write operations in the memory circuit to correct stored data. This can reduce the likelihood that additional errors that make the data uncorrectable (e.g., a second bit flip in a memory that supports correction of single-bit flips but not multiple-bit flips) will occur.

[0023] In some embodiments, the device is configured to track and record non-correctable errors and correctable errors (e.g., via a separate table structure) and can take various actions based on various error thresholds. In some embodiments, the device includes a memory cache and a memory cache controller, and the memory cache controller is configured to track errors. In some embodiments, both correctable and non-correctable errors are also tracked at the source of the error (e.g., in certain processor clusters and their caches). The memory cache controller can aggregate error information and trigger various signals in response to specific thresholds. The various disclosed techniques can enable potential identification problems (e.g., a threshold count associated with a specific physical address can indicate a bad DRAM cell).

[0024] In various embodiments, the disclosed techniques can advantageously improve memory reliability in devices where server-grade memory reliability techniques are impractical. Overview of Memory System

[0025] FIG. 1 is a block diagram showing an exemplary memory system according to some embodiments. In the illustrated embodiment, system 100 includes a memory controller circuit 101 and a plurality of memory circuits 104-106 (note that in other embodiments, any of various numbers of memory circuits can be implemented, and the memory circuits can include various different numbers of banks per circuit). In the illustrated embodiment, memory controller circuit 101 is configured to communicate with memory circuit 104 via bus 108.

[0026] In the illustrated embodiment, the memory control circuit 101 receives access requests 109 via a plurality of virtual channels 110. In some embodiments, the virtual channels carry different types of requests and have different quality of service requirements. Requests from a particular agent may be sent via a particular virtual channel, or the agent may be configured to send requests via multiple different virtual channels. In some embodiments, as discussed in more detail below, the virtual channels include a real-time channel, a low-latency channel, and a bulk (or best-effort) channel.

[0027] In the illustrated embodiment, the memory controller circuit 101 includes a queue circuit 102, an arbitration circuit 103, and a priority tracking structure 111. In some embodiments, the queue circuit 102 is configured to queue received requests. In some embodiments, the arbitration circuit 103 is configured to select which requests are permitted to access a particular memory bank 107. In some embodiments, the arbitration circuit 103 is configured to use the information in the priority tracking structure 111 to determine which requests to permit.

[0028] In some embodiments, arbitration circuit 103 is configured to implement a category-based arbitration scheme. In some embodiments, a category value is assigned to each virtual channel (e.g., C0 to C3 in some embodiments, although other embodiments may implement any of a variety of numbers of categories). Arbitration circuit 103 can assign the category of each bank to each virtual channel. In some embodiments, each virtual channel starts with C3 for each bank, and arbitration circuit 103 is configured to prioritize the C3 channel over other channels. The least-recently-used (LRU) scheme can be used to select from among virtual channels having the same category for a bank. In some embodiments, a particular low category, such as C1 or C0, is always assigned to a particular low-priority virtual channel.

[0029] When a virtual channel wins arbitration and is permitted access to a particular bank, in some embodiments, memory controller 101 decrements the category of that bank (e.g., from C3 to C2, or from C2 to C1). In some embodiments, when a virtual channel is reduced below a particular level (e.g., below C2) for each bank for which the virtual channel has a request, memory controller 101 is configured to increment all categories for that virtual channel by one level (e.g., from C2 to C3). Note that when discussing "each" memory bank of a set of multiple memory banks herein, the disclosed techniques can be applied to the set of memory banks, but it should be noted that they are not necessarily applied to all memory banks within a device or system. For example, other memory banks within the same device may be controlled by other memory controllers or devices.

[0030] In some embodiments, the memory controller 101 also executes a credit system to permit a certain number of requests per virtual channel for a given read or write order, based on, for example, the bandwidth requested for, or allocated to, different virtual channels. This credit system can affect which virtual channels actually send requests to the arbitration circuit 103 during a given order.

[0031] In various embodiments, the category-based arbitration scheme can provide fair access to a given bank from among a plurality of virtual channels while cycling among the banks to avoid delays associated with accessing the same bank in quick succession.

[0032] Memory circuits 104 - 106 each include a plurality of banks 107a - n in the illustrated embodiments. Memory circuits 104 - 106 can be implemented using any of a variety of suitable memory technologies. Memory circuits 104 - 106 may need to be refreshed periodically, for example, if implemented as dynamic random access memory (DRAM). Further, it can be efficient to spread access requests across different banks, for example, because there can be a delay between successive accesses to different pages of the same bank. Thus, generally speaking, the arbitration circuit 103 attempts to permit access to one of a set of banks that have not been accessed within a threshold time interval. Exemplary override for propagating contamination indicators

[0033] FIG. 2 is a block diagram showing a memory controller circuit configured to force an uncorrectable error when writing to a memory circuit according to some embodiments. In the illustrated embodiment, the computing system includes a memory controller circuit 101 and a memory circuit 104. Note that these circuits are manufactured separately and can be connected during the assembly of the computing device. The memory controller circuit 101 includes a control circuit 230. The memory circuit 103 includes a write link ECC check circuit 220, a circuit 250 configured to write a data / parity combination for an uncorrectable error (UE) in the write link, a CE correction circuit 245, an error correction code (ECC) calculation circuit 255, and cells 260A to 260N.

[0034] The write link ECC check 220 includes, in the illustrated embodiment, a circuit configured to check the parity of the write data transmitted from the memory controller 101. For example, the circuit 220 can generate a parity value based on the received data and check that it matches the received parity value. As shown, the circuit 220 can indicate whether the data transmitted via the link exhibits an uncorrectable error (UE), a correctable error (CE), or no error (NE). In the case of a correctable write link error, the CE correction circuit 245 can correct the error. The circuit 255 is configured to generate ECC information for error-free data or data with corrected CE and store the data and parity information in the memory cells 260. Elements 245 and element 255 are shown using dashed lines and can be omitted in some embodiments. As shown, the memory circuit 104 is configured to store data and parity information in a given memory cell 260 (note that the data and parity information can be stored using sideband techniques or inline techniques depending on the memory technology of the memory circuit 104).

[0035] In the illustrated embodiment, circuit 250 is configured to handle irreparable errors on the write link. In particular, circuit 250 is configured to write to cell 260 a data parity combination that causes a UE when the cell is later read (e.g., by on-chip ECC checking discussed below with reference to FIG. 3). The data value and parity value may or may not match the actual data and parity received from the memory controller circuit via the link. The data and parity values written may be vendor specific and the original irreparable data need not be stored.

[0036] In some embodiments, control circuit 230 is configured to override the write link ECC and force irreparable errors in contaminated data. For example, control circuit 230 may write via the link a combination of data and parity that causes the write link ECC check circuit 220 to intentionally detect a UE. This can propagate a contamination indicator for the data that will remain when a UE subsequently occurs upon reading to the location. Note that the contaminated data may be corrupted in another circuit (e.g., a cache, a link between a processor and another element, etc.) and by tracking this corrupted data, inappropriate use of the corrupted data can be avoided. In this scenario, memory controller circuit 101 may not need to consider the actual value of the corrupted data.

[0037] In other embodiments, control circuit 230 may use other techniques to override the link ECC. For example, rather than providing a data / parity combination that exhibits a UE, control circuit 230 may assert a signal indicating an override, and memory circuit 104 may write a data / parity combination to cell 260 in response to detecting the override signal.

[0038] FIG. 3 is a block diagram showing an exemplary read link and on-chip error detection circuit according to some embodiments. In the illustrated embodiment, memory controller 101 includes a read link ECC check circuit 315, and memory circuit 104 includes a CE correction circuit 345 and a check ECC circuit 355.

[0039] In some embodiments, read link ECC check circuit 325 is configured to generate and check parity information similar to write link ECC check circuit 220. In some embodiments, an on-chip error or a read link error can be detected and reported by read link ECC check circuit 325, as discussed in detail below.

[0040] Check ECC circuit 355, in the illustrated embodiment, is configured to read data and parity information of memory cells, generate a parity value based on this data, and confirm that the parity values match. CE correction circuit 345, in the illustrated embodiment, is configured to correct the CE detected by circuit 355. The UE can be reported via a decode status flag (DSF) transmitted via the same interface as the link parity information. In some embodiments, the decode status flag enables indication of whether memory circuit 104 has detected an error in the memory cells. Thus, memory circuit 104 can indicate, via the DSF for a given location, a corrected CE, UE, or no error. Note that various elements of the device (e.g., SoC components) can similarly detect and correct CE.

[0041] For a UE being read, the memory controller circuit 101 may mark the data as contaminated. Similarly, the UE can detect in various circuits of the device and can bring about a contamination indication for the data in a circuit that supports such an indication. In the case of CE, the memory controller circuit 101 can trigger a demand scrub operation, as discussed in detail below with reference to FIG. 5.

[0042] For the purposes of explanation, various error detection and correction techniques discussed in this specification are included, but it should be noted that it is not intended to limit the scope of the present disclosure. In other embodiments, any of various suitable ECC schemes or parity schemes can be implemented. Generally speaking, in the context of an ECC scheme that supports correction of a CE having up to N incorrect bits, errors on more than N bits can correspond to the UE. Similarly, although separate parity lines and data lines are shown, in other embodiments, any of various suitable link interfaces can be implemented and these fields can share the interface.

[0043] In some cases, it should be noted that the memory circuit 104 can be the original source of an irreparable error that caused the contamination indicator in the memory controller 101. This can raise the question of whether the contamination indicator will propagate when data is rewritten to a known bad cell. If this is a soft error or a temporary memory error, the contaminated irreparable error can propagate when the data is rewritten to the cell. If this is a hard or permanent memory circuit error, there are two possibilities. First, the rewrite to the cell may still store the data / parity combination corresponding to the irreparable error, and the propagation of the contamination indicator is safe. Second, the cell may eventually store the data / parity combination corresponding to a repairable error, so there is a possibility that the contamination indication will not propagate when the cell is read. In this scenario, since the error in the memory circuit 104 initially caused the contamination indication, the overall error problem can be addressed by software, and the memory controller circuit 101 may have notified when generating the original contamination indication. Additionally, a hard memory failure can also be detected during a zeroing operation where all zeros are written to the memory location. Any detection technique can enable the operating system to take the page offline to avoid further errors, for example, due to hard memory errors or permanent memory errors. When the operating system takes the page offline in response to an irreparable error and is quite pessimistic, the irreparable error initiated by the memory is very unlikely to cause a failure that propagates the contamination indicator. Example of Propagation of Contamination Indicator

[0044] Overriding the link ECC is an example of contamination indicator propagation, but it should be noted that the contamination indicator can propagate across various circuit elements and overall operations, as discussed in detail below.

[0045] FIG. 4 is a block diagram showing an exemplary circuit within, for example, a SoC configured to propagate a contamination indication. In the illustrated example, the system includes a memory controller 101, a memory circuit 104, a memory cache controller 410, a fabric 420, and an agent 440.

[0046] Various agents, the memory cache controller 410, and the memory controller circuit 101 communicate via the fabric 420. In some embodiments, the fabric may include a field (e.g., a bit) for a contamination indicator of data transmitted via the fabric. This may enable circuits to propagate contamination indicators via the fabric. In other embodiments, the fabric 420 may not include a dedicated field for a contamination indicator, but various circuits may encode a contamination indicator within data transmitted via the fabric for decoding by a receiving circuit.

[0047] The memory cache controller 410 may control a memory cache that may be the cache furthest from one or more processors within the cache / memory hierarchy (e.g., one or more lower-level L1, L2, L3 caches, etc. may exist). The memory cache (not shown) may be configured to write evicted data to the memory circuit 104 and read data for cache misses from the memory circuit 104. The memory cache controller 410 may be configured to detect damaged data within the memory cache and mark that data as contaminated. The memory cache controller 410 may also be configured to maintain a contamination indicator for data damaged elsewhere before being stored in the memory cache.

[0048] Memory controller 101 can also generate a contamination indicator for data based on a match between the data's address and a channel address mask. This can enable, for example, the intentional insertion of various types of errors for debugging purposes, and errors (including CE and UE) can be injected when receiving data from memory or writing data to memory. Masking can be made to trigger over a range of addresses. This can be important for testing considering that CE is quite rare and UE is even rarer. Thus, error injection can facilitate the testing of various memory reliability features.

[0049] Memory controller 101 can also include a write queue field for tracking contamination indicators. Memory controller 101 can perform various operations on queued accesses to improve efficiency. For example, memory controller 101 can transfer write data from the write queue to a read queue entry for the same location and avoid accessing memory circuit 104 for reading. As another example, memory controller 101 can merge accesses to improve efficiency, avoid hazards (such as WAW, WARAW, etc.), or do both. In some embodiments, memory controller circuit 101 is configured to properly maintain contamination indicators through such operations.

[0050] Agent 440 can be various circuits such as a processor, a graphics processor, an I / O controller, etc. Agent 440 can similarly generate or maintain a contamination indicator for the data being processed.

[0051] Consider the following exemplary paths that data can take through the system. A data block can be flagged as contaminated by the memory cache controller 410 based on an error in the memory cache. The contamination indicator can be communicated to the memory controller 101 via the fabric 420 along with the write of the data to memory. The memory controller 101 can combine the contamination indicator with any contamination indicator generated due to a channel address mask (e.g., by indicating contamination if any contamination indicator is set). The memory controller 101 can propagate the contamination indicator to the write queue circuit along with the write data. In the case of any write-read transfer from a write queue entry to a read queue, the memory controller 101 can likewise propagate the contamination indicator. In the case of any access merge operation, the memory controller 101 can likewise propagate any contamination indicator for the merged data to the merge operation. Due to a write link override, data corresponding to an irreparable error may be stored in the memory cell. When read later, the memory controller circuit 101 can mark the data as contaminated in response to the detection of a DSF value for the irreparable error, and the contamination indicator can be propagated to various circuits within the system.

[0052] With respect to a memory controller circuit that maintains dedicated information regarding which memory cells are contaminated or a dedicated field within the memory cells to track this information, the disclosed technique can advantageously reduce the area and power consumption within the memory cache controller while accurately propagating the contamination indicator. Overview and Limitations of LPDDR5 Memory

[0053] The various techniques discussed herein may be particularly relevant in the context of LPDDR5 memory circuits, but it should be noted that similar techniques can be used with various memory technologies. In general, LPDDR5 memory can provide good performance for various applications (e.g., mobile devices) with relatively low power consumption. This memory technology and these applications may not be able to incorporate various memory reliability features implemented in other contexts such as server applications that incorporate large-scale redundancy and ECC functions. The following discussion presents some LPDDR5 features that may be relevant to the present disclosure.

[0054] The fifth generation of Low-Power Double Data Rate (LPDDR) SDRAM technology was first released in the first half of 2019. It inherits from the previous generation of LPDDR4 / 4X and provides speeds up to 6400 Mbps (1.5 times faster). Additionally, by implementing several power-saving developments, LPDDR5 can reduce power by up to 20% compared to the previous generation. LPDDR5 can provide a link ECC scheme, a scalable clocking architecture, multiple frequency-set points (FSPs), decision feedback equalization (DFE) to mitigate inter-symbol interference (ISI), a write X function, a flexible bank architecture, and in-line on-chip ECC. LPDDR5 systems typically do not provide server-level reliability features such as single-device data correction (SDDC), memory mirroring and redundancy, demand scrubbing, patrol scrubbing, data contamination, redundant links, clock and power monitoring / redundancy / failover, CE separation, online sparing with automatic failover, double device data correction (DDDC), etc. Exemplary demand scrub circuit

[0055] In some embodiments, the memory circuit 104 is configured to detect correctable errors in the memory cell data and correct the errors before providing the read data to the memory controller 101. However, the incorrect data may remain uncorrected in the memory cell. The likelihood of uncorrectable errors in such data may increase. For example, if the system is configured to correct single-bit errors but cannot correct multiple-bit errors (or more generally, errors exceeding a threshold number of bit errors), data that already exhibits correctable errors may be more likely to be further corrupted to exhibit uncorrectable errors.

[0056] Accordingly, in some embodiments, the memory controller circuit 101 is configured to execute a demand scrub to cause the memory circuit 104 to correct the data stored in the memory cells. The memory circuit 104 can support one or more types of write operations to efficiently perform the correction.

[0057] FIG. 5 is a block diagram showing an exemplary demand scrub circuit according to some embodiments. In the illustrated embodiment, the memory controller 101 includes a demand scrub circuit 510, and the demand scrub circuit 510 includes a snoop circuit 520 and a correction CE circuit 530. The remaining elements of FIG. 5 may be configured as described above for elements numbered the same as in the previous drawings.

[0058] In the illustrated embodiment, the demand scrub circuit 510 is configured to detect corrected errors from the memory circuit 104 and trigger the memory circuit 104 to correct the errors. Specifically, in the illustrated embodiment, the snoop circuit 510 is configured to snoop (monitor) the DSF status of the read operation executed by the memory controller 101. When the memory circuit 104 detects and corrects a CE, the DSF associated with the data indicates that the CE has been corrected. The DSF is an example of an encoding that can be used for LPDDR5, but is not intended to limit the scope of the present disclosure. Generally, the snoop circuit can utilize any of various suitable fields to determine when the memory circuit 104 has corrected an error in a memory cell without updating the memory cell to the corrected value. In some embodiments, the snoop circuit 520 collects the DRAM channel addresses corresponding to each detected CE.

[0059] In response to the detection of a CE, the snoop circuit 520 notifies the correction CE circuit 530, which triggers an internal correction within the memory circuit 104. In the illustrated example, the trigger is a fully masked partial write to the location presenting the CE, which causes an internal read / correction / write to that memory cell within the memory circuit 104 (without changing the correct value of the data). Generally, the memory circuit 104 can support commands such as a fully masked partial write operation that indicates reading a location, correcting the CE of that location, and writing the corrected value back to that location.

[0060] The disclosed technique may enable simplification of the memory circuit 104 for a memory circuit with built-in scrubbing while still providing a demand scrub function in some scenarios.

[0061] In some embodiments, multiple demand scrub corrections at the same location may indicate bad memory cells, and the operating system may take the corresponding page offline. However, in the case of temporary errors or soft errors, the demand scrub techniques discussed herein can reduce the rate at which CEs in the memory circuit cells become UEs.

[0062] The demand scrub function may be programmable, for example, to disable demand scrub. In some embodiments, demand scrub may not be performed in one or more modes in which the DSF is disabled. In some embodiments, the status of demand scrub may be locked so that it cannot be changed after boot.

[0063] Note that the demand scrub operation can be arbitrated with other access operations by the memory controller circuit 101. In some embodiments, the demand scrub operation has a relatively low quality of service (QoS) level or class compared to one or more other types of traffic, thereby reducing or avoiding interference with QoS for that traffic. In some situations, the demand scrub operation may be deferred. In some embodiments, the snoop circuit can track information regarding multiple CEs at once, but can only allow a threshold number of demand scrub operations to be in-flight at a given time (e.g., 1).

[0064] In some embodiments, the demand scrub circuit 510 includes a transfer progress counter that accumulates over time, and may increase the priority of the demand scrub operation when the transfer progress counter reaches a threshold.

[0065] In some embodiments, the demand scrub circuit 510 includes a timeout timer that can start when a demand scrub write is placed into the write queue and enforce the write order when the timeout timer reaches a threshold. The demand scrub circuit 510 can also disable demand scrubbing in response to certain operating conditions, such as when the write queue already has a threshold number of valid entries.

[0066] In some embodiments, the data associated with a demand scrub write is not software accessible (e.g., the data is internally read, corrected, and written to the memory circuit 104). In some embodiments, the demand scrub operation is not controlled by software but is fully hardware controlled (e.g., the snoop circuit 520 and the correction CE circuit 530 can operate according to a finite state machine).

[0067] In some embodiments, the demand scrub circuit 510 is configured to record demand scrub operations. For example, the demand scrub circuit 510 can include software accessible configuration registers that indicate the count of DSFs with CE status (which can be maintained independently for different lanes), the count of demand scrub writes that completed successfully, and the count of demand scrub writes that were deferred. These counters can be reset to zero by software, or both, at reset. In some embodiments, the counters are only available in a debug operation mode. As used herein, the term "software" broadly refers to program instructions executed by one or more processors and includes user applications, firmware, operating systems, and the like. Exemplary Error Tracking Techniques

[0068] FIG. 6 is a block diagram showing an exemplary memory cache controller configured to track and record correctable and uncorrectable errors and output software visible signaling, according to some embodiments. In the illustrated embodiment, memory cache controller 410 includes an uncorrectable error (UE) logger 610 and a correctable error (CE) tracker 620.

[0069] In the illustrated embodiment, UE logger 610 is configured to record detected memory errors and track certain information (e.g., physical address, error source, client identifier, etc., discussed in detail below). In the illustrated embodiment, UE logger 610 is specifically configured to record detected uncorrectable memory errors. In some embodiments, UE logger 610 tracks the source of uncorrectable errors. In some embodiments, UE logger 610 does not aggregate addresses and is not content addressable.

[0070] In the illustrated embodiment, CE tracker 620 is configured to record detected memory errors and track certain information (e.g., physical address, count of errors at that address, client identifier, etc.). In the illustrated embodiment, CE tracker 620 is specifically configured to record detected correctable memory errors. In some embodiments, CE tracker 620 implements a count field indicating the number of correctable errors that have occurred corresponding to a given physical address. In some embodiments, CE tracker 620 aggregates addresses and is content addressable.

[0071] In the illustrated embodiment, the memory cache controller 410 is configured to generate software visible signal(s). These signals can notify the software of the tracker / logger content that a threshold regarding the content has been met, or generally, can indicate to the software that certain actions (e.g., clearing an entry, marking data as contaminated, taking a page offline, etc.) may need to be taken.

[0072] Note that in other embodiments, the device may implement the disclosed recording / tracking circuitry elsewhere, in addition to or instead of the memory cache controller 410. However, tracking in the memory cache controller 410 can be particularly advantageous since the memory cache controller can operate using physical memory channel addresses. This information may not be available to other circuits, and thus, tracking in the memory cache controller can provide detailed information to the software without the need to transmit this information to other circuit elements.

[0073] Generally, the disclosed tracking structure advantageously can provide software with various useful information not available in conventional implementations, whereby the software can take appropriate corrective actions when errors are detected.

[0074] FIG. 7 is a diagram showing an exemplary UE logger data structure configured to record irreparable memory errors according to some embodiments. In the illustrated embodiment, the exemplary UE logger data structure 610 includes a valid field, a physical address field, a client identifier field, and an error source field.

[0075] In the illustrated embodiment, the valid field indicates whether the data entry is valid. In some embodiments, all entries within the UE logger data structure 610 are initially set to invalid.

[0076] In the illustrated embodiment, the physical address field contains memory address information regarding a data entry that enables the data bus to access a particular memory cell of the memory. This information can be particularly useful when the memory cell is the source of an error.

[0077] In the illustrated embodiment, the client identifier field identifies the client circuit within the SoC that previously accessed the data. For example, this field can indicate the fabric identifier of the client to the communication fabric.

[0078] In the illustrated embodiment, the error source field contains address information regarding a data entry that identifies the source of a memory error. Non-limiting exemplary error sources that can be encoded include UEs from DRAM reads, memory cache read data with uncorrectable errors (based on error checks or contamination indicators), or snoop response contaminated data (e.g., when a snoop to another cache determines that another cache controls the location and marks the data as contaminated).

[0079] In some embodiments, when there are no free entries in the UE logger and a UE is detected, an overflow signal (e.g., a bit) is asserted. In some embodiments, the overflow bit can be sticky and persistent until it is cleared (e.g., via a write-1-to-clear operation). Software can initiate a correction action based on the overflow signal to reduce the risk of damage associated with not being able to record subsequent UEs.

[0080] In some embodiments, software can invalidate an entry, for example, via a write-1-to-clear operation, after reading the entry from the UE logger data structure 610.

[0081] FIG. 8A is a diagram showing an exemplary CE tracker data structure configured to track correctable memory errors, according to some embodiments. In the illustrated embodiment, the exemplary CE tracker data structure 620 includes a valid field, a physical address field, a client identifier field, and a count field.

[0082] The valid field, the physical address field, and the client identifier field can track information similar to that described above in the context of the UE logger data structure 610. In some embodiments, the CE tracker 620 utilizes a content addressable memory (CAM) structure in which at least a portion of the physical address is used as a tag to determine whether there is a hit on a valid entry and increment its count, as discussed below with reference to FIG. 9.

[0083] In the illustrated embodiment, the count field indicates the number of correctable errors detected for each physical address within the interval since its entry was last cleared.

[0084] FIG. 8B is a block diagram showing an exemplary memory cache controller configured to track CE errors and output a signal based on whether a particular threshold is met or exceeded. In the illustrated embodiment, the memory cache controller 410 includes a CE tracker 620 and outputs a first signal corresponding to a valid occupancy threshold and a second signal corresponding to a count threshold.

[0085] In the illustrated embodiment, the control circuit is configured to assert a signal indicating the valid occupancy threshold when the number of valid entries in the CE tracker 620 meets the threshold. Note that "meeting" the threshold can correspond to being equal to the threshold or exceeding the threshold (e.g., having a value one step greater or one step less than the threshold) in different implementations.

[0086] In the illustrated embodiment, the control circuit is configured to assert a signal indicating the count threshold when the count field of a particular physical address in the CE tracker 620 reaches a value that satisfies the count threshold.

[0087] Software can perform various corrective actions based on these signals, which can include stopping a particular activity when a valid occupancy threshold is met or accessing one or more CE tracker entries when the count threshold is met. Exemplary techniques for allocating and deallocating CE tracker entries

[0088] FIG. 9 is a flowchart showing an exemplary method for allocating a new CE. The method shown in FIG. 9 can be used, among other things, in conjunction with any of the computer circuits, systems, devices, elements, or components disclosed herein. In various embodiments, some of the illustrated method elements may be executed simultaneously, may be executed in an order different from that shown, or may be omitted. Additional method elements may be executed as needed.

[0089] At 910, in the illustrated embodiment, the control circuit (e.g., of the memory cache controller 410) receives a new CE.

[0090] At 920, in the illustrated embodiment, the control circuit determines whether the new CE hits or misses in the CE tracker. If it hits, the flow proceeds to 950; if it misses, the flow proceeds to 930.

[0091] At 930, in the illustrated embodiment, in the case of a miss in the CE tracker, the control circuit allocates an entry in the CE tracker for the new CE and initializes its count (e.g., to 1 or a default value).

[0092] At 940, in the illustrated embodiment, the control circuit determines whether the occupancy threshold is met (e.g., if the number of valid entries in the CE tracker meets the occupancy threshold after entries are assigned at 930). If so, the control circuit asserts a signal indicating that the valid occupancy threshold has been met.

[0093] In some embodiments, in response to the signal, software can take a snapshot of the visible valid entries, clear the entries, and clear the space in the CE tracker. In some embodiments, when there are no free entries in the CE tracker, new CEs may not be tracked. Note that in some situations, an entry may not be visible to software. For example, the control circuit may allow software to access all or some of the entries only after one of the disclosed thresholds is hit.

[0094] At 950, in the illustrated embodiment, in response to a hit in the CE tracker, the control circuit increments the count value for the hit entry and updates the client identifier of that entry to the latest client associated with the error. In other embodiments, the client identifier field may track multiple client identifiers, and the control circuit may add the latest client identifier to the list of identifiers.

[0095] At 960, in the illustrated embodiment, the control circuit determines whether the count threshold is met due to the increment at 950. If so, the control circuit asserts a signal indicating that the count threshold has been met. In some embodiments, such a signal can warn software of a potential bad DRAM cell, whereby software can take various actions such as taking the page containing the cell offline.

[0096] FIG. 10 is a flow diagram illustrating an exemplary method for deallocating CE tracker entries. The method shown in FIG. 10 can be used, among other things, with any of the computer circuits, systems, devices, elements, or components disclosed herein. In various embodiments, some of the illustrated method elements may be executed simultaneously, may be executed in an order different from that shown, or may be omitted. Additional method elements may be executed as necessary.

[0097] At 1010, in the illustrated embodiment, the control circuit determines whether the CE tracker is accessible by software. If so, the flow proceeds to 1020. Otherwise, the control circuit may not perform further operations.

[0098] At 1020, in the illustrated embodiment, upon verifying that the CE tracker is accessible by software, the control circuit reads one or more entries. In some embodiments, a protocol is initiated to take a snapshot of all visible valid entries within the CE tracker structure.

[0099] At 1030, in the illustrated embodiment, the control circuit determines whether to deallocate one or more entries within the CE tracker. In some embodiments, deallocation is performed by software, for example, using a write-1-to-clear mechanism.

[0100] In some embodiments, deallocation of one or more entries within the CE tracker is at the discretion of the software. The software has the option not to deallocate the entries. According to some embodiments, the software can move the CE tracker information to another data structure to make space available within the CE tracker. This can be useful, for example, in situations where there are a fairly large number of unique CE addresses or when a threshold is reduced. Exemplary method

[0101] FIG. 11 is a flow diagram illustrating an exemplary method for tracking damaged data according to some embodiments. The method shown in FIG. 11 can be used, among other things, with any of the computer circuits, systems, devices, elements, or components disclosed herein. In various embodiments, some of the illustrated method elements may be executed simultaneously, may be executed in an order different from that shown, or may be omitted. Additional method elements may be executed as needed.

[0102] At 1110, in the illustrated embodiment, the memory controller circuit communicates with the memory circuit via an interface. The memory circuit may implement both link error correction and on-die error correction. In some embodiments, the memory circuit supports interface error detection (e.g., write link ECC) that causes a combination of data and parity to be written to the target memory location for detected uncorrectable write interface errors, and this combination corresponds to uncorrectable errors.

[0103] At 1120, in the illustrated embodiment, the memory controller circuit arbitrates between requests to access the memory circuit from a request agent circuit, including a first request to write first data to a first location within the memory circuit.

[0104] At 1130, in the illustrated embodiment, the memory controller circuit maintains a damage indicator for a data block, including a first damage indicator indicating that the first data has been determined to be damaged. In some embodiments, one of the agent circuits is configured to generate the first damage indicator, for example, based on a detected UE.

[0105] In some embodiments, a device that includes a memory controller circuit is configured to maintain a corruption indicator through a plurality of operations that include any combination of the following operations. That is, propagation of the corruption indicator after merging one or more requests to resolve a hazard, propagation of the corruption indicator for a write-read transfer operation from a write queue, conversion of the corruption indicator to a non-correctable write interface error, communication of the corruption indicator from a memory cache controller circuit to the memory controller circuit, and propagation of the corruption indicator determined based on an address mask.

[0106] At 1140, in the illustrated embodiment, the memory controller circuit transmits a combination of data and parity for a first data block that causes the memory circuit to detect a non-correctable write interface error.

[0107] At 1150, in the illustrated embodiment, the memory controller circuit reads a memory location following a write to a first request and generates a corruption indicator for the read data in response to a report of a non-correctable error from the memory circuit for the read data.

[0108] In some embodiments, the demand scrub circuit detects corrected errors indicated by a memory circuit where incorrect data remains stored in the memory cells of the memory circuit, and in response to the detection of the corrected errors, initiates a demand scrub write operation to the memory circuit that causes an internal read, error correction of correctable errors, and writing of the corrected data to the memory circuit. In some embodiments, the write operation is a fully masked partial write operation to the detected DRAM address of the corrected error. In some embodiments, the demand scrub circuit is configured to record in one or more software-accessible registers the number of detected correctable errors and the number of successful demand scrub writes. In some embodiments, the detection of the corrected error is based on a decoded status flag reported by the memory circuit indicating whether the provided data has no errors, has correctable errors, or has uncorrectable errors.

[0109] In some embodiments, the memory circuit includes an error circuit configured to verify parity information of error-free write data for a write operation and correct detected correctable errors associated with the interface. In some embodiments, the memory circuit includes an error circuit configured to correct detected errors associated with a read location for a read operation and report detected uncorrectable errors associated with the read location via the interface.

[0110] FIG. 12 is a flow diagram illustrating an exemplary method for tracking the number of detected correctable errors according to some embodiments. The method shown in FIG. 12 can be used, among other things, in conjunction with any of the computer circuits, systems, devices, elements, or components disclosed herein. In various embodiments, some of the illustrated method elements may be executed simultaneously, may be executed in an order different from that shown, or may be omitted. Additional method elements may be executed as needed.

[0111] At 1210, in the illustrated embodiment, data operated on by one or more processors within the memory cache is cached.

[0112] At 1220, in the illustrated embodiment, the number of detected correctable errors associated with each of a plurality of locations is tracked using a plurality of tracking circuit entries.

[0113] At 1230, in the illustrated embodiment, in response to detecting the number of correctable errors for a particular location, a signal identifying the particular location is generated for one or more processors.

[0114] In some embodiments, a signal identifying a particular location is asserted to indicate that a count threshold has been hit. This signal can warn software that there is a page potentially having defective DRAM that may be near a failure.

[0115] In some embodiments, an alert signal is generated in response to the number of valid entries within a tracking circuit entry matching or exceeding an occupancy threshold. In some embodiments, in response to matching or exceeding the occupancy threshold, software is enabled to access one or more tracking circuit entries.

[0116] In some embodiments, one or more of the tracking circuit entries can be deallocated in response to software signaling.

[0117] In some embodiments, a plurality of circuit entries include respective client identifier fields indicating a client associated with a given correctable error. In some embodiments, detected UEs associated with each of a plurality of locations of data are tracked using a plurality of UE tracking circuit entries.

[0118] In some embodiments, a UE trace circuit entry includes a source field that identifies the source of a given UE. In some embodiments, the source field is configured to encode a source that includes at least a memory error, a memory cache error, and a source of a snoop response. In some embodiments, a plurality of UE trace circuit entries are not tagged, and a plurality of trace circuit entries are tagged with at least a portion of an address for a given location.

[0119] In some embodiments, the device is configured to maintain a corruption indicator for a data block, the corruption indicator indicating that the data block has been determined to be corrupted. Exemplary Device

[0120] Referring now to FIG. 13, a block diagram illustrating an exemplary embodiment of device 1300 is shown. In some embodiments, the elements of device 1300 may be included within a system-on-chip. In some embodiments, device 1300 may be included in a battery-powered mobile device. Thus, power consumption by device 1300 can be an important design consideration. In the illustrated embodiment, device 1300 includes fabric 1310, compute complex 1320, input / output (I / O) bridge 1350, cache / memory controller 1345, graphics unit 13135, and display unit 1365. In some embodiments, in addition to or instead of the illustrated components, device 1300 may include other components (not shown) such as a video processor encoder and decoder, an image processing element or recognition element, a computer vision element, etc.

[0121] Fabric 1310 may include various interconnections, buses, MUXes, controllers, etc., and may be configured to facilitate communication between various elements of device 1300. In some embodiments, portions of Fabric 1310 may be configured to implement various different communication protocols. In other embodiments, Fabric 1310 may implement a single communication protocol, and elements coupled to Fabric 1310 may internally convert from a single communication protocol to other communication protocols.

[0122] In the illustrated embodiment, compute complex 1320 includes bus interface unit (BIU) 1325, cache 1330, and cores 1335 and 1340. In various embodiments, compute complex 1320 may include various numbers of processors, processor cores, and caches. For example, compute complex 1320 may include 1, 2, or 4 processor cores, or any other suitable number. In one embodiment, cache 1330 is a set associative L2 cache. In some embodiments, cores 1335 and 1340 may include internal instruction and / or data caches. In some embodiments, a coherence unit (not shown) within Fabric 1310, cache 1330, or other locations within device 1300 may be configured to maintain coherence between various caches of device 1300. BIU 1325 may be configured to manage communication between compute complex 1320 and other elements of device 1300. Processor cores such as cores 1335 and 1340 may be configured to execute instructions of a particular instruction set architecture (ISA) that may include operating system instructions and user application instructions.

[0123] The cache / memory controller 1345 may be configured to manage the transfer of data between the fabric 1310 and one or more caches and / or memories. For example, the cache / memory controller 1345 may be coupled to the L3 cache, which may in turn be coupled to the system memory. In other embodiments, the cache / memory controller 1345 may be directly coupled to the memory. In some embodiments, the cache / memory controller 1345 may include one or more internal caches.

[0124] As used herein, the term "coupled" can indicate one or more connections between elements, and the coupling may include intervening elements. For example, in FIG. 13, the graphics unit 1375 may be described as "coupled" to the memory via the fabric 1310 and the cache / memory controller 1345. In contrast, in the illustrated embodiment of FIG. 13, since there are no intervening elements, the graphics unit 1375 is "directly coupled" to the fabric 1310.

[0125] The graphics unit 1375 may include one or more processors, such as one or more graphics processing units (GPUs). The graphics unit 1375 can receive graphics-oriented instructions, such as OPENGL (registered trademark), Metal, or DIRECT3D (registered trademark) instructions. The graphics unit 1375 may execute specialized GPU instructions or perform other operations based on the received graphics-oriented instructions. The graphics unit 1375 may generally be configured to process large blocks of data in parallel, may construct an image in a frame buffer for output to a display, and the display may be included in the device or may be a separate device. The graphics unit 1375 may include conversion, lighting, triangle, and rendering engines in one or more graphics processing pipelines. The graphics unit 1375 can output pixel information for a displayed image. In various embodiments, the graphics unit 1375 may include a programmable shader circuit that can include highly parallel execution cores configured to execute a graphics program, which may include pixel tasks, vertex tasks, and compute tasks (which may or may not be graphics-related).

[0126] The display unit 1365 may be configured to read data from a frame buffer and provide a stream of pixel values for display. In some embodiments, the display unit 1365 can be configured as a display pipeline. Additionally, the display unit 1365 may be configured to blend multiple frames to generate an output frame. Further, the display unit 1365 may include one or more interfaces (e.g., MIPI (registered trademark) or embedded display port (eDP)) for coupling to a user display (e.g., a touch screen or an external display).

[0127] The I / O bridge 1350 may include various elements configured to implement, for example, Universal Serial Bus (USB) communication, security, audio, and / or low-power always-on functions. The I / O bridge 1350 may also include interfaces such as, for example, Pulse Width Modulation (PWM), General Purpose Input / Output (GPIO), Serial Peripheral Interface (SPI), and Inter-Integrated Circuit (I2C). Various types of peripheral devices and devices may be connected to the device 1300 via the I / O bridge 1350.

[0128] In some embodiments, the device 1300 may include a network interface circuit (not explicitly shown) that can be connected to the fabric 1310 or the I / O bridge 1350. The network interface circuit may be configured to communicate via various networks that can be wired, wireless, or both. For example, the network interface circuit may be configured to communicate via a wired local area network, a wireless local area network (e.g., via WiFi), or a wide area network (e.g., the Internet or a virtual private network). In some embodiments, the network interface circuit is configured to communicate via one or more cellular networks using one or more wireless access technologies. In some embodiments, the network interface circuit is configured to communicate using device-to-device communication (e.g., Bluetooth® or WiFi Direct), etc. In various embodiments, the network interface circuit may provide the device 1300 with connectivity to various types of other devices and networks.

[0129] The various elements of FIG. 13 can utilize the disclosed techniques. For example, the memory cache controller 410, the memory controller circuit 101, or both can be included in element 1345. The fabric 1310 can support a damage indicator. Various agent circuits such as the graphics unit 1375, the compute complex 1320, etc. can detect data contamination and propagate a contamination indicator. Advantageously, the disclosed techniques can improve memory reliability in various embodiments. Exemplary Uses

[0130] Referring now to FIG. 14, various types of systems are shown that can include any of the circuits, devices, or systems described above. A system or device 1400 that can incorporate or otherwise utilize one or more of the techniques described herein can be used in a wide range of areas. For example, the system or device 1400 can be used as part of the hardware of a system such as a desktop computer 1410, a laptop computer 1420, a tablet computer 1430, a cellular or mobile phone 1440, or a television 1450 (or a set-top box connected to a television).

[0131] Similarly, the disclosed elements can be used in wearable devices 1460 such as smartwatches or health monitoring devices. A smartwatch can implement various different functions in many embodiments, such as access to email, cellular service, a calendar, health monitoring, etc. A wearable device can also be designed to perform only health monitoring functions such as monitoring a user's vital signs, performing epidemiological functions such as contact tracing, and providing communication to emergency medical services. Other types of devices are also contemplated, including glasses or helmets designed to provide a computer-generated reality experience, such as devices worn around the neck, devices implantable in the human body, and those based on augmented and / or virtual reality.

[0132] System or device 1400 may also be used in a variety of other contexts. For example, system or device 1400 may be utilized in the context of a server computer system such as a dedicated server or shared hardware that executes a cloud-based service 14130. Further, system or device 1400 may be implemented in a wide range of dedicated everyday devices including devices 1480 commonly found in the home such as refrigerators, thermostats, security cameras, etc. The interconnection of such devices is often referred to as the "Internet of Things" (IoT). The elements may also be implemented in various forms of transportation. For example, system or device 1400 may be used in control systems, guidance systems, entertainment systems, etc. of various types of vehicles 1490.

[0133] The uses shown in FIG. 14 are merely illustrative and are not intended to limit the potential future uses of the disclosed system or device. Other exemplary uses include, but are not limited to, portable game devices, music players, data storage devices, unmanned aerial vehicles, etc. Exemplary computer-readable media

[0134] The present disclosure has been described in more detail above with respect to various exemplary circuits. The present disclosure is intended to cover not only embodiments including such circuits, but also computer-readable storage media containing design information specifying such circuits. Accordingly, the present disclosure is intended to support claims that cover not only devices including the disclosed circuits, but also storage media that specify circuits in a format recognized by a manufacturing system configured to generate hardware (e.g., integrated circuits) including the disclosed circuits. Claims for such storage media are intended to cover, for example, entities that generate circuit designs but do not themselves manufacture the designs.

[0135] FIG. 15 is a block diagram showing an exemplary non-transitory computer-readable storage medium storing circuit design information according to some embodiments. In the illustrated embodiment, semiconductor manufacturing system 1520 is configured to process design information 1515 stored in non-transitory computer-readable medium 1510 and manufacture integrated circuit 1530 based on design information 1515.

[0136] The non-transitory computer-readable storage medium 1510 may include any of various suitable types of memory devices or storage devices. The non-transitory computer-readable storage medium 1510 may be an installation medium, e.g., a CD-ROM, floppy disk, or tape device, computer system memory or random access memory such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, flash, magnetic media such as a hard drive, or an optical storage device, non-volatile memory, a register, or other similar types of memory elements. The non-transitory computer-readable storage medium 1510 may also include other types of non-transitory memory, or combinations thereof. The non-transitory computer-readable storage medium 1510 may include two or more memory media that may exist at different locations, e.g., different computer systems connected through a network.

[0137] Design information 1515 can be specified using any of a variety of suitable computer languages, including but not limited to hardware description languages such as VHDL, Verilog, SystemC, SystemVerilog, RHDL, M, and MyHDL. Design information 1515 can be used by semiconductor manufacturing system 1520 to manufacture at least a portion of integrated circuit 1530. The format of design information 1515 can be recognized by at least one semiconductor manufacturing system 1520. In some embodiments, design information 1515 may also include one or more cell libraries that specify the synthesis, layout, or both of integrated circuit 1530. In some embodiments, the design information is specified, in whole or in part, in the form of a netlist that specifies cell library elements and their connectivity. Design information 1515 may or may not contain sufficient information to manufacture the corresponding integrated circuit by itself. For example, design information 1515 may specify the circuit elements to be manufactured, but not their physical layout. In this case, design information 1515 may need to be combined with layout information to actually manufacture the specified circuit.

[0138] Integrated circuit 1530 can include, in various embodiments, one or more custom macrocells such as memories, analog or mixed-signal circuits, etc. In such cases, design information 1515 may include information related to the macrocells included. Such information includes, but is not limited to, circuit diagram capture databases, mask design data, behavioral models, and device or transistor level netlists. As used herein, the mask design data may be formatted according to Graphics Data System (GDSII), or any other suitable format.

[0139] The semiconductor manufacturing system 1520 may include any of a variety of suitable elements configured to manufacture integrated circuits. This may include, for example, depositing semiconductor materials (e.g., on a wafer, which may include masking), removing materials, changing the shape of the deposited materials, modifying the materials (e.g., by doping the materials or changing the dielectric constant using ultraviolet treatment), and the like. The semiconductor manufacturing system 1520 may also be configured to perform various tests on the circuits manufactured for proper operation.

[0140] In various embodiments, the integrated circuit 1530 is configured to operate according to a circuit design specified by the design information 1515, which may include performing any of the functions described herein. For example, the integrated circuit 1530 may include any of the various elements shown in FIGS. 1-8 or FIG. 13. Further, the integrated circuit 1530 may be configured to perform the various functions described herein in conjunction with other components. Further, the functions described herein may be performed by a plurality of connected integrated circuits.

[0141] As used herein, the phrase "design information specifying the design of a circuit configured to" does not mean that the circuit in question must be fabricated in order for the element to be satisfied. Rather, this phrase indicates that the design information describes a circuit that, when fabricated, is configured to perform the indicated action or includes the specified components. ***

[0142] This disclosure includes references to a group of "one embodiment" or "embodiments" (e.g., "some embodiments" or "various embodiments"). Embodiments are different implementations or examples of the disclosed concepts. References to "one embodiment", "an embodiment", "a particular embodiment", etc. do not necessarily refer to the same embodiment. A number of possible embodiments including those specifically disclosed, as well as modifications or alternatives within the spirit or scope of this disclosure are contemplated.

[0143] This disclosure can discuss potential advantages that can arise from the disclosed embodiments. Not all implementations of these embodiments necessarily exhibit any or all of the potential advantages. Whether an advantage is realized for a particular implementation depends on many factors, some of which are outside the scope of this disclosure. In fact, there are many reasons why an implementation within the scope of the claims may not exhibit some or all of the disclosed advantages. For example, a particular implementation may include other circuitry outside the scope of this disclosure that, in combination with one of the disclosed embodiments, invalidates or reduces one or more of the disclosed advantages. Additionally, sub-optimal design implementation of a particular implementation (e.g., implementation techniques or tools) can also invalidate or reduce the disclosed advantages. Even assuming skilled implementation, the realization of advantages can still depend on other factors such as the environmental circumstances in which the implementation is deployed. For example, the input supplied to a particular implementation can prevent one or more of the problems addressed in this disclosure from occurring in a particular opportunity, and as a result, the benefits of the solution may not be realized. Considering the existence of possible external factors of this disclosure, it is clearly intended that any potential advantages described herein should not be construed as limitations of the claims that must be met to demonstrate infringement. Rather, the identification of such potential advantages is intended to illustrate the types of improvements available to designers who benefit from this disclosure. The fact that such advantages are permissively described (e.g., a particular advantage is described as "can occur") is not intended to convey doubt as to whether such advantages can actually be realized, but rather to recognize the technical reality that the realization of such advantages often depends on additional factors.

[0144] Unless otherwise specified, the embodiments are non-limiting. That is, even if only a single example is described with respect to a particular feature in the disclosed embodiments, it is not intended to limit the scope of the claims made based on the present disclosure. The disclosed embodiments are intended to be illustrative rather than limiting in the absence of a contrary description in the present disclosure. The above description is intended to enable claims that cover not only the disclosed embodiments but also alternatives, modifications, and equivalents that will be apparent to those skilled in the art who benefit from the present disclosure.

[0145] For example, the features of the present application can be combined in any suitable manner. Accordingly, new claims can be formulated during the examination procedure of the present application (or an application claiming priority to the present application) for any such combination of features. In particular, referring to the appended claims, the features from the dependent claims can be combined, as appropriate, with the features of other dependent claims, including claims that depend on other independent claims. Similarly, the features from each independent claim can be combined as appropriate.

[0146] Accordingly, the appended dependent claims can be drafted such that each depends on a single other claim, but additional dependencies are contemplated. Any combination of features in the dependent claims consistent with the present disclosure is contemplated and can be claimed in the present application or another application. In summary, the combinations are not limited to those specifically recited in the appended claims.

[0147] Claims drafted in one format or statutory type (e.g., apparatus) are also intended to support corresponding claims in another format or statutory type (e.g., method), as appropriate. ***

[0148] Since this disclosure is a legal document, various terms and phrases may be subject to administrative and judicial interpretation. The following paragraphs, as well as the definitions provided throughout this disclosure, publicly notify how they are to be used in interpreting the claims made based on this disclosure.

[0149] References to singular items (i.e., nouns or noun phrases preceded by "a", "an", or "the") are intended to mean "one or more" unless the context specifically indicates otherwise. Thus, references to "an item" in the claims do not exclude additional instances of the item without context. "A plurality of" items refers to a set of two or more items.

[0150] The word "may" is used herein in the sense of permission (i.e., having the possibility, being possible), and not in the sense of obligation (i.e., not being mandatory).

[0151] The terms "comprising" and "including" and their forms are open-ended and mean "including but not limited to".

[0152] When the term "or" is used in this disclosure with respect to a list of alternatives, it will generally be understood to be used in an inclusive sense unless the context specifically indicates otherwise. Thus, the listing of "x or y" is equivalent to "x or y, or both", and thus includes 1) x but not y, 2) y but not x, and 3) both x and y. On the other hand, the phrase "either x or y, but not both" makes it clear that "or" is being used in an exclusive sense.

[0153] The enumeration of "w, x, y, z, or any combination thereof", or "at least one of... w, x, y, and z" is intended to cover all possibilities including single elements up to the total number of elements in the set. For example, in the case of the set [w, x, y, z], these expressions cover any single element of the set (e.g., w but not x, y, or z), any two elements (e.g., w and x but not y or z), any three elements (e.g., w, x, and y but not z), and all four elements. Thus, the phrase "at least one of... w, x, y, and z" refers to at least one element of the set [w, x, y, z], thereby covering all possible combinations in the list of these elements. This phrase should not be interpreted as requiring that there be at least one instance of w, at least one instance of x, at least one instance of y, and at least one instance of z.

[0154] In the present disclosure, various "labels" may precede a noun or noun phrase. Unless otherwise explicitly indicated in the context, the various labels used for features (e.g., "first circuit", "second circuit", "specific circuit", "given circuit", etc.) refer to different examples of the feature. Further, when applied to a feature, the labels "first", "second", and "third" do not, unless otherwise specified, imply any type of order (e.g., spatial, temporal, logical, etc.).

[0155] As used herein, the phrase "based on" is used to describe one or more factors that affect a determination. This term does not exclude the possibility that additional factors may affect the determination. That is, the determination may be based only on the specified factors, or may be based on the specified factors as well as other unspecified factors. Consider the phrase "determine A based on B". This phrase identifies that B is a factor used to determine A or that affects the determination of A. This phrase does not exclude the possibility that the determination of A may also be based on some other factor, such as C. This phrase is intended to cover embodiments in which A is determined based only on B. As used herein, the phrase "based on" is synonymous with the phrase "at least partially based on".

[0156] The phrases "in response to" and "responsive to" describe one or more factors that trigger an effect. This phrase does not exclude the possibility that additional factors may affect or otherwise trigger the effect, either together with the specified factor or independently of the specified factor. That is, the effect may be in response only to these factors, or may be in response to the specified factors as well as other unspecified factors. Consider the phrase "perform A in response to B". This phrase identifies that B is a factor that triggers the performance of A or a particular result for A. This phrase does not exclude the possibility that the performance of A may also be in response to some other factor, such as C. This phrase also does not exclude the possibility that performing A may be in response to both B and C. This phrase is intended to cover embodiments in which A is performed only in response to B. As used herein, the phrase "responsive to" is synonymous with the phrase "at least partially responsive to". Similarly, the phrase "in response to" is synonymous with the phrase "at least partially in response to". ***

[0157] Within this disclosure, various physical entities (which may be variously referred to as "units", "circuits", other components, etc.) may be described or claimed as being "configured" to perform one or more tasks or operations. This expression of an [entity] "configured" to [perform one or more tasks] is used herein to refer to a structure (i.e., something physical). More specifically, this expression is used to indicate that the structure is arranged to perform one or more tasks during operation. A structure may be described as being "configured" to perform some task even if the structure is not currently operating. Thus, an entity described or characterized as being "configured" to perform some task refers to something physical, such as a device, a circuit, a processor unit, and a system having a memory storing program instructions executable to perform the task. This phrase is not used herein to refer to something intangible.

[0158] In some cases, various units / circuits / components can be described herein as performing a set of tasks or operations. Even if not specifically described, it is understood that those entities are "configured" to perform those tasks / operations.

[0159] The term "configured to" is not intended to mean "configurable to". For example, an unprogrammed FPGA is not considered to be "configured" to perform a particular function. However, this unprogrammed FPGA may be "configurable" to perform that function. After appropriate programming, the FPGA can then be said to be "configured" to perform a particular function.

[0160] For purposes of the U.S. patent application based on this disclosure, reciting in a claim that a structure is "configured" to perform one or more tasks is not expressly intended to invoke 35 U.S.C. § 112(f) with respect to that claim element. If Applicant desires to invoke 35 U.S.C. § 112(f) during the examination of a U.S. patent application based on this disclosure, it will be by using "means for" performing the function to describe the claim element.

[0161] This disclosure may describe various "circuits." These circuits or "circuit configurations" comprise hardware including various types of circuit elements such as combinational logic, clocked storage devices (e.g., flip-flops, registers, latches, etc.), finite state machines, memories (e.g., random access memories, embedded dynamic random access memories), programmable logic arrays, and the like. The circuits may be custom designed or obtained from standard libraries. In various implementations, the circuit configuration can include digital components, analog components, or a combination of both, as needed. A particular type of circuit may generally be referred to as a "unit" (e.g., a decoder unit, an arithmetic logic unit (ALU), a functional unit, a memory management unit (MMU), etc.). Such a unit also refers to a circuit or circuit configuration.

[0162] The disclosed circuits / units / components and other elements shown in the drawings and described herein include hardware elements such as those described in the preceding paragraphs. In many cases, the internal arrangement of the hardware elements within a particular circuit can be specified by describing the function of that circuit. For example, a particular "decoder unit" can be described as performing the function of "processing the opcode of an instruction and routing the instruction to one or more of a plurality of functional units," which means that the decoder unit is "configured" to perform this function. This description of the function is sufficient to imply to those of ordinary skill in the computer art a set of possible structures for the circuit.

[0163] In various embodiments, as discussed in the previous paragraph, circuits, units, and other elements can be defined by the functions or operations they are configured to perform. The arrangement of such circuits / units / components relative to each other, and the way they interact, ultimately generates a microarchitecture definition of the hardware that is either manufactured within an integrated circuit or programmed into an FPGA, forming a physical implementation form of the microarchitecture definition. Thus, the microarchitecture definition is recognized by those skilled in the art as a structure from which many physical implementation forms can be derived, and all of those implementation forms belong to a broader structure described by the microarchitecture definition. That is, one skilled in the art presented with the microarchitecture definition provided according to the present disclosure can implement the structure by coding the description of the circuits / units / components into a hardware description language (HDL) such as Verilog or VHDL using ordinary techniques without undue experimentation. HDL descriptions are often expressed in a manner that appears to be functional. However, to those skilled in the art, this HDL description is the way used to convert the structure of a circuit, unit, or component to the next level of implementation detail. Such HDL descriptions can take the form of behavioral level code (typically not synthesizable), register transfer language (RTL) code (typically synthesizable as opposed to behavioral level code), or structural code (e.g., a netlist specifying logic gates and their connections). The HDL description may be synthesized against a library of cells designed for a given integrated circuit manufacturing technology, modified for timing, power, and other reasons, and result in a final design database that can be sent to a foundry to generate masks and ultimately manufacture the integrated circuit. Some hardware circuits or parts thereof can also be custom designed with a circuit diagram editor and incorporated into the integrated circuit design along with the synthesized circuits. The integrated circuit can further include transistors and other circuit elements (e.g., passive elements such as capacitors, resistors, inductors, etc.), as well as interconnects between the transistors and circuit elements.Some embodiments can implement a plurality of integrated circuits integrally connected to realize a hardware circuit, and / or, in some embodiments, individual elements can be used. Alternatively, the HDL design may be integrated into a programmable logic array such as a field programmable gate array (FPGA), or implemented on an FPGA. This separation between the design of this circuit group and the subsequent low-level implementation of these circuits generally results in a scenario where, since this process is performed at different stages of the circuit implementation process, the circuit or logic designer does not specify any specific set of structures for the low-level implementation form other than an explanation of how the circuit is configured.

[0164] The fact that many different low-level combinations of circuit elements can be used to implement the same specification of a circuit results in a number of equivalent structures for that circuit. As described above, these low-level circuit implementation forms can vary depending on changes in manufacturing technology, the foundry selected to manufacture the integrated circuit, the library of cells provided for a particular project, and the like. In many cases, the selection made by different design tools or methods for generating these different implementation forms can be arbitrary.

[0165] Furthermore, for a given embodiment, it is common for a single implementation form of a particular functional specification of a circuit to include a large number of devices (e.g., millions of transistors). Therefore, due to this absolute amount of information, it goes without saying that there are an enormous number of equivalent possible implementation forms, and it is unrealistic to fully enumerate the low-level structures used to implement a single embodiment. For this reason, the present disclosure uses functional omissions commonly used in the industry to describe the structure of the circuit.

Claims

1. An apparatus comprising: one or more agent circuits and a memory controller circuit, wherein the memory controller circuit communicates with a memory circuit via an interface, the memory circuit supporting error detection for the interface that causes a combination of data and parity to be written to a target memory location for a detected uncorrectable write interface error, the combination corresponding to an uncorrectable error, communicates with the memory circuit, arbitrates among requests to access the memory circuit from a requesting agent circuit, including a first request to write first data to a first location within the memory circuit, maintains a damage indicator for a data block including a first damage indicator indicating that the first data has been determined to be damaged, transmits, via the interface, a combination of data and parity for the first data block that causes the memory circuit to detect an uncorrectable write interface error, reads the memory location subsequent to the write for the first request and generates a damage indicator for the read data in response to a report of an uncorrectable error from the memory circuit for the read data, the apparatus being configured as such.

2. The apparatus according to claim 1, wherein one of the agent circuits is configured to generate the first damage indicator.

3. The apparatus according to claim 1, configured to maintain a damage indicator through a plurality of operations including propagation of the damage indicator after merging one or more requests to resolve hazards, propagation of the damage indicator for write-read transfer operations from a write queue, and conversion of the damage indicator to a forced uncorrectable write interface error.

4. The apparatus according to claim 3, wherein the plurality of operations further include communication of the damage indicator from the memory controller circuit by a memory cache controller circuit and propagation of the damage indicator determined based on an address mask.

5. The memory controller circuit detects a corrected error indicated by the memory circuit in which incorrect data remains stored in a memory cell of the memory circuit, ​ ​ ​ ​ In response to detecting a corrected error, initiate a demand scrub write operation to the memory circuit that causes an internal read, error correction of the correctable error, and writing of the corrected data to the memory circuit. The apparatus of claim 1, comprising a demand scrub circuit configured as such. **Claim 6** The apparatus of claim 5, wherein the write operation is a fully masked partial write operation to the detected DRAM address of the corrected error. **Claim 7** The demand scrub circuit is configured to record, in one or more software-accessible registers, the number of detected correctable errors, and the number of successful demand scrub writes, as described in claim 5. **Claim 8** The detection of the corrected error is based on a decoding status flag reported by the memory circuit indicating whether the provided data has no errors, has correctable errors, or has uncorrectable errors, as described in claim 5. **Claim 9** The memory circuit of the apparatus of claim 1 implements both link error correction and on-die error correction. **Claim 10** Further comprising the memory circuit, wherein the memory circuit in the case of a write operation, verifies parity information of the error-free write data, corrects the detected correctable errors associated with the interface, in the case of a read operation, corrects the detected errors associated with the read location, and includes an error circuit configured to report the detected uncorrectable errors associated with the read location via the interface, as described in claim 1. **Claim 11** The memory circuit and a central processing unit configured to access the memory circuit via the memory controller circuit, and a network interface circuit, as described in claim 1. **Claim 12** A method comprising communicating with a memory circuit via an interface by a memory controller circuit, wherein the memory circuit supports error detection for the interface that causes a write of a combination of data and parity to a target memory location for detected uncorrectable write interface errors, the combination corresponding to uncorrectable errors, and communicating with the memory circuit. Arbitrating between requests to access the memory circuit from a requesting agent circuit, including a first request to write first data to a first location within the memory circuit by the memory controller circuit; Maintaining a damage indicator for a data block, including a first damage indicator indicating that the first data has been determined to be damaged, by the memory controller circuit; Transmitting, via the interface, a combination of data and parity for the first data block that causes the memory controller circuit to detect an uncorrectable write interface error in the memory circuit; Following the write for the first request, the memory controller circuit reads the memory location and generates a damage indicator for the read data in response to a report of an uncorrectable error from the memory circuit for the read data. A method including:

13. The method of claim 12, further comprising maintaining a damage indicator through a plurality of operations by an apparatus including the memory controller circuit, the plurality of operations including: Propagation of the damage indicator after merging one or more requests to resolve hazards; Propagation of the damage indicator for write-read transfer operations from a write queue; Conversion of the damage indicator to a forced uncorrectable write interface error.

14. Detecting a corrected error indicated by the memory circuit in which incorrect data is still stored in a memory cell of the memory circuit by a demand scrub circuit; In response to detecting the corrected error, starting a demand scrub write operation to the memory circuit, wherein the demand scrub circuit causes an internal read, error correction of the correctable error, and writing of the corrected data to the memory circuit. The method of claim 12, further comprising:

15. The method of claim 14, wherein the write operation is a fully masked partial write operation to the detected DRAM address of the corrected error.

16. A non-transitory computer-readable medium storing design information that specifies at least a portion of a hardware integrated circuit in a format recognized by a semiconductor manufacturing system configured to use the design information to generate a circuit according to the design, the design information specifying that the circuit includes one or more agent circuits and a memory controller circuit, wherein the memory controller circuit communicates with a memory circuit via an interface, the memory circuit supporting error detection for the interface that causes a combination of data and parity to be written to a target memory location for a detected non-correctable write interface error, the combination corresponding to a non-correctable error, and communicates with the memory circuit, arbitrates between requests to access the memory circuit from a requesting agent circuit, the requests including a first request to write first data to a first location within the memory circuit, maintains a damage indicator for a data block including a first damage indicator indicating that the first data has been determined to be damaged, transmits, via the interface, a combination of data and parity for the first data block that causes the memory circuit to detect a non-correctable write interface error, subsequent to the write for the first request, reads the memory location and generates a damage indicator for the read data in response to a report of a non-correctable error from the memory circuit. A non-transitory computer-readable medium configured as such. **Claim 17** The non-transitory computer-readable medium according to claim 16, wherein one of the agent circuits is configured to generate the first damage indicator. **Claim 18** The circuit is configured to maintain a damage indicator through a plurality of operations, the plurality of operations including propagation of the damage indicator after merging one or more requests to resolve hazards, propagation of the damage indicator for write-read transfer operations from a write queue, conversion of the damage indicator to a forced non-correctable write interface error, communication of the damage indicator to the memory controller circuit by a memory cache controller circuit The non - transitory computer - readable medium according to claim 16, including the propagation of a damage indicator determined based on an address mask.

19. The memory controller circuit detects a corrected error indicated by the memory circuit in which incorrect data is still stored in the memory cells of the memory circuit and in response to detecting the corrected error, starts a demand scrub write operation to the memory circuit that causes an internal read, error correction of the correctable error, and writing of the corrected data to the memory circuit. The non - transitory computer - readable medium according to claim 16, including a demand scrub circuit configured as described above.

20. The detection of the corrected error is based on a decoding status flag reported by the memory circuit indicating whether the provided data has no error, has a correctable error, or has an uncorrectable error, for the non - transitory computer - readable medium according to claim 19.

21. An apparatus, one or more processors configured to execute program instructions, a memory cache and a memory cache controller circuit, wherein the memory cache controller circuit caches data operated on by the one or more processors in the memory cache, uses a plurality of trace circuit entries to track the number of detected correctable errors associated with each of a plurality of locations of data processed by the memory cache controller circuit, and is configured to generate a signal to the one or more processors that identifies the specific location in response to detecting a threshold number of correctable errors for the specific location.

22. The memory cache controller circuit is further configured to generate an alert signal in response to the number of valid entries in the trace circuit entries matching or exceeding an occupancy threshold, for the apparatus according to claim 21.

23. The memory cache controller circuit is further configured to enable software to access one or more trace circuit entries in response to matching or exceeding the occupancy threshold, for the apparatus according to claim 22.

24. The memory cache controller circuit The apparatus according to claim 23, further configured to release one or more of the tracking circuit entries in response to software signaling. **Claim 25** The apparatus according to claim 21, wherein the plurality of tracking circuit entries each include a respective client identifier field indicating a client associated with a given correctable error. **Claim 26** The memory cache controller circuit The apparatus according to claim 21, further configured to track detected uncorrectable errors associated with respective locations of data processed by the memory cache controller circuit using a plurality of uncorrectable error (UE) tracking circuit entries. **Claim 27** The apparatus according to claim 26, wherein the UE tracking circuit entry includes a source field identifying a source of a given UE. **Claim 28** The apparatus according to claim 27, wherein the source field is configured to encode a source including at least a memory error, a memory cache error, and a source of a snoop response. **Claim 29** The apparatus according to claim 26, wherein the plurality of UE tracking circuit entries are not tagged, and the plurality of tracking circuit entries for detected correctable errors are tagged with at least a portion of an address for a given location. **Claim 30** The apparatus according to claim 26, configured to maintain a corruption indicator for a data block indicating that the data block has been determined to be corrupt. **Claim 31** The apparatus a central processing unit, a display, and a network interface circuit, and is a computing device further comprising the apparatus according to claim 21. **Claim 32** A method comprising: caching, by a memory cache controller circuit, data operated on by the one or more processors in the memory cache; tracking, using a plurality of tracking circuit entries, a number of detected correctable errors associated with respective locations of data processed by the memory cache controller circuit; and generating, in response to detecting a threshold number of correctable errors for a particular location, a signal identifying the particular location to the one or more processors. **Claim 33** The method according to claim 32, further comprising generating, by the memory cache controller circuit, an alert signal to software indicating a potential future occupancy problem in response to the number of valid entries in the trace circuit entry matching or exceeding an occupancy threshold.

34. The method according to claim 33, further comprising enabling software to access one or more trace circuit entries in response to matching or exceeding the occupancy threshold.

35. The method according to claim 32, further comprising tracking detected uncorrectable errors associated with respective locations of data processed by the memory cache controller circuit using a plurality of uncorrectable error (UE) trace circuit entries.

36. A non-transitory computer-readable storage medium storing design information specifying at least a portion of a hardware integrated circuit in a format recognized by a semiconductor manufacturing system configured to use the design information to generate a circuit according to a design, the design information specifying that the circuit comprises: one or more processors configured to execute program instructions; a memory cache, and a memory cache controller circuit; wherein the memory cache controller circuit is configured to: cache data operated on by the one or more processors in the memory cache; track the number of detected correctable errors associated with respective locations of data processed by the memory cache controller circuit using a plurality of trace circuit entries; generate a signal to the one or more processors identifying the specific location in response to detecting a threshold number of correctable errors for the specific location.

37. The memory cache controller circuit is further configured to: generate an alert signal to software indicating a potential future occupancy problem in response to the number of valid entries in the trace circuit entry matching or exceeding an occupancy threshold, the non-transitory computer-readable medium according to claim 36.

38. The memory cache controller circuit is configured to: The non - transitory computer - readable medium according to claim 37, further configured to enable software to access one or more trace circuit entries in response to matching or exceeding the occupancy threshold. **Claim 39** The memory cache controller circuit The non - transitory computer - readable medium according to claim 36, further configured to use a plurality of irreparable error (UE) trace circuit entries to track detected irreparable errors associated with respective locations of data processed by the memory cache controller circuit. **Claim 40** The non - transitory computer - readable medium according to claim 39, wherein the plurality of UE trace circuit entries are not tagged, and the plurality of trace circuit entries are tagged with at least a portion of an address for a given location.

Citation Information

Patent Citations

  • A system and method for tracking error data in a storage device.

    JP2012532372A

  • Memory module error tracking

    US20170372799A1

  • Memory system with error detection

    US20210350870A1

  • Memory system and memory operation program

    WO2022018950A1