Systems, devices, and methods for lockstep corrected error reporting, data poisoning, and potential recovery mechanisms

By introducing circuit system to the processor to identify and handle corrected errors, the platform reset problem caused by incorrect comparison in dynamic lockstep mode is solved, achieving longer uptime and lower costs.

CN120179439APending Publication Date: 2025-06-20INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411644221.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-18
Filing Date
2024-11-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In dynamic lockstep mode, corrected errors can lead to false comparisons, which triggers software platform reset, resulting in high cost and long downtime, especially in data center environments.

Method used

By introducing a circuit system into the processor, error comparisons caused by corrected errors occurring in the core pair involved in the lock step operation are identified and notified to the software entity. The circuit system is configured to treat the corrected error as a recoverable error rather than an uncorrected error, thereby avoiding unnecessary platform resets.

Benefits of technology

Reduces error explosion radius, extends processor uptime, especially in data center environments, and avoids corrected errors due to lockstep operations being promoted to uncorrected errors, thereby reducing the frequency and cost of platform resets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179439A_ABST
    Figure CN120179439A_ABST
Patent Text Reader

Abstract

Systems, devices, and methods for lockstep corrected error reporting, data poisoning, and potential recovery mechanisms are disclosed. In one embodiment, an apparatus includes a first core and a second core to execute instructions; and an interface circuit coupled to the first core and the second core. In a lockstep mode in which a first core and a second core are configured to execute redundantly, the interface circuitry is configured to: identify a miscomparison between the first core and the second core, the miscomparison being caused by a corrected error in one of the first core or the second core; and indicating the miscomparison as a recoverable error. Other embodiments are described and claimed.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Many modern processors provide mechanisms for reducing silent data errors, which may occur due to single-bit flips somewhere in the signal processing path. One such technique is the Dynamic Lockstep Mode (DLSM), which is an enhanced reliability, availability, and serviceability (RAS) capability that can be selectively enabled by allowing system software to place logical processors in lockstep mode and to exit lockstep mode, for protecting high-integrity applications, containers, virtual machines (VMs), and subroutines. In DLSM, two cores are placed in an operating mode in which they execute in lockstep fashion, cycle by cycle, to execute the same instructions in a given cycle. When in this mode, any functional differences between them are detected, allowing for a much higher error detection rate. This error detection triggers an error indication that may cause a software platform reset, which is a very expensive and time-consuming process, especially in a data center context. Brief Description of the Drawings

[0002] Figure 1 is a block diagram of a portion of a processor according to an embodiment.

[0003] Figure 2 is a block diagram of a portion of a processor according to another embodiment.

[0004] Figure 3 is a flowchart of a method according to an embodiment.

[0005] Figure 4 is a flowchart of a method according to another embodiment.

[0006] Figure 5 Illustrates an example computing system.

[0007] Figure 6 Illustrates a block diagram of an example processor according to an embodiment.

[0008] Figure 7 is a block diagram of a processor core according to an embodiment. Detailed Description

[0009] In a processor with lockstep capabilities, when the common output of two functional units (e.g., two cores in a core pair) diverges, a miscomparison is detected. This miscomparison can be indicated via a machine check event.

[0010] In various embodiments, the processor is equipped with circuitry for identifying a miscomparison caused by a corrected error that occurs in a core pair involved in a lockstep operation. In addition, the circuitry is configured to notify a software entity about the miscomparison. More specifically, in some cases, the circuitry may notify the software via a recoverable error indication in a machine check status register (e.g., Software Recoverable Action Required (SRAR)).

[0011] This indication allows the notified software to take appropriate action since an SRAR error is an uncorrected error from which the software can attempt to recover. Once the software has performed a recovery action, continued execution is possible. In many cases, the action initiated by the software may be far less severe than the response to an uncorrected error that would have occurred in the absence of the embodiments. That is, the software can choose to handle the recoverable error in a more graceful manner than for an uncorrected error. For example, the software can recover from the error without causing a platform reset.

[0012] In other words, the embodiments help reduce the blast radius of errors and contribute to increased uptime, particularly in processors configured for data centers. The processor circuitry described herein can detect when a miscomparison between common results determined by two (or more) cores is due to a corrected error. In this way, the embodiments avoid elevating the severity of a corrected error due to lockstep operation to an uncorrected error, thereby achieving more platform uptime, particularly for large-scale data centers.

[0013] Core-wide corrected errors can be detected at any time, including during lockstep operation. While the embodiments are applicable to various forms of lockstep operation, the discussion herein focuses on such errors that occur during DLSM. Considering the timing implications of the correction, these errors generally result in timing involvement. Such errors are typically recorded in the core-wide (core-internal) architectural state (such as a machine check status register or other model-specific register) for consumption by software. Both this timing involvement and this core-wide error logging cause miscomparisons in cores executing in lockstep. In the absence of the embodiments, when a DLSM miscomparison occurs, an uncorrected error event is triggered, which requires a software platform reset, as described above, and a software platform reset is a costly event. In other words, in the absence of the embodiments, the default DLSM behavior has the practical effect of elevating the severity of a core-wide corrected error to an uncorrected error.

[0014] Now refer to Figure 1, shown is a block diagram of a portion of a processor according to an embodiment. As Figure 1 shown, processor 100 includes a plurality of cores. In the illustrated implementation, a pair of cores 1101, 1102 are depicted, which together form a core pair that can be placed in a redundant mode (such as DLSM). It should be understood that core 110 can operate independently outside of the lockstep mode. When placed in the lockstep mode, core 1101 can be configured as the active core, and core 1102 can be configured as the shadow core for redundantly executing the same instruction stream in a lockstep manner so that the execution results can be compared within interface circuit 120.

[0015] As Figure 1 shown, interface circuit 120 includes a plurality of comparators, which include a plurality of signal comparators 130 0-N . Each such signal comparator 130 is configured to compare common results from core 110, that is, signals that are routed outside of the core periphery of these cores. It should be understood that signal comparator 130 can be used to compare the corresponding outputs of the cores to ensure output matching. In the case of a mismatch, a miscomparison is identified. Generally, a miscomparison may cause DLSM to be deactivated.

[0016] As Figure 1 further illustrated, core 110 includes logic circuits 1151, 1152 configured as logical OR gates. As illustrated, each logical OR gate 115 receives inputs from various functional units of a given core. More specifically, these inputs are corrected error signals that are output from the functional units when an error (such as error correction coding (ECC) or other such errors) has been corrected within the core. In operation, whenever logic circuit 115 receives an active corrected error signal from the functional unit, it sends an active corrected error indication signal to corrected error comparator 135 within interface circuit 120. Thus, logic gate 115 operates to logically OR the core-wide corrected errors together into a single signal that is provided to corrected error comparator 135.

[0017] In this way, one or more core-wide corrected error detections are logically ORed into a single signal that is brought to the module level, that is, interface circuit 120 at the core pair boundary. Within interface circuit 120, comparator 135 is regarded as a special type of miscomparison. When software enables this architecturally visible capability and the first miscomparison belongs to the type of corrected error, the comparator error is recorded as recoverable in, for example, a machine check status register or other model-specific register.

[0018] Although in some implementations,Figure 1 The processor hardware described in [[ID=]] and the corrected error handling discussed herein can be configured as the default processor behavior, but in other cases, the circuitry and its operation can be enabled as an opt-in feature. In such cases, when software opts in, the circuitry operates to treat miscomparisons caused by the detection of corrected errors as recoverable errors rather than uncorrected errors.

[0019] In a particular embodiment, the processor identifies support for DLSM via a CPU identifier (e.g., CPUID(0x7).ECX(0x1).EDX[12:12] in an x86 processor), which provides DLSM MSR accessibility. In turn, the opt-in ability for the corrected error handling described herein can be set in a given configuration storage device (e.g., an MSR). In one embodiment, a capabilities register can include a field for indicating support for the opt-in feature, as follows: IA32_DLSM_CAPABILITIES[CORRECTED_MISCO_SEVERITY] (IA32_DLSM capabilities [corrected miscomparison severity]). In turn, when DLSM is activated to opt in to this ability, software can set the fields of the configuration register. Specific exemplary registers are described below.

[0020] Accordingly, embodiments provide an architectural mechanism for software to opt in to a mode for treating corrected errors that cause miscomparisons as recoverable. In an embodiment, this opt-in feature, when activated, can report miscomparisons due to corrected errors via a recoverable (e.g., SRAR) machine check signature. Software can determine an appropriate course of action in response to this machine check signature.

[0021] One or more embodiments can also utilize data poisoning to enable graceful resolution of miscomparisons between cores operating in lockstep mode. In data poisoning, an indication of an uncorrected error is carried with the data affected by the uncorrected error. This mechanism allows software suppression of uncorrected errors on the data path. To this end, a poison indicator (which can be a single bit that indicates an uncorrected error when set) extends from within a cache structure within a core to main memory. Overall, when an uncorrected error is detected at that address, the error is marked as poisonous, and the poison indication follows the data wherever it goes in the system. If the data is consumed by software, the fault contains the error and allows software to have an opportunity to recover from the fault by only terminating the affected software flow.

[0022] Embodiments provide a hardware capability for writing data from a core pair that has encountered a miscomparison to be marked as poisonous. This is true even if the data write itself is the first source of the miscomparison. This capability allows the impact of the miscomparison to be contained on the data path. In this way, software can reduce the blast radius of a core pair that has encountered a miscomparison to only the software running on that core pair.

[0023] Now refer to Figure 2 , a block diagram of a portion of a processor according to another embodiment is shown. The processor 200 is implemented similar to Figure 1 's system 100 (with the same reference numerals, although in the "200" series). Note that in Figure 2 , the additional comparator 230 d is implemented as a data path comparator. In the embodiments herein, the data path comparator 230 d is configured to compare the output data from each core 210 to identify whether a mismatch occurs. When data poisoning according to an embodiment is enabled, a mismatch indicated by the data path compared to 230 d triggers marking the given data as poisonous. This poison indicator can flow through the machine with the data. Thus, as shown in Figure 2 , the write data that has been associated with the mismatch in the data path comparator 230 d can be provided to the memory hierarchy with the poison indicator. In Figure 2 , the bus interface unit 250 may include a second-level cache 255, and when a data path mismatch is identified, the write data can be stored in the cache 255 with the set poison indicator.

[0024] As shown in Figure 2 , the BIU 250 hardware marks all stores from the core pair where the miscomparison occurred as poisonous until the DLSM deactivation in response to the miscomparison is complete. This arrangement protects the hardware on the address path that might otherwise introduce silent data corruption into the address space of the OS, host, or other guest. Although shown at this high level in the embodiments of Figure 2 , many variations and alternatives are possible.

[0025] One example use case for this poison-based error reporting can be for "shutdownable" software similar to virtual machine (VM) guest workloads. In such a use case, a hypervisor such as a virtual machine monitor (VMM) manages data access such that miscomparisons can be attributed to a known set of shutdownable VM guests.

[0026] Although in some implementations, Figure 2The processor hardware described in [[ ]] and the data poisoning handling discussed herein can be configured as the default processor behavior, but in other cases, the circuit system and its operation can be enabled as an opt-in feature. In such cases, when the software opts in, the circuit system operates to mark the stores from the core pairs that have encountered miscomparisons as poisonous.

[0027] In one or more embodiments, an additional opt-in implements the hardware recording of miscomparisons as recoverable rather than uncorrected. When enabled, the hardware intercepts data stores from active cores and marks them as poisonous if these data stores occur when or after a miscomparison has been detected, e.g., until the DLSM deactivation is complete.

[0028] The opt-in ability for data poisoning as described herein can be set in a given configuration storage device (e.g., MSR). In one embodiment, a capabilities register can store fields for providing an opt-in feature for poisonous handling of miscomparisons as follows: IA32_DLSM_CAPABILITIES[POISON_MISCO] (IA32_DLSM Capabilities [Poison Miscomparison]) = 1 enumerates the ability for poisonous suppression for core pairs for which miscomparisons have been detected. Further, when activating DLSM to opt in to this ability, the software can set fields in the configuration register.

[0029] Another ability indicator IA32_DLSM_CAPABILITIES[SRAR_MISCO] (IA32_DLSM Capabilities [SRAR Miscomparison]) = 1 enumerates the ability to mark miscomparisons as recoverable errors. When activating DLSM to opt in to this ability, the software can set another field in the configuration register. In an embodiment, when the software chooses to set SRAR_MISCO, the software can set POISON_MISCO.

[0030] Now referring to Table 1, shown is an example capabilities register for DLSM operation, including fields for identifying the presence of the features described herein. Table 1

[0031] Now referring to Table 2, shown is an example configuration register for setting various behaviors for DLSM operation according to an embodiment.

[0032] In one or more embodiments, the command register may be a thread - scope register, and in a particular embodiment may be enumerated as the IA32_DLSM_CMD register, and may have the fields and definitions shown in Table 2 below:. Table 2

[0033] Now referring to Table 3, an illustration of mis - comparison handling according to an embodiment is shown. More specifically, Table 3 illustrates the impact of mis - comparison based on configuration settings in a configuration register according to an embodiment. Table 3

[0034] As shown in Table 3, when the CORRECTED_MISCO_SEVERITY field is set when DLSM is activated, a mis - comparison first detected with a corrected error signaling event causes an SRAR error to be logged. Additionally, even if POISON_MISCO is set, this type of mis - comparison does not cause a poison to be asserted. In one or more embodiments, if the first detected mis - comparison is due to a corrected error detection event, CORRECTED_MISCO_SEVERITY only affects the type of error logged. Other types of mis - comparisons result in uncorrectable errors.

[0035] As further shown in Table 3, when the POISON_MISCO field is set when DLSM is activated, this poisoning mechanism causes hardware poisoning of any data writes (both stores and evictions) from the core pair after the core pair has detected a mis - comparison, thereby implementing suppression against address and data corruption. However, if the first detected mis - comparison is due to a corrected error signaling event and CORRECTED_MISCO_SEVERITY is set, the data will not be marked as poisoned even if POISON_MISCO is set. This is because the machine state and workload can potentially be recoverable without workload termination. This can avoid unnecessarily marking data as poisonous.

[0036] When the SRAR_MISCO field is set when activated in the DLSM, if any type of miscomparison is detected, a recoverable error is logged. This ability and this type of error require software action to recover and can allow the software to avoid a thermal reset when such miscomparisons are detected. For this error type, the software is responsible for determining whether recovery is possible and what recovery action (if any) should occur. If the software expects address and data suppression for the core pair that detects the miscomparison, the software can enable POISON_MISCO and SRAR_MISCO when activating the DLSM.

[0037] To recover properly from such miscomparisons, the software can ensure that the error is confined to a specific scope of the software (e.g., confined to a VM or application) and terminate it. However, other VMs or applications can remain active. If a miscomparison is detected in other non-terminable software scopes (e.g., detected during VMM or OS execution), the error can be considered non-recoverable.

[0038] When SRAR occurs because a corrected error was the cause of the first miscomparison detected and CORRECTED_MISCO_SEVERITY is enabled, the corrected error can be logged in both the active and shadow core Machine Check Architecture (MCA) blocks. When this type of miscomparison is detected, the Machine Check Status Register can indicate that a restart of execution using the context of the interrupt is possible. When the miscomparison is detected, the comparison stops and the DLSM deactivation begins. In some embodiments, the software can choose to attempt to reconstruct the lockstep mode and continue execution at the execution point where the active core deactivated the lockstep.

[0039] When the software chooses to enable SRAR_MISCO, it may wish to enable POISON_MISCO. If the software enables POISON_MISCO, it can operate to clear the poisoned lines generated by the DLSM deactivation. When this type of miscomparison is detected, the Machine Check Status Register can indicate that a restart of execution using the interrupted context is possible. If the software has determined that this recoverable miscomparison is confined to a terminable software domain (e.g., a VM guest or application), the system software can terminate that software domain but can continue execution in other contexts (e.g., other VM guests or applications). If the system software has determined that this recoverable miscomparison is associated with a non-terminable software domain (e.g., the VMM), the error can be considered non-recoverable.

[0040] Thus, embodiments can suppress the miscomparison effect by using data poisoning. That is, instead of logging the miscomparison as an uncorrected error that triggers a platform reset, the miscomparison is logged as a recoverable error that allows the software to terminate the affected workload and then continue execution.

[0041] Now referring to Figure 3 , a flowchart of a method according to an embodiment is shown. Figure 3 Method 300 is a method for providing a corrected error report when in a lockstep mode according to an embodiment. Method 300 may be performed by a hardware circuit system that includes both core internal circuit systems and core external circuit systems, such as interface circuit systems between cores and / or additional core external circuit systems (such as bus interface units). In some implementations, method 300 may be performed by the hardware circuit system alone and / or in combination with firmware and / or software.

[0042] Method 300 begins when the core is configured into a lockstep mode, in which the core redundantly executes the same instructions (block 310). For purposes of discussion, assume the lockstep mode is DLSM. During execution, it is determined whether a core error (i.e., an error within the core scope) has occurred (diamond 315). If not, continued redundant code execution occurs at block 310. When an error is identified, it is determined whether the error has been corrected (diamond 320). If not, the uncorrected error is recorded in, for example, a machine check register or other model-specific register.

[0043] Still referring to Figure 3 , conversely, if the error is a corrected error, control passes to block 330, where the corrected error is recorded and a corrected error signal is sent to a corrected error comparator such as that shown above in Figure 1 . Next, control passes to block 340, where various core peripheral outputs may be sent to an associated comparator of an interface circuit that couples the active core and the shadow core of the core pair.

[0044] Next, in the interface circuit, at diamond 350, it is determined whether a miscomparison has occurred at the corrected error comparator. Note that the miscomparison when identified is thus an identification of an error within the core pair scope. If so, control passes adjacent to diamond 360 to determine whether an opt-in behavior for the corrected error miscomparison is enabled. This determination may be based on a setting in a model-specific register (e.g., a setting in the miscomparison severity field of the DLSM command register). If this opt-in behavior is enabled, control passes to block 365, where the miscomparison is recorded as a recoverable error. For example, in one embodiment, an SRAR indicator may be set within a given machine check status register.

[0045] Still referring to Figure 3, if there is no miscomparison by the corrected error comparator, control passes sequentially to diamond 370 to determine if there is a miscomparison at the signal comparator. If not, further redundant code execution occurs at block 310. If such a miscomparison is identified, control passes to block 380, where the miscomparison can be recorded as an uncorrected error. Thereafter, control passes to block 390, where steps are taken to deactivate the lockstep mode. After a recoverable error record has been executed at block 365, control also passes to block 390 for lockstep mode deactivation. Although shown at this high level in the Figure 3 embodiment, many variations and alternatives are possible.

[0046] Now referring to Figure 4 , a flowchart of a method according to another embodiment is shown. Figure 4 Method 400 is a method for providing data poisoning while in a lockstep mode according to an embodiment. Method 400 may be executed by the same hardware circuitry (and / or firmware and / or software) as discussed above in Figure 3 .

[0047] Method 400 begins when the core is configured into a lockstep mode, in which the core redundantly executes code (block 410). For purposes of discussion, again assume the lockstep mode is DLSM. During execution, core peripheral outputs (e.g., outputs from various functional units) are sent to a comparator in the interface circuitry of the coupled core pair (block 420). Next, it is determined whether a miscomparison is detected in the data path in the interface circuitry, which may be identified by a miscomparison in one or more data path comparators (diamond 430). If not, continued redundant code execution occurs at block 410.

[0048] When a data path miscomparison is identified, it is determined whether it is due to a corrected error (diamond 435). If not, control passes to diamond 440, where it is determined whether opt-in behavior for data poisoning is enabled. This determination may be based on settings in a model-specific register (e.g., settings in the poison miscomparison enable field of the DSLM command register). If this opt-in behavior is enabled, control passes to block 460, where write data from the core is marked as poisonous and sent to the memory hierarchy. Next, at block 470, steps may be taken to deactivate the lockstep mode.

[0049] Still referring to Figure 4, in contrast, if it is determined at diamond 440 that data poisoning has not been enabled, control passes to block 450, where the written data can be sent directly to the memory hierarchy without any data being marked as poisonous. Thereafter, control passes again to block 470 for lockstep mode deactivation.

[0050] Finally, further reference is made to Figure 4 , if the miscomparison detected on the data path is due to a corrected error miscomparison (as determined at diamond 435), control passes to diamond 445. At diamond 445, it is determined whether the opt-in behavior for the corrected error miscomparison is enabled (which can be determined, for example, based on model-specific register settings). If this opt-in behavior is enabled, control passes to block 450, where, since the corrected error reporting is enabled, the data is sent directly to the memory hierarchy without any poison marking. Otherwise, when this opt-in behavior is not enabled, control passes to diamond 440 (discussed above) to determine whether to mark the data as poisonous. Although shown at this high level in the Figure 4 embodiment of, many variations and alternatives are possible.

[0051] As discussed above, the programmable behavior described herein can be enabled by system software such as a virtual machine monitor (VMM) or other hypervisor, an OS, firmware, etc. To enable or disable the corrected errors and / or data poisoning described herein, such privileged software can write to an MSR such as a command register.

[0052] Figure 5 Illustrated example computing system. Multiprocessor system 500 is a system provided with interfaces and includes multiple processors or cores, which include a first processor 570 and a second processor 580 coupled via an interface 550 such as a point-to-point (P-P) interconnect, a fabric, and / or a bus. In some examples, the first processor 570 and the second processor 580 are homogeneous. In some examples, the first processor 570 and the second processor 580 are heterogeneous. Although example system 500 is shown as having two processors, the system can have three or more processors, or can be a single-processor system. In some examples, the computing system is a SoC. In any case, system 500 includes interface circuitry as described herein for performing DLSM as described herein and identifying at least some miscomparisons as recoverable errors and / or poisoned write data.

[0053] Processors 570 and 580 are shown as including integrated memory controller (IMC) circuitry 572 and 582, respectively. Processor 570 also includes interface circuits 576 and 578; similarly, second processor 580 includes interface circuits 586 and 588. Processors 570, 580 may exchange information via interface circuits 578, 588 through interface 550. IMCs 572 and 582 couple processors 570, 580 to respective memories, namely memory 532 and memory 534, which may be part of the main memories locally attached to the respective processors.

[0054] Processors 570, 580 may each utilize interface circuits 576, 594, 586, 598 to exchange information with a network interface (NW I / F) 590 via respective interfaces 552, 554. Network interface 590 (e.g., one or more of an interconnect, bus, and / or fabric, which is a chipset in some examples) may optionally exchange information with coprocessor 538 via interface circuit 592. In some examples, coprocessor 538 is a dedicated processor, such as a high-throughput processor, a network or communication processor, a compression engine, a graphics processor, a general purpose graphics processing unit (GPGPU), a neural-network processing unit (NPU), an embedded processor, and so on.

[0055] A shared cache (not shown) may be included in either processor 570, 580 or outside both processors but connected to these processors via an interface (e.g., a P-P interconnect) such that: if a processor is placed in a low-power mode, the local cache information of either or both processors may also be stored in the shared cache.

[0056] The network interface 590 may be coupled to the first interface 516 via the interface circuit 596. In some examples, the first interface 516 may be an interface such as a Peripheral Component Interconnect (PCI) interconnect, a PCI Express interconnect, or another I / O interconnect. In some examples, the first interface 516 is coupled to a power control unit (PCU) 517, which may include circuitry, software, and / or firmware to perform power management operations regarding the processors 570, 580, and / or the coprocessor 538. The PCU 517 provides control information to a voltage regulator (not shown) such that the voltage regulator generates an appropriate regulated voltage. The PCU 517 also provides control information to control the generated operating voltage. In various examples, the PCU 517 may include various power management logic units (circuitry) to perform hardware-based power management. Such power management may be fully controlled by the processor (e.g., controlled by various processor hardware and may be triggered by workload and / or power constraints, thermal constraints, or other processor constraints), and / or the power management may be performed in response to an external source (e.g., a platform or a power management source or system software).

[0057] The PCU 517 is illustrated as existing as a separate logic from the processor 570 and / or the processor 580. In other cases, the PCU 517 may execute on one or more given cores in the core (not shown) of the processor 570 or 580. In some cases, the PCU 517 may be implemented as a microcontroller (dedicated or general-purpose) or other control logic that is configured to execute its own dedicated power management code (sometimes referred to as P-code). In still other examples, the power management operations to be performed by the PCU 517 may be implemented external to the processor, such as by a separate power management integrated circuit (PMIC) or another component external to the processor. In still other examples, the power management operations to be performed by the PCU 517 may be implemented within the BIOS or other system software.

[0058] Various I / O devices 514 and a bus bridge 518 may be coupled to a first interface 516, and the bus bridge 518 couples the first interface 516 to a second interface 520. In some examples, one or more additional processors 515 are coupled to the first interface 516, such as a coprocessor, a high throughput many integrated core (MIC) processor, a GPGPU, an accelerator (such as a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array (FPGA), or any other processor. In some examples, the second interface 520 may be a low pin count (LPC) interface. Various devices may be coupled to the second interface 520, such devices including, for example, a keyboard and / or a mouse 522, a communication device 527, and a storage circuitry 528. The storage circuitry 528 may be one or more non-transitory machine-readable storage media as described below, such as a disk drive or other mass storage device, which may include instructions / code and data 530. Additionally, audio I / O 524 may be coupled to the second interface 520. Note that other architectures are possible in addition to the point-to-point architecture described above. For example, a system such as the multiprocessor system 500 may implement a multi-drop interface or other such architecture instead of the point-to-point architecture.

[0059] Example core architectures, processors, and computer architectures.

[0060] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, the implementation of these cores can include: 1) general-purpose in-order cores, for general computing purposes; 2) high-performance general-purpose out-of-order cores, for general computing purposes; 3) specialized cores, mainly for graphics and / or scientific (throughput) computing purposes. The implementation of different processors can include: 1) a CPU, including one or more general-purpose in-order cores for general computing purposes and / or one or more general-purpose out-of-order cores for general computing purposes; and 2) a coprocessor, including one or more specialized cores mainly for graphics and / or scientific (throughput) computing purposes. These different processors result in different computer system architectures, which can include: 1) the coprocessor and the CPU on separate chips; 2) the coprocessor and the CPU on separate dies within the same package; 3) the coprocessor and the CPU on the same die (in this case, such a coprocessor is sometimes referred to as specialized logic, such as integrated graphics and / or scientific (throughput) logic, or as a specialized core); and 4) a system-on-chip (SoC), which can be included on the same die as the described CPU (sometimes referred to as (one or more) application cores or (one or more) application processors), the above-mentioned coprocessor, and additional functionality. An example core architecture will be described next, followed by a description of example processors and computer architectures.

[0061] Figure 6 A block diagram of an example processor and / or SoC 600 is illustrated, which can have one or more cores and have an integrated memory controller. The processor 600 illustrated by the solid-line block diagram has a single core 602(A), a system agent unit circuitry 610, and a set of one or more interface controller unit circuitries 616, while the optionally added dashed-line block diagram illustrates an alternative processor 600 as having multiple cores 602(A)-(N), a set of one or more integrated memory control unit circuitries 614 in the system agent unit circuitry 610, specialized logic 608, and a set of one or more interface controller unit circuitries 616. Note that the processor 600 can be Figure 5 one of the processors 570 or 580 or the coprocessors 538 or 515.

[0062] Thus, different implementations of the processor 600 may include: 1) a CPU, where the dedicated logic 608 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores, not shown), and the cores 602(A)-(N) are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of both); 2) a coprocessor, where the cores 602(A)-(N) are a large number of dedicated cores mainly for graphics and / or scientific (throughput) purposes; and 3) a coprocessor, where the cores 602(A)-(N) are a large number of general-purpose in-order cores. Thus, the processor 600 may be a general-purpose processor, a coprocessor, or a special-purpose processor, such as a network or communication processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput integrated many-core (MIC) coprocessor (including 30 or more cores), an embedded processor, and so on. The processor may be implemented on one or more chips. The processor 600 may be part of one or more substrates and / or may be implemented on one or more substrates using any of a variety of process technologies, such as complementary metal oxide semiconductor (CMOS), bipolar CMOS (BiCMOS), P-type metal oxide semiconductor (PMOS), or N-type metal oxide semiconductor (NMOS).

[0063] The memory hierarchy includes one or more levels of cache unit circuitry 604(A)-(N) within cores 602(A)-(N), a set of one or more shared cache unit circuitry 606, and external memory (not shown) coupled to the set of integrated memory controller unit circuitry 614. The set of one or more shared cache unit circuitry 606 may include one or more intermediate-level caches, such as a second level (L2), third level (L3), fourth level (L4), or other level of cache, such as a last level cache (LLC), and / or combinations thereof. Although in some examples interface network circuitry 612 (e.g., a ring interconnect) provides an interface to dedicated logic 608 (e.g., integrated graphics logic), the set of shared cache unit circuitry 606, and system agent unit circuitry 610, alternative examples use any number of known techniques to provide an interface to these units. In some examples, coherence is maintained between one or more of the circuitry in the shared cache unit circuitry 606 and the cores 602(A)-(N). In some examples, interface controller unit circuitry 616 couples these cores 602 to one or more other devices 618, such as one or more I / O devices, storage devices, one or more communication devices (e.g., wireless networks, wired networks, etc.), and so on.

[0064] In some examples, one or more of the cores 602(A)-(N) have multithreading capabilities. System agent unit circuitry 610 includes those components that coordinate and operate the cores 602(A)-(N). System agent unit circuitry 610 may include, for example, power control unit (PCU) circuitry and / or display unit circuitry (not shown). The PCU may be (or may include) the logic and components required to regulate the power states of the cores 602(A)-(N) and / or dedicated logic 608 (e.g., integrated graphics logic). The display unit circuitry is used to drive one or more externally connected displays. At least the dedicated logic 608 includes interface circuitry 609, which may perform core-pair-based analysis of redundant execution and, in some cases and depending on configuration settings, identify miscomparisons due to corrected errors as recoverable errors and / or identify generated write data as poisonous, as described herein. Of course, similar circuitry may be located throughout the processor 600, including within the cores 602, system agent unit 610, and shared cache unit 606.

[0065] The cores 602(A)-(N) may be homogeneous with respect to an instruction set architecture (ISA). Alternatively, the cores 602(A)-(N) may be heterogeneous with respect to the ISA; that is, a subset of the cores 602(A)-(N) may be capable of executing one ISA, while other cores may be capable of executing only a subset of that ISA or of executing another ISA.

[0066] Figure 7 A processor core 790 is shown, the processor core 790 including a front-end unit circuitry 730 coupled to an execution engine unit circuitry 750, both of which are coupled to a memory unit circuitry 770. The core 790 may be a reduced instruction set architecture computing (RISC) core, a complex instruction set architecture computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, the core 790 may be a specialized core, such as, for example, a network or communication core, a compression engine, a coprocessor core, a general purpose computing graphics processing unit (GPGPU) core, a graphics core, and the like.

[0067] The front-end unit circuit system 730 may include a branch prediction circuit system 732, which is coupled to an instruction cache circuit system 734, which is coupled to an instruction translation lookaside buffer (TLB) 736, which is coupled to an instruction fetch circuit system 738, which is coupled to a decoding circuit system 740. In one example, the instruction cache circuit system 734 is included in the memory unit circuit system 770 rather than the front-end circuit system 730. The decoding circuit system 740 (or decoder) may decode an instruction and generate one or more micro-operations, microcode entry points, micro-instructions, other instructions, or other control signals as output, which are decoded from the original instruction, or otherwise reflect the original instruction, or are derived from the original instruction. The decoding circuit system 740 may also include an address generation unit (AGU, not shown) circuit system. In one example, the AGU uses the forwarded register ports to generate LSU addresses and may further perform branch forwarding (e.g., immediate offset branch forwarding, LR register branch forwarding, etc.). Various different mechanisms may be utilized to implement the decoding circuit system 740. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLAs), microcode read only memories (ROMs), etc. In one example, the core 790 includes a microcode ROM (not shown) or other medium that stores microcode for certain macro instructions (e.g., in the decoding circuit system 740 or otherwise within the front-end circuit system 730). In one example, the decoding circuit system 740 includes a micro-operation (micro-op) or operation cache (not shown) to save / cache the decoded operations, microtags, or micro-operations generated during the decoding or other stages of the processor pipeline 700. The decoding circuit system 740 may be coupled to a rename / allocator unit circuit system 752 in the execution engine circuit system 750.

[0068] The execution engine circuitry 750 includes a rename / allocator unit circuitry 752 that is coupled to a retirement unit circuitry 754 and a set of one or more scheduler circuitry systems 756. The scheduler circuitry system 756 represents any number of different schedulers, including reservation stations, a central instruction window, and so on. In some examples, the (one or more) scheduler circuitry systems 756 may include an arithmetic logic unit (ALU) scheduler / scheduling circuitry, an ALU queue, an address generation unit (AGU) scheduler / scheduling circuitry, an AGU queue, and so on. The (one or more) scheduler circuitry systems 756 are coupled to the (one or more) physical register file circuitry systems 758. Each of the (one or more) physical register file circuitry systems 758 represents one or more physical register files, and different physical register files among these physical register files store one or more different data types, such as scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points, status (e.g., an instruction pointer, i.e., the address of the next instruction to be executed), and so on. In one example, the (one or more) physical register file circuitry systems 758 include a vector register unit circuitry, a write mask register unit circuitry, and a scalar register unit circuitry. These register units may provide architected vector registers, vector mask registers, general-purpose registers, and so on. The (one or more) physical register file circuitry systems 758 are coupled to the retirement unit circuitry 754 (also referred to as a retirement queue) to illustrate various ways that can be used to implement register renaming and out-of-order execution (e.g., using the (one or more) reorder buffer(s) (ROB) and the (one or more) retirement register files; using the (one or more) future heaps, the (one or more) history buffers, and the (one or more) retirement register files; using a register map and a pool of registers; and so on). The retirement unit circuitry 754 and the (one or more) physical register file circuitry systems 758 are coupled to the (one or more) execution clusters 760. The (one or more) execution clusters 760 include a set of one or more execution unit circuitry systems 762 and a set of one or more memory access circuitry systems 764. The (one or more) execution unit circuitry systems 762 may perform various arithmetic, logical, floating-point, or other types of operations (e.g., shifts, additions, subtractions, multiplications) on various types of data (e.g., scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points). While some examples may include several execution units or execution unit circuitry systems dedicated to specific functions or function sets, other examples may include only one execution unit circuitry system or multiple execution units / execution unit circuitry systems that perform all functions.(One or more) scheduler circuitry 756, (one or more) physical register file circuitry 758, and (one or more) execution clusters 760 are shown as potentially multiple because some examples create separate pipelines for certain types of data / operations (e.g., scalar integer pipelines, scalar floating point / tight integer / tight floating point / vector integer / vector floating point pipelines, and / or memory access pipelines, each having its own scheduler circuitry, (one or more) physical register file circuitry, and / or execution cluster—and in the case of a separate memory access pipeline, in some examples only the execution cluster of that pipeline has (one or more) memory access unit circuitry 764). It should also be understood that in cases where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution while the rest are in-order.

[0069] In some examples, the execution engine unit circuitry 750 may perform load / store unit (LSU) address / data pipelining to an advanced microcontroller bus (AMB) interface (not shown), as well as address phase and write-back, data phase load, store, and branch.

[0070] A set of memory access circuitry 764 is coupled to a memory unit circuitry 770 system that includes data TLB circuitry 772, which is coupled to data cache circuitry 774, which is coupled to a level 2 (L2) cache circuitry 776. In one example, the memory access circuitry 764 may include load unit circuitry, store address unit circuitry, and store data unit circuitry, each of which is coupled to the data TLB circuitry 772 in the memory unit circuitry 770. Instruction cache circuitry 734 is further coupled to the level 2 (L2) cache circuitry 776 in the memory unit circuitry 770. In one example, the instruction cache 734 and the data cache 774 are combined into a single instruction and data cache (not shown) in the L2 cache circuitry 776, a level 3 (L3) cache circuitry (not shown), and / or main memory. The L2 cache circuitry 776 is coupled to one or more other levels of cache and ultimately to main memory.

[0071] As Figure 7As further shown, portions of the processor core 790 can include circuitry for performing the following operations: identifying certain core - range errors as corrected errors, and triggering a notification to external - to - core circuitry (such as the interface circuitry described herein) when such corrected errors occur during a DLSM or other lock - step mode. To this end, the execution engine 750 includes logic circuitry 751, which can be a full - core OR logic for sending a corrected - error signal outside the core in response to an indication of a corrected error occurring at any of various locations within the core. Although not shown for ease of illustration, it should be understood that the memory unit 770 can also include such logic circuitry to identify the corrected errors described herein.

[0072] The following examples relate to further embodiments.

[0073] In one example, a device includes: a first core for executing instructions; a second core for executing instructions, where, in a lock - step mode, the first core and the second core are configured to execute in a lock - step manner; and an interface circuit coupled to the first core and the second core, where, in the lock - step mode, the interface circuit is configured to: identify a mis - comparison between the first core and the second core, the mis - comparison being caused by a corrected error in one of the first core or the second core; and indicate the mis - comparison as a recoverable error.

[0074] In the example, the interface circuit is further configured to: in response to a first value in a mis - comparison severity field of a model - specific register, indicate the mis - comparison as a recoverable error.

[0075] In the example, the interface circuit is configured to: in response to a second value in a mis - comparison severity field of a model - specific register, indicate the mis - comparison as an uncorrected error.

[0076] In the example, the device is configured to deactivate the lock - step mode in response to the mis - comparison.

[0077] In the example, the first core includes a first logic circuit configured to: receive a corrected - error indication from at least one functional circuit; and in response to the corrected - error indication, send a corrected - error signal to the interface circuit.

[0078] In the example, the interface circuit includes a corrected - error comparator configured to: receive the corrected - error signal; and in response to the corrected - error signal, indicate a mis - comparison caused by the corrected error.

[0079] In the example, the interface circuit includes a plurality of signal comparators, each of the plurality of signal comparators being configured to compare a result from the first core with a redundant result from the second core.

[0080] In an example, the interface circuit includes at least one data path comparator, and in response to a second miscomparison between a first core and a second core detected by the at least one digital path comparator, the apparatus is configured to: write at least one data to a memory hierarchy; and when a poison field of a model-specific register has a first value for enabling a poison indicator flag, mark the at least one data with the poison indicator.

[0081] In an example, the interface circuit includes a bus interface that is configured to: mark at least one data with a poison indicator; and send the at least one data and the poison indicator to a memory hierarchy.

[0082] In an example, in response to the poison indicator, the system software is configured to: clear at least one data from the memory hierarchy; and in a first software domain in which a second miscomparison occurs during its termination while a second software domain is maintained active.

[0083] In an example, in response to a data path miscomparison between a first core and a second core, the apparatus is configured to write at least one data to a memory hierarchy and is configured to: when a poison field of a model-specific register has a first value for enabling a poison indicator flag, mark the at least one data with the poison indicator; and when the poison field of the model-specific register has a second value for disabling the poison indicator flag, not mark the at least one data with the poison indicator.

[0084] In another example, a method includes: when a first core and a second core are operating in a lockstep mode, in a processor circuit system of a processor, identifying a mismatch between first data output from the first core and first redundant data output from the second core; determining whether the processor is enabled to report a toxic miscomparison; and in response to identifying the mismatch and determining that the processor is enabled to report a toxic miscomparison, associating a poison indicator with the first data and sending the first data and the poison indicator to a memory hierarchy.

[0085] In an example, the method further includes: in response to identifying the mismatch, deactivating the lockstep mode.

[0086] In an example, the method further includes: determining whether the processor is enabled to report miscomparison severity; and in response to identifying the mismatch and determining that the processor is enabled to report a toxic miscomparison and miscomparison severity, in response to the processor circuit system identifying a corrected error occurring in at least one of the first core or the second core, not associating the poison indicator with the first data and sending the first data to the memory hierarchy.

[0087] In an example, the method further includes: accessing a field of a model-specific register to determine whether the processor is enabled for a toxic miscomparison report, the field being written by system software.

[0088] In an example, the method further includes: in response to identifying a mismatch, disabling the lockstep mode; clearing the first data from the memory hierarchy; and terminating a first software domain during which the mismatch occurred.

[0089] In an example, the method further includes: maintaining at least one other software domain active while terminating the first software domain.

[0090] In another example, a computer-readable medium includes instructions for performing a method as in any of the above examples.

[0091] In a further example, a computer-readable medium includes data for use by at least one machine to fabricate at least one integrated circuit to perform a method as in any of the above examples.

[0092] In an even further example, a device includes means for performing a method as in any of the above examples.

[0093] In yet another example, a system includes a processor and a dynamic random access memory coupled to the processor. The processor can include: a first register having a first field for storing one or more first bits for enabling or disabling a corrected error indication for a miscomparison due to a corrected error in the lockstep mode; a first core for executing instructions; a second core for executing instructions, wherein, in the lockstep mode, the first core and the second core are for executing redundantly; and interface circuitry coupled to the first core and the second core, wherein the interface circuitry is for: identifying a miscomparison between the first core and the second core due to a corrected error in one of the first core or the second core; and indicating the miscomparison as a recoverable error when the one or more first bits enable the corrected error indication.

[0094] In an example, the first register further includes a second field for storing one or more second bits for indicating enabling or disabling of data poisoning for a miscomparison in the lockstep mode.

[0095] In an example, in response to a miscomparison in a lockstep mode, an interface circuit system is configured to: mark write data from a first core with a poison indicator; and send the write data and the poison indicator to a dynamic random access memory when one or more second bits enable data poisoning.

[0096] It should be understood that various combinations of the above examples are possible.

[0097] Note that the terms "circuit" and "circuit system" may be used interchangeably herein. As used herein, these terms, as well as the term "logic", are used to refer, individually or in any combination, to analog circuit systems, digital circuit systems, hardwired circuit systems, programmable circuit systems, processor circuit systems, microcontroller circuit systems, hardware logic circuit systems, state machine circuit systems, and / or any other type of physical hardware component. Embodiments may be used in many different types of systems. For example, in one embodiment, a communication device may be arranged to perform the various methods and techniques described herein. Of course, the scope of the present invention is not limited to communication devices, and instead, other embodiments may relate to other types of apparatus for processing instructions, or one or more machine-readable media that include instructions that, when executed on a computing device, cause the device to perform one or more of the methods and techniques described herein.

[0098] Embodiments can be implemented in code and stored on a non-transitory storage medium having instructions stored thereon that can be used to program a system to execute the instructions. Embodiments can also be implemented in data and stored on a non-transitory storage medium that, if used by at least one machine, causes the at least one machine to fabricate at least one integrated circuit to perform one or more operations. Further embodiments can be implemented in a computer-readable storage medium including information that, when fabricated in a system-on-a-chip (SOC) or other processor, configures the SOC or other processor to perform one or more operations. The storage medium can include, but is not limited to: any type of disk, including floppy disks, optical disks, solid state drives (SSDs), compact disk read-only memories (CD-ROMs), compact disk rewritable (CD-RWs), and magneto-optical disks; semiconductor devices such as read-only memories (ROMs), random access memories (RAMs) such as dynamic random access memories (DRAMs) and static random access memories (SRAMs), erasable programmable read-only memories (EPROMs), flash memories, electrically erasable programmable read-only memories (EEPROMs); magnetic or optical cards; or any other type of medium suitable for storing electronic instructions.

[0099] Although the present disclosure has been described with reference to a limited number of implementations, those skilled in the art who benefit from the present disclosure will appreciate numerous modifications and variations therefrom. The appended claims are intended to cover all such modifications and variations.

Claims

1. A device comprising: a first core, the first core being used to execute instructions; a second core, the second core being configured to execute instructions, wherein in lockstep mode, the first core and the second core are configured to execute in lockstep; and an interface circuit coupled to the first core and the second core, wherein, in the lockstep mode, the interface circuit is to: identify a miscomparison between the first core and the second core, the miscomparison being due to a corrected error in one of the first core or the second core; and indicate the miscomparison as a recoverable error.

2. The device according to claim 1, wherein: The interface circuit is further configured to indicate the miscompare as the recoverable error in response to a first value in a miscompare severity field of a model specific register.

3. The device as claimed in claim 2, wherein: The interface circuit is to indicate the miscompare as an uncorrected error in response to a second value in the miscompare severity field of the model specific register.

4. The device according to claim 1, wherein: The means is for disabling the lockstep mode in response to the miscomparison.

5. The device according to claim 1, wherein: The first core includes a first logic circuit configured to: receive a corrected error indication from at least one functional circuit; and send a corrected error signal to the interface circuit in response to the corrected error indication.

6. The device according to claim 5, wherein: The interface circuit includes a corrected error comparator configured to: receive the corrected error signal; and, in response to the corrected error signal, indicate the miscomparison due to the corrected error.

7. The device according to claim 6, wherein: The interface circuit includes a plurality of signal comparators, each of the plurality of signal comparators being used to compare a result from the first core with a redundant result from the second core.

8. The device of claim 1, wherein: The interface circuit includes at least one data path comparator, and in response to a second miscomparison between the first core and the second core detected by the at least one digital path comparator, the device is used to: write at least one data to a memory hierarchy; and mark the at least one data with a poison indicator when a poison field of a model-specific register has a first value for enabling a poison indicator mark.

9. The device of claim 8, wherein: The interface circuit includes a bus interface for: marking the at least one data with the poison indicator; and sending the at least one data and the poison indicator to the memory hierarchy.

10. The device of claim 8, wherein: In response to the poison indicator, system software is to: flush the at least one data from the memory hierarchy; and terminate the first software domain during which the second miscompare occurred, while a second software domain is maintained active.

11. The device according to any one of claims 1 to 10, wherein: In response to a data path miscompare between the first core and the second core, the apparatus is configured to write at least one data to a memory hierarchy, and to: marking the at least one data with a poison indicator when the poison field of the model specific register has a first value for enabling the poison indicator marking; as well as When the poison field of the model specific register has a second value for disabling the poison indicator marking, the at least one data is not marked with the poison indicator.

12. A method comprising: identifying, in processor circuitry of a processor, a mismatch between first data output from the first core and first redundant data output from the second core when the first core and the second core are operating in a lockstep mode; determining whether the processor is enabled for poison miscompare reporting; as well as In response to identifying the mismatch and determining that the processor is enabled for the poisoned miscompare reporting, a poison indicator is associated with the first data, and the first data and the poison indicator are sent to a memory hierarchy.

13. The method of claim 12, wherein: The method further includes, in response to identifying the mismatch, disabling the lockstep mode.

14. The method of claim 12, wherein: The method further comprises: determining whether the processor is enabled for miscompare severity reporting; and In response to identifying the mismatch and determining that the processor is enabled to perform the poisoned miscompare reporting and the miscompare severity reporting, in response to the processor circuitry identifying that a corrected error has occurred in at least one of the first core or the second core, not associating the poison indicator with the first data, and sending the first data to the memory hierarchy.

15. The method of claim 12, wherein: The method further includes accessing a field of a model specific register to determine whether the processor is enabled for the poisoned miscompare reporting, the field being written by system software.

16. The method of claim 12, wherein: The method further comprises: responsive to identifying the mismatch, deactivating the lockstep mode; clearing the first data from the memory hierarchy; and The first software domain during which the mismatch occurred is terminated.

17. The method of claim 16, wherein: The method further includes maintaining at least one other software domain active while terminating the first software domain.

18. A computer program product comprising instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 12 to 17.

19. A system comprising: A processor, the processor comprising: a first register having a first field for storing one or more first bits for enabling or disabling a corrected error indication for a miscompare due to a corrected error in a lockstep mode; a first core, the first core being used to execute instructions; a second core, the second core configured to execute instructions, wherein in the lockstep mode, the first core and the second core are configured to execute redundantly; and an interface circuit system coupled to the first core and the second core, wherein the interface circuit system is to: identify a miscomparison between the first core and the second core, the miscomparison being due to a corrected error in one of the first core or the second core; and indicate the miscomparison as a recoverable error when the one or more first bits enable the corrected error indication; and A dynamic random access memory is coupled to the processor.

20. The system of claim 19, wherein: The first register further includes a second field for storing one or more second bits for indicating enabling or disabling of data poisoning for miscomparison in the lockstep mode.

21. The system of claim 20, wherein: In response to the miscompare in the lockstep mode, the interface circuit system is to: mark write data from the first core with a poison indicator; and send the write data and the poison indicator to the dynamic random access memory when the one or more second bits enable the data poisoning.

22. A device comprising: A first core device for executing instructions; a second core device for executing instructions, wherein in a lockstep mode, the first core device and the second core device are configured to execute in lockstep; as well as An interface device, wherein, in the lockstep mode, the interface device is used to: identify a miscomparison between the first core device and the second core device, the miscomparison being caused by a corrected error in one of the first core device or the second core device; and indicate the miscomparison as a recoverable error.

23. The apparatus of claim 22, wherein: The interface device is further configured to indicate the miscompare as the recoverable error in response to a first value in a miscompare severity field of a model specific register.

24. The apparatus of claim 23, wherein: The interface device is configured to indicate the miscompare as an uncorrected error in response to a second value in the miscompare severity field of the model specific register.