Delayed error handling

Through the delay error handling mechanism, the system stops the container running when an uncorrectable error is encountered, and waits for other containers to reach a recoverable state to deal with the error, which solves the data loss and system crash caused by uncorrectable errors, and improves the availability and reliability of the system.

CN108984329BActive Publication Date: 2025-06-24INTEL CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201810546266.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-05-31
Filing Date
2018-05-31
Publication Date
2025-06-24
Estimated Expiration
2038-05-31

AI Technical Summary

Technical Problem

In multi-containerized systems, non-correctable errors in one container can cause hierarchical failures and even lead to data loss across containers, especially in unexpected interactions between lower priority containers and higher priority containers.

Method used

Provides a delayed error handling mechanism that allows the container to stop running when an uncorrectable error is encountered instead of immediately performing catastrophic error recovery. The system will wait for other containers to reach a recoverable state and then perform error handling, including notifying other containers to find the recoverable state or prepare for error handling.

Benefits of technology

By delaying error handling, the system can avoid data loss and system crashes caused by uncorrectable errors, improve system availability and reliability, and reduce the performance impact of error recovery on other containers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN108984329B_ABST
    Figure CN108984329B_ABST
Patent Text Reader

Abstract

A computing device includes: a hardware platform including a processor and a memory; and a System Management Interrupt (SMI) handler; a first logic configured to provide, via the hardware platform, a first container and a second container; and a second logic configured to: detect an uncorrectable error in the first container; generate a degraded system state in response to the detection; provide a degraded state message to the SMI handler; instruct the second container to find a recoverable state; determine that the second container has entered the recoverable state; and initiate a recovery operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of cloud computing, and more specifically but not exclusively, to a system and method for handling latency errors. Background Art

[0002] Modern computing practices have abandoned dedicated hardware computing and moved towards "network as a device". Modern networks can include data centers that host large numbers of general-purpose hardware server devices, which are contained in, for example, server racks and controlled by a hypervisor. Each hardware device can run one or more instances of virtual devices such as workload servers or virtual desktops. Brief Description of the Drawings

[0003] The present disclosure is best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be emphasized that, according to standard practice in the industry, the various features are not necessarily drawn to scale and are for illustrative purposes only. Where scale is explicitly or implicitly shown, it only provides an illustrative example. In other embodiments, for the sake of clarity in discussion, the dimensions of the various features may be increased or decreased arbitrarily.

[0004] Figure 1 is a network-level diagram of a cloud service provider (CSP) according to one or more examples of the present specification.

[0005] Figure 2 is a block diagram of a data center according to one or more examples of the present specification.

[0006] Figure 3 Illustrates a block diagram of a central processing unit according to one or more examples of the present specification.

[0007] Figure 4 is a block diagram of a data center computing architecture according to one or more examples of the present specification.

[0008] Figure 5 is a block diagram illustrating how the recovery of uncorrectable errors affects multiple containers according to one or more examples of the present specification.

[0009] Figure 6a –6b is a signal flow diagram of a method for performing latency error handling according to one or more examples of the present specification. Detailed Description

[0010] The following disclosure provides many different embodiments or examples for implementing the different features of the present disclosure. Specific examples of components and arrangements are described below to simplify the present disclosure. These are of course merely examples and are not intended to be limiting. Additionally, the present disclosure may repeat reference numerals and / or letters in various examples. Such repetition is for simplicity and clarity purposes and does not inherently indicate a relationship between the various embodiments and / or configurations being discussed. Different embodiments may have different advantages, and not necessarily any embodiment requires a particular advantage.

[0011] In a modern data center, very high computing densities can be achieved. For example, a batch of high-performance computing platforms can be aggregated into a blade chassis or a compute sled, and the chassis can then consume one or more slots in a rack chassis. A rack with several high-density computing nodes of this type can thus host dozens or hundreds of cores in a single rack with a capacity of, for example, 42 U or similar.

[0012] Software engineering techniques can target each core adopting such an architecture to run a single thread in a multithreaded process. A single application can have multiple threads and, thus, can consume multiple processor cores. One or more additional cores can also be dedicated to providing the operating system and / or other support software.

[0013] In some cases, to save the overhead of providing a separate operating system for each discrete application, a single operating system can run many "containers" while also maintaining some logical separation between the applications. These containers can share low-level operating system resources but, on the other hand, can be isolated from each other.

[0014] Such an architecture supports providing computing, storage, communication, and acceleration resources that can be provided in a structurally connected data center.

[0015] The advantage of the containers as described above is that they provide a modular and flexible rack-scale implementation for Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS).

[0016] One challenge in such multi-containerized systems is that a single uncorrectable error in one container can cause a failure of the hierarchical structure and, in some cases, can cause the underlying operating system to fail, thereby causing data loss across other containers. This can be particularly challenging in cases where a lower-priority container may encounter an error and may thereby cause a higher-priority container to fail. In some examples, the situation may be exacerbated because lower-priority containers may have a less robust programming model, while higher-priority containers may be more "solid" and more robust. Thus, the less robust lower-priority containers may cause unexpected interactions with the more robust higher-priority containers.

[0017] For example, an implementation may include one container providing an email server and another container providing a high-availability database server. If the low-priority email server encounters an uncorrectable error (such as a corrupted memory location), the operating system error handling routine may need to be fully restarted to ensure memory integrity. Unfortunately, this full restart will not only affect the low-priority email server, but will also affect the high-priority high-availability database server. Additionally, if a failure occurs while the database server is performing a critical operation (such as a database write operation), the failure may actually result in the corruption of one or more records in the database itself.

[0018] Although this problem can be partially avoided by only providing homogeneous containers on a single operating system (such as only providing other database servers for database servers in a single operating system), this strategy may affect the advantages of containerized computing. Additionally, such errors cannot be completely avoided because even very robust applications may encounter errors. Thus, if multiple containers are each running very robust database servers and if one of those database servers encounters an uncorrectable error, it will cause all database servers to fail.

[0019] Examples of such uncorrectable errors include errors in the memory subsystem RAS stack (hardware, firmware, and / or software). Attempts to recover from such uncorrectable errors may include enhancing MCA generation to the base firmware model.

[0020] Therefore, it is beneficial to provide a system that can more gracefully recover from uncorrectable errors. In particular, instead of immediately going to an error handler routine when an application running in a single container encounters an uncorrectable error, it is possible to implement deferred error handling. This deferred error handling is feasible because although containers may share underlying operating system services, they generally do not share memory pages or other resources. Therefore, the fact that a memory page may be corrupted or inaccessible for one container should not affect another container. Thus, instead of immediately going to a catastrophic error recovery that may cause other containers to fail and result in data loss, deferred error handling can be implemented. With deferred error handling, other containers can be notified to look for a "recoverable state" or, on the other hand, to prepare for error handling. As used throughout this specification, a "recoverable state" is a state in which the workload of a node is completed, minimized, or reduced and, therefore, the risk of data loss or data corruption is also eliminated, minimized, or reduced. Compared to the active state of a container, the recoverable state can be a relatively quiescent state. For example, if the container is a database driver, the recoverable state can be a state in which it no longer accepts new incoming database connections and all outstanding operations have been completed and committed. In the case of a web server, the recoverable state can be a state in which it no longer accepts incoming HTTP connections and all outstanding transactions have been reasonably disposed of. In the case of a computing node (e.g., for large-scale parallel computing), the recoverable state can be a state in which it no longer accepts incoming computing transactions and the existing transactions have been completed and output. In another example, the recoverable state can include a state in which the container can be migrated to a new hardware platform with minimal data loss, in which case the container can be migrated before error recovery occurs.

[0021] Although error recovery is deferred, the container that encountered the error itself can be stopped. Since it has encountered an uncorrectable error, it may not be able to continue computing or processing. However, in a flexible data center that employs software-defined networking and network function virtualization, it is often feasible to spawn a new instance of the service to handle any additional workload that was lost due to the loss of one of the failed containers.

[0022] Note that, as disclosed herein, "finding" a recoverable state does not necessarily require a "degraded" (but not stopped) container to immediately stop receiving incoming transactions. To avoid a data center service crash, a degraded container can continue to receive new incoming connections or transactions while the data center is under higher load, and can wait until the load gradually decreases to stop receiving new incoming connections or transactions. However, since stopped containers are consuming resources that cannot be allocated elsewhere, error handling may not be indefinitely postponed. Instructions to find a recoverable state can include a timeout. If the container does not reach its optimal recoverable state before the timeout expires, error handling can be processed in any way such that resources consumed by the stopped node can be brought back into circulation on the data center.

[0023] The ultimate response to an uncorrectable error can depend on the system capabilities. For example, the response can include shutdown and restart or error recovery. In some cases where error recovery is disabled or unavailable, the response can typically be a shutdown of the computing resources, followed by error harvesting, and then restart. Although this error recovery may seem transparent from the perspective of the operating system, it can in fact have a severe performance impact on the containers that are currently executing. This is because, for recovery, several hardware subunits may have to be reprogrammed and restored, and transactions restarted in some cases. This means that the pre-error state may be lost and may have to be rebuilt by software re-execution. Additionally, sometimes the high-level recovery routines are very complex and may require several system management routines to be completed. For example, mirror failover may require all system address mappings to be updated to reflect the failed memory as the main memory, which can take a significant amount of time.

[0024] In another example, memory sparing is used, where the engine copies all previously saved data in the primary record to the secondary record and then marks the secondary record as the primary record. Again, this may require several iterations of the SMI flow to be completed, regardless of whether the system has recovered. In any case, whether the operating system treats the error as non-recoverable or recoverable, other tasks (such as apps or containers) may see data loss or performance degradation or both.

[0025] Therefore, it is beneficial to have the following situation: provide a delayed error recovery mechanism that waits to attempt error recovery until other apps or containers are in a more suitable (recoverable) state so that they will not lose data.

[0026] In an embodiment, the system can have an interface for receiving instructions regarding error recovery options, such as from a user or a coordinator. For example, a system administrator can operate a user interface for the coordinator, where she defines error recovery policies in the coordinator, including the error recovery actions to be taken, the maximum timeout for certain types of containers, parameters for how a container can find a recoverable state when receiving a degradation notice, and how the coordinator can help, such as by steering traffic away from the degraded container and / or generating new instances of the services provided by the stopped container (either autonomously or in response to a request from the hardware platform hosting the stopped container).

[0027] Embodiments of this specification support recovery from critical errors, such as memory errors and cache errors. As non-limiting examples, supported error types can include error-correcting code (ECC) memory, PCIe data errors, and other examples of "poisonous" data. In some embodiments, a distinction can be made between deferrable errors and other errors that should not be deferred (such as parity errors on a buffer). Accordingly, certain embodiments of this specification identify such errors and recover them immediately, regardless of the deferrable error mechanism.

[0028] Error suppression can also be a consideration in embodiments of this specification. It may be desirable to ensure that, in the case of a data error, other data is not corrupted. Accordingly, data isolation can be used. This can include preventing write-back from a corrupted memory I / O, such as a hard disk drive or other external data source.

[0029] While a deferred error recovery is pending, other apps and containers can continue to operate gracefully until they reach a good "stopping point". This can include allowing other containers to complete program execution or continue program execution for a reasonable period of time. Some embodiments can notify the scheduler not to schedule more tasks until the error is disposed of. Accordingly, not only are the containers allowed to remain running, but measures can be taken to ensure that additional workload is removed from it so that it can reach a recoverable state.

[0030] In certain embodiments, a timeout can also be specified for the case where a container runs longer than expected by the system. In other words, an unaffected container can be requested to start reducing its workload and reach a recoverable state. This can be issued by the operating system or firmware. Once this request is issued, the operating system or firmware can set a timeout after which, if the container has not notified that it has reached a graceful recovery state, it can be forced to shut down or power off. This can be similar to a maintenance mode shutdown request. Accordingly, although the application is given the opportunity to reach a recoverable state, it is not given unlimited time to do so, as that could affect the ability for the ultimate error recovery to occur.

[0031] The operating system is also given the ability to suspend the offending application, making further data from that application unavailable. Thus, the operating system scheduler has the ability to ensure that there are no further calls to the offending application. The state of the offending application can be kept available for later debugging and reference.

[0032] In some embodiments, the system agent firmware may also notify the system administrator or other management tasks that the system is running in a degraded mode. This can take the form of, for example, issuing a notification to a data center coordinator. This ensures that the data center coordinator begins to direct traffic away from the offending platform so that it can reach a recoverable state. For example, the coordinator may instruct the load balancer not to dispatch the degraded container to any new load balancing buckets and may also instruct the load balancer to gradually redirect the buckets from the degraded container to other containers. This allows the degraded container to gradually shut down so that it can reach a recoverable state.

[0033] Beneficially, when a critical error in another container related to memory and other data paths and subsystems is encountered, delaying error handling allows other containers to avoid data loss or data corruption. Embodiments may also create an architectural dump of the default app or container, which can subsequently be replayed in recovery mode or software to allow for possible failure analysis and later debugging.

[0034] Systems and methods for delaying error handling will now be described more specifically with reference to the accompanying drawings. It should be noted that throughout the drawings, certain reference numerals may be repeated to indicate that a particular device or block is identical or substantially identical between the various drawings. However, this is not intended to imply any particular relationship between the various disclosed embodiments. In some examples, a class of elements may be represented by a particular reference numeral ("widget 10"), while individual species or examples of that class may be represented by numbers with hyphens ("first particular widget 10-1" and "second particular widget 10-2").

[0035] Figure 1 is a network-level diagram of a network 100 of a cloud service provider (CSP) 102 according to one or more examples of the present specification. As a non-limiting example, the CSP 102 can be a traditional enterprise data center, an enterprise "private cloud", or a "public cloud", providing services such as infrastructure as a service (IaaS), platform as a service (PaaS), or software as a service (SaaS).

[0036] The CSP 102 can provide a certain number of workload clusters 118, which can be clusters of individual servers, blade servers, rack-mounted servers, or any other suitable server topologies. In this illustrative example, two workload clusters 118-1 and 118-2 are shown, with each workload cluster providing rack-mounted servers 146 in a chassis 148.

[0037] Each server 146 can host an independent operating system and provide server functionality, or the servers can be virtualized, in which case they can be under the control of a virtual machine manager (VMM), hypervisor, and / or coordinator, and can host one or more virtual machines, virtual servers, or virtual appliances. These server racks can be co-located in a single data center or can be located in different geographical data centers. According to a contract agreement, some servers 146 can be specifically dedicated to certain enterprise customers or tenants, while other servers 146 can be shared.

[0038] Various devices in the data center can be connected to each other via a switching fabric 170, which can include one or more high-speed routing and / or switching devices. The switching fabric 170 can provide both "north-south" traffic (e.g., traffic to and from a wide area network (WAN), such as the Internet) and "east-west" traffic (e.g., traffic across the data center). Historically, north-south traffic has accounted for the majority of network traffic, but as web services have become more complex and distributed, east-west traffic volume has increased. In many data centers, east-west traffic now accounts for the majority of traffic.

[0039] In addition, as the capabilities of each server 146 increase, the traffic volume can further increase. For example, each server 146 can provide multiple processor slots, each slot accommodating a processor with four to eight cores and sufficient memory for the cores. Thus, each server can host many VMs, with each VM generating its own traffic.

[0040] To accommodate the large amount of traffic in the data center, a high-performance switching fabric 170 can be provided. The switching fabric 170 is illustrated in this example as a "flat" network, where each server 146 can have a direct connection to a top-of-rack (ToR) switch 120 (e.g., a "star" configuration), and each ToR switch 120 can be coupled to a core switch 130. This two-tier flat network architecture is shown only as an illustrative example. In other examples, other architectures can be used, as non-limiting examples, such as a three-tier star or leaf-spine (also known as "fat tree" topology) based on the "Clos" architecture, a hub-and-spoke topology, a mesh topology, a ring topology, or a 3-D mesh topology.

[0041] The structure itself can be provided via any suitable interconnection. For example, each server 146 can include a fabric interface such as an Intel® Host Fabric Interface (HFI), a Network Interface Card (NIC), or other host interface. The host interface itself can be coupled to one or more processors via an interconnection or bus such as PCI, PCIe, or the like, and in some cases, such an interconnection bus can be considered part of the fabric 170.

[0042] The interconnection technology can be provided by a single interconnection or a hybrid interconnection, where PCIe provides on-chip communication, 1 Gb or 10 Gb copper Ethernet provides a relatively short connection to the ToR switch 120, and optical cables provide a relatively long connection to the core switch 130. As a non-limiting example, interconnection technologies include Intel® OmniPath™, TrueScale™, Ultra Path Interconnect (UPI) (formerly known as QPI or KTI), STL, Fibre Channel, Ethernet, Fibre Channel over Ethernet (FCoE), InfiniBand, PCI, PCIe, or optical fibre, etc. Some of these interconnection technologies will be more suitable for certain deployments or functions than others, and it is the job of one of ordinary skill in the art to select a suitable fabric for an immediate application.

[0043] However, it should be noted that although high-end fabrics such as OmniPath™ are provided here for illustration, more generally, the fabric 170 can be any suitable interconnection or bus for a particular application. In some cases, this can include traditional interconnections such as Local Area Networks (LANs), Token Ring Networks, Synchronous Optical Networks (SONETs), Asynchronous Transfer Mode (ATM) Networks, wireless networks (such as WiFi and Bluetooth), "Plain Old Telephone System" (POTS) interconnections, etc. It is also explicitly contemplated that in the future, new network technologies will emerge to supplement or replace some of those listed here, and any such future network topologies and technologies can be part of or form the fabric 170.

[0044] In certain embodiments, the fabric 170 can provide communication services at various "layers", as initially outlined in the OSI seven-layer network model. In contemporary practice, the OSI model is not strictly followed. Generally speaking, layers 1 and 2 are often referred to as the "Ethernet" layers (but in large data centers, Ethernet has often been replaced by newer technologies). Layers 3 and 4 are often referred to as the Transmission Control Protocol / Internet Protocol (TCP / IP) layers (which can be further subdivided into TCP and IP layers). Layers 5 - 7 can be referred to as the "application layer". These layer definitions are presented as a useful framework but are intended to be non-limiting.

[0045] Figure 2 is a block diagram of data center 200 according to one or more examples of this specification. In various embodiments, data center 200 can be the same data center as data center 100, or can be a different data center. Additional views are provided in Figure 1 to illustrate different aspects of data center 200. Figure 2

[0046] In this example, fabric 270 is provided to interconnect various aspects of data center 200. Fabric 270 can be the same as fabric 170 of Figure 1 , or can be a different fabric. As described above, fabric 270 can be provided by any suitable interconnect technology. In this example, Intel® OmniPath™ is used as an illustrative and non-limiting example.

[0047] As illustrated, data center 200 includes many logical elements that form multiple nodes. It should be understood that each node can be provided by a physical server, a group of servers, or other hardware. Each server can be running one or more virtual machines suitable for its application.

[0048] Node 0 208 is a processing node that includes processor socket 0 and processor socket 1. The processor can be, for example, an Intel® Xeon™ processor with multiple cores (such as 4 or 8 cores). Node 0 208 can be configured to provide network or workload functions, such as by hosting multiple virtual machines or virtual appliances.

[0049] On-board communication between processor socket 0 and processor socket 1 can be provided by on-board uplink 278. This can provide a very high-speed, short-length interconnect between the two processor sockets, enabling virtual machines running on node 0 208 to communicate with each other at very high speeds. To facilitate this communication, a virtual switch (vSwitch) can be provided on node 0 208, and the virtual switch (vSwitch) can be considered part of fabric 270.

[0050] Node 0 208 is connected to fabric 270 via fabric interface 272. Fabric interface 272 can be any suitable fabric interface as described above, and in this particular illustrative example, can be an Intel® host fabric interface for connecting to an Intel® OmniPath™ fabric. In some examples, a channel can be opened for communication with fabric 270, such as by providing UPI tunneling via OmniPath™.

[0051] ​Because data center 200 can provide many functions that were provided on-board in previous generations in a distributed manner, a high-performance fabric interface 272 can be provided. The fabric interface 272 can operate at speeds of several gigabits per second and, in some cases, can be tightly coupled with node 0 208. For example, in some embodiments, the logic for the fabric interface 272 is integrated directly with the processors on the system-on-chip. This provides ultra-high-speed communication between the fabric interface 272 and the processor socket without the need for an intermediate bus device that might introduce additional latency into the fabric. However, this does not imply that embodiments that provide the fabric interface 272 on a traditional bus will be excluded. On the contrary, it is explicitly anticipated that, in some examples, the fabric interface 272 can be provided on a bus such as a PCIe bus, which is a serialized version of PCI and provides higher speeds than traditional PCI. Throughout data center 200, the various nodes can provide different types of fabric interfaces 272, such as on-board fabric interfaces and plug-in fabric interfaces. It should also be noted that certain blocks in the system-on-chip can be provided as intellectual property (IP) blocks that can be "dropped" into an integrated circuit as modular units. Thus, the fabric interface 272 can be derived from such IP blocks in some cases.

[0052] Note that, in a "network as a device" manner, node 0 208 can provide limited on-board memory or storage, or no on-board memory or storage. Instead, node 0 208 can rely primarily on distributed services such as memory servers and networked storage servers. On-board, node 0 208 can provide only enough memory and storage to boot the device and communicate with the fabric 270. Because of the ultra-high speeds of modern data centers, this distributed architecture is possible and can be beneficial because there is no need to over-provision resources for each node. Instead, a large amount of high-speed or dedicated memory can be provided dynamically among many nodes such that each node can access a large amount of resources, but those resources are not idle when that particular node does not need them.

[0053] In this example, node 1 memory server 204 and node 2 storage server 210 provide the operating memory and storage capabilities for node 0 208. For example, memory server node 1 204 can provide Remote Direct Memory Access (RDMA), whereby node 0 208 can access the memory resources on node 1 204 in a DMA fashion via fabric 270, similar to the way it would access its own on-board memory. The memory provided by memory server 204 can be traditional memory, such as volatile Double Data Rate type 3 (DDR3) Dynamic Random Access Memory (DRAM), or it can be a more unique type of memory, such as Persistent Fast Memory (PFM), like Intel® 3D Crosspoint™ (3DXP), which operates at the same speed as DRAM but is non-volatile.

[0054] Similarly, instead of providing on-board hard disks for node 0 208, storage server node 2 210 can be provided. Storage server 210 can provide Networked Block of Disks (NBOD), PFM, Redundant Array of Independent Disks (RAID), Redundant Array of Independent Nodes (RAIN), Network Attached Storage (NAS), optical storage devices, tape drives, or other non-volatile memory solutions.

[0055] Thus, when performing its designated functions, node 0 208 can access the memory from memory server 204 and store the results on the storage device provided by storage server 210. Each of these devices is coupled to fabric 270 via fabric interface 272, and fabric interface 272 provides the fast communication that enables these technologies.

[0056] As a further illustration, node 3 206 is also depicted. Node 3 206 also includes fabric interface 272 and two processor sockets internally connected by an uplink. However, unlike node 0 208, node 3 206 includes its own on-board memory 222 and storage device 250. Thus, node 3 206 can be configured to primarily perform its functions on-board and may not need to rely on memory server 204 and storage server 210. However, in suitable cases, similar to node 0 208, node 3 206 can supplement its own on-board memory 222 and storage device 250 with distributed resources.

[0057] The basic building blocks of the various components disclosed herein may be referred to as “logic elements”. Logic elements may include hardware (including, for example, software programmable processors, ASICs, or FPGAs), external hardware (digital, analog, or mixed signal), software, interactive software, services, drivers, interfaces, components, modules, algorithms, sensors, components, firmware, microcode, programmable logic, or objects that can cooperate to perform logical operations. Additionally, some logic elements are provided by a tangible non-transitory computer-readable medium having executable instructions stored thereon for instructing a processor to perform a certain task. As a non-limiting example, such non-transitory media may include, for example, hard drives, solid state memories or disks, read only memories (ROMs), persistent fast memories (PFM) (e.g., Intel® 3D Crosspoint™), external storage devices, redundant arrays of independent disks (RAID), redundant arrays of independent nodes (RAIN), network attached storage devices (NAS), optical storage devices, tape drives, backup systems, cloud storage devices, or any combination of the foregoing. Such media may also include instructions written into an FPGA or encoded in the hardware of an ASIC or processor.

[0058] Figure 3 FIG. is a block diagram of a central processing unit (CPU) 312 according to certain embodiments. Although the CPU 312 depicts a particular configuration, the cores and other components of the CPU 312 may be arranged in any suitable manner. The CPU 312 may include any processor or processing device, such as a microprocessor, an embedded processor, a digital signal processor (DSP), a network processor, an application processor, a coprocessor, a system on a chip (SOC), or other device for executing code. In the depicted embodiment, the CPU 312 includes four processing elements (cores 330 in the depicted embodiment), which may include asymmetric or symmetric processing elements. However, the CPU 312 may include any number of processing elements, which may be symmetric or asymmetric.

[0059] Examples of hardware processing elements include: thread units, thread slots, threads, process units, contexts, context units, logical processors, hardware threads, cores, and / or any other element capable of holding the state of a processor (such as, an execution state or an architectural state). In other words, in one embodiment, a processing element represents any hardware capable of independently associating with code (such as, software threads, operating systems, applications, or other code). A physical processor (or processor socket) typically represents an integrated circuit, which may include any number of other processing elements, such as cores or hardware threads.

[0060] A core may represent logic on an integrated circuit that is capable of maintaining an independent architectural state, where each independently maintained architectural state is associated with at least some dedicated execution resources. A hardware thread may represent any logic on an integrated circuit that is capable of maintaining an independent architectural state, where the independently maintained architectural states share access to execution resources. A physical CPU may include any suitable number of cores. In various embodiments, a core may include one or more out-of-order processor cores or one or more in-order processor cores. However, each core may be individually selected from any type of core, such as a native core, a software-managed core, a core adapted to execute a native instruction set architecture (ISA), a core adapted to execute a translated ISA, a co-designed core, or other known cores. In a heterogeneous core environment (i.e., asymmetric cores), some form of translation, such as binary translation, may be used to schedule or execute code on one or both cores.

[0061] In the depicted embodiment, core 330A includes an out-of-order processor having a front-end unit 370 that is configured to fetch incoming instructions, perform various processing (e.g., caching, decoding, branch prediction, etc.), and pass the instructions / operations forward to an out-of-order (OOO) engine. The OOO engine performs further processing on the decoded instructions.

[0062] The front-end 370 may include a decoding module that is coupled to the fetch logic to decode the fetched elements. In one embodiment, the fetch logic includes individual sequencers associated with the thread slots of core 330. Generally, core 330 is associated with a first ISA that defines / specifies the instructions that may be executed on core 330. Machine code instructions that are part of the first ISA often include a portion of the instruction (referred to as an opcode) that references / specifies the instruction or operation to be executed. The decoding module may include circuitry that identifies these instructions from their opcodes and passes the decoded instructions in a pipeline for processing as defined by the first ISA. In one embodiment, the decoder of core 330 identifies the same ISA (or a subset thereof). Alternatively, in a heterogeneous core environment, the decoder of one or more cores (e.g., core 330B) may identify a second ISA (a subset of the first ISA or a different ISA).

[0063] In the depicted embodiment, the out-of-order engine includes an allocation unit 382 for receiving decoded instructions from the front-end unit 370 (the decoded instructions may have one or more microinstructions or uin the form of ops), and allocate them to appropriate resources (such as registers, etc.). Next, the instructions are provided to the reservation station 384, which reserves resources and schedules their execution on one of the multiple execution units 386A - 386N. There can be various types of execution units, including, for example, arithmetic logic units (ALUs), load and store units, vector processing units (VPUs), floating - point execution units, etc. The results from these different execution units are provided to the re - order buffer (ROB) 388, which takes the out - of - order results and restores them to the correct program order.

[0064] In the depicted embodiment, both the front - end unit 370 and the out - of - order engine 380 are coupled to different levels of the memory hierarchy. Specifically shown is the instruction - level cache 372, which in turn is coupled to the mid - level cache 376, which in turn is coupled to the last - level cache 395. In one embodiment, the last - level cache 395 is implemented in the on - chip (sometimes referred to as non - core) unit 390. The non - core 390 can communicate with the system memory 399, which in the illustrated embodiment is implemented via embedded DRAM (eDRAM). The various execution units 686 within the OOO engine 380 communicate with the first - level cache 374, which also communicates with the mid - level cache 376. Additional cores 330B–330D can also be coupled to the last - level cache 395.

[0065] In a particular embodiment, the non - core 390 can be located in a voltage domain and / or frequency domain separate from that of the core. That is, the non - core 390 can be powered by a supply voltage different from the supply voltage used to power the core, and / or can operate at a frequency different from the operating frequency of the core.

[0066] The CPU 312 can also include a power control unit (PCU) 340. In various embodiments, the PCU 340 can control the supply voltage and operating frequency applied to each core (on a per - core basis) and to the non - core. The PCU 340 can also instruct the core or non - core to enter an idle state (where no voltage and clock are supplied) when not executing a workload.

[0067] In various embodiments, the PCU 340 may detect one or more stress characteristics of hardware resources such as cores and non-cores. The stress characteristics may include an indication of the amount of stress applied to the hardware resources. As an example, the stress characteristic may be the voltage or frequency applied to the hardware resource; the power level, current level, or voltage level sensed at the hardware resource; the temperature sensed at the hardware resource; or other suitable measurements. In various embodiments, when sensing the stress characteristics at a particular moment, multiple measurements of the particular stress characteristic may be performed (e.g., at different locations). In various embodiments, the PCU 340 may detect the stress characteristics at any suitable interval.

[0068] In various embodiments, the PCU 340 is a component discrete from the core 330. In a particular embodiment, the PCU 340 operates at a clock frequency different from the clock frequency used by the core 630. In some embodiments where the PCU is a microcontroller, the PCU 340 executes instructions according to an ISA different from the ISA used by the core 330.

[0069] In various embodiments, the CPU 312 may further include non-volatile memory 350 for storing stress information associated with the core 330 or non-core 390 (such as stress characteristics, incremental stress values, accumulated stress values, stress accumulation rates, or other stress information) such that the stress information is retained when power is lost.

[0070] Figure 4 is a block diagram of a data center computing architecture according to one or more examples of the present specification. Architecture 400 illustrates the correlation of various components in a data center such that error handling can affect more than one container.

[0071] In this example, the blade chassis 404 includes blades 404-1 to 404-n. The blade chassis 404 may be implemented as a multi-u module 408, and the multi-u module 408 includes computing modules 408-1 to 408-n. The computing modules 408 may be loaded into drawers 412. The drawers 412 may be installed in one or more slots of a rack 416. And the rack 416 may be part of a replaceable data center enclosure 420.

[0072] The blades 404-1 to 404-n may provide multiple resources such as processors 424, memories 428, fabric interfaces 432, and storage containers 436. The fabric interface 432 may couple the computing nodes to a fabric module 440. The fabric module 440 may provide switch ports 448 and be connected to a VLAN 444. The VLAN 444 may include many VLAN ports 452.

[0073] The storage container 436 may include an iSCSI target 438 and one or more logical drives 456, each of which is hosted on a physical drive 460.

[0074] As discussed above, various containers running on an operating system instance may isolate many of these resources from one another in a single computing node.

[0075] However, the error recovery paradigm in such a system may be based on the machine check architecture. In this case, when an error is consumed, the machine check exception is immediately disposed of. For recoverable errors, the corresponding task is terminated by the operating system, and the OS may then isolate the memory from further use to avoid future errors. In the case of non-recoverable errors, a cold restart is used to bring the entire machine down to resume nominal system operation.

[0076] In the case of such non-recoverable errors, when bringing the entire machine down, a firmware-first model may be applied to mitigate some issues, but firmware-first may not necessarily solve data loss. Additionally, there may be minimal hardware error harvesting when tasks are terminated or the machine is restarted. This may provide minimal insight into application-level fault analysis. For example, the Microsoft bug check 0x124 named Windows Hardware Error Architecture (WHEA)_uncorrectable_error does not provide information other than the WHEA record.

[0077] In contrast, with deferred error handling, the faulty memory region can be immediately isolated for both recoverable and non-recoverable error scenarios. This optimizes operational costs by providing the option for operators to plan services. It also improves manageability by improving task scheduling control (such as machine development and operations (DevOp)). Additionally, it can provide predictive fault analysis of memory type or manufacturing.

[0078] Regarding data loss, the system described here removes the containers affected by the error and saves the machine check context for future services. This allows the remaining containers to continue execution without immediate interruption, thus reducing data loss and increasing system uptime and availability. Regarding fault analysis, in addition to error records, the current solution may also save task-specific records or contexts in memory. This allows additional post-processing or postmortem fault analysis.

[0079] Figure 5 is a block diagram illustrating how the recovery of uncorrectable errors according to one or more examples of this specification affects multiple containers. In Figure 5In the example, blade system 502 hosts computing platform 503, and computing platform 503 provides the hardware for the data center process. Operating system 512 runs on blade system 502. Operating system 512 hosts multiple tasks or applications 508 and also provides two containers, namely container 504-1 and container 504-2.

[0080] Hardware platform 503 may include various components to provide hardware services for the software components of the system. This may include multiple cores 522-1 to 522-6. Cores 522 may access one or more DRAM modules, such as DRAM 516-1 and DRAM 516-2. Memory controllers 526-1 and 526-2 may provide hardware control for DRAM modules 516. One or more levels of cache 530 may also be provided. Input / Output Controller (IOCTL) 534 may provide input and output operations. System agent 538 may provide firmware for, e.g., detecting errors and recovering from errors and other system services. System agent 538 may interface with one or more I / O modules 546, directly as in the case of I / O module 546-2, or via Platform Controller Hub (PCH) 542.

[0081] It should be understood that blade system 502 is provided only as a non-limiting and illustrative example of a hardware platform 503 that may provide computing services. Many other configurations are possible, and this specification is not intended to be limited to the example of blade system 502 or any other particular hardware platform.

[0082] For example, when container 504-1 is accessing a block of memory within DRAM 516-1, an error may occur. The block of memory may be specifically partitioned and dedicated to container 504-1 and may thus not be accessible to container 504-2. While performing its operations and accessing DRAM 516-1, container 504-1 may encounter an uncorrectable error. This may be, for example, the result of faulty programming that causes an uncorrectable software error, or it may be the result of a hardware failure within DRAM 516-1 (such as a bad or damaged memory block).

[0083] Subsequently, container 504-2 may be accessing a completely separate memory block, which may be located on DRAM 516-1 or may be in a completely separate DRAM, such as DRAM 516-2. Since containers 504-1 and 504-2 have separately partitioned memory blocks, a memory error encountered by container 504-1 does not directly affect container 504-2. However, since the error on container 504-1 is an uncorrectable error, it may not be possible to resume the functionality of container 504-1 without restarting blade system 502. When blade system 502 is restarted, the memory can be checked, and if the error is the result of a hardware error, in some cases, the bad memory block can be removed from the cycle so that it is not addressable, and blade system 502 can then continue to operate normally.

[0084] However, if blade system 502 is immediately powered down by system agent 538 when container 504-1 encounters an uncorrectable error, container 504-2 is also immediately powered down. Thus, while the memory error encountered by container 504-1 does not directly affect container 504-2, the cold restart required to resume container 504-1 does affect container 504-2. Additionally, directly cold restarting without any warning to container 504-2 can cause container 504-2 to lose data and / or corrupt data.

[0085] If container 504-2 is providing a higher-priority function than container 504-1, this also means that the higher-priority function is terminated or stopped as a result of an error in the relatively lower-priority function. As discussed above, container 504-1 can be an email server, while container 504-2 can be a high-availability database server. In this case, powering down blade system 502 without warning to container 504-2 can cause data loss or database corruption in container 504-2.

[0086] Therefore, in some embodiments, it is beneficial for container 504-2 to be notified that a cold restart will be necessary and then wait for container 504-2 to reach a good "stopping point" before powering down blade system 502, rather than immediately restarting blade system 502. As described herein, this notification instructs container 504-2 to seek a recoverable state. Note that, as a non-limiting example, the notification is provided for container 504-2. In other embodiments, system agent 538 can simply monitor container 504-2 and autonomously determine when container 504-2 has reached a recoverable state. Note also that other network elements, such as a coordinator or controller, can help container 504-2 seek a recoverable state, such as by instructing a load balancer to direct less traffic to container 504-2.

[0087] Once the blade system 502 reaches a state where all containers are in a suitable state for restart, the system agent 538 can perform its normal error handling up to and including a cold restart of the blade system 502. Meanwhile, the container 504-1 can be stopped and made unavailable, but the loss of the container 504-1 can often be mitigated in a data center, especially one that provides software-defined networking and orchestration capabilities, by simply creating a new instance of the function provided by the container 504-1. In some cases, if the container 504-2 provides a function that is very intolerant of interruption, a new instance of the function of the container 504-2 can be created before the blade system 502 is restarted, and a handover can be completed between the two instances so that the restart is relatively seamless.

[0088] FIGS. 6A–6B are signal flow diagrams of a method for performing latency error handling according to one or more examples of the present specification. The example of FIGS. 6A–6B illustrates the interaction between an operating system 608, system firmware 606 that may provide a system agent, CPU hardware 604, and CPU microcode 602. For illustrative purposes of an embodiment, the logic presented here is divided into separate blocks. However, it should be understood that the division shown here is merely illustrative, and the functions provided in one block can often be moved to another block in many instances. In particular, the division between the microcode 602 and the hardware 604 in the CPU is often a matter of design optimization, and it is often feasible to move functions between one side or the other in different embodiments. Additionally, the functions provided by the system firmware 606 can often be moved to the microcode 602 or provided elsewhere.

[0089] In the illustrated example, starting at block 662 in the operating system 608 in FIG. 6B, a memory location performs a memory access on a memory page. This access flows from the off-page connector to the microcode 602, where at block 610, a data collector unit (DCU) receives the data access request and determines that this is a “poisonous” access request. In other words, the memory at this location is corrupted or otherwise inaccessible, and the memory access cannot continue.

[0090] In block 614, in response to receiving an indication that poisonous data has been received, the microcode can take appropriate action, such as triggering a “memory corruption” event in the microcode.

[0091] In block 618, in response to a memory corruption event, the microcode 602 may issue an error signal. In one example, the error signal is a page fault with an error code of 0x20 or some other "invalid" indicator. In another example, the page fault may have a null page error code, or a specific designated toxic code may be used. The purpose of this specific code may be to trigger a segmentation fault on the operating system or to replay a dump of the container's architectural state as needed. This is similar to the situation during a switch in SMM or from real mode to protected mode.

[0092] The page fault may be sent to the operating system 608 via the off-page connector B. In block 664, the operating system 608 may trigger a segmentation fault. A core dump may occur, and the application may be removed from execution. The architectural state may be saved for future replay. Advantageously, with a core dump, a debugger can be used to analyze the program flow to determine what caused the memory error and to avoid future memory errors.

[0093] In block 668, on the operating system 608, other containers on the hardware platform may continue to execute.

[0094] In block 694, a task switch to a thread going to another container may be performed.

[0095] Returning to block 618, within the microcode 602, a degraded state is issued to the system agent.

[0096] In block 622, the microcode 602 may trigger a "degraded" state for the system agent of the firmware 606. The degraded state may be used to notify resources on the platform itself and other resources in the data center that the containers on this hardware platform should now be considered "degraded".

[0097] In one example, the degraded state is triggered using a new degraded state register. Exemplary implementations of the degraded state register may include the following:

[0098] Timer [63:07] Counter [6:3] Delay Disposal Complete [2:2] EN / MCE / DHC [1:1] MCE Trigger [0:0] Offset for triggering an event. The delay state can be retained until the counter has reached its maximum value The firmware has completed the delay error handling #timg# Setting this bit triggers the next sequence (e.g., error handling can continue). Enable Machine Check Exception (MCE) upon completion of the delay disposal Now trigger MCE (related to debugging scenarios).

[0099] In block 626, the microcode 602 may notify the system agent of the degraded state.

[0100] For example, within the hardware 604, in block 630, the system agent receives the degraded state signal and may trigger an SMI for initial degraded response handling. The SMI may be issued via the off-page connector C to the SMI handler in the firmware.

[0101] Within the hardware 604, the system agent may then start a counter for the next SMI generation. This counter may be used to ensure that no other container takes too long to reach a recoverable state and unnecessarily delays the recovery of the hardware platform.

[0102] Accordingly, in decision block 638, the hardware 604 continuously monitors the timeout to see if it has expired. As long as the timeout has not expired, the system agent continues to wait for the container to reach a graceful exit point.

[0103] If the timer expires without other containers reaching an appropriate exit point, the process proceeds to off-page connector D.

[0104] Returning to FIG. 6B, from off-page connector C, the SMI handler in the system firmware 606 receives an SMI that notifies it of the degraded state.

[0105] In block 646, the SMI handler evaluates whether the system is in a degraded state. If the system has been identified as being in a degraded state, it can notify the operating system of a planned shutdown or maintenance cycle. The operating system can then notify other containers that they should start working towards a reset point. The operating system can also notify other data center components (such as the data center coordinator) so that the coordinator can take measures to reduce the workload on the degraded container. For example, if the email container encounters an error and the database driver is running on a separate container on the same hardware platform, when the operating system notifies the coordinator that the database driver is running in a degraded state, the coordinator can start diverting as much traffic as possible away from that database driver instance. For example, if sufficient resources are available, the coordinator can spawn a new instance of the database driver. The coordinator can also instruct the load balancer not to dispatch any new traffic buckets to the degraded instance of the database driver. To further reduce the load on the database driver, the coordinator can gradually instruct the load balancer to re-dispatch existing buckets to other instances. This can allow the degraded instance of the database driver to reach a state where it does not receive any new incoming traffic. It can then gracefully complete the processing of any outstanding requests and, once it has no outstanding requests, it is in a state where it can gracefully shut down and can notify the system agent of this situation.

[0106] In block 650, if the Machine Check Architecture (MCA) or Enhanced MCA (EMCA) is enabled, the SMI handler can collect appropriate logs; the SMI handler can also notify, for example, a top-of-rack switch or Remote Monitoring and Management (RMM) of the degraded system state. This can be in addition to or an alternative to the notification provided by the operating system 608.

[0107] In decision block 654, the system firmware 606 receives a hard timeout limit from the hardware 604. In a loop, the SMI handler can continuously check whether other containers have reached a recoverable state, such as whether the system agent or the system firmware has received notification that other containers are no longer handling new incoming traffic. Alternatively, a hard timeout may occur.

[0108] Any one of these events ultimately triggers the SMI handler 606 to set the delayed disposal completion bit.

[0109] Once the delayed disposal completion bit is set, control flows back to FIG. 6A via the off-page connector E. In block 642, the delayed disposal completion bit triggers the final error handling performed by the SMI or MCE.

[0110] Once the final error handling has been triggered, in block 698, the error is disposed of according to the existing error disposal method as described herein.

[0111] The foregoing has outlined the features of several embodiments so that those skilled in the art may better understand the various aspects of the present disclosure. Those skilled in the art should understand that they can readily use the present disclosure as a basis for designing or modifying other processes and structures for carrying out the same purposes and / or achieving the same advantages as the embodiments introduced herein. Those skilled in the art should also realize that such equivalent constructs do not depart from the spirit and scope of the present disclosure, and that they can make various changes, substitutions, and alterations herein without departing from the spirit and scope of the present disclosure.

[0112] All or part of any of the hardware elements disclosed herein can be readily provided in a system-on-chip (SoC) including a central processing unit (CPU) package. An SoC represents an integrated circuit (IC) that integrates the components of a computer or other electronic system into a single chip. Thus, for example, a client device or a server device can be provided wholly or partially in an SoC. The SoC can include digital, analog, mixed-signal, and radio frequency functions, all of which can be provided on a single chip substrate. Other embodiments can include a multi-chip module (MCM), where multiple chips are located within a single electronic package and are configured to interact closely with each other through the electronic package.

[0113] It should also be noted that in some embodiments, some components may be omitted or combined. In a general sense, the arrangements depicted in the drawings may be more logical in their representation, while the physical architecture may include various arrangements, combinations, and / or mixtures of these elements. It must be noted that countless possible design configurations can be used to achieve the operational objectives outlined herein. Therefore, the associated infrastructure has countless alternative arrangements, design choices, device possibilities, hardware configurations, software implementations, and equipment options.

[0114] In a general sense, any suitably configured processor can execute any type of instruction associated with data to implement the operations detailed herein. Any processor disclosed herein can transform an element or article (e.g., data) from one state or thing to another. In operation, a storage device can store information in any suitable type of tangible non-transitory storage medium (e.g., random access memory (RAM), read only memory (ROM), field programmable gate array (FPGA), erasable programmable read only memory (EPROM), electrically erasable programmable ROM (EEPROM), etc.), software, hardware (e.g., processor instructions or microcode), or, where appropriate and based on specific needs, in any other suitable component, device, element, or object. Additionally, based on specific needs and implementations, information tracked, sent, received, or stored in a processor can be provided in any database, register, table, cache, queue, control list, or storage structure, and all such databases, registers, tables, caches, queues, control lists, or storage structures can be referenced at any suitable time frame. Any memory or storage element disclosed herein should be construed, as appropriate, to be included within the broad terms "memory" and "storage device". The non-transitory storage medium herein is expressly intended to include any non-transitory dedicated or programmable hardware configured to provide the disclosed operations or to cause a processor to execute the disclosed operations.

[0115] Computer program logic for implementing all or part of the functionality described herein is implemented in various forms, including but in no way limited to source code form, computer-executable form, machine instructions or microcode, programmable hardware, and various intermediate forms (e.g., forms produced by assemblers, compilers, linkers, or locators). In an example, source code includes a series of computer program instructions implemented in various programming languages or in a hardware description language, such as object code, assembly language, or a high-level language (such as OpenCL, FORTRAN, C, C++, JAVA, or HTML) for use with various operating systems or operating environments, or a hardware description language such as Spice, Verilog, and VHDL. Source code can define and use various data structures and communication messages. Source code can have a computer-executable form (e.g., via an interpreter), or source code can be transformed (e.g., via a converter, assembler, or compiler) into a computer-executable form or into an intermediate form (such as bytecode). Where appropriate, any of the foregoing can be used to build or describe a suitable discrete or integrated circuit, whether sequential, combinatorial, state machine, or otherwise.

[0116] In one exemplary embodiment, any number of circuits of the drawings may be implemented on a board of an associated electronic device. The board may be a general circuit board that may house various components of the internal electronic system of the electronic device and further provide connectors for other peripheral devices. Any suitable processor and memory may be suitably coupled to the board based on specific configuration requirements, processing needs, and computing designs. It should be noted that, using many of the examples provided herein, interactions may be described in terms of two, three, four, or more electrical components. However, this description is for clarity and illustrative purposes only. It should be understood that the system may be incorporated or reconfigured in any suitable manner. According to similar design alternatives, any of the illustrated components, modules, and elements of the drawings may be combined in various possible configurations, all of which configurations fall within the broad scope of this specification.

[0117] Those skilled in the art can determine many other changes, substitutions, variations, alterations, and modifications, and this disclosure is intended to cover all such changes, substitutions, variations, alterations, and modifications that fall within the scope of the appended claims. To assist the United States Patent and Trademark Office (USPTO) and additionally assist any reader of any patent issued on this application in interpreting the appended claims herein, the applicant wishes to note that the applicant: (a) does not intend any of the appended claims to invoke 35 U.S.C. § 112, paragraph 6 (6) (pre-AIA) or paragraph (f) (post-AIA) of the same section as it exists on the filing date of the application herein, unless the words "means for" or "step for" are expressly used in a particular claim; and (b) does not intend any statement in the specification to limit this disclosure in any way that is not expressly reflected in the appended claims.

[0118] Exemplary implementations

[0119] The following examples are provided for illustration.

[0120] Example 1 includes a computing device comprising: a hardware platform including a processor and a memory; and a system management interrupt (SMI) handler; first logic configured to provide a first container and a second container via the hardware platform; and second logic configured to: detect an uncorrectable error in the first container; in response to the detection, generate a degraded system state; provide a degraded state message to the SMI handler; instruct the second container to find a recoverable state; determine that the second container has entered a recoverable state; and initiate a recovery operation.

[0121] Example 2 includes the computing device of Example 1, wherein the second logic is further configured to set a timeout and initiate the recovery operation after the timeout expires.

[0122] Example 3 includes the computing device of Example 1 and further includes a fabric interface; wherein the second logic also provides a degradation notification to the controller.

[0123] Example 4 includes the computing device of Example 3, wherein the second logic also requests the controller to generate a new instance of the service provided by the first container.

[0124] Example 5 includes the computing device of Example 1, wherein the recoverable state includes a state in which the second container can be migrated with minimal data loss.

[0125] Example 6 includes the computing device of Example 5, wherein the second logic is further configured to migrate the second container.

[0126] Example 7 includes the computing device of any one of Examples 1 - 6 and further includes an operating system, and the operating system is further configured to perform a core dump of the first container.

[0127] Example 8 includes the computing device of Example 7, wherein the operating system is configured to receive machine check architecture (MCA) record information from the processor.

[0128] Example 9 includes the computing device of Example 7, wherein the second logic is further configured to notify the operating system of the degradation state of the device.

[0129] Example 10 includes the computing device of any one of Examples 1 - 6, wherein the second logic further includes a configuration interface configured to receive configuration options.

[0130] Example 11 includes one or more tangible non - transitory computer - readable media having instructions stored thereon for providing logic for: providing a system management interrupt (SMI) handler; providing a first container and a second container; detecting an uncorrectable error in the first container; in response to the detection, generating a degraded system state; providing a degradation status message to the SMI handler; instructing the second container to find a recoverable state; determining that the second container has entered a recoverable state; and starting a recovery operation.

[0131] Example 12 includes the one or more tangible non - transitory computer - readable media of Example 11, wherein the logic is further configured to set a timeout and start the recovery operation after the timeout expires.

[0132] Example 13 includes the one or more tangible non - transitory computer - readable media of Example 11, wherein the logic is further for providing a degradation notification to the controller via a fabric interface.

[0133] Example 14 includes the one or more tangible non-transitory computer-readable media of Example 13, wherein the logic is further configured to request a controller to generate a new instance of a service provided by a first container.

[0134] Example 15 includes the one or more tangible non-transitory computer-readable media of Example 11, wherein the recoverable state includes a state in which a second container can be migrated with minimal data loss.

[0135] Example 16 includes the one or more tangible non-transitory computer-readable media of Example 15, wherein the logic is further configured to migrate a second container.

[0136] Example 17 includes the one or more tangible non-transitory computer-readable media of any one of Examples 11-16, wherein the logic is further configured to provide an operating system that is configured to perform a core dump of a first container.

[0137] Example 18 includes the one or more tangible non-transitory computer-readable media of Example 17, wherein the operating system is configured to receive machine check architecture (MCA) record information from a processor.

[0138] Example 19 includes the one or more tangible non-transitory computer-readable media of Example 17, wherein the logic is further configured to notify the operating system of a degraded state of the device.

[0139] Example 20 includes the one or more tangible non-transitory computer-readable media of any one of Examples 11-16, wherein the second logic further includes a configuration interface that is configured to receive configuration options.

[0140] Example 21 includes a computer-implemented method for providing delayed error handling, the method comprising: providing a system management interrupt (SMI) handler; providing a first container and a second container; detecting an uncorrectable error in the first container; generating a degraded system state in response to the detection; providing a degraded state message to the SMI handler; instructing the second container to find a recoverable state; determining that the second container has entered the recoverable state; and starting a recovery operation.

[0141] Example 22 includes the method of Example 21, further comprising: setting a timeout and starting the recovery operation after the timeout expires.

[0142] Example 23 includes the method of Example 21, further comprising: providing a degradation notification to a controller via a structural interface.

[0143] Example 24 includes the method of Example 23, further comprising: requesting a controller to generate a new instance of a service provided by the first container.

[0144] Example 25 includes the method of Example 21, wherein the recoverable state includes a state in which the second container can be migrated with minimal data loss.

[0145] Example 26 includes the method of Example 26, further comprising: migrating the second container.

[0146] Example 27 includes the method of any one of Examples 21-26, further comprising: providing an operating system configured to perform a core dump of the first container.

[0147] Example 28 includes the method of Example 27, wherein the operating system is configured to receive machine check architecture (MCA) record information from a processor.

[0148] Example 29 includes the method of Example 27, further comprising: notifying the operating system of a degraded state of the device.

[0149] Example 30 includes the method of any one of Examples 21-26, further comprising: providing a configuration interface configured to receive configuration options.

[0150] Example 31 includes a device comprising means for performing the method of any one of Examples 21-30.

[0151] Example 32 includes the device of Example 31, wherein the means for performing the method comprises a processor and a memory.

[0152] Example 33 includes the device of Example 32, wherein the memory includes machine-readable instructions that, when executed, cause the device to perform the method of any one of Examples 21-30.

[0153] Example 34 includes the device of any one of Examples 31-33, wherein the device is a computing system.

[0154] Example 35 includes at least one tangible non-transitory computer-readable medium comprising instructions that, when executed, implement the method or device as shown in any one of Examples 21-34.

Claims

1. A computing device, comprising: A hardware platform including a processor and a memory; And A system management interrupt (SMI) handler; A first logic configured to provide a first container and a second container via the hardware platform; And A second logic configured to: Detect an uncorrectable error in the first container; Generate a degraded system state in response to the detection; Provide a degraded state message to the SMI handler; Instruct the second container to find a recoverable state in response to the degraded system state; Determine that the second container has entered a recoverable state; And Initiate a recovery operation.

2. The computing device according to claim 1, wherein the second logic is further configured to set a timeout and initiate the recovery operation after the timeout expires.

3. The computing device according to claim 1, further comprising a fabric interface; wherein the second logic further provides a degradation notification to a controller.

4. The computing device according to claim 3, wherein the second logic further requests the controller to generate a new instance of a service provided by the first container.

5. The computing device according to claim 1, wherein the recoverable state includes a state in which the second container can be migrated with minimal data loss.

6. The computing device according to claim 5, wherein the second logic is further configured to migrate the second container.

7. The computing device according to any one of claims 1-6, further comprising: An operating system, the operating system is further configured to perform a core dump of the first container.

8. The computing device according to claim 7, wherein the operating system is configured to receive machine check architecture (MCA) record information from the processor.

9. The computing device according to claim 7, wherein the second logic is further configured to notify the operating system of the degraded state of the device.

10. The computing device according to any one of claims 1-6, wherein the second logic further includes a configuration interface configured to receive configuration options.

11. One or more tangible non-transitory computer-readable media having instructions stored thereon for providing logic for: Providing a system management interrupt (SMI) handler; Providing a first container and a second container; Detecting an uncorrectable error in the first container; Generating a degraded system state in response to the detection; Providing a degraded state message to the SMI handler; Instructing the second container to find a recoverable state in response to the degraded system state; Determining that the second container has entered a recoverable state; and Initiating a recovery operation.

12. The one or more tangible non-transitory computer-readable media according to claim 11, wherein the logic is further configured to set a timeout and initiate the recovery operation after the timeout expires.

13. The one or more tangible non-transitory computer-readable media according to claim 11, wherein the logic is further for providing a degradation notification to a controller via a fabric interface.

14. The one or more tangible non-transitory computer-readable media according to claim 13, wherein the logic is further for requesting the controller to generate a new instance of a service provided by the first container.

15. The one or more tangible non-transitory computer-readable media of claim 11, wherein the recoverable state includes a state in which the second container can be migrated with minimal data loss.

16. The one or more tangible non-transitory computer-readable media of claim 15, wherein the logic is further configured to migrate the second container.

17. The one or more tangible non-transitory computer-readable media of any one of claims 11-16, wherein the logic is further configured to provide an operating system, the operating system being configured to perform a core dump of the first container.

18. The one or more tangible non-transitory computer-readable media of claim 17, wherein the operating system is configured to receive machine check architecture (MCA) record information from the processor.

19. The one or more tangible non-transitory computer-readable media of claim 17, wherein the logic is further configured to notify the operating system of a degraded state.

20. A computer-implemented method for providing delayed error handling, comprising: providing a system management interrupt (SMI) handler; providing a first container and a second container; detecting an uncorrectable error in the first container; generating a degraded system state in response to the detection; providing a degraded state message to the SMI handler; instructing the second container to find a recoverable state in response to the degraded system state; determining that the second container has entered a recoverable state; and initiating a recovery operation.

21. The method according to claim 20, further comprising: setting a timeout and initiating the recovery operation after the timeout expires.

22. The method according to claim 20, further comprising: providing a degradation notification to a controller via a structural interface.

23. The method according to claim 22, further comprising: requesting the controller to generate a new instance of a service provided by the first container.

24. The method of claim 20, wherein the recoverable state includes a state in which the second container can be migrated with minimal data loss.

25. An apparatus for providing delayed error handling, comprising: means for providing a system management interrupt (SMI) handler; means for providing a first container and a second container; means for detecting an uncorrectable error in the first container; means for generating a degraded system state in response to the detection; means for providing a degraded state message to the SMI handler; means for instructing the second container to find a recoverable state in response to the degraded system state; means for determining that the second container has entered a recoverable state; and means for initiating a recovery operation.

26. The apparatus according to claim 25, further comprising: means for setting a timeout and initiating the recovery operation after the timeout expires.

27. The device according to claim 25, further comprising: means for providing a degradation notification to a controller via a structural interface.

28. The apparatus according to claim 27, further comprising: means for requesting the controller to generate a new instance of a service provided by the first container.

29. The apparatus of claim 25, wherein the recoverable state includes a state in which the second container can be migrated with minimal data loss.

30. A computer program product having instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 20-24.

Citation Information

Patent Citations

  • Methods and systems for handling software operations associated with startup and shutdown of handheld devices

    US20090117889A1

  • Methods and Apparatus for Handling Errors Involving Virtual Machines

    US20090144579A1