Selective endpoint isolation for self-healing in cache and memory coherent systems
By isolating faulty hardware by dynamically updating routing tables in the coherent CPU system, the system crash caused by hardware failure is resolved, the system's self-healing capability during failures is achieved, and efficient user and resource support is maintained.
Patent Information
- Application Number
- CN202080095343.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-30
- Filing Date
- 2020-12-14
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2040-12-14
AI Technical Summary
In a coherent CPU system, a hardware failure causes the entire system to crash. Existing technologies can only generate a fault notification and are unable to keep the system operating normally before the hardware is repaired and replaced.
The coherent mesh fabric and routing engine dynamically update the routing table to selectively isolate faulty hardware components, allowing the rest of the system to continue operating normally.
In the event of a hardware failure, the system can continue to support the majority of users and processing resources until the failed component is repaired or replaced.
Smart Images

Figure CN115039085B_ABST
Abstract
Description
Background Art
[0001] As demand for cloud-based storage and computing services grows rapidly, so too does the need for technologies that can quickly scale hardware in existing data centers. Traditionally, hardware expansion has been achieved by adding resources to "scale up," such as by adding power or capacity to a data center. However, more recent solutions aim to "scale out" or "scale out" to support higher performance levels, throughput, and redundancy for advanced fault tolerance without increasing cost and / or the total amount of hardware (e.g., without increasing the number of servers, drives, etc.). Architectures that enable this type of horizontal scaling are sometimes referred to as "hyperscale." Summary of the Invention
[0002] According to one implementation, a cache and memory coherence system includes a plurality of processing chips, each hosting a different subset of a shared memory space. One or more routing tables define access routes between logical addresses of the shared memory space and endpoints, each corresponding to a selected one of the plurality of processing chips. A coherent mesh structure physically couples each of the plurality of processing chips together and is configured to execute routing logic for updating the routing table(s) in response to identification of a faulty hardware component hosted by a first processing chip in the system. The update to the routing table effectively removes from the routing table(s) all access routes having an endpoint corresponding to the first processing chip hosting the faulty hardware component.
[0003] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0004] Other implementations are also described and listed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] Figure 1A An example coherent CPU system is shown that implements self-healing logic for selectively isolating a region of the system upon detection of a faulty hardware component.
[0006] Figure 1B Shows the update Figure 1A The CPU coheres one or more routing tables in a system to selectively isolate the effect of an area of the system.
[0007] Figure 2A Another example coherent CPU system is shown that implements self-healing logic for selectively isolating a region of the system upon detection of a faulty hardware component.
[0008] Figure 2B Shows the update Figure 1A The CPU coheres one or more routing tables in a system to selectively isolate the effect of an area of the system.
[0009] Figure 3 An exemplary architecture of a coherent CPU system that implements self-healing logic for selectively isolating a system endpoint upon detection of a hardware component failure is shown.
[0010] Figure 4 Example operations for implementing self-healing logic within a cache and memory coherent multi-chip system are shown.
[0011] Figure 5 An example schematic diagram illustrating a processing device suitable for implementing aspects of the disclosed technology is illustrated. DETAILED DESCRIPTION
[0012] Hyperscale architectures that allow memory, cache, and / or I / O coherence between processors are gaining popularity as a way to increase the number of users that can be supported by existing hardware and, as a secondary effect, to de-duplicate storage (and thereby reduce data costs). A multiprocessor system is said to be "coherent" or have a coherent CPU architecture when its multiple CPUs share a single memory space so that data can be read or written to any memory location at any CPU. Coherent CPU architectures rely on the coherence of both caches and memory (e.g., each processor of the system knows where the most recent version of data is, even if that data is in a local cache hosted by another processor).
[0013] In a coherent CPU architecture, problems can arise and, in fact, be amplified when a shared hardware component fails. For example, if multiple processors are configured to read and write data from the same memory component, a failure of that memory component can lead to a system-wide hang. Similarly, a failure in a single processing component or on a shared bus can have similar consequences, ultimately disrupting more users and / or suspending more system processes than would otherwise occur in a traditional (non-coherent) CPU architecture.
[0014] Although some coherent systems include modules that enable self-monitoring of hardware health to more quickly detect potential hardware failure issues, these systems only generate notifications about recommended replacements and repairs. In the time between a hardware failure and component repair and / or replacement, the entire CPU-coherent system (e.g., a server with multiple systems on a chip (SoCs) each supporting different processes and / or users) may become inoperable.
[0015] The techniques disclosed herein allow these coherent CPU systems to at least partially "self-heal" by logically isolating the portion of the system that includes the faulty component(s) in response to detecting and identifying the relative location of the faulty (e.g., malfunctioning, failing, or failing) hardware component(s). In one implementation, this isolation of the faulty hardware allows the remaining healthy subset of CPUs in the memory and cache coherent system to continue to function nominally, such as by continuing to service user requests and execute processes as long as these requests and processes do not require access to resources within the isolated portion of the system. In many cases, this isolation allows the system to perform at a high percentage of its nominal "health" level, such as by continuing to support a significant subset (e.g., 50%, 75%, or more) of users connected to the coherent CPU system (e.g., a server) and / or by continuing to operate a large portion (e.g., 50%, 75%, or more) of the total processing and storage resources nominally.
[0016] Figure 1A An example CPU-coherent system 100 is shown that implements self-healing logic for selectively isolating an area of the system after detecting a faulty hardware component. The CPU-coherent system 100 includes multiple processing chips coupled together by a coherent mesh structure 130. The coherent mesh structure 130 can generally be understood to encompass the physical interconnects and logical architecture that allow each processing chip to collectively operate as a memory-coherent and cache-coherent system so that all processing entities in the system can read and write data to any memory location and know the location of the most recent version of any data stored in the system, even if the data resides in a local cache hosted by a different processing chip.
[0017] The example CPU-coherent system 100 includes four processing chips, labeled IC_1, IC_2, IC_3, and IC_4. Each of these different processing chips hosts memory (e.g., memory 108, 110, 112, and 114) and includes at least one CPU (e.g., CPU 104, 116, 118, and 120). In one implementation, each of the processing chips (IC_1, IC_2, IC_3, and IC_4) is a system-on-chip (SoC) that includes multiple processors, including, for example, one or more CPUs and one or more graphics processing units (GPUs). As used herein, a processing chip is referred to as "hosting" memory if it physically includes the memory or is directly coupled to the memory so that data routed to the memory does not flow across any other processing chip between the processing chip hosting the memory and the memory itself.
[0018] exist Figure 1AIn one implementation, four different processing chips are coupled to the same printed circuit board assembly (PCBA), such as to operate together as part of the same server (e.g., a blade server sometimes referred to as a "blade"). In other implementations, one or more of these processing chips are on different PCBAs, but are still physically close to the rest of the processing chips in the memory and cache coherent system to allow extremely fast chip-to-chip volatile memory access (a critical feature for the operation of CPU-coherent systems). For example, the processing chips can be on different PCBAs within two or more servers that are physically close together and integrated within the same server or in the same data storage center.
[0019] The memory 108, 110, 112, and 114 hosted by each processing chip can be understood to include at least volatile memory and, in many implementations, both volatile and non-volatile memory. In one implementation, each of the processing chips (IC_1, IC_2, IC_3, and IC_4) hosts a different subset of system memory that is collectively mapped to the same logical address space. For example, the non-volatile memory space hosted by each processing chip can be mapped to a logical address range used by the host when reading and writing data to the CPU-coherent system 100. In one implementation, each logical address of the host addressing scheme is mapped to only a single non-volatile memory location in the system (e.g., a location hosted by a single processing chip among the four processing chips). For example, if the host's logical address range is represented by the alphanumeric sequence AZ, each different processing chip can host a different subset of that sequence (e.g., IC_1 can host AG, while IC_2 hosts HN, IC_3 hosts OT, and IC_4 hosts UZ). Additionally, the memory of each processing chip may include volatile memory and one or more local caches where data may sometimes be updated locally without immediately updating the corresponding non-volatile memory locations.
[0020] The system's coherent mesh fabric 130 provides the physical and logical infrastructure that allows each of the four processing chips to operate collectively as a cache, memory, and I / O coherent system by: (1) providing direct data links between the processing chips and each memory resource in the system, and (2) providing a coherence scheme that allows memory and resources hosted by any processing chip to be accessed and used (e.g., read from and write to) by any processor anywhere on the coherent mesh fabric 130 diagram.
[0021] In physical terms, the coherent mesh fabric 130 can be understood as including a silicon-level data bus (e.g., physical interconnect 132) that provides a direct path between each processing chip and each other processing chip in the CPU coherent system. In addition, the coherent mesh fabric 130 includes routing logic and routing control electronics for executing this logic. The routing logic and routing control electronics are collectively represented in FIG1 by routing engine 122. Although routing engine 122 is shown as being central (for conceptual simplicity), it should be understood that the logic and hardware resources employed by the routing engine can actually be distributed among the various processing chips, as shown in FIG1. Figure 2A and 2B shown.
[0022] In different implementations, the physical interconnect 132 can take various forms. In one implementation, the physical interconnect 132 is a Peripheral Component Interconnect Express (PCIe) physical interface that provides a physical data link between each pair of processing chips in the system. For example, the coherent mesh fabric 130 can be a memory interconnect using the PCIe interface standard or other interface standards (e.g., Serial ATA (SATA), USB, Serial Attached SCSI (SAS), FireWire, Accelerated Graphics Port (AGP), Peripheral Component Interconnect (PCI)) to facilitate fast and direct communication between each pair of processing chips in the CPU coherent system 100.
[0023] Routing engine 122 implements a coherence scheme via coherent mesh structure 130 that allows each processing entity coupled to coherent mesh structure 130 to be aware of the most recent version of all system data. While the exact coherence scheme utilized may vary depending on the implementation, routing engine 122 implements this scheme at least in part by managing and dynamically updating one or more system routing tables. By way of example and not limitation, a zoomed-in view 136 of routing engine 126 shows example routing tables 134 at two different points in time, t1 and t2, to illustrate dynamic updates used to isolate a portion of the system, as discussed in more detail below.
[0024] The routing table 134 provides a mapping between each logical address in the system and a physical "endpoint", where each physical endpoint identifies a selected one of the different processing chips that stores the most recent version of the data corresponding to the logical address. In one implementation, the routing engine 122 manages the routing table with respect to each processing chip. In this case, the routing table 134 for each different processing chip may list only the logical addresses that are not hosted by that chip. For example, the routing table for processing chip "IC_3" may list the logical addresses corresponding to data storage locations hosted by IC_1, IC_2, and IC_4. In other implementations, the routing table for each processing chip provides a complete mapping between every logical address in the system's memory space and every corresponding endpoint (e.g., the processing chip that hosts the memory storing the associated data).
[0025] When any processing chip in the system receives a host request to read or write data, the routing engine of that processing chip determines the endpoint location for storing the data. For example, if IC_3 receives a request to read data at logical address "ABC", the routing engine 122 consults the appropriate routing table (e.g., routing table 134) and determines that the most recent version of the data is stored in memory hosted by IC_4. In this example, the data at logical address "ABC" may be stored in non-volatile memory hosted by IC_4, or instead stored in volatile memory on IC_4, such as a local cache hosted by IC_4. The routing engine 122 directs the read control signal to IC_4 along the appropriate data link of the physical interconnect 132 and uses logic stored locally on the receiving chip to further direct the read control signal to the appropriate physical storage location, whether that location is in volatile memory (e.g., cache) or in non-volatile memory.
[0026] One consequence of the coherence scheme implemented in CPU-coherent system 100 is that a hardware failure occurring anywhere within the system has the potential to crash the entire CPU-coherent system 100. For example, if a DRAM component fails in memory 110 of chip IC_2, the processing components on each processing chip may freeze (e.g., hang indefinitely) in response to receiving and attempting to process the next read or write command to the failed DRAM component. This may result in a system-wide failure, as all four processing chips may freeze one at a time until the entire CPU-coherent system 100 is hung and requires a reboot.
[0027] While there are scenarios where a memory failure (e.g., DRAM) can be resolved upon restart by simply remapping the data previously loaded into the failed DRAM to spare volatile memory, there are scenarios where it is not possible to remap the logical memory space simply to exclude specific physical memory components due to, for example, memory constraints (e.g., lack of sufficient spare memory and / or the non-swappable nature of the memory hardware components) or other reasons. For example, some systems require memory mapping to DRAM dual in-line memory modules (DIMMs) at boot time, while other systems require memory mapping to multiple surface-mounted memory channels (e.g., DRAM soldered to a board) and / or mapping to high-bandwidth memory (HBM memory) integrated as part of the processing chip. When these types of components fail, memory remapping may not be a viable option.
[0028] In other scenarios, hardware failures affect parts of the silicon die and non-memory chip components of the bus. For example, a failed solder joint or fan could cause thermal damage to the processor. In these scenarios, proper isolation of the subsystem hosting the failed component (as described below) can allow the system to continue operating as expected.
[0029] 1 , the CPU-coherent system 100 includes a board management controller 138 that continuously monitors system event data and health indicators to record information that may be about hardware component faults and failures in an error log 140. For example, the board management controller 138 may attempt to periodically (e.g., every 10 ms) poll each processing chip and immediately log a time-stamped event if any processing chip becomes unresponsive.
[0030] When a partial or system-level failure occurs due to a failed or failing hardware component, the board management controller 138 can restart the entire system and, upon restart, analyze the error log 140 to determine the most likely location of the failed (e.g., failed, failing, or malfunctioning) component that caused the system-level failure. In the example where the failed hardware component is the DRAM in the memory 110 of IC_2, the board management controller 138 analyzes the error log 140 to determine that IC_2 is the endpoint associated with the failed component.
[0031] Error log 140 can be used as a basis for generating notifications to system administrators that a particular chip or server needs to be repaired or replaced. In a system lacking coherent mesh structure 130, the entire CPU coherent system may remain offline (inoperative) after detecting a failed component until the system administrator is actually able to perform the necessary maintenance to repair or replace the failed component. However, in the presently disclosed system, coherent mesh structure 130 includes logic for selectively isolating an endpoint (e.g., a processing chip) that is the host of a failed or failing hardware component. This host endpoint of a failed hardware component is hereinafter referred to as a "failed endpoint."
[0032] In one implementation, isolation of a failed endpoint is achieved by selectively updating all system routing tables (e.g., routing table 134) to remove all access routes directed to the physical location hosted by the failed endpoint. For example, if a DRAM component fails in memory 110 of IC_2, the board management controller 138 can analyze the error log 140 upon reboot to identify IC_2 as the failed endpoint. The board management controller 138 can then instruct the coherent mesh structure 130 to update the routing tables of each of the other processing chips (e.g., IC_1, IC_2, and IC_4) to remove all access routes with endpoints corresponding to the failed endpoint.
[0033] exist Figure 1A In a basic example, updates to routing tables in the system are illustrated by the changes shown relative to routing table 134 between times t1 and t2. At time t1, routing table 134 includes all logical addresses hosted by any processing chip endpoint. Between times t1 and t2, routing engine 126 receives and executes an instruction from board management controller 138 that causes all addresses associated with the failed endpoint to be removed. This can, for example, cause each of the remaining processing chips in the system (e.g., IC_1, IC_3, and IC_4) to be unable to service addresses having access routes mapped to the failed endpoint. The removal of all access routes to the failed endpoint allows the remaining endpoints to continue operating nominally and to service requests that are not directly dependent on the failed endpoint. In other words, the failed endpoint is taken offline and the remaining endpoints do not hang indefinitely because they no longer have the ability to process requests to the failed endpoint.
[0034] Figure 1BThe practical effect of isolating a dead endpoint (e.g., IC_2) in a CPU-coherent system 100 by updating one or more routing tables (as described above) is shown. Although the physical structure of the coherent mesh structure 130 remains unchanged, the dead endpoint is logically isolated so that other endpoints in the system (e.g., IC_1, IC_3, and IC_4) can no longer "see" the dead endpoint. All routing tables in the system have been updated to remove references to the logical addresses hosted by the dead endpoint. During the period that these removed logical addresses are missing from the system's routing tables, the remaining endpoints remain active and perform nominally, regardless of the fact that the system may not immediately (or within a period of time, hours, days, or more) reload the data associated with the removed addresses. For example, the CPU-coherent system 100 may not immediately (or ever) replace the removed routes with new routes to replace the physical location where the data associated with the removed addresses is stored.
[0035] It is noteworthy that the system mapping of logical addresses to physical memory space can be maintained unchanged by the above-mentioned updates to the physical routing table. That is, although the routing table (e.g., the table that provides inter-chip memory access) is updated to remove the subset of routes mapped to the failed endpoints, the actual logical-to-physical block mapping maintained on each processing chip remains unchanged.
[0036] In the event that any remaining active endpoint receives a read or write request for an address previously associated with the failed endpoint (e.g., IC_2), the routing engine processing the read / write request may return an "address not found" indicator. In some implementations, the system may, at this time, initiate a contingency protocol for locating data previously stored at the failed endpoint. For example, the system may use the address itself or other information in the read / write command to locate a backup copy of the data associated with the address and, in some scenarios, load the data from the backup location into an available location with memory 108, 112, or 114. However, all system endpoints except the failed endpoint continue to perform nominally during the period after the failed endpoint is isolated and before the data associated with the failed endpoint is restored from a redundant (backup) storage location elsewhere on the system.
[0037] Figure 2AAnother example CPU-coherent system 200 is shown that implements self-healing logic for selectively isolating an area of the system after detecting a faulty hardware component. The CPU-coherent system 200 includes a server blade 202 that includes a plurality of SoCs 204, 206, 208, and 210. Each of the SoCs 204, 206, 208, and 210 hosts memory, including volatile memory such as RAM or DRAM (e.g., VMem 212) and non-volatile memory such as one or more hard disk drive assemblies (HDAs), solid-state drives (SSDs), and the like (e.g., non-volatile storage banks 214). In addition, each SoC may include multiple individual processors (CPUs and GPUs) that are cache and memory coherent with each other and with all other processing entities in the server blade 202, coupled to each other via a coherent mesh fabric 230 that includes physical links between each pair of SoCs in the system and their associated resources, as well as coherent logic and hardware used to execute this logic (e.g., routing engines 222, 224, 226, and 228) to route requests to different endpoints in the system.
[0038] although Figures 1A-1B The examples of both 100 and 2A include four processing chips in the exemplary CPU-coherent systems 100 and 200, but it should be understood that the techniques disclosed herein are applicable to memory and cache-coherent systems having any number of interconnected chips (on the same PCBA or on different PCBAs in physical proximity (e.g., the same shelf or rack in a data center)).
[0039] By way of example and not limitation, server blade 202 is shown as a game server. Each of the different SoCs 204, 206, 208, and 210 acts as a host for an associated group of online game players (e.g., user groups 216, 218, 220, or 222, respectively). In this example, server blade 202 stores multiple games in non-redundant locations on respective non-volatile memory storage media. For example, SoC 204 may store games, such as Candy and Words With While SoC 206 stores other games, such as Call of and Grand Theft Each of these games can be stored in a single non-volatile memory location within the system. Due to the memory and cache coherent architecture of the CPU coherent system 200, each of the SoCs 204, 206, 208, and 210 is able to load any game stored in the system's non-volatile memory into local volatile memory, thereby allowing all of the different user groups 216, 218, 220, and 222 to play the game regardless of which non-volatile drive permanently stores the game. For example, It may actually be stored in the SSD of the non-volatile storage 214, but loaded into volatile memory (e.g., DRAM) hosted by the SoC 210 whenever one of the users in the user group 222 requests to load a new instance of the game.
[0040] Coherent mesh fabric 230 includes a data bus comprised of a physical interconnect 232, which in one implementation is a PCIe physical interface that provides data links between the various SoCs and between each SoC and each data storage resource in the system. For example, physical interconnect 232 gives SoC 204 direct access to memory hosted by SoC 206, SoC 208, and SoC 210, so that SoC 204 can read or update data at any system memory location without requiring the processor at the associated endpoint to take any action.
[0041] In addition to the physical interconnect 232, the coherent mesh structure 230 also includes Figure 2A 3. Routing logic and processing hardware represented in FIG3 as routing engines 234, 236, 238, and 240. In one implementation, coherent mesh fabric 304 is a memory interconnect with control electronics that use the PCIe interface standard to facilitate fast and direct communication between each pair of processing chips in CPU coherent system 300.
[0042] Server blade 202 includes a board management controller 242 that monitors the health and performance of various components in the system and creates a log file (not shown) that records actions that indicate potential hardware failures. For example, board management controller 242 can continuously record information related to the function of buses (e.g., physical interconnect 232 and routing engines 234, 236, 238, and 240), all system memory components, and all system processing components (e.g., each system CPU, GPU). Upon detecting any potential hardware component failure at a given endpoint (e.g., "failed endpoint 244," which includes SoC 206 and all memory and hardware hosted by SoC 206), board management controller 242 updates the log file to indicate the physical location of the potentially failed hardware component.
[0043] In the example shown, the volatile memory component 246 hosted by the SoC 206 experiences a sudden failure. This failure may affect the system in different ways, depending on what the failed drive is being used for in the system. For example, if the volatile memory component 246 stores game If there is no active copy of the game, the other system endpoints (e.g., SoC 204, SoC 208, and SoC 210) may experience a hang (system freeze) the next time they try to access the game.
[0044] In one implementation, the board management controller 242, in response to detecting a potential hardware failure, initiates a restart of the server blade 202. Upon restart, the board management controller 242 analyzes log data and, based on the analysis, determines that the failed component is hosted by the SoC 206. In response to identifying the endpoint hosting the failed component, the board management controller 242 instructs the coherent mesh fabric 230 (e.g., routing engines 234, 236, 238, and 240) to update its respective routing tables to remove all access routes to the endpoint corresponding to the SoC hosting the failed component.
[0045] In another implementation, the analysis of the log files and the dynamic rerouting are performed without taking the entire system offline. For example, the board management controller 242 can detect the failed volatile memory component 246 immediately when its host SoC (206) experiences a detectable error due to a hardware failure. At this point in time, the board management controller can dynamically (e.g., without restarting the CPU coherent system 200) instruct the coherent mesh structure 230 to update all system routing tables to remove the access route to the endpoint corresponding to the failed endpoint 244.
[0046] Figure 2B Shows the update Figure 1AThe CPU coherence system of one or more routing tables is used to selectively isolate one of the SoCs hosting the failed hardware component. In the example shown, routing engines 234, 238, and 240 have deleted from their respective routing tables all routes with endpoints mapped to the physical storage location hosted by failed endpoint 244. In one implementation, these access routes are completely erased from the routing tables without implementing alternative routes to the data hosted by failed memory component 246 in the event of a hardware failure. Due to the erasure of the routes, coherent mesh structure 232 can no longer determine where the latest version of the data is stored. Since the routing table of coherent mesh structure 232 no longer includes the route mapped to failed endpoint 244, the remaining active processing entities hosted by SoC 204, SoC 208, and SoC 210 no longer have a way to direct read or write requests for logical addresses routed through the routing table to failed endpoint 244 in the event of a hardware failure. This allows the remaining SoCs 204 , SoC 208 , and SoC 210 to continue operating at limited or near-nominal capacity until such time as the failed volatile memory component 246 can be replaced.
[0047] It is worth noting that the isolation of the dead endpoint 244 may render a significant portion of the system data temporarily unavailable to users connected to other system endpoints. In some implementations, the server blade 202 is configured to automatically initiate a contingency operation for reloading the data previously hosted by the dead endpoint 244 to an alternative system storage location when needed. For example, the next time the SoC 204 receives a request to join a game When a request is received for a new instance of the game, the SoC 204 may determine that the game cannot be found on the server blade 202 and begin searching for a backup copy of the data in other network locations. When a redundant copy of the game is located (e.g., on another server, a backup database, etc.), the game is then loaded into the non-volatile memory of one of the still active endpoints, where it can be accessed by all remaining active SoCs.
[0048] When driver 246 initially fails, the sudden failure may have the consequence of disrupting the connections of all users in user group 218 connected to SoC 206. For example, users in user group 218 may be temporarily disconnected from server blade 202 and added to one of user groups 216, 220, and 222 hosted by a different system endpoint while attempting to reestablish a connection. Alternatively, the users may connect to the gaming system through an entirely different server blade providing the same gaming service upon reconnection.
[0049] It should be understood that the gaming service is intended to represent only one of many different possible use cases for the technology disclosed herein. For example, similar applications can be implemented for cloud computing service providers that provide data storage and computing resources to cloud service providers. For example, thousands to millions of server blades similar to server blade 202 can be present in a data server center to support non-gaming workloads. In this system, each server can include multiple processor cores distributed on different chips that share memory and cache resources with each other. The ability to isolate the endpoint of a managed fault component allows the processing cores on other endpoints to remain active and continue to execute the nominal workload until the fault component can be serviced or replaced.
[0050] Similarly, the above-described isolation techniques can also be used in systems where multiple SoCs are coupled together to create virtual machines. For example, there is a scenario where multiple SoCs are configured to pool resources together to increase the processing and computing power available to end-user(s) in some way. In one implementation, a virtual machine is created by pooling the GPU processing power on two different SoCs together, effectively converting an 80-core GPU into a 160-core GPU. If one of the two SoCs of a virtual machine experiences a hardware failure, the other SoC can continue to attempt to use the resources on the failed chip. In this example, the coherent mesh structure 230 can reconfigure the routing table to remove access routes with logical addresses mapped to the failed chip, thereby reducing the effective computing power of the virtual machine from 160 cores back to 80 cores, but still allowing the active 80 cores to perform nominal computing operations.
[0051] Figure 3 An exemplary architecture of a CPU-coherent system 300 is shown, which implements self-healing logic for selectively isolating system endpoints after detecting a hardware component failure. View A shows a server blade 302 having four different SoCs 306, 308, 310, and 312 interconnected via a coherent mesh fabric 304. View B shows a zoomed-in view of SoC 310. It will be appreciated that the other SoCs 306, 310, and 312 may have the same or similar architecture.
[0052] SoC 308 is shown in view B as including four CPUs 320, 322, 324, and 326, each with eight processing cores. In one example implementation, multiple of these four different CPUs are configured to pool resources to provide a virtual machine experience to the end user. In addition to the multiple different CPUs, SoC 310 includes a graphics processing unit (GPU) 328, as well as hardware and software (e.g., various drivers), as well as encoders, decoders, I / O components, and buses (represented by functional block 330).
[0053] SoC 310 further includes multiple connections for optional coupling to DRAM (e.g., DRAM 314, 316). In addition, SoC 310 includes some of the hardware and executable logic represented by coherent mesh structure 304 in view A. Because coherent mesh structure 304 does not physically reside in its entirety on SoC 310, view B references coherent mesh structure 304a, which will be understood as references to a subset of coherent mesh structure 304. Coherent mesh structure 304a includes physical data links to memory and processing resources on each of the other SoCs (e.g., SoCs 306, 308, and 312). In addition, coherent mesh structure 304a includes a routing engine 332 that locally manages routing table 334 according to a system-level coherence scheme. The routing engine 332 dynamically updates the access routes in the routing table 334 so that the routing table 334 maps each logical address in the system to the physical endpoint (e.g., SoC 306, 308, 310, or 312) storing the latest version of the associated data, whether the data resides in volatile memory or non-volatile memory. Aspects of the routing table 334 not explicitly described herein can be the same or similar to those described above with reference to the routing table 134 of FIG. 1 .
[0054] In one implementation, the coherent mesh fabric 304 is a 1.2 to 1.8 GHz communication bus that uses PCIe interface links to provide connectivity between resources of each pair of endpoints in the CPU-coherent system 300. For example, the coherent mesh fabric 304 can be a bus that incorporates a controller at each endpoint to manage data flow across the physical links of the bus. In other implementations, the coherent mesh fabric 304 can take other physical forms, including, for example, any silicon bus that uses PCIe or other interface standards to provide direct communication between system endpoints.
[0055] Figure 4 Example operations 400 for implementing self-healing logic within a cache and memory coherent multi-chip system are shown. According to one implementation, a cache and memory coherent system includes multiple processing chips that each host a different subset of a shared memory space. Monitoring operations 402 monitor the health of various components in the cache and memory coherent system and populate a log file with event information that may indicate a hardware failure. For example, monitoring operations 402 may include sending test signals to multiple different processing components and recording time-stamped response information. Determining operations 404 determine whether and when a system event occurs that meets predefined emergency criteria. The predefined emergency criteria may be met, for example, when the system event is an unresponsive CPU or other event that indicates a CPU crash, a system hang, etc.
[0056] Until the detected system event meets the emergency criteria, monitoring operation 402 continues to monitor the system and populate the log file. When determining operation 404 determines that the detected event does meet the emergency criteria, parsing operation 406 parses the log file to identify the faulty hardware component that caused the system event. In some implementations, the system may be restarted and the log file parsed in response to the restart. Parsing operation 406 identifies a selected processing chip from among a plurality of processing chips in the system that hosts the faulty hardware component that caused the system event.
[0057] In response to identifying the selected processing chip as the host of the failed hardware component, a routing table update operation 408 updates routing tables throughout the cache and memory coherent system to remove (e.g., delete) all access routes defined by these tables that have endpoints corresponding to the selected processing chip hosting the failed hardware component. For example, a subset of the system's logical block address space may be temporarily erased from the system, as evidenced by the logical addresses ceasing to correspond to any defined physical storage locations. Removing these logical addresses from the system routing tables allows the remaining processing chips (e.g., excluding the selected chip) to continue operating as normal, while the selected processing chip remains isolated (e.g., invisible to other processing chips for system routing operations).
[0058] Figure 5 An example schematic diagram of a processing device 500 suitable for implementing various aspects of the disclosed technology is illustrated. The processing device 500 can, for example, represent a user device that interfaces with a cache and memory coherent system or, instead, represent a device that includes multiple processing chips operating as a cache and memory coherent system. The processing device includes one or more processor units 502, memory device(s) 504, a display 506, and other interfaces 608 (e.g., buttons). The processor unit(s) 502 can each include one or more CPUs, GPUs, etc.
[0059] Memory 504 generally includes both volatile memory (e.g., RAM) and non-volatile memory (e.g., flash memory). Operating system 510 (such as Microsoft Operating system, Microsoft The Windows Phone operating system or a specific operating system designed for gaming devices) may reside in memory 504 and be executed by processor unit(s) 502, but it will be appreciated that other operating systems may be employed.
[0060] One or more applications 512 are loaded into the memory 604 and executed by the processor unit(s) 602 on the operating system 610. The applications 512 may receive input from various local input devices such as a microphone 534, input accessories 535 (e.g., a keypad, mouse, stylus, touchpad, game pad, steering wheel, joystick), and a camera 532 (e.g., to provide scene footage for multiple object trackers). In addition, the applications 512 may receive input from one or more remote devices (such as remotely located smart devices) by communicating with such devices over a wired or wireless network using further communication transceivers 530 and antennas 538 to provide network connectivity (e.g., mobile phone networks, ). The processing device 500 may also include one or more storage devices 528 (eg, non-volatile storage). Other configurations may also be employed.
[0061] Where the processing device 500 operates a multi-chip cache and memory coherent system, the processor 502 may be distributed across different chips (e.g., SoCs) interconnected by a coherent mesh structure (not shown) including the processors described herein. Figures 1A-1B , 2A-2B or 3-4.
[0062] The processing device 500 further includes a power supply 516, which is powered by one or more batteries or other power sources and provides power to the other components of the processing device 500. The power supply 516 may also be connected to an external power source (not shown) that overrides or recharges the internal batteries or other power sources.
[0063] Processing device 500 may include various tangible computer-readable storage media and intangible computer-readable communication signals. Tangible computer-readable storage can be embodied by any available media accessible by processing device 600 and includes both volatile and non-volatile storage media, and removable and non-removable storage media. Tangible computer-readable storage media does not include intangible and transient communication signals, but rather includes volatile and non-volatile, removable and non-removable storage media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Tangible computer-readable media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVDs) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other tangible medium that can be used to store the desired information and that can be accessed by processing device 600. In contrast to tangible computer-readable storage media, intangible computer-readable communication signals may embody computer-readable instructions, data structures, program modules, or other data using a modulated data signal such as a carrier wave or other signal transmission mechanism. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, intangible communication signals include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency, infrared, and other wireless media.
[0064] Some embodiments may include an article of manufacture. The article of manufacture may include a tangible storage medium (memory device) for storing logic. Examples of storage media may include one or more types of processor-readable storage media capable of storing electronic data, including volatile or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, etc. Examples of logic may include various software elements, such as software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, operation segments, methods, processes, software interfaces, application program interfaces (APIs), instruction sets, computing codes, computer codes, code segments, computer code segments, text, values, symbols, or any combination thereof. For example, in one implementation, an article of manufacture may store executable computer program instructions that, when executed by a computer, cause the computer to perform methods and / or operations according to the described various implementations. Executable computer program instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. Executable computer program instructions can be implemented according to a predefined computer language, method or syntax for instructing a computer to perform a specific operation segment. These instructions can be implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language.
[0065] An example cache and memory coherence system disclosed herein includes a plurality of processing chips, each processing chip hosting a different subset of a shared memory space and one or more routing tables defining access routes between logical addresses of the shared memory space and endpoints corresponding to a selected one of the plurality of processing chips. The system further includes a coherent mesh structure physically coupling each of the plurality of processing chips together. The coherent mesh structure is configured to execute routing logic for updating the one or more routing tables in response to detecting a faulty hardware component and identifying a first processing chip hosting the faulty hardware component. The updating of the one or more routing tables effectively removes all access routes having endpoints corresponding to the first processing chip.
[0066] In an example cache and memory coherent system according to any of the foregoing systems, the coherent mesh fabric includes a communication bus using a Peripheral Component Interconnect Express (PCIe) physical interface. In another example cache and memory coherent system according to any of the foregoing systems, the processing chips are system-on-chip (SoCs) disposed on a same printed circuit board assembly (PCBA).
[0067] In yet another cache and memory coherent system of any of the foregoing systems, the system includes a controller stored in the memory and configured to analyze system log information to identify a first processing chip hosting a failed hardware component.
[0068] In another example cache and memory coherence system of any of the foregoing systems, the cache and memory coherence system is further configured to pool resources on two or more of the processing chips together to provide a virtual machine experience to the user.
[0069] In yet another example cache and memory coherent system according to any of the preceding systems, the coherent mesh fabric executes routing logic for updating the one or more routing tables in response to a restart of the cache and memory coherent system.
[0070] An example method disclosed herein includes analyzing system log information to identify a location of a failed hardware component within a cache and memory coherent system, the cache and memory coherent system including a plurality of processing chips each hosting a different subset of a shared memory space. The method further provides updating one or more routing tables to remove all access routes that map a logical address of the shared memory space to an endpoint corresponding to a first processing chip corresponding to the location of the failed hardware component.
[0071] In a further example method of any of the preceding methods, updating the one or more routing tables includes removing an access route having an endpoint corresponding to the first processing chip without adding a new access route to a logical address identified by the removed access route.
[0072] In a further example method of any of the foregoing methods, the plurality of processing chips are coupled together via a coherent mesh fabric including a communication bus using a Peripheral Component Interconnect Express (PCIe) physical interface.
[0073] In yet another example method of any of the foregoing methods, the processing chip is a system on a chip (SoC) disposed on a same printed circuit board assembly (PCBA).
[0074] In a further example method of any of the foregoing methods, the method further includes generating a log file including system log information indicative of a potentially failed hardware component hosted by any of the plurality of processing chips.
[0075] In yet another example method of any of the foregoing methods, updating the one or more routing tables includes updating a routing table stored on each of a plurality of different processing chips in the plurality of processing chips.
[0076] In a further example method of any of the foregoing methods, the coherent mesh fabric executes logic for updating the one or more routing tables in response to a restart of the cache and memory coherent system.
[0077] An example tangible computer-readable storage medium disclosed herein encodes computer-executable instructions for performing a computer process, the computer process comprising: analyzing system log information to identify a location of a faulty hardware component within a cache and memory coherent system, the cache and memory coherent system comprising a plurality of processing chips each hosting a different subset of a shared memory space. The method further provides, in response to determining that the location of the faulty hardware component corresponds to a first processing chip among the plurality of processing chips, updating one or more routing tables to remove all access routes that map logical addresses of the shared memory space to endpoints corresponding to the first processing chip.
[0078] In another example tangible computer-readable storage medium of any of the foregoing storage media, a coded computer process provides for updating a routing table for facilitating communication between a plurality of processing chips coupled together via a coherent mesh fabric including a communication bus using a Peripheral Component Interconnect Express (PCIe) physical interface.
[0079] In yet another example tangible computer readable storage medium of any of the foregoing storage media, the routing logic facilitates communication between processing chips that are systems on a chip (SoCs) disposed on a same printed circuit board assembly (PCBA).
[0080] In yet another example tangible computer readable storage medium of any of the foregoing storage media, updating the one or more routing tables includes removing an access route having an endpoint corresponding to the first processing chip without adding a new access route to a logical address identified by the removed access route.
[0081] In yet another example tangible computer readable storage medium of any of the foregoing storage media, updating the one or more routing tables includes updating a routing table stored on each of a plurality of different ones of the plurality of processing chips.
[0082] In yet another example tangible computer readable storage medium of any of the foregoing storage media, the logic for updating the one or more routing tables is executed in response to a restart of the cache and memory coherent system.
[0083] An example system disclosed herein includes a device for analyzing system log information to identify the location of a failed hardware component within a cache and memory coherent system, the cache and memory coherent system including a plurality of processing chips each hosting a different subset of a shared memory space. The system further includes a device for updating one or more routing tables to remove all access routes that map a logical address of the shared memory space to an endpoint corresponding to a first processing chip corresponding to the location of the failed hardware component.
[0084] The logical operations described herein are implemented as logical steps in one or more computer systems. The logical operations may be implemented as: (1) a sequence of processor-implemented steps executed in one or more computer systems; and (2) interconnected machine or circuit modules within one or more computer systems. The implementation is a matter of choice depending on the performance requirements of the computer systems being utilized. Accordingly, the logical operations that make up the various implementations described herein may also be referred to as operations, steps, objects, or modules. Furthermore, it should be understood that the logical operations may be performed in any order unless expressly stated or the claim language inherently requires a particular order. The above description, examples, and data, together with the accompanying drawings, provide a comprehensive description of the structure and use of the exemplary implementations.
Claims
1. A cache and memory coherence system comprising: a plurality of processing chips each hosting a different subset of the shared memory space, each hosted subset of the shared memory space corresponding to a logical address range; one or more routing tables that define access routes between logical addresses of the shared memory space and endpoints that each correspond to a selected subset of the shared memory space hosted by a respective processing chip of the plurality of processing chips; as well as a coherent mesh fabric physically coupling each of the plurality of processing chips together, the coherent mesh fabric configured to execute routing logic for updating the one or more routing tables in response to detecting a failed hardware component and identifying a first processing chip among the plurality of processing chips that hosts the failed hardware component, the updating of the one or more routing tables effectively removing all access routes having endpoints corresponding to a first logical address range hosted by the first processing chip.
2. The cache and memory coherent system of claim 1 , wherein the coherent mesh structure executes the routing logic to remove an access route having an endpoint corresponding to the first processing chip without adding a new access route to a logical address identified by the removed access route.
3. The cache and memory coherence system of claim 1, wherein the coherent mesh fabric comprises a communication bus using a Peripheral Component Interconnect Express (PCIe) physical interface.
4. The cache and memory coherent system of claim 1, wherein the processing chip is a system on a chip (SoC) disposed on a same printed circuit board assembly (PCBA).
5. The cache and memory coherence system of claim 1 , further comprising a controller stored in the memory and configured to analyze system log information to identify the first of the plurality of processing chips hosting the failed hardware component.
6. The cache and memory coherence system of claim 1, wherein the cache and memory coherence system is further configured to pool resources on two or more of the processing chips together to provide a virtual machine experience to a user.
7. The cache and memory coherent system of claim 1, wherein the coherent grid fabric executes routing logic for updating the one or more routing tables in response to a restart of the cache and memory coherent system.
8. A method comprising: analyzing system log information to identify a location of a failed hardware component within a cache and memory coherent system, the cache and memory coherent system including a plurality of processing chips each hosting a different subset of a shared memory space, each hosted subset of the shared memory space corresponding to a logical address range; as well as In response to determining that the location of the failed hardware component corresponds to a first processing chip among the plurality of processing chips, one or more routing tables defining access routes between logical addresses of the shared memory space and endpoints that each correspond to a selected subset of the shared memory space hosted by a corresponding processing chip among the plurality of processing chips are updated to remove all access routes that map logical addresses of the shared memory space to endpoints corresponding to a first logical address range hosted by the first processing chip.
9. The method of claim 8, wherein updating the one or more routing tables comprises removing an access route having an endpoint corresponding to the first processing chip without adding a new access route to a logical address identified by the removed access route.
10. The method of claim 8, wherein the plurality of processing chips are coupled together via a coherent mesh fabric comprising a communication bus using a Peripheral Component Interconnect Express (PCIe) physical interface.
11. The method of claim 8, wherein the processing chip is a system on chip (SoC) disposed on a same printed circuit board assembly (PCBA).
12. The method of claim 8, further comprising: A log file is generated that includes the system log information, the system log information indicating a possibly failed hardware component hosted by any of the plurality of processing chips.
13. The method of claim 8, wherein updating the one or more routing tables comprises updating a routing table stored on each of a plurality of different ones of the plurality of processing chips.
14. The method of claim 10, wherein the coherent mesh fabric executes logic for updating the one or more routing tables in response to a restart of the cache and memory coherent system.
15. One or more tangible computer-readable storage media encoding computer-executable instructions for performing a computer process comprising: analyzing system log information to identify a location of a failed hardware component within a cache and memory coherent system, the cache and memory coherent system including a plurality of processing chips each hosting a different subset of a shared memory space, each hosted subset of the shared memory space corresponding to a logical address range; as well as In response to determining that the location of the failed hardware component corresponds to a first processing chip among the plurality of processing chips, one or more routing tables defining access routes between logical addresses of the shared memory space and endpoints that each correspond to a selected subset of the shared memory space hosted by a corresponding processing chip among the plurality of processing chips are updated to remove all access routes that map logical addresses of the shared memory space to endpoints corresponding to a first logical address range hosted by the first processing chip.
16. The one or more tangible computer-readable storage media of claim 15, wherein the plurality of processing chips are coupled together via a coherent mesh fabric comprising a communication bus using a Peripheral Component Interconnect Express (PCIe) physical interface.
17. The one or more tangible computer-readable storage media of claim 15, wherein the processing chip is a system on a chip (SoC) disposed on a same printed circuit board assembly (PCBA).
18. The one or more tangible computer-readable storage media of claim 15, wherein updating the one or more routing tables comprises removing an access route having an endpoint corresponding to the first processing chip without adding a new access route to a logical address identified by the removed access route.
19. The one or more tangible computer-readable storage media of claim 15, wherein updating the one or more routing tables comprises updating a routing table stored on each of a plurality of different ones of the plurality of processing chips.
20. The one or more tangible computer-readable storage media of claim 15, wherein updating the one or more routing tables is performed in response to a restart of the cache and memory coherent system.
21. A computer system comprising means for performing the method of any one of claims 8-14.