Flexible High-Availability Computing with a Parallel Configurable Structure
Through the parallel interconnect structure and resource adapter mechanism, the computing module and I/O module logic are bound to an equivalent PC server, which solves the problems of inefficient resource utilization and insufficient fault tolerance in the existing technology, and realizes efficient and flexible computing resource management and failover.
Patent Information
- Application Number
- CN202111261399.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-30
- Filing Date
- 2021-10-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-10-28
AI Technical Summary
Existing custom configuration technologies have failed to optimally configure in providing efficient use, resource utilization and fault tolerance of computing functionality, especially in the PCIe standard, where fault tolerance support is lacking in link failure and resource utilization is inefficient.
The parallel interconnect structure and resource adapter (RA) mechanism are adopted to logically bind the computing module and I/O module logic to equivalent PC servers through carrier switching structures such as PCIe, InfiniBand or Gen-Z. The structure manager dynamically manages resource configuration and failover, providing high availability and fault tolerance.
It realizes efficient utilization and fault tolerance of computing resources, supports dynamic resource management and failover, and improves the flexibility and reliability of the system.
Smart Images

Figure CN115269476B_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] Current custom configuration techniques generally involve selecting which components (e.g., processors, memories, input / output (I / O) interfaces, storage devices) are connected within a physical chassis. For example, the interconnection can be provided via a Peripheral Component Interconnect (PCI) compliant bus to which components can be connected. The PCI standard is available from the PCI Special Interest Group (PCI-SIG) located in Beaverton, Oregon, USA. By providing a standard interface, various components are interconnected to provide a desired system configuration. However, these architectures are not optimally configured in terms of providing efficient use of computing functionality, resource utilization, and / or fault tolerance. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Embodiments of the present invention are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which like reference numerals refer to similar elements.
[0003] Figure 1 is a block diagram of one embodiment of a composable computing environment having a single parallel interconnect structure.
[0004] Figure 2 is a block diagram of one embodiment of a composable computing environment having a stacked interconnect structure.
[0005] Figure 3 is a logical representation of one embodiment of an architecture that can utilize a parallel interconnect structure to provide a flexible configurable system.
[0006] Figure 4 is a conceptual diagram of one embodiment of an infrastructure resource pool that provides composable resources from which resource instances or complete devices can be bound together and configured to create a logical server that operates as a traditional server.
[0007] Figure 5 is a block diagram of one embodiment of a trusted composable computing architecture for trusted composable infrastructure resources having intelligent modules, simple modules, and a parallel interconnect structure from which server modules can expand their I / O capabilities or create logical servers.
[0008] Figure 6 is a block diagram of one embodiment of an infrastructure resource pool that provides composable infrastructure resources in a hardened security environment where logical server resource bindings are connected cryptographically via a parallel interconnect structure.
[0009] Figure 7 is an example of a parallel interconnect structure packet frame header that can be used with the composable architectures described herein for a PCIe parallel interconnect structure built using a standard PCIe switch fabric.
[0010] Figure 8 It is a flowchart of an embodiment of a technique for configuring computing infrastructure resources in an environment that provides composable resources.
[0011] Figure 9 It is a flowchart of an embodiment of a technique for configuring computing resources in an environment that provides composable resources.
[0012] Figure 10 It is a conceptual diagram of an embodiment of an infrastructure resource pool that provides composable infrastructure resources configured as a set of logical servers.
[0013] Figure 11 It is a conceptual diagram of a logical server configured to utilize peer - to - peer communication on a parallel interconnect structure of a composable architecture. Detailed Description
[0014] In the following description, many specific details are set forth. However, embodiments of the present invention may be practiced without these specific details. In other instances, well - known structures and techniques have not been shown in detail in order not to obscure the understanding of this description.
[0015] As described in more detail below, an efficient and configurable computing environment, referred to as a composable infrastructure architecture, can be provided that is modular, scalable, and highly available. This composable infrastructure architecture is not constrained by distance and utilizes a parallel interconnect structure that can provide parallel interconnect and advanced fault tolerance.
[0016] In various embodiments, the composable architecture described herein can provide the functionality of logically binding simple components (e.g., computing modules (e.g., (multiple) processors, memory)) to I / O modules (e.g., Peripheral Component Interconnect Express or “PCIe”, Graphics Processing Unit or “GPU”, Network Interface Card or “NIC”, storage controller, extended memory, Non - Volatile Memory Express or “NVMe”) across a carrier - switching fabric (e.g., PCIe, InfiniBand or “IB”, Generation Z or “GenZ”), such that the logically - bound components (partially or fully) function as an equivalent PC server (logical server).
[0017] As described in more detail below, this functionality can be provided using a single interface (e.g., an environment Application Programming Interface (API)), which can provide the ability to control environment functionality (e.g., composition, binding, Software License Agreement (SLA) enforcement) through a single interface and facilitate the use of Resource Adapters (RAs) that provide interfaces to individual modules within the computing environment. Thus, functionally, the composable environment can be viewed as a single configurable device accessible by a single interface.
[0018] In some embodiments, multiple parallel interconnect structures can be utilized to provide a single integrated structure that can provide resiliency, redundancy, and / or parallelism beyond that of the individual parallel interconnect structures alone. The individual physical interconnections can be managed by their own managers and protocols and can be fully addressable such that each RA can be reached by other RAs. Thus, an RA can be considered "multi-homed" to each of the parallel interconnect structures in the parallel interconnect structure to provide an interface to each interconnect structure. Accordingly, the physical interconnect structures can be managed by their own native managers and protocols, and an RA can be reached by any other RA in the environment and can provide hardware failover functionality. In some embodiments, the individual physical interconnect managers can be clustered to provide a single interface to the integrated structure.
[0019] As described in more detail below, RAs can communicate with and bind to each other using an appropriate addressing scheme for the underlying physical fabric (e.g., Ethernet can utilize MAC addresses, and PCIe can utilize bus device functions). A single management fabric (and corresponding interface) can be used to communicate with the various RAs on how to bind together within a composable environment.
[0020] These mechanisms and structures can allow a system administrator (or other entity) to describe desired functionality at a relatively high level using an API, and the underlying management entities and supporting hardware within the environment can provide the desired functionality. Further, when conditions or requirements change, the configuration of the environment can be updated / changed using an API.
[0021] In various embodiments, the composable architectures described herein can provide the functionality of logically binding simple components (e.g., compute modules (e.g., (multiple) processors, memory)) to IO modules (e.g., Peripheral Component Interconnect Express or "PCIe", Graphics Processing Unit or "GPU", Network Interface Card or "NIC", storage controller, extended memory, Non-Volatile Memory Express or "NVMe") across carrier switching fabrics (e.g., PCIe, InfiniBand or "IB", Generation Z or "GenZ") such that the logically bound components (partially or fully) act as, for example, an equivalent PC server (logical server).
[0022] In various embodiments described herein, a parallel interconnect structure based on a Peripheral Component Interconnect (PCI) compliant interface can provide a configurable, fault-tolerant interconnect between multiple resources and compute nodes. Subsequent PCI standards can also be used. For example, related standards such as Compute Express Link (CXL) can be used with the techniques and architectures described herein. CXL is an open standard high-speed processor-to-device and processor-to-memory interconnect standard provided by the Compute Express Link Consortium of Beaverton, Oregon. Similarly, the Gen-Z interconnect standard can be supported. The Gen-Z standard is provided by the Gen-Z Consortium.
[0023] In the described embodiments, enhanced functionality can be achieved via the use of one or more Resource Adapters (RAs - e.g., illustrated as Application-Specific Integrated Circuit or “ASIC” modules). In one example embodiment, an RA can be configured in one of three modes: 1) Compute RA (C-RA); 2) Target RA (T-RA); and 3) Stacked RA (S-RA). A Compute RA (e.g., Compute Resource Adapter) or a Target RA (e.g., Target I / O Resource Adapter) can provide a hardware protocol link that bridges between module local interconnects (e.g., PCIe, CXL) and a carrier switching fabric (e.g., PCIe, IB, GenZ).
[0024] In various embodiments described herein, each RA provides a single control surface (i.e., programming interface) that enhances the carrier switching fabric capabilities, such as using a PCIe switch, into a hardware resilient (e.g., hardware failover, i.e., without a higher-level management software to orchestrate the failover) structure on multiple isolated and independent paths to provide a parallel interconnect structure. Additionally, an RA may be able to enhance the carrier switching fabric capabilities to provide end-to-end fabric service quality of service (QoS) with cryptographic isolation between, for example, bound and logically bound resources (e.g., logical servers that can be linked to an external cryptographic key management service).
[0025] Further, in other embodiments, binding can be extended to standard intelligent infrastructure resources such as traditional PC servers or storage arrays, thereby providing flexible I / O expansion or peer-to-peer communication across the parallel interconnect structure. For example, each isolated fabric can be independently managed by its own fabric manager but orchestrated together as a clustered set to provide a single unified interface and experience.
[0026] Due to structural independence, the structure types or structural latencies do not need to be the same. Thus, a parallel structure only requires structural node addressability (e.g., to the RA) to be structurally reachable. In some embodiments, when a parallel interconnect structure grows or expands to provide scalability, a stacked RA (S-RA) can be used to bridge between two parallel interconnect structure domains. That is, for example, a structural branch in a parallel interconnect structure A domain is connected to a structural branch in a parallel interconnect structure B domain using an S-RA to address resolution and quality of service (QoS) considerations within the target parallel structure domain.
[0027] In some configurations, all parallel interconnect structure domains along with the same structural node addressability to the RA domain can be referred to as an infrastructure resource pool, which can be independent of the number of parallel interconnect structures. Due to the use of programmatic binding, all resources can be managed in a real-time on-demand manner, thus supporting modern intelligent applications that require binding new resources to them (computing resources) or creating new computing resources (logical server nodes). Additionally, unused resources can be released by the intelligent applications back to the idle pool for reallocation.
[0028] For example, the described modularity can be provided because input / output (I / O) components do not need to be constrained within the physical devices they support, computing nodes, resources, switches, and routers can all be dynamically added to the architecture, and various components can be interconnected over any distance (e.g., chip-to-chip, board-to-board, via copper cables, via optical fibers, via the Internet). In various embodiments, application-transparent alternate path partitioning can provide independence and fault isolation. Further, a single management footprint can be provided.
[0029] Conceptually, the architecture can be considered as PCI-based intermediate management of I / O devices, including, for example, NVMe memory devices and / or PCIe memory devices. Further, the architecture described herein can be used to provide fault tolerance for non-fault-tolerant devices and redundancy for PCIe-based systems (or other interconnect architectures that do not provide native redundancy). That is, the architecture described herein provides I / O resiliency in systems that do not natively support fault tolerance and I / O resiliency.
[0030] For example, PCIe is based on a point-to-point topology with separate serial links connecting each device to the host. However, the PCIe standard does not provide native support for fault tolerance in response to link failures. Further, providing fault tolerance by providing redundant links results in inefficient use of resources. Thus, the techniques and mechanisms described herein can allow mature, widely deployed technologies such as PCIe to be leveraged to provide an improved, more efficient, and more advanced architecture.
[0031] In some embodiments, a first switching fabric routes packets between compute resources and I / O resources. The compute resources each have a compute resource adapter (C-RA) to configure a destination identifier with a corresponding first frame identifier and a source identifier with a corresponding second frame identifier for the header of a packet. The I / O resources each have a target resource adapter (T-RA) to manage packets based on the header. A second switching fabric also routes packets between the compute resources and the I / O resources.
[0032] One or more fabric managers can be coupled to the first and second switching fabrics to route packets in parallel through the first and second switching fabrics when the packet transmission failure rate is below a preselected failure rate, and to dynamically reconfigure the routing of packets on the first and second switching fabrics in response to the packet transmission failure rate exceeding the preselected failure rate. The (s) fabric manager(s) can also be used to provide more complex available resource allocation by supporting and managing a composable computing environment.
[0033] Figure 1 is a block diagram of an embodiment of a composable computing environment having a single parallel interconnect fabric. Figure 1 Examples of configure one or more logical servers from the resources of a composable computing environment; however, this is only one example of the type of configuration that can be supported.
[0034] In Figure 1 example, a composable resource manager 120 (which can be a cluster of composable resource managers) can provide management connectivity for the control of the (s) parallel interconnect fabric(s) 110 and the (s) compute module(s) 150 and the (s) target module(s) 160. In some implementations, the composable resource manager 120 can also have management connectivity (e.g., 154, 164) to resource adapters (RAs) in the compute and target modules (e.g., 152, 162), but in Figure 1 example, the RAs 152 and 162 are controlled via the parallel interconnect fabric (PIF) 110 through an API from a fabric management module (e.g., 130, 135, 138). The parallel interconnect fabric 110 can include any number of fabric fabrics (e.g., 112, 114, 116) having fabric management modules (e.g., 130, 135, 138).
[0035] In Figure 1In the example, the user interface is represented by the API / GUI / Script 125 of the composable explorer. In one embodiment, the composable explorer to API / GUI / script interface can provide information, health status, and control of the entire topology of the parallel interconnect structure 110 to the attached computing and target modules (e.g., 150, 160). In one embodiment, the composable explorer API 125 will include information, health status, and control of the PIF 110, and can include information, health status, and control of each RA connected to the PIF 110, similar to Figure 1 the example case.
[0036] In one embodiment, the composable explorer 120 directly manages the RA modules (e.g., 152, 162) via one or more APIs specific to the information, health status, and control of different RAs. In one embodiment, the composable explorer 120 can also associate each RA with a computing module (e.g., 150) or a target module (e.g., 160). In Figure 1 the example, the composable explorer 120 has management connections to the computing module 150 and the target module 160, and the APIs on these connections provide information about module identification, module inventory (including the identification of the fabric RAs), and other resources within the module.
[0037] An example of a management entity that can provide information, health status, and control in the computing module 150 or the target module 160 is the baseboard management controller or BMC. In one embodiment, the composable explorer 120 can discover the (multiple) parallel interconnect structures 110 through management connections (e.g., 122, 126, 128) to the fabric managers (e.g., 130, 135, 138), including, for example, the complete connection topology to all RAs.
[0038] The composable explorer discovers the association of each resource adapter with the parallel interconnect structure and the host computing or target module through the (multiple) fabric managers or through management connections to each computing or target module management API. In one embodiment, the composable explorer 120 can also discover the resources within each computing or target module to build a complete resource inventory and parallel interconnect structure topology, and can store this information in the database 140, memory, or some combination thereof.
[0039] In one embodiment, the user interface 125 of the composable resource manager 120 is capable of querying information and health status regarding the overall topology, or requesting a change to the (multiple) resources of any computing module. The composable resource manager 120 is capable of responding to a requested resource addition or reduction or a response that the request cannot be satisfied or an error that occurs during operation by identifying the (multiple) resources that can be reconfigured on the parallel interconnect fabric to meet the request and completing the reconfiguration.
[0040] In Figure 1 the example of, one or more logical servers 170 can be configured as described above. Figure 1 An example of a logical server configuration is a PCIe-based configuration; however, as discussed above, various other protocols can be supported. Each logical server can include one or more processing resources 172 (from one or more processor modules 150), a logical PCIe switch 175 (from the parallel interconnect fabric 110), and any number of logical PCIe devices 180, 184, 188 (from one or more target modules 160).
[0041] Figure 2 is a block diagram of one embodiment of a composable computing environment having a stacked interconnect fabric. Figure 2 The example of configures one or more logical servers from the resources of the composable computing environment; however, this is just one example of the types of configurations that can be supported.
[0042] Figure 2 The composable computing architecture includes three types of RAs (described in more detail below): a Compute Resource Adapter (C-RA), a Target Resource Adapter (T-RA), and a Stack Resource Adapter (S-RA). In addition to the various types of resource adapters, one or more fabric modules and corresponding fabric managers can provide composable interconnectivity.
[0043] At a high level, in one embodiment, Figure 2 the architecture can utilize end-to-end credit-based flow control (e.g., between one or more compute resources and one or more target resources). Figure 2 The example embodiment provides three example types of parallel interconnect fabrics: 1) Parallel Interconnect Fabric 1 (260), which provides an interconnect between one or more compute resources and another parallel interconnect fabric; 2) Parallel Interconnect Fabric 2 (262), which provides an interconnect between other parallel interconnect fabrics; and 3) Parallel Interconnect Fabric 3, which provides an interconnect between one or more target resources and another parallel interconnect fabric. Any number of any of these types of parallel interconnect fabrics can be supported. Compared with Figure 1 the non-stacked architecture of Figure 2 the illustrated stacked architecture provides more functionality and greater flexibility.
[0044] A variety of distances can be supported by providing interconnections at the chip-to-chip level, board-to-board level, copper cable, optical fiber, or any combination thereof. Thus, the dynamic addition and removal of end nodes, switches, and / or routers can be supported.
[0045] In various embodiments, the stacking links (such as 202, 204, 206, 208) can be any type of interconnection link, such as copper, optical fiber, and can be utilized between S-RAs (such as 253, 254, 255, 256, 265, 266). That is, the stacking links can be a different type of interconnection from the interconnection structure links.
[0046] In Figure 2 an example, the composable resource manager 215 can be a cluster of composable resource managers and has management connectivity for controlling the parallel interconnection structures 260, 262, and 264 and the shown (multiple) computing modules 250 and intelligent target modules 260 (such as 280, 281, 282, 283, 284, 285, 288, 289). In some implementations, the composable resource manager 215 can also have management connectivity to the RAs in the computing module 250 (such as 252) and the target module 260 (such as 262), but in this example, the RA is controlled through the API from the fabric management module through the parallel interconnection structures 260, 262, 264.
[0047] In one embodiment, the user interface 210 can be the API / GUI / script of the composable resource manager 215. In one embodiment, the composable resource manager to the API / GUI / script interface provides information, health status, and topology of any number of parallel interconnection structures (such as 260, 262, 264) to all attached computing modules 250 and target modules 260. In one embodiment, the composable resource manager API to the (multiple) parallel interconnection structure managers can include information, health status, and control of the structure, and can include information, health status, and control of each RA connected to the structure.
[0048] In one embodiment, the composable resource manager 215 can directly manage the RA modules (such as 250, 260), which can utilize APIs specific to the different RAs for information, health status, and control. In one embodiment, the composable resource manager 215 can be used to associate each RA with a computing module or a target module.
[0049] In this example, the composable resource manager 215 can have management connections to the (multiple) compute modules 250 and the (multiple) target modules 260, and the APIs on these connections can be used to provide information such as module identification, module inventory (including identification of the fabric RAs), and other resources within the module. Examples of management entities that can provide information, health status, and control within a compute module or target module are the baseboard management controller or BMC.
[0050] In one embodiment, the composable resource manager 215 can be used to discover a parallel interconnect fabric (such as 260, 262, 264) through management connections (such as 280 to 285) with fabric managers 290, 292, 294, 295, 297, 298, including the complete connection topology to all RAs. In one embodiment, the composable resource manager 215 discovers the association of each RA with the parallel interconnect fabrics 260, 262, 264 and the hosting compute module or target module through the (multiple) fabric managers or through management connections to each compute module or target module management API.
[0051] In one embodiment, the composable resource manager 215 also discovers the resources within each compute module or target module in order to build a complete resource inventory and parallel interconnect fabric topology, and this information can be stored in the database 220, memory, or some combination thereof. In one embodiment, the user interface 210 of the composable resource manager 215 can query information and health status about the entire topology, for example, or request a change to the (multiple) resources of any compute module. In one embodiment, the composable resource manager 215 can respond to a requested resource addition or reduction or a response that the request cannot be satisfied or an error that occurs during operation by identifying the (multiple) resources that can be reconfigured on the parallel interconnect fabrics 260, 262, 264 to meet the request and successfully complete the reconfiguration.
[0052] In various embodiments, each RA (i.e., C-RA or T-RA) is mapped to its own interconnect fabric address domain. Each S-RA bridges transactions between interconnect fabric address domains and is thus a member of two interconnect fabric address domains.
[0053] Multiple resource modules can be interconnected via an interconnect fabric. In various embodiments, a resource module can have multiple RAs that can provide different types of resources to a compute module. Figure 2 Examples can include, for example, intelligent I / O modules that can include multifunctional devices, intelligent I / O modules that can include multi-host devices, I / O modules that can include an interface for receiving pluggable resources, or memory modules. In other embodiments, different types of resources and / or different quantities of resources can be supported.
[0054] Figure 2 The example architecture can provide several advantages over current architecture strategies. At a high level, the architecture described herein utilizes a composable structure that enables platform disaggregation to facilitate modular computing resources and modular I / O resources that can be dynamically configured to meet various needs. Thus, I / O scaling can be provided to meet the changing needs of the computing resources. Further, since the I / O resources can be shared sequentially or concurrently among the computing resources, I / O utilization can be more efficient. In some embodiments, peer-to-peer connections between resource modules can provide more efficient utilization of the computing node resources.
[0055] In an example embodiment, the I / O module can include one or more dual-port NVMe devices that can be coupled to two RAs. Thus, the dual-port functionality of the NVMe devices can be used within the composable computing architecture. Various embodiments and configurations of the NVMe devices are described in more detail below.
[0056] In various embodiments, the fabric manager can control the traffic flow through the interconnect fabric to utilize parallel paths as much as possible when sufficient capacity is available, and reroute traffic and reduce parallelism when one or more links are unavailable. In some embodiments, the functionality provided by the fabric manager is supported by the encapsulation and additional information provided by various resource adapters. In an example embodiment, each frame type includes one or more interconnect fabrics with corresponding fabric managers.
[0057] The architecture described herein can provide hardware failover for PCIe cards and NVMe drives through the interconnect fabric. Moreover, the architecture can provide quality of service (QoS) control for fabric bandwidth and throughput. The example architecture defines QoS as a "unit of consumption" of shared resources, shared concurrently (simultaneously) or sequentially (one at a time) as part of a logical server. For the interconnect fabric, the sharing is concurrent, and the unit of consumption is based on the use of end-to-end traffic control credits for the fabric. In one embodiment, the use of end-to-end traffic control credits limits each RA from injecting new transactions or network packets into the interconnect fabric until the available credits have been accumulated (acknowledgment packets) for any outstanding transactions, network packets previously inserted into the network.
[0058] In one embodiment, for a T-RA, the QoS "consumption unit" further includes a target IO resource device, or the order of mapping of the target I / O device as any of all (e.g., PCIe devices (all PCIe device PFs physical functions)), PCIe adapter devices, PCIe GPU devices, or PCIe NVMe devices. In addition to the sequential mapping of the target device, the T-RA is also capable of concurrently mapping the target I / O resources, i.e., the structural binding in some parts of the IO resources / devices (e.g., one PCIe PF physical function in a PCIe multifunction device or a PCIe PF virtual function in a PCIe device supporting many VFs within its (multiple) PFs).
[0059] In some embodiments, for all T-RAs, when the same target resource is also bound to the same C-RA logic server, the target I / O is either bound to the C-RA that created the logic server for peer transactions or is bound to another T-RA. In one embodiment, a logic server hosting a hypervisor software stack can simply map the processor I / O memory management unit (IOMMU) to the C-RA to group the T-RA bindings of the C-RA such that within the same logic server, they are hardware isolated between VM instances. Additionally, each RA QoS can be allocated based on the logic server type, category, and support for IO target resource bindings for peer transactions, including transactions (network packets) flowing through the S-RA stack nodes. Moreover, the architecture can provide a provable infrastructure to enhance security.
[0060] In various embodiments, the RA provides a physical interface between resources (such as computing, memory modules, NVMe devices) and a PCIe-based interconnect fabric. In addition to providing encapsulation (and possibly other data formatting) services, the RA can also include a state machine that supports flow control functionality and failover functionality by further encapsulating PCI-based traffic. Similarly, a compute resource adapter (C-RA) provides RA functionality for compute resources such as, for example, processor cores.
[0061] In some embodiments, multiple parallel connection fabrics can be provided, which can be used to provide a parallel interconnect fabric when fully functional and provide failover capabilities when some parts of the interconnect fabric are non-functional. That is, when multiple parallel paths with full functionality can be provided between compute modules and resources, but when a part of one of the fabrics in the fabric fails, the overall interconnect configuration can be dynamically rearranged to compensate for the failure without using a primary backup type of functionality.
[0062] Figure 3is a logical representation of an embodiment of an architecture that can utilize a parallel interconnect structure to provide a flexible configurable system. Figure 3 Examples of Figure 3 include a small number of components; however, more complex systems with any number of components can be supported.
[0063] In Figure 3 an example of
[0064] node 312 includes input / output (I / O) target resource devices (such as 320, 321) and computing resource nodes (such as 323, 324, 325) coupled to switch 322. The I / O devices can be, for example, storage devices, memory systems, user interfaces, and the computing nodes can be one or more processors (or processor cores) providing processing functionality. Switch 322 is used to interconnect I / O devices 320 and 321 and computing devices 323, 324, and 325. In some embodiments, these components can be co-located with node 312; however, in other embodiments, these components can be geographically distributed. Figure 3 Node 314 is similarly configured with switch 335 that interconnects I / O devices (such as 330, 331, 332) and computing devices (such as 333, 334). Node 318 is an I / O node having a plurality of I / O devices 360, 361, 362, 363, 364 interconnected by switch 365. Node 316 is a control node including controller 352 and storage medium 353, which can be used to control
[0065] the architecture of
[0066] Figure 3The illustrated example modular and composable architecture enables modularity because I / O resources need not be constrained within any specific physical system. Further, for example, consolidation of storage, cluster services, and / or I / O devices can be provided. The example modular and composable architecture also provides high availability by allowing application transparent alternate paths, independence partitioning, fault isolation, and a single management footprint.
[0067] Figure 4 is a conceptual diagram of an embodiment of an infrastructure resource pool that provides composable resources from which resource instances or complete devices can be bound together and configured to create a logical server that operates as a traditional server. Figure 4 Provides an example use case for mapping physical and virtual functions provided by a compute module and a target module to implement a logical server architecture. The figure shows a physical infrastructure and a logical infrastructure with a parallel structure, where the parallel structure is invisible and only the mapped resources are visible to the server.
[0068] The infrastructure resource pool 400 includes the various physical infrastructure modules discussed above in Figure 2 In the example embodiment of Figure 4 , this includes a compute module 410 having one or more processing cores 412 and corresponding C-RAs 414, respectively. For example, the processing core 412 can be connected to the C-RA 414 using the CXL or PCIe protocol. The processing core 412 can be any type of processing resource, such as a central processing unit (CPU), a graphics processing unit (GPU), etc.
[0069] The infrastructure resource pool 400 also includes a target module 430 having one or more target / IO resources 434 that provide physical functions (PF) and corresponding T-RAs 432, respectively. For example, the resource 434 can be connected to the T-RA 414 using the CXL or PCIe protocol. The processing core 412 can be any type of resource that can be used to support the processing core 412, such as a memory module, a smart I / O module, etc.
[0070] Interconnect structures 421 and 425 having associated structure managers 420 and 426, respectively, can be used to provide a configurable interconnect between the C-RA 414 and the T-RA 432, which provides an interconnect between the compute module 410 and the target module 430. In some embodiments, the interconnect structure is a PCIe-based interconnect structure. In alternative embodiments, the interconnect structure is a CXL- or Gen-Z-based interconnect.
[0071] In various embodiments, using the interconnect structures 421 and 425 with structure managers 420 and 426, and the modules from the infrastructure resource pool 400, various computing architectures can be composed.Figure 4 The example of Figure 4 illustrates a set of logical servers (e.g., 450, 452, 454). Although the
[0072] illustrated configuration is of the same logical device type (i.e., logical server); the described functionality is not limited to multiple identical logical device types. Thus, in alternative embodiments, many different concurrent logical device types can be configured. Figure 4 In the example configuration of
[0073] each logical server includes at least one processor resource 460, which is some or all of the computing modules 410 (i.e., processors 412 and corresponding C-RAs 414) from the infrastructure resource pool 400. Some or all of the interconnect fabrics 421 and 425 can be configured (via fabric managers 420 and 426) to provide a logical PCIe switch 465. In alternative embodiments, e.g., based on CXL or Gen-Z, the logical switch will be based on the appropriate protocol.
[0074] Figure 5 is a block diagram of an embodiment of a trusted composable computing architecture for trusted composable infrastructure resources, which have intelligent modules (e.g., servers, storage arrays), simple modules (e.g., CPUs, PCIe adapters, NVMe drives), and parallel interconnect fabrics from which traditional server modules (intelligent modules) can expand their IO capabilities or create logical servers. Figure 5 The architecture of Figure 5 can be a fully provable infrastructure and can provide complete encryption and key management isolation services.
[0075] Figure 5The example of [description] illustrates how compatibility and security can be added to the system at each of the management endpoints including, for example, a composable resource manager, one or more fabric managers, one or more resource adapters, and one or more baseboard management controllers. This compatibility and security is illustrated using the "bowtie" symbol at each endpoint. The "bowtie" indicates that the endpoint provides authentication, attestation, and encryption key management services.
[0076] Figure 5 The illustrated architecture utilizes the expansion fabrics 505 and 508 that operate as described above. The expansion fabrics 505 and 508 can be PCIe-based or alternatively CXL- or Gen-Z-based fabrics. The fabric managers 512 and 518 can be used to configure the connections within the expansion fabrics 505 and 508 respectively to provide the composable functionality described herein. The resource managers 510 and 516 can provide resource configuration management functionality for one or more compute resources and / or one or more target resources. In alternative embodiments, any number of expansion fabrics, fabric managers, and resource managers can be supported. In some embodiments, the fabric managers 512 and 518 can provide identification, enumeration, and discovery, and resource pool authentication services. Figure 5 The configuration of [description] can provide high availability management via the resource manager.
[0077] Figure 5 The illustrated architecture can utilize compute resources (e.g., 540) including one or more processor resources 542 and one or more C-RAs 546. In some embodiments, the one or more C-RAs 546 can provide identification and / or consumption unit authentication services. In some embodiments, one or more intelligent compute resources 520 can be included. The intelligent compute resources include at least one app 522 that provides a degree of intelligence to the intelligent compute resources 520 compared to the compute resources 540. In addition to the app 522, the intelligent compute resources 520 can also include one or more processor resources 524 and one or more corresponding C-RAs 526. The intelligent compute resources 520 can also include an authentication agent 528 to provide identification and / or consumption unit authentication services for the intelligent compute resources 520.
[0078] Figure 5Various embodiments of the illustrated composable computing architecture can include a simple expansion resource 580 having expansion resources 582 and T-RAs 584. The expansion resources can include, for example, SmartIO modules, NVMe modules, etc. In some embodiments, the (multiple) C-RAs 584 can provide identity and / or consumption unit authentication services. In some embodiments, one or more intelligent expansion resources 560 can be included. The intelligent expansion resource includes at least one processing element 568, which provides a certain degree of intelligence to the intelligent expansion resource 560 compared to the simple expansion resource 580. In addition to the processing element 568, the intelligent expansion resource 560 can include one or more expansion resources 564 and one or more corresponding T-RAs 562. The intelligent expansion resource 560 can also include an authentication agent 566 to provide identity and / or consumption unit authentication services for the intelligent expansion resource 560.
[0079] Generally, Figure 5 the architecture diagram illustrates how security and attestation functionality layers can be provided to composable architecture modules. This additional functionality can support enumeration and discovery functionality, including, for example, traversing the interconnect (e.g., link by link in a PCIe embodiment). Further, resource binding (e.g., RA fabric binding (host to target, target to target)) and I / O target binding (e.g., logical memory ranges, logical NVMe devices, dual-path NVMe devices to a single host, physical PCIe adapters) can be supported.
[0080] By Figure 5 the architecture provides authentication functionality that can utilize mutual authentication and a fully encrypted structured data path that provides cryptographic isolation between the bindings / logical servers resulting in secure management and control. Further, hardware failover for triggering events or software control (e.g., scheduling events) can be supported. Controller failure recovery can be software-controlled using path verification and error condition clearing.
[0081] Figure 6 is a block diagram of an embodiment of an infrastructure resource pool that provides composable infrastructure resources in a hardened security environment, where logical server resource bindings are cryptographically connected via a parallel interconnect fabric. In Figure 6 the example, the infrastructure resource 600 can include any number of compute modules 610, interconnect modules 640, and target modules 660.
[0082] Similar to the configurations discussed above, the (multiple) compute modules 610 can include one or more processors 612 and corresponding C-RAs 614. The interconnect module 640 can include one, two, or more interconnect fabrics 642 and corresponding fabric managers 645. Figure 6The example target module(s) 660 can include one or more T-RAs 662 and corresponding switches 664 to provide access to various PCI slots 666. Other target module configurations providing different functionality can also be supported.
[0083] In Figure 6 the illustrated architecture, the C-RA 614 and T-RA 662 are used to provide PCIe host packet compatibility and encapsulation / decapsulation functionality. In various embodiments, the C-RA 614 is used to receive logical server PCIe packets 620 into encapsulated packets 650. The encapsulated packets 650 are sent via the interconnect module 640 to the T-RA 662, which also provides compatibility and decapsulation functionality. Traffic can also flow in the opposite direction using the same procedures and protocols.
[0084] In one embodiment, the encapsulated packet 650 includes a tail 652, a payload 654 (e.g., a PCIe packet 620), and a message header 656. In one embodiment, the message header 656 includes a vendor ID, a requester ID (e.g., a compute RAID, a target RA ID), a target ID, traffic control credits, and transaction ordering requirements. In alternative embodiments, different message header configurations can be used.
[0085] Thus, Figures 1 to 6 the combination of the various elements and corresponding descriptions illustrated can have various levels of security, encryption, cryptographic isolation with logical server binding, attestation, etc.
[0086] Figure 7 is an example of a parallel interconnect structure packet frame header that can be used with the composable architecture described herein for a PCIe parallel interconnect structure built using standard PCIe switches. Figure 7 An example is an example of a PCIe transaction layer packet (TLP) header that utilizes vendor-reserved bytes to support the composable structures and composable computing architectures described herein; however, other protocols can use other header formats.
[0087] In Figure 7 the example format, bytes 0 to 11 of the TLP header (815) can be as defined in the PCIe standard. In various embodiments, bytes 12 to 15 of the header (825) can be used to support the composable architecture described herein for this purpose, as bytes 12 to 15 are currently reserved by the standard for vendor use. Thus, as the standard evolves, the exact configuration of the header may change, but the functionality can be maintained by leveraging the portions reserved for vendor use.
[0088] In various embodiments, the information in bytes 12 through 15 of the TLP header can be utilized by one or more RAs within a composable architecture. In one embodiment, bytes 12 and 13 of the header are used for the destination identifier (840) and the frame identifier (845). Similarly, bytes 14 and 15 of the header are used for the source identifier (850) and the frame identifier (855). Other header structures can be used to provide similar functionality.
[0089] In various embodiments, transactions across the interconnect fabric address domains are updated by allowing S-RA frames with a new S-RA ID (header bytes 6+7... requester ID) and a destination ID (header bytes 10+11... bus and device number) to flow independently between different interconnect fabric address domains. When the bus number and device number (bytes 10+11) match the destination ID (bytes 14+15), the packet has reached its final destination, i.e., the C-RA or T-RA attached to the destination interconnect fabric address domain.
[0090] In various embodiments, each fabric can have a corresponding subnet, and frame IDs (e.g., 745, 755) can be used to identify frames and thus identify the corresponding address domains. In some embodiments, frame IDs can be used to route packets using device IDs that can be reused. In some embodiments, routing tables and / or address translation tables can be used for traffic between frames.
[0091] Figure 8 is a flowchart of an embodiment of a technique for configuring computing infrastructure resources in an environment that provides composable resources. Figure 8 An example process describes how a composable resource manager (e.g., Figure 2 215 in Figure 2 can control one or more fabric managers (e.g., 290, 292, 294, 295, 297, 298 in
[0092] to use a composable fabric to provide computing customized to API requirements, in one embodiment.
[0093] In one embodiment, a system manager (or other component, such as one or more fabric managers) can evaluate the environment and determine whether computing resources that meet the requested requirements are available at block 810. If computing resources are available, then at block 810, the system manager (or other component(s)) can allocate computing resources from a resource pool at block 870. When the computing resources are allocated, at block 870, a response can be provided via a system API at block 880. In one embodiment, the response indicates that the requested computing resources have been successfully allocated. In an alternative embodiment, additional information can also be provided.
[0094] If computing resources are not available, then at block 810, the system manager (or other component(s)) can determine whether the requested requirements can be met by configuring available resources from the resource pool at block 820. If not, a failure response can be provided via an API at block 890. If the requested requirements can be met using the resource pool, a fabric API and / or a fabric manager can be used to connect the desired resources (such as processors, memories, storage devices, networks) using RAs and other elements discussed herein at block 830.
[0095] If the fabric manager(s) successfully connect the desired resources, then at block 840, a response can be provided via the fabric API indicating that the computing resources have been allocated at block 850. If the fabric manager(s) are not successful in connecting the desired resources, then at block 840, a failure response can be provided via the fabric API indicating that the computing resources have not been allocated at block 855.
[0096] When a response is received via the fabric API, the system manager can determine at block 860 whether the resource connection was successful. If not, a failure response can be provided via an API at block 890. If the resource connection was successful, then at block 860, the response can be provided via the system API at block 880. In one embodiment, the response indicates that the requested computing resources have been successfully allocated. In an alternative embodiment, additional information can also be provided.
[0097] Figure 9 is a flowchart of one embodiment of a technique for configuring computing resources in an environment that provides composable resources. Figure 9 The flow of Figure 8 provides an alternative example of the flow of Figure 9 In the example of Figure 5 one or more resource managers (such as 510 and 516 in Figure 5512 and 518) in it to provide customized computing for API requirements using a composable structure. In this example, the (multiple) resource managers communicate with other fabric switches and C-RA and T-RA to provide the desired configuration.
[0098] In one embodiment, the environment provides (or includes) a resource API that supports the creation of new computing resources, which can include, for example, (multiple) processors, memory, storage, network, and / or accelerator resources. In alternative embodiments, additional and / or different resources can be supported. In one embodiment, one or more computing specifications are received via the resource API in block 900.
[0099] In one embodiment, a resource manager (or other component, such as one or more fabric managers) can evaluate the environment and determine whether computing resources that meet the requested requirements are available in block 910. If computing resources are available, then in block 910, the resource manager (or the (multiple) other components) can allocate computing resources from a resource pool in block 950. When the computing resources are allocated, in block 950, a response can be provided via the resource API in block 960. In one embodiment, the response indicates that the requested computing resources have been successfully allocated. In alternative embodiments, additional information can also be provided.
[0100] If computing resources are not available, then in block 910, the resource manager (or the (multiple) other components) can determine whether the requested requirements can be met by configuring available resources from the resource pool in block 920. If not, a failure response can be provided via the resource API in block 970. If the requested requirements can be met using the resource pool, the resource manager can be used to connect the desired resources (such as (multiple) processors, memory, storage, network) using RA and other elements discussed herein in block 930.
[0101] Figure 10 is a conceptual diagram of one embodiment of an infrastructure resource pool that provides composable infrastructure resources configured as a set of logical servers. Figure 10 An example use case is provided for mapping the physical and virtual functions provided by a computing module and a target module to implement a logical server architecture. Compared with Figure 3 the example of Figure 10 the example provides a more complex and flexible example.
[0102] The infrastructure resource pool 1000 includes various physical infrastructure modules, such as computing resource modules 1010 and 1020 and target resource modules 1040, 1050, and 1060. Any number of computing modules and target modules can be supported. In Figure 10In an example embodiment, this includes a computing resource module 1010 having at least one processing core 1011 and a corresponding C-RA 1012 (labeled RA-1), and a computing resource module 1020 having processing cores 1022 and 1024 and corresponding C-RAs 1026 and 1028 (labeled RA-4 and RA-5, respectively). For example, the processing cores can be connected to the C-RAs using the CXL or PCIe protocol.
[0103] The infrastructure resource pool 1000 also includes target resource modules 1040, 1050, and 1060, each having one or more resources 1045, 1055, and 1065, respectively. For example, the resources can be connected to the T-RAs using the CXL or PCIe protocol. In Figure 10 the example, the target resource module 1040 includes a single T-RA 1041 (labeled RA-2), a PCIe switch 1042, and a plurality of PCIe slots 1045 coupled to the switch 1042. Similarly, the target module 1060 includes a single T-RA 1061 (labeled RA-7) and a multi-functional PCIe device 1064 shown as a smart resource 1065. The target module 1050 includes two T-RAs 1051 and 1052 (labeled RA-3 and RA-6, respectively), which can be used to access a dual-port NVMe device 1055 through switches 1053 and 1054.
[0104] Interconnect structures 1031 and 1035 having associated structure managers 1030 and 1036, respectively, can be used to provide configurable interconnections between various RAs, which provide interconnections between the computing resource modules 1010 and 1020 and the target resource modules 1040, 1050, and 1060. In some embodiments, the interconnect structure is a PCIe-based interconnect structure. In an alternative embodiment, the interconnect structure is a CXL- or Gen-Z-based interconnect.
[0105] In various embodiments, using the interconnect structures 1031 and 1035 having structure managers 1030 and 1036 and the modules from the infrastructure resource pool 1000, various computing architectures can be composed. Figure 10 The example illustrates logical server pairs (e.g., 1070, 1080). Other configurations can also be supported.
[0106] In Figure 10 the illustrated example configuration, the logical server 1070 can include (a plurality of) processing cores 1011, and the logical server 1080 can include processing cores 1022 and 1024. The C-RA 1012 (RA-1) is configured to provide the functionality of a logical PCIe switch 1072.
[0107] The portion of T-RA 1041 (RA-2) can be configured to provide the functionality of logical PCIe switch 1073, which can be used as an interface to some or all of the PCIe slots 1045 (labeled as PCIe devices 1076). Portions of the interconnect structure 1031 and portions of the interconnect structure 1035 and T-RA 1051 (RA-3) can be configured to provide the functionality of logical PCIe switch 1074, which can be used as an interface to some or all of the NVMe devices 1055 (labeled as NVMe devices 1077). Portions of the interconnect structure 1031 and portions of the interconnect structure 1035 and T-RA 1061 (RA-7) can be configured to provide the functionality of logical PCIe switch 1075, which can be used as an interface to the multi-function PCIe device 1064 (labeled as multi-function device 1078).
[0108] In Figure 10 In the illustrated example configuration, the logical server 1080 can include processing cores 1022 and 1024. Portions of C-RA 1026 (RA-4), C-RA 1028 (RA-5), portions of the interconnect structure 1031, and portions of the interconnect structure 1035 can be configured to provide the functionality of logical PCIe switch 1083 and logical PCIe switch 1084.
[0109] Portions of the interconnect structure 1031 and portions of the interconnect structure 1035 and T-RA 1041 (RA-2) can be configured to provide the functionality of logical PCIe switch 1085, which can be used as an interface to some or all of the PCIe slots 1045 (labeled as PCIe devices 1090). Portions of the interconnect structure 1031 and portions of the interconnect structure 1035 and T-RA 1061 (RA-7) can be configured to provide the functionality of logical PCIe switch 1086, which can be used as an interface to some or all of the multi-function devices 1064 (labeled as multi-function devices 1091).
[0110] Portions of the interconnect structure 1031 and portions of the interconnect structure 1035 and T-RA 1041 (RA-2) can be configured to provide the functionality of logical PCIe switch 1087, which can be used as an interface to the PCIe slot 1045 (labeled as PCIe device 1092). Portions of the interconnect structure 1031 and portions of the interconnect structure 1035, T-RA 1051 (RA-3), and T-RA 1052 (RA-6) can be configured to provide the functionality of logical PCIe switch 1089, which can be used as an interface to some or all of the NVMe devices 1055 (labeled as NVMe devices 1093 and 1094).
[0111] Figure 11It is a conceptual diagram of a logic server configured to utilize peer - to - peer communication on a composable architecture parallel interconnection structure. Figure 11 An example use case is provided for mapping physical and virtual functions provided by computing modules and target modules to implement a logic server architecture with internal peer - to - peer connections. Compared with the Figure 10 example, Figure 11 the example provides a more complex and flexible example.
[0112] The infrastructure resource pool 1100 includes various physical infrastructure modules, such as computing resource modules 1110 and 1120 and target resource modules 1140, 1150, 1160, and 1165. Any number of computing modules and target modules can be supported. In the Figure 11 example embodiment, this includes computing module 1110 having one processing core 1111 and corresponding C - RA 1112 (labeled RA - 1) and computing module 1120 having processing cores 1122 and 1124 and corresponding C - RAs 1126 and 1128 (labeled RA - 4 and RA - 5 respectively). For example, the processing cores can be connected to the C - RA using the CXL or PCIe protocol.
[0113] The infrastructure resource pool 1100 also includes target modules 1140, 1150, 1160, and 1165, each having one or more resources 1145, 1155, 1164, and 1168 respectively. For example, the resources can be connected to the T - RA using the CXL or PCIe protocol. In the Figure 11 example, target module 1140 includes a single T - RA 1141 (labeled RA - 2), a PCIe switch 1142, and a plurality of PCIe slots 1145 coupled to the switch 1142. Target module 1150 includes two T - RAs 1151 and 1152 (labeled RA - 3 and RA - 6 respectively), which can be used to access a dual - port NVMe device 1155 through switches 1153 and 1154.
[0114] Similarly, target module 1160 includes a single T - RA 1161 (labeled RA - 7), a multi - function PCIe device 1162 shown as intelligent resource 1164. Target module 1165 includes T - RA 1166 (labeled RA - 8), a CXL memory controller 1167, and one or more memory modules 1168.
[0115] Structure managers 1130 and 1136 can be used to provide configurable interconnections between various RAs, which provide interconnections between computing modules 1110 and 1120 and target modules 1140, 1150, 1160, and 1165. In some embodiments, the interconnection structure is a PCIe-based interconnection structure. In alternative embodiments, the interconnection structure is a CXL- or Gen-Z-based interconnection.
[0116] In various embodiments, using the interconnection structures 1131 and 1135 with structure managers 1130 and 1136 and modules from the infrastructure resource pool 1100, various computing architectures can be composed. Figure 11 The example of [the figure] illustrates a logical server 1170 with internal peer functionality 1169. Other configurations can also be supported.
[0117] In Figure 11 In the illustrated example configuration, the logical server 1170 can include processing cores 1122 and 1124. Portions of C-RA 1126 (RA-4), C-RA 1128 (RA-5), the interconnection structure 1131, and the interconnection structure 1135 can be configured to provide the functionality of logical PCIe switches 1183 and 1184.
[0118] Portions of the interconnection structure 1131 and portions of the interconnection structure 1135 and T-RA 1141 (RA-2) can be configured to provide the functionality of a logical PCIe switch 1185, which can serve as an interface to some or all of the PCIe slots 1145 (labeled as PCIe devices 1190). Portions of the interconnection structure 1131 and portions of the interconnection structure 1135 and T-RA 1161 (RA-7) can be configured to provide the functionality of a logical PCIe switch 1186, which can serve as an interface to some or all of the multifunction devices 1164 (labeled as multifunction devices 1191).
[0119] Portions of the interconnection structure 1131 and portions of the interconnection structure 1135 and T-RA 1141 (RA-2) can be configured to provide the functionality of a logical PCIe switch 1187, which can serve as an interface to the PCIe device 1145 (labeled as PCIe device 1192). Portions of the interconnection structure 1131 and portions of the interconnection structure 1135, T-RA 1151 (RA-3), and T-RA 1152 (RA-6) can be configured to provide the functionality of a logical PCIe switch 1189, which can serve as an interface to some or all of the NVMe devices 1155 (labeled as NVMe devices 1193 and 1194).
[0120] Portions of the interconnect structure 1131 and portions of the interconnect structure 1135 and the T-RA 1141 (RA-2) can be configured to provide the functionality of a logical PCIe switch 1185, which can serve as an interface to some or all of the PCIe slots 1145 (labeled as PCIe devices 1190). Portions of the interconnect structure 1131 and portions of the interconnect structure 1135 and the T-RA 1161 (RA-7) can be configured to provide the functionality of a logical PCIe switch 1186, which can serve as an interface to some or all of the multi-function devices 1164 (labeled as multi-function devices 1191).
[0121] Portions of the interconnect structure 1131 and portions of the interconnect structure 1135 and the T-RA 1141 (RA-2) can be configured to provide the functionality of a logical PCIe switch 1187, which can serve as an interface to the PCIe device 1145 (labeled as PCIe device 1192). Portions of the interconnect structure 1131 and portions of the interconnect structure 1135, the T-RA 1151 (RA-3) and the T-RA 1152 (RA-6) can be configured to provide the functionality of a logical PCIe switch 1189, which can serve as an interface to some or all of the NVMe devices 1155 (labeled as NVMe devices 1193 and 1194).
[0122] In some embodiments, in addition to the configurable processor-to-resource connections, peer-to-peer connections within the logical server 1170 can also be supported. For example, the resource manager 1136 can be used to provide a path through the interconnect structure 1135 between the multi-function device 1119 and the memory device 1198 via the logical PCIe switches 1186 and 1183 and the memory controller 1167. This functionality can support the processor 1122. A similar configuration (not explicitly illustrated in Figure 11 can use the PCIe logical switches 1186 and 1184 and the memory controller 1167 to provide (to the memory device 1199) for the processor 1124.
[0123] References to "one embodiment" or "an embodiment" in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the invention. The appearances of the phrase "in one embodiment" in various places in this specification are not necessarily all referring to the same embodiment.
[0124] Although the invention has been described in terms of multiple embodiments, those skilled in the art will recognize that the invention is not limited to the described embodiments, but can be modified and changed within the spirit and scope of the appended claims to be practiced. Therefore, this description is to be considered illustrative and not restrictive.
Claims
1. A composable computing platform, comprising: A first parallel interconnection structure, including a first plurality of structural branches that provide parallel interconnection via the first parallel interconnection structure; And A second parallel interconnection structure, including a second plurality of structural branches that provide parallel interconnection via the second parallel interconnection structure; Wherein: A first structural branch among the first plurality of structural branches and a first structural branch among the second plurality of structural branches are communicatively connected to each other and include a first switching structure, A second structural branch among the first plurality of structural branches and a second structural branch among the second plurality of structural branches are communicatively connected to each other and include a second switching structure, and The first switching structure and the second switching structure are isolated from each other and provide independent paths for routing packets.
2. The composable computing platform according to claim 1, further comprising: Computing resources, the computing resources including a computing resource adapter C-RA, and the C-RA communicatively connects the computing resources to the first plurality of structural branches in the first parallel interconnection structure.
3. The composable computing platform according to claim 1, further comprising: Input / output I / O resources, the I / O resources including a target resource adapter T-RA, and the T-RA communicatively connects the I / O resources to the first plurality of structural branches in the first parallel interconnection structure.
4. The composable computing platform according to claim 1, wherein: The first parallel interconnection structure further includes a first plurality of stacked resource adapters S-RA; The first structural branch among the first plurality of structural branches is communicatively connected to a first S-RA among the first plurality of S-RA; And The first S-RA among the first plurality of S-RA is communicatively connected to a first S-RA in the second parallel interconnection structure.
5. The composable computing platform according to claim 4, wherein: The second parallel interconnection structure further includes a second plurality of S-RA, the second plurality of S-RA including the first S-RA in the second parallel interconnection structure; and The first structural branch among the second plurality of structural branches is communicatively connected to the first S-RA among the second plurality of S-RA, such that the first structural branch among the second plurality of structural branches is communicatively connected to the first structural branch among the first plurality of structural branches.
6. The composable computing platform according to claim 1, wherein: A first plurality of structure managers are coupled to the first switching structure; And A second plurality of structure managers are coupled to the second switching structure.
7. The composable computing platform according to claim 6, further comprising: A composable resource manager that discovers the first parallel interconnection structure and the second parallel interconnection structure through management connections with the first plurality of structure managers and the second plurality of structure managers.
8. The composable computing platform according to claim 7, further comprising a user interface, the user interface being an API / GUI / script of the composable resource manager.
9. The combinable computing platform according to claim 7, wherein the structure manager is configured to: Receive I / O or memory resource requirements for a computing workload for a host device; and map I / O resources to individual computing resources based on the received requirements; wherein the first parallel interconnect structure and the second parallel interconnect structure provide a communication connection between the I / O resources and the computing resources.
10. A combinable computing platform, comprising: A first parallel interconnect structure, including a first plurality of structure branches providing parallel interconnect via the first parallel interconnect structure; wherein: The first structure branch of the first plurality of structure branches includes a first switching structure, and The second structure branch of the first plurality of structure branches includes a second switching structure, and The first switching structure and the second switching structure are isolated from each other and provide independent paths for routing packets; and A second parallel interconnect structure, including a second plurality of structure branches providing parallel interconnect via the second parallel interconnect structure; wherein: The first structure branch of the second plurality of structure branches includes the first switching structure, and The second structure branch of the second plurality of structure branches includes the second switching structure.
11. The combinable computing platform according to claim 10, wherein: The first structure branch of the first plurality of structure branches is communicatively connected to the first structure branch of the second plurality of structure branches; and The second structure branch of the first plurality of structure branches is communicatively connected to the second structure branch of the second plurality of structure branches.
12. The combinable computing platform according to claim 11, wherein: The first parallel interconnect structure further includes a first plurality of S-RAs; The first structure branch of the first plurality of structure branches is communicatively connected to the first S-RA of the first plurality of S-RAs; and The first S-RA of the first plurality of S-RAs is communicatively connected to the first S-RA in the second parallel interconnect structure.
13. The combinable computing platform according to claim 12, wherein: The second parallel interconnect structure further includes a second plurality of S-RAs, the second plurality of S-RAs including the first S-RA in the second parallel interconnect structure; and The first structure branch of the second plurality of structure branches is communicatively connected to the first S-RA of the second plurality of S-RAs, such that the first structure branch of the second plurality of structure branches is communicatively connected to the first structure branch of the first plurality of structure branches.
14. The combinable computing platform according to claim 10, wherein: A first plurality of structure managers are coupled to the first switching structure; A second plurality of structure managers are coupled to the second switching structure.
15. The combinable computing platform according to claim 14, further comprising: A combinable resource manager that discovers the first parallel interconnect structure and the second parallel interconnect structure through management connections with the first plurality of structure managers and the second plurality of structure managers.
16. The combinable computing platform according to claim 15, wherein the structure manager is configured to: Receive I / O or memory resource requirements for a computing workload for a host device; and Map I / O resources to individual computing resources based on the received requirements; wherein the first parallel interconnect structure provides a communication connection between the I / O resources and the computing resources.
17. The composable computing platform according to claim 10, wherein the first structural branch and the second structural branch in the first plurality of structural branches are communicatively disconnected.
18. The composable computing platform according to claim 10, wherein the first parallel interconnect structure and the second parallel interconnect structure are built using a switching mechanism.
19. A composable computing platform, comprising: A composable resource manager that discovers a parallel interconnect structure through an administrative connection; and The parallel interconnect structure includes a plurality of structural branches that provide parallel interconnect via the parallel interconnect structure; wherein: The first structural branch of the plurality of structural branches includes a first switching structure, and The second structural branch of the plurality of structural branches includes a second switching structure, and The first switching structure is isolated from the second switching structure and provides independent paths for routing packets; and A second parallel interconnect structure, including a second plurality of structural branches that provide parallel interconnect via the second parallel interconnect structure; wherein: The first structural branch of the second plurality of structural branches includes the first switching structure, and The second structural branch of the second plurality of structural branches includes the second switching structure.
Citation Information
Patent Citations
Specifying disaggregated compute system
CN108885554A
Resource management for peripheral component interconnect-express domains
CN109032974A