Apparatus and method for trusted access to extended memory through multiple potential heterogeneous compute nodes

By dynamically managing memory domains through CXL interconnection and switches, the flexibility and performance issues of memory consistency management in multi-processor systems are resolved, fine-grained memory access control and efficient consistency management in multi-tenant environments are achieved, and system performance and isolation capabilities are improved.

CN120654277APending Publication Date: 2025-09-16INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510304596.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-15
Filing Date
2025-03-14
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve flexible and efficient memory consistency management in multi-processor systems, resulting in increased communication volume and performance degradation. In particular, it is difficult to achieve fine-grained memory domain isolation and consistency control in cloud computing and multi-tenant environments.

Method used

Dynamically manage memory domains through CXL interconnects and switches, allowing applications to dynamically configure and update consistency states, leveraging CXL.cache and CXL.mem semantics for fine-grained memory access control, optimizing communication traffic with telemetry and QoS mechanisms, and supporting independent consistency domain management in multi-tenant environments.

Benefits of technology

It realizes flexible memory consistency management in multi-processor systems, reduces communication volume, improves system performance, supports memory domain isolation and consistency control in multi-tenant environments, and reduces system complexity and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654277A_ABST
    Figure CN120654277A_ABST
Patent Text Reader

Abstract

Apparatuses and methods for trusted access to extended memory through a plurality of potential heterogeneous compute nodes are disclosed. A plurality of roots of trust are described under a super root of trust (SROT). One embodiment comprises: a core of a host processor; a host processor memory subsystem to provide access to a host processor memory; a home agent to provide access by the core to the host processor memory subsystem and the extension memory subsystem; and a source address decoder to decode the memory requests generated from the plurality of cores to determine whether the memory requests are to be directed to the host processor memory subsystem or the extended memory subsystem. The host-based security circuitry encrypts and decrypts a memory request directed to the extended memory subsystem, the host-based security circuitry performing encryption and decryption based on a second key stored in a cache maintained by the host-based security circuitry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to computer systems and, more particularly, to apparatus and methods for providing trusted access to extended memory across multiple, potentially heterogeneous computing nodes. Background Art

[0002] A processor or a collection of processors executes instructions from an instruction set (e.g., an instruction set architecture (ISA)). An instruction set is the programming-related part of a computer's architecture and generally includes native data types, instructions, register architecture, addressing modes, memory architecture, interrupt and exception handling, and external input and output (I / O). BRIEF DESCRIPTION OF THE DRAWINGS

[0003] Figure 1 is a block diagram of a portion of a data center architecture according to an embodiment.

[0004] Figure 2 is a block diagram of a switch according to an embodiment.

[0005] Figure 3 is a flowchart of a method according to an embodiment.

[0006] Figure 4 is a flow chart of a method according to another embodiment.

[0007] Figure 5 The diagram illustrates an example implementation of a computing system including a host processor and an accelerator coupled by a link.

[0008] Figure 6 The diagram illustrates an example implementation of a computing system that includes two or more interconnected processor devices.

[0009] Figure 7 The diagram illustrates a representation of example ports of a device comprising a layered stack.

[0010] Figures 8A-8C is a simplified block diagram illustrating an example implementation of an interface utilizing an example adapter.

[0011] Figure 9 The figure illustrates one embodiment of an expansion memory device coupled to a host processor.

[0012] Figure 10 The diagram illustrates an apparatus for determining a path for a memory request.

[0013] Figures 11A-11BOne embodiment of switching circuitry including a microcontroller and a lookup table for allocating and de-allocating logical expansion memory devices is shown.

[0014] Figure 12 The figure illustrates a method according to an embodiment of the invention.

[0015] Figure 13 The figure shows a specific implementation of an embodiment of the present invention.

[0016] Figures 14A-14B The diagram shows a specific implementation that is operable with both Type 2 and Type 3 CXL devices. DETAILED DESCRIPTION

[0017] In various embodiments, a system may have a memory space that can be dynamically configured to include multiple independent memory domains, each of which can be dynamically created and updated. In addition, each of these independent memory domains can be dynamically controlled to be consistent or inconsistent, and can be dynamically updated to switch consistency states. To this end, a switching circuit system within the system (such as a switch that couples multiple processors, devices, memories, etc.) can be configured to dynamically assign memory ranges to given memory domains. In addition, when such a memory domain is indicated as having a consistent state, the switch can maintain and implement a consistency mechanism. Thus, the switching circuit system can dynamically handle incoming memory requests differently depending on whether the request is directed to a consistent memory domain or a non-consistent memory domain. In addition, the switching circuit system can handle consistency operations differently depending on, for example, the traffic conditions in the system. For example, consistent memory domains can be allocated and can be associated with one or more fallback rules to provide different consistency mechanisms to be used when there is a high traffic condition.

[0018] Although the embodiments are not limited in this respect, example cloud-based edge architectures may communicate using interconnects and switches in accordance with the Compute Express Link (CXL) specification, such as the CXL 1.1 specification or any future version, modification, variation, or alternative of the CXL specification. Furthermore, while the example embodiments described herein incorporate CXL-based technologies, the embodiments may be used with other coherent interconnect technologies, such as the IBM XBus protocol, the Nvidia NVLink protocol, the AMD Infinity Fabric (IF) protocol, the Cache Coherent Interconnect for Accelerators (CCIX) protocol, or the Coherent Accelerator Processor Interface (OpenCAPI).

[0019] Many systems provide a single coherent memory domain so that all computing devices (e.g., multiple processor sockets) and additional devices (such as accelerators, etc.) are in the same coherent domain. Such a configuration can be beneficial for enabling shared computation and shared memory across processors. However, increasing the number of coherence agents also increases the amount of coherence traffic. For example, adding four processor sockets to a system to convert it from a 4-socket system to an 8-socket system increases the amount of coherence traffic by a factor of three, which can adversely affect latency, and a higher number of sockets can increase the amount of traffic even further. This is especially true when additional devices and accelerators are also considered, which can be part of this single coherent memory domain.

[0020] Thus, embodiments can dynamically and finely granularly control memory coherence. In embodiments, protocols based on shared coherence domains can communicate over the CXL interconnect in a flexible and scalable manner. As a result, multiple servers or racks can communicate in memory semantics using CXL.cache or CXL.mem semantics via a CXL switch. Using embodiments, applications can dynamically and independently achieve coherence using CXL.cache semantics.

[0021] When coherency is disabled for a memory device attached via a CXL link, the memory device can be locally non-coherent. As an example, an attached accelerator with attached memory or an attached memory expansion card can be: (1) configured in "device biased" mode and not coherent with any other entity and used exclusively by the device; or (2) configured in "host biased" mode and globally coherent with the rest of the platform.

[0022] In cloud server implementations such as multi-tenant data centers, the system may have multiple consistency domains, such as per-tenant consistency. As an example, each of a plurality (potentially a large number of different tenants) may be associated with a memory domain (or multiple memory domains). Note that these separate memory domains may be isolated from each other such that a first tenant assigned to a first memory domain cannot access a second memory domain assigned to a second tenant (and a second tenant assigned to a second memory domain cannot access the first memory domain assigned to the first tenant). In other cases, there may be a more flexible relationship between tenants and memory domains. In an embodiment, consistency domains are managed on a per-tenant basis.

[0023] An example implementation may be associated with a database server or database management system configured to run on a cloud-based architecture. In such a system, multiple nodes may be implemented, where at least some of the nodes have a segment called primary storage that does not require consistency because it is read-only. This primary storage may consume a large percentage (e.g., 50%) of the total memory capacity used by the database. While other parts of the database may require consistency for specific transactions, embodiments may provide fine-grained, flexible mechanisms within the application to define consistency requirements. Thus, the dynamic and flexible approach provided in embodiments differs from static, pre-defined, hard partitioning at the node level or memory area level.

[0024] To achieve this arrangement, embodiments provide mechanisms that expose to applications or other requestors the ability to dynamically configure and update the coherence state and other aspects of memory domains. For example, when allocating a memory region that does not require coherence (such as main storage), an application may specify a memory allocation request as follows: cxl-mmap([A,B],allocate,800GB,NULL <coherence>,NULL <call-back>)(cxl-mmap([A,B],allocate,800GB,empty <consistency>,empty <callback>)). With this example memory allocation request, the requester provides information about the memory range request type (allocation request), the amount of space requested, and indicators of the consistency state and callback information (neither of which is active in this particular request).

[0025] However, when allocating a memory region to be used for a transaction, an application can specify consistency and further define the entities that are allowed to access that consistent memory domain (e.g., in terms of process address space identifiers (PASIDs), such as PASID2, PASID3, and PASID5). This is illustrated with the following memory allocation request: cxl-mmap([C,D],allocate,100GB,PASID2,PASID3,PASID5,NULL <call-back>)(cxl-mmap([C,D],allocate,100GB,PASID2,PASID3,PASID5,empty <callback>)). Note that, in addition, memory domains can be associated with tenant IDs, which in turn can be mapped to one or more PASIDs to provide tenant-by-tenant consistency. Note that in some implementations, a "tenant" can be defined as one instance of all processes. Embodiments may enable definition of consistency domains as one of two options: (1) ID (tenant ID), which includes a set of PASIDs; and (2) PASID granularity (which can be identified by tenant ID and PASID).

[0026] Now, you can also turn off consistency after the transaction is complete by using the same memory allocation request, but using a modification indicator instead of an allocation indicator as follows: cxl-mmap([C,D],modify,100GB,NULL <coherence>,NULL <call-back>)(cxl-mmap([C,D],modify,100GB,NULL <consistency>,NULL <callback>)). The same mechanism can be used to open consistency later, for example, only update the consistency for PASID5 as follows: cxl-mmap([C,D],modify,100GB,PASID5,NULL <call-back>)(cxl-mmap([C,D],modify,100GB,PASID5,empty<callback>)).

[0027] As further shown above, these memory allocation and update requests can include an extension called "callback", which can be used to specify CXL-based callback rules. These rules can provide fallback operations for handling consistency if one or more links are saturated. This is similar to the fallback mechanism for locking, for example, in which if the lock is not acquired, another code path or option is taken. As an example, if a switch generates a callback signal indicating that the interconnect is saturated due to a coherence operation, the callback option can require the use of the software multi-phase commit protocol to achieve consistency: cxl-mmap([C,D],modify,100GB,PASID5,CALL-BACK CODEPATH*swcommitprotocol(C,D,PASID5))(cxl-mmap([C,D],modify,100GB,PASID5,CALL-BACK CODEPATH*swcommitprotocol(C,D,PASID5))).

[0028] Another option for callback could be quality of service, where if the interconnect is saturated, a given PASID (e.g., PASID 2) receives high priority / dedicated switch credits as follows (e.g., PASID 2 is performing the main consistency-required operations, while PASID 3 and PASID 5 are just collecting statistics or doing garbage collection): cxl-mmap([C,D],modify,100GB,PASID5,CALL-BACK QOS PASID 2)

[0029] Now refer to Figure 1 , which is a block diagram of a portion of a data center architecture according to an embodiment. Figure 1 As shown in , system 100 can be a collection of components implemented as one or more servers in a data center. As shown, system 100 includes a switch 110, such as a CXL switch according to an embodiment. In other implementations, switch 110 can be another type of coherent switch. However, note that in all cases, switch 110 is implemented as a coherent switch rather than an Ethernet-type switch. Through switch 110 acting as a fabric, various components including one or more central processing units (CPUs) 120 and 160, one or more special function units such as a graphics processing unit (GPU) 150, and a network interface circuit (NIC) 130 can communicate with each other. More specifically, each of these devices can be implemented as one or more integrated circuits that perform functions that communicate with other functions in other devices via one of multiple CXL communication protocols. For example, CPU 120 can communicate with NIC 130 via the CXL.io communication protocol. Furthermore, CPUs 120 and 160 can communicate with GPU 150 via the CXL.mem communication protocol. Furthermore, as an example, CPUs 120, 160 may communicate with each other and from CPU 160 to GPU 150 via the CXL.cache communication protocol. Switch 110 may include control circuitry that allows for dynamic allocation and updating (including coherence state) of different memory domains for devices and applications or services. For example, different processes may request coherence across certain memory ranges, while other processes may not require coherence at all.

[0030] like Figure 1 As further shown in FIG, the system memory can be formed by various memory devices. In the embodiment shown, the pooled memory 160 is coupled to the switch 110. Various components can access the pooled memory 160 via the switch 110. In addition, multiple portions of the system memory can be directly coupled to specific components. As shown, the memory device 170 0-3 Distributed such that each region is directly coupled to the corresponding CPU 120 , 160 , NIC 130 , and GPU 150 .

[0031] like Figure 1 As further shown in FIG, various coherent and non-coherent memory domains may be maintained within memory 170 in response to memory allocation requests issued by processes. Figure 1 The embodiments are shown at this high level, but many variations and alternatives are possible.

[0032] Via an interface according to an embodiment, software (e.g., a system stack) enables dynamic specification of these types of memory domains. In an embodiment, a memory domain consists of a set of memory regions having an address range, a PASID list associated with the memory domain, and a consistency type (e.g., consistency, non-consistency, read-only, etc.). Device-level (e.g., GPU and CPU) memory domains can also be defined. In other cases, a memory domain can be mapped to a single address range, where a tenant can have multiple memory domains.

[0033] Circuitry within the switch can implement the aforementioned coherence domains. To do so, the circuitry can be configured to intercept snoop and other CXL.cache flows and determine whether they need to traverse the switch. If not, it returns a corresponding CXL.cache response to inform the snoop requester that the address is not hosted on the target platform or device for the request.

[0034] Note that the dynamic coherent memory domains as described herein can be implemented without any modifications to any coherence agent, such as a caching agent (CA) in a CPU.

[0035] Now refer to Figure 2 , which is a block diagram of a processor according to an embodiment. Figure 2 As shown in FIG, switch 200 includes various circuit systems, including ingress circuitry 212 and egress circuitry 219, for receiving incoming requests via ingress circuitry 212 and for sending outgoing communications via egress circuitry 219. For the purposes of describing the dynamic consistency mechanisms herein, switch 210 further includes a configuration interface 214, which can expose the capabilities herein to applications, including the ability to dynamically instantiate and update coherent memory domains. To determine whether an incoming request is for a coherent domain, coherence circuitry 220 can utilize information in a system address decoder 218, which can decode an incoming system address in a request.

[0036] like Figure 2 As further shown in the inset of FIG, coherence circuitry 220 includes caching agent (CA) circuitry 222, which can perform snooping and other coherence processing. More specifically, when control circuitry 224 determines that a request will be processed coherently, it can register CA circuitry 222 to perform coherence processing. This determination can be based at least in part on information maintained by telemetry circuitry 226, which can track the amount of traffic passing through the system, including interconnect bandwidth levels.

[0037] like Figure 2 As further shown in FIG, a rule database 230 is provided within the switch 210, which can store information about different memory domains. As shown, the rule database 230 includes multiple entries, each entry being associated with a given memory domain. As shown, each entry includes multiple fields, including a rule ID field, a memory range field, a PASID list field, a device list field, a callback field, and a consistency status field. These various fields can be populated in response to a memory allocation request and can be further updated in response to additional requests for updates, etc.

[0038] Embodiments can be applicable to multi-tenant usage in cloud and edge computing, as well as cloud-native applications with many microservices that do not have global consistency. For further illustration, multiple independent CXL consistency domains associated with different tenants can be isolated in system memory. For example, an application can deploy a container or virtual machine that specifies the following domains: Domain 1 - VMs A, B, C = compute devices S1, S2, S3, A3 sharing memory range [x, y] Domain 2 - VM D, E = computing devices S3, S4, S5, A4 sharing the memory range [z, t] Domain 3 - Shared memory between VMs C and D - All computing devices Application A generates monitoring @X1[x,y], and the CXL switch only monitors S1, S2, S3, and A3.

[0039] like Figure 2 As shown in [1], these different memory domains shared across platforms are not consistent across all computing devices. For each memory range, a set of targets to be monitored is specified, such as shown in domains 1, 2, and 3 above. In addition, some areas of memory may be read-only, such as the main storage of a database, which may account for a large percentage of memory capacity usage. For such defined areas, no monitoring or consistency is required.

[0040] With this arrangement, switch 210 can provide consistent quality of service (QoS) between and within coherence domains. In this way, switch 210 exposes an interface that can be used by: (1) infrastructure owners to specify what coherence QoS (in terms of priority or coherence transactions per second) is associated with each coherence domain; and (2) coherence domain owners to specify the QoS level associated with coherence flows between each of the domain's participants.

[0041] Active telemetry coherence saturation awareness is implemented via telemetry circuitry 226. This allows the software stack to be aware of how access to different objects within a coherence domain may be experiencing performance degradation. In an embodiment, telemetry circuitry 226 can track saturation of various paths between each of the domain's participants and various objects, and notify each of them based on provided monitoring rules.

[0042] In an embodiment for implementing monitoring and quality of service flows, switch 210 may include a content addressable memory (CAM)-type structure that can be tagged with object IDs to track access and apply QoS enforcement. To this end, system address decoder 216 tracks different objects and maps coherence requests (such as read requests) to the object. Thus, upon a particular coherence request, switch 210 can use SAD 216 to discover which coherence domain and object it belongs to; identify the implemented and specified QoS; and determine when to process the request. Note that if it is determined that the request has not yet been processed, the request can be stored in a queue. When processing the request, if the domain is coherent, it can proceed. If it is not coherent, switch 210 can perform a "fake" flow and respond to the originator with the response expected when the target does not have a row. Further, switch 210 sends the request directly to the target via egress circuitry 219. As an example, when forging a flow, the switch can return a global observation signal (e.g., ACK GO) (indicating to the originator that no one has the row).

[0043] The switch 210 can allow registration of new coherence domains via a configuration interface 214. In one embodiment, the interface allows specification of the identity of the address domain; and the memory ranges belonging to the memory domain. The assumption here is that the physical memory range (from 0..N) is mapped to all the different addressable memories in the system; the interface also enables specification of the elements within the memory domain, a list of process address IDs (PASIDs) belonging to the memory domain, and optionally, a list of devices within the memory domain. The configuration interface 214 can further enable changing or removing memory domains.

[0044] Coherence circuitry 220 can be configured to intercept CXL.cache requests and determine whether to intercept them. To this end, control circuitry 224 can use system address decoder 218 to identify, for each request, whether there is any coherence domain mapped to a specific address space that matches the memory address in the request. If no coherence domain is found, the request exits egress circuitry 219 toward its final destination.

[0045] If one or more domains are found, then for each of them, coherence circuitry 220 may check whether the PASID included in the request maps to that domain. If so, the request exits egress circuitry 219 toward the final destination. If not, coherence circuitry 220 may discard the snoop or memory CXL.cache request. Coherence circuitry 220 implements a coherence response corresponding to that particular CXL.cache request. For example, the response may be invalid.

[0046] Now refer to Figure 3 , which shows a flow chart of a method according to an embodiment. Figure 3 As shown in , method 300 is a method for generating and updating memory attributes in response to a memory allocation request. Thus, method 300 can be performed by switching circuitry, such as coherence circuitry within a switch according to an embodiment. Thus, method 300 can be performed by hardware circuitry, firmware, software, and / or a combination thereof.

[0047] As shown, method 300 begins by receiving a memory allocation request in a switch (block 310). By way of example, an application such as a VM, a process, or any other software entity may issue the request, which may include various information. Although embodiments are not limited in this respect, example information in the request may include memory range information, coherence status, address space identifier information, and the like.

[0048] Next, control passes to diamond 320 where it is determined whether an entry already exists in the memory domain table for the memory range of the memory allocation request. If not, control passes to block 330 where an entry in the table may be generated. As an example, an entry may include the one described above with respect to Figure 2 Otherwise, if it is determined that the entry already exists, control passes to block 340 where the entry may be updated. For example, the consistency state may be changed, such as making a consistency domain a non-consistent domain (such as after a transaction completes), deleting a memory domain (such as when an application terminates), etc. Although in Figure 3 The embodiments are shown at this high level, but many variations and alternatives are possible.

[0049] Now refer to Figure 4 , which shows a flow chart of a method according to another embodiment. Figure 4 As shown in , method 400 is a method for handling incoming memory requests in a switch. Therefore, method 400 can be performed by various circuit systems within the switch. Therefore, method 400 can be performed by hardware circuit systems, firmware, software, and / or a combination thereof.

[0050] Method 400 begins by receiving a memory request in a switch (block 410). For the purposes of discussion, assume that the memory request is for reading data. The read request includes the address at which the requested data is located. Next, at block 420, a memory domain table may be accessed based on the address of the memory request, for example, to identify an entry in the table associated with the memory domain that includes the address.

[0051] At diamond 425, a determination is made as to whether the memory request is for a coherent memory domain. This determination may be based on the coherence status indicator present in the coherence status field of the associated entry in the memory domain table. If not, control passes to block 430 where the memory request is forwarded to the destination location without further processing within the switch because the request is directed to a non-coherent domain.

[0052] Still refer to Figure 4 If it is determined that the request is for a coherent memory domain, control passes to diamond 440 to determine whether the memory request is associated with a snoop. This determination can be based on whether the request is for a read, in which case snoop processing can be performed. Other memory requests, such as write requests, can be handled directly without snoop processing (block 445).

[0053] Next, control passes to diamond 450 to determine whether the snoop process is allowed. This determination can be based on one or more system parameters, such as the interconnect status. If it is determined that the snoop process is not allowed, such as in the presence of high interconnect traffic, control passes to block 460. At block 460, the memory request can be handled based on the callback information. More specifically, the relevant entry in the memory domain table can be accessed to determine a fallback mechanism that can be used to handle the snoop process. In this manner, reduced interconnect traffic can be achieved.

[0054] Still refer to Figure 4 If it is determined at diamond 450 that the snoop process is allowed, control passes to box 470 where a snoop process is performed to determine the presence and status of the requested data in various distributed caches and other memory structures. Next, at box 480, the memory request can be handled based on the snoop results. For example, when it is determined that the latest copy of the data is valid, the read request can be executed. Alternatively, in the case where dirty data is indicated, the dirty data can be used to provide read completion. Although in Figure 4 The embodiments are shown at this high level, but many variations and alternatives are possible.

[0055] Various interconnect architectures and protocols can utilize the concepts discussed in this article. As computing systems and performance requirements advance, improvements to interconnect structures and link implementations are under development, including interconnections of components based on or utilizing PCIe or other traditional interconnect platforms. In one example, Compute Express Link (CXL) has been developed, which provides improved high-speed CPU-to-device and CPU-to-memory interconnection, which is designed to accelerate next-generation data center performance and other applications. CXL maintains memory consistency between the CPU memory space and the memory on the attached device, which allows resource sharing to obtain higher performance, reduced software stack complexity and lower overall system cost, as well as other example advantages. CXL enables communication between a host processor (e.g., CPU) and a set of workload accelerators (e.g., graphics processing units (GPUs), field programmable gate arrays (FPGAs), tensor and vector processor units, machine learning accelerators, dedicated accelerator solutions, and other examples). In fact, CXL is designed to provide a standard interface for high-speed communication because accelerators are increasingly used to supplement CPUs to support emerging computing applications such as artificial intelligence, machine learning, and other applications.

[0056] CXL links can be low-latency, high-bandwidth discrete or packaged uplinks that support dynamic protocol multiplexing of coherence, memory access, and input / output (I / O) protocols. In other applications, CXL links can enable accelerators to access system memory as a caching agent and / or host system memory (among other examples). CXL is a dynamic multi-protocol technology designed to support a variety of accelerators. CXL provides a rich set of protocols, including I / O semantics similar to PCIe (CXL.io), cache protocol semantics (CXL.cache), and memory access semantics (CXL.mem) via discrete or packaged uplinks. Based on the specific accelerator usage model, all CXL protocols or only a subset of the protocols can be enabled. In some implementations, CXL can be built on the well-established, widely adopted PCIe infrastructure (e.g., PCIe 5.0), leveraging the PCIe physical and electrical interfaces to provide advanced protocols in areas including I / O, memory protocols (e.g., allowing the host processor to share memory with the accelerator device), and coherence interfaces.

[0057] Steering Figure 5 , a simplified block diagram 500 illustrating an example system utilizing a CXL link 550 is shown. For example, the link 550 can interconnect a host processor 505 (e.g., a CPU) to an accelerator device 510. In this example, the host processor 505 includes one or more processor cores (e.g., 515a-515b) and one or more I / O devices (e.g., 518). Host memory (e.g., 560) can be provided with the host processor (e.g., on the same package or die). The accelerator device 510 can include accelerator logic 520 and, in some implementations, can include its own memory (e.g., accelerator memory 565). In this example, the host processor 505 can include circuitry for implementing coherency / cache logic 525 and interconnect logic (e.g., PCIe logic 530). CXL multiplexing logic (e.g., 555a-555b) may also be provided to enable multiplexing of CXL protocols (e.g., I / O protocols 535a-535b (e.g., CXL.io), cache protocols 540a-540b (e.g., CXL.cache), and memory access protocols 545a-545b (CXL.mem)) so that data for any one of the supported protocols (e.g., 535a-535b, 540a-540b, 545a-545b) can be sent in a multiplexed manner over the link 550 between the host processor 505 and the accelerator device 510.

[0058] In some implementations, the Flex Bus TM The ports can be used in conjunction with CXL-compliant links to flexibly adapt the device for interconnection with a variety of other devices (e.g., other processor devices, accelerators, switches, memory devices, etc.). A FlexBus port is a flexible, high-speed port that is statically configured to support either PCIe or CXL links (and potentially links supporting other protocols and architectures). The FlexBus port allows designs to choose between providing native PCIe protocol or CXL over a high-bandwidth, out-of-encapsulation link. The selection of the protocol applied at the port can occur via auto-negotiation during boot time and is based on the device plugged into the slot. The FlexBus uses PCIe electrical, making it compatible with PCIe retimers and conforming to the standard PCIe form factor for add-in cards.

[0059] Go to Figure 6 , shows (in a simplified block diagram 600) an example of a system that utilizes flexible bus ports (e.g., 635-640) to implement CXL (e.g., 615a-615b, 650a-650b) and PCIe links (e.g., 630a-630b) to couple various devices (e.g., 510, 610, 620, 625, 645, etc.) to a host processor (e.g., CPU 505, 605). In this example, the system may include interconnected processors via inter-processor links 670 (e.g., utilizing UltraPath Interconnect (UPI), Infinity Fabric, TM (Infinity Fabric TM ) or other interconnection protocol). Each host processor device 505, 605 can be coupled to a local system memory block 560, 660 (e.g., a double data rate (DDR) memory device), which is coupled to the corresponding host processor 505, 605 via a memory interface (e.g., a memory bus or other interconnection).

[0060] As described above, CXL links (e.g., 615a, 650b) can be used to interconnect various accelerator devices (e.g., 510, 610). Accordingly, corresponding ports (e.g., flexible bus ports 635, 640) can be configured (e.g., selecting CXL mode) to enable establishing CXL links and interconnecting corresponding host processor devices (e.g., 505, 605) to accelerator devices (e.g., 510, 610). As shown in this example, flexible bus ports (e.g., 636, 639) or other similarly configurable ports can be configured to implement general-purpose I / O links (e.g., PCIe links) 630a-630b instead of CXL links to interconnect host processors (e.g., 505, 605) to I / O devices (e.g., smart I / O devices 620, 625, etc.). In some implementations, the memory of the host processor 505 can be connected to the memory expander device (e.g., 645) of the host processor(s) 505, 605, for example, through the memory (e.g., 565, 665) of the connected accelerator device (e.g., 510, 610) or via corresponding CXL links (e.g., 650a-650b) implemented on the flexible bus ports (637, 638), among other example implementations and architectures.

[0061] Figure 7 is a simplified block diagram illustrating an example port architecture 700 (e.g., a flexible bus) for implementing a CXL link. For example, the flexible bus architecture can be organized into multiple layers to implement multiple protocols supported by the port. For example, a port can include transaction layer logic (e.g., 705), link layer logic (e.g., 710), and physical layer logic (e.g., 715) (e.g., implemented in whole or in part in circuitry). For example, the transaction (or protocol) layer (e.g., 705) can be subdivided into transaction layer logic 725, which implements a basic PCIe transaction layer 755 and a CXL transaction layer enhancement 760 (for CXL.io) of the PCIe transaction layer 755, and logic 730 for implementing cache (e.g., CXL.cache) and memory (e.g., CXL.mem) protocols for the CXL link. Similarly, link layer logic 735 can be provided to implement a basic PCIe data link layer 765 and a CXL link layer (for CXL.io) representing an enhanced version of the PCIe data link layer 765. The CXL link layer 710 may also include cache and memory link layer enhancement logic 740 (eg, for CXL.cache and CXL.mem).

[0062] continue Figure 7 In some examples, the CXL link layer logic 710 can provide an interface to CXL arbitration / multiplexing (ARB / MUX) logic 720 that interleaves traffic from two logical streams (e.g., PCIe / CXL.io and CXL.cache / CXL.mem), as well as other example implementations. During link training, the transaction layer and the link layer are configured to operate in either PCIe mode or CXL mode. In some instances, the host CPU can support an implementation of either PCIe or CXL mode, while other devices such as accelerators can support only CXL mode, as well as other examples. In some implementations, a port (e.g., a flexible bus port) can utilize a physical layer 715 that is based on a PCIe physical layer (e.g., PCIe electrical PHY 750). For example, the flexible bus physical layer can be implemented as a fused logical physical layer 745 that can operate in either PCIe mode or CXL mode based on the results of alternative mode negotiation during the link training process. In some implementations, the physical layer can support multiple signaling rates (e.g., 8GT / s, 16GT / s, 32GT / s, etc.) and multiple link widths (e.g., x16, x8, x4, x2, x1, etc.). In PCIe mode, the link implemented by port 700 can fully conform to native PCIe features (e.g., as defined in the PCIe specification), while in CXL mode, the link supports all features defined for CXL. Accordingly, the flexible bus port can provide a point-to-point interconnect that can electrically transmit native PCIe protocol data or dynamic multi-protocol CXL data over PCIe to provide I / O, coherence, and memory protocols, among other examples.

[0063] The CXL I / O protocol, CXL.io, provides a non-coherent load / store interface for I / O devices. The transaction types, transaction grouping format, credit-based flow control, virtual channel management, and transaction ordering rules in CXL.io can conform to all or part of the PCIe definition. The CXL cache coherence protocol, CXL.cache, defines interactions between devices and hosts as multiple requests, each with at least one associated response message and, optionally, a data transfer. This interface consists of three channels in each direction: request, response, and data.

[0064] The CXL memory protocol, CXL.mem, is the transaction interface between the processor and the memory and uses the physical and link layers of CXL when communicating across the die. CXL.mem can be used for multiple different memory attachment options, including when the memory controller is located in the host CPU, when the memory controller is located in an accelerator device, or when the memory controller is moved to a memory buffer chip. CXL.mem can be applied to transactions involving different memory types (e.g., volatile, persistent, etc.) and configurations (e.g., flat, hierarchical, etc.), as well as other example features. In some implementations, the host processor's consistency engine can provide an interface to the memory using CXL.mem requests and responses. In this configuration, the CPU consistency engine is considered a CXL.mem master, and the Mem device is considered a CXL.mem slave. The CXL.mem master is the agent responsible for obtaining CXL.mem requests (e.g., read, write, etc.), and the CXL.mem slave is the agent responsible for responding to CXL.mem requests (e.g., data, completion, etc.). When the slave device is an accelerator, the CXL.mem protocol assumes the presence of a device coherency engine (DCOH). This agent is assumed to be responsible for implementing coherency-related functions, such as snooping the device cache based on CXL.mem commands and updating metadata fields. In implementations where metadata is backed by device-attached memory, the metadata can be used by the host to implement a coarse snoop filter for the CPU socket, among other example uses.

[0065] In some implementations, an interface may be provided to couple circuitry or other logic (e.g., an intellectual property (IP) block or other hardware element) that implements a link layer (e.g., 710) to circuitry or other logic (e.g., an IP block or other hardware element) that implements at least a portion of a physical layer (e.g., 715) of a protocol. For example, an interface based on the Logical PHY Interface (LPIF) specification is used to define a common interface between a link layer controller, module, or other logic and a module that implements a logical physical layer ("logical PHY," or "logPHY") to facilitate interoperability, design, and verification reuse between one or more link layers and a physical layer of an interface for physical interconnection, such as in Figure 7 In the example of . Additionally, as in Figure 7 In some implementations, each block (e.g., 715, 720, 735, 740) in a multi-protocol implementation may provide an interface to another block via a separate LPIF interface (e.g., 780, 785, 790). In the case of bifurcation support, each bifurcation port may also have its own separate LPIF interface, among other examples.

[0066] While the examples discussed herein may refer to the use of a link layer logical PHY interface based on LPIF, it should be understood that the details and principles discussed herein may be equally applicable to non-LPIF interfaces. Similarly, while some examples may refer to the use of a common link layer logical PHY interface to couple a PHY to a controller implementing CXL or PCIe, other link layer protocols may utilize such an interface as well. Similarly, while some references may be made to a flexible bus physical layer, other physical layer logic may also be employed in some implementations and utilize a common link layer logical PHY interface, such as discussed herein, and other example variations within the scope of the present disclosure.

[0067] Advances in multi-chip packaging (MCP) technology allow multiple silicon dies to be included in the same package. High-density, low-latency die-to-die interconnects optimized for short distances can have very low bit error rates (BER) (e.g., better than 1×10 -18 As a result, these interconnects typically omit the overhead of the serializer / deserializer (SERDES) circuitry and the synchronization associated with packet tracking transmissions, and also omit the overhead of a complex link training and status state machine (LTSSM) in the logic PHY.

[0068] A variety of different protocols (e.g., CXL, PCIe, Ultra Path Interconnect (UPI), In-Die Interconnect (IDI), etc.) would benefit from a common logical PHY interface to enable the use of die-to-die interconnects, where the common logical PHY interface (or adapter) acts as a transport mechanism that abstracts the handshakes used for initialization, power management, and link training. For example, a conventional logical PHY implementation may require a customized handshake with the conventional logical PHY for each different protocol. In an improved implementation, an adapter circuit system may be provided to implement a common logical PHY that allows upper protocol layers (e.g., link layer) to be transported over a variety of different die-to-die building blocks. The adapter may enable the transport of raw bit streams through a die-to-die interface that uses a subset of a common link layer to PHY interface protocol (e.g., LPIF). Potentially, any die-to-die electrical interface may utilize such an interface by providing such an adapter. In some implementations, the adapter can utilize a subset of defined common link layer to PHY interfaces (such as LPIF) with which existing link layer controllers have been configured to interoperate (e.g., LPIF for PCIe / flexbus / R-Link logPHY, etc.), among other example uses and advantages.

[0069] Go to Figures 8A-8C , simplified block diagrams 800a-800c are shown illustrating example implementations of interfaces utilizing an adapter block (e.g., 805) to facilitate implementation of a common interface between various link layer blocks and various die-to-die PHY blocks. An example adapter (e.g., 805) may be provided to terminate, recondition, and gear data to be transmitted over a die-to-die interface 815. For example, two or more dies may be provided on a package, with a die-to-die interconnect 815 (e.g., implemented as a high-bandwidth die-to-die PHY IP block on the same package) serving as an interface between the two dies on the package. The adapter 805 may be provided with state machine logic to support and transition between a simplified set of states for a die-to-die environment. Additionally, the adapter may include logic for defining efficient sideband channels (e.g., 830) for various handshakes used to initiate a link, provide power management, and facilitate state transitions, among other example features. Such an adapter device can provide a universal and protocol-agnostic die-to-die interface including an in-band data channel (e.g., 835) and a sideband channel (e.g., 830). The adapter can be used to support a simplified logic PHY, which can result in lower latency and lower power for the interface, while enabling substantially greater bandwidth per millimeter of die shoreline (or edge) due to the dense I / O possible for the die-to-die interface. Additionally, the adapter can support the simultaneous use of multiple protocols to allow such bandwidth to be scalable, among other example advantages.

[0070] like Figures 8A-8C As shown in , various implementations may utilize an example adapter 805 (e.g., based on a common link layer to PHY interface (such as LPIF)) that may be provided to implement a defined interface capable of transmitting and receiving common link layer to PHY interface data (e.g., LPIF data) via a die-to-die PHY 815. The link layer may have one or more functional pipes (e.g., 820, 825), each corresponding to an implemented protocol. In some implementations, such as Figure 8A As shown in , traffic from each pipe 820, 825 can interface to a single shared LPIF adapter 805. Additionally, in some implementations, such as in Figure 8A In the example of , an ARB / MUX 720 can be instantiated between a multi-protocol link layer (e.g., including link layer controllers 820, 825) and an adapter block 805, where the ARB / MUX 720 arbitrates between traffic from different pipes (e.g., 820, 825) to drive to the adapter 805. The adapter can terminate and / or recondition data from the common link layer to the PHY interface for transmission over the die-to-die PHY (e.g., 815). The adapter 805 can also coordinate various handshakes with the link layer for power management (PM) and clock gating. The adapter 805 can also perform handshakes with the remote die when applicable for error / reset / power management propagation, as well as other example features. In short, as in Figures 8A-8C In the example of FIG, using an LPIF-based adapter 805 allows a link layer element that already or natively supports LPIF to seamlessly connect to a logical PHY for PCIe / CXL or an LPIF adapter for die-to-die transport, where the LPIF adapter implements a very lightweight logical PHY for die-to-die communication.

[0071] The LPIF adapter 805 is used to send raw protocol streams over a multi-die interface for die-to-die operations. In some implementations, the die-to-die PHY 815 is implemented as a simplified high-density die-to-die PHY that achieves lower latency and power performance than a conventional die-to-die PHY, among other examples. The LPIF adapter 805 may include digital logic for providing an interface to the PHY 815. The adapter 805 may implement a substantially simplified logical PHY for die-to-die transmission. The LPIF adapter 805 may facilitate handshakes according to the LPIF interface when transmitting raw data bits. The LPIF adapter 805 may implement a sideband channel 830 to the PHY 815 to exchange adapter-to-adapter handshakes. These handshakes may also be accomplished via the primary band by assigning specific packets / microslices that are unique to the LPIF adapter (and not used by the protocol), among other example implementations.

[0072] Figure 8A An example implementation of an LPIF adapter 805 is shown that provides an interface between an ARB / MUX device 720 and a die-to-die PHY 815. The ARB / MUX circuitry 720 may include an LPIF interface to each of a plurality of link layer function pipes (provided by corresponding logic (e.g., 820, 825)). The ARB / MUX circuitry 720 may additionally include a single LPIF interface for coupling to a single LPIF adapter (e.g., 805). In other implementations, the ARB / MUX circuitry may be omitted, such as in Figure 8B As shown in the example. Figure 8B In the example of , multiple link layer controllers (e.g., 820, 825) may be provided that are not multiplexed to the same LPIF adapter but instead provide interfaces to dedicated adapter instances (e.g., 805, 805') that provide interfaces and logical PHYs between the link layer pipes and the die-to-die PHY 815. Figure 8A-8B In an example, there is a single cluster of die-to-die PHYs (eg, an atomic unit or cluster of die-to-die PHYs such that all signals within the cluster are natively synchronized).

[0073] For example, multiple LPIF adapters (e.g., 805, 805') may be provided to facilitate higher bandwidth applications (e.g., for sending parallel transfers of CXL.io and CXL.cache / CXL.mem data). Figure 8C In some embodiments, the bandwidth may be doubled (or otherwise multiplied) by providing multiple die-to-die PHY interfaces (e.g., 815, 815'). Respective adapters (e.g., 805, 805') may provide corresponding logical PHYs and interfaces between link layer controllers (e.g., 820, 825) and corresponding die-to-die PHY blocks (e.g., 815, 815'). In other examples, a single adapter may be provided as a logical PHY implementation and interface to multiple die-to-die PHY blocks (e.g., 815, 815'). In this example, the die-to-die PHY blocks (e.g., 815, 815') may be coupled (e.g., using sideband or other channels (e.g., 830)) to enable additional synchronization between the die-to-die PHY blocks (e.g., 815, 815'), as well as other example features and implementations. Synchronization between different clusters of die-to-die PHYs (e.g., 815, 815') illustrates that the LPIF data width does not need to be tightly coupled to the die-to-die cluster width. Figures 8A to 8C The examples shown in are simplified examples and represent only a small portion of potential implementations that may utilize an adapter that provides an interface to a die-to-die PHY, such as described herein.

[0074] Various signals and interactions can be defined between the adapter device and the link layer elements. In addition, in some implementations, a sideband channel (e.g., 830) can be defined to perform auxiliary communication with a remote die (e.g., a corresponding adapter on the remote die). In some instances, a unique microchip / packet can be assigned to communicate with the LPIF adapter of the remote die's LPIF adapter. Communication between adapters on interconnected dies (e.g., sideband handshakes or dedicated packets) can be used, for example, for link startup and operation. The adapter can utilize a subset of the signals defined in the link layer to PHY interface (e.g., LPIF) to implement an adapter lightweight logic PHY. For example, some signals can be omitted because, in some die-to-die applications, once the link is operational, there is no need to retrain (recover) the link. In addition, the scope of the adapter can be expanded to include it using the mechanisms defined by the corresponding link layer to PHY interface. Signal transmission between the adapter and the PHY block (e.g., die-to-die PHY) can be flexible and / or implementation-specific, wherein the adapter is configured to communicate according to a specific die-to-die PHY design.

[0075] In some implementations, the LPIF adapter and the corresponding PHY block can scale with data width using a single or multiple instantiations. In this case, synchronization across multiple instantiations can be implemented in the PHY. The data transmission rate ratio (gear ratio) can also be scaled based on the PHY implementation. This can allow bridge cores operating at different frequencies, as well as other example applications. The example LPIF adapter can support serialization / deserialization, or simple throttling logic to ensure that no data is lost when transmitting to different frequencies. In some implementations, if necessary, the adapter can also implement, include or otherwise instantiate a clock crossing FIFO queue. The back pressure to the link layer can be controlled, for example, using one of the signals defined in the corresponding link layer to PHY interface (e.g., pl_trdy in LPIF), as well as other examples. In addition, in some implementations, error correction can be supported by the adapter, for example, for additional link protection to ensure specific BER requirements, as well as other examples. Error correction can be implemented in either (or both) of the PHY block and the adapter, as well as other example features and implementations.

[0076] In some implementations, each instance of the link layer to PHY interface on a particular die can operate at the same clock frequency and within the same power domain. If different clock frequencies or clock sources are used, additional FIFOs can be used for clock crossing. If different power domains are used, voltage isolation can be utilized, among other example features. In addition, the link layer logic and the corresponding adapter block can be within the same reset (RESET) domain. The secondary side adapter clock (e.g., the PHY clock) can be derived from the same phase-locked loop (PLL) circuit as the primary side clock (LPIF clock). In some implementations, the adapter can be configured such that some or all portions of the adapter are in an always-ON power domain to enable wakeup from a low power state (e.g., via sideband or mainband communication), among other example features.

[0077] In some implementations, it may be desirable to maintain the security of information for a virtual machine, as well as isolation from the hypervisor / virtual machine monitor (VMM) and from other (one or more) virtual machines. Certain processors (e.g., a system-on-a-chip (SoC) including a processor) include hardware that assists with this isolation. A securely isolated VM is sometimes referred to as a "trust domain" (TD), a "trusted VM" (TVM), or a "secure VM." For the remainder of this disclosure, these terms are used interchangeably to refer to any VM or guest protected by hardware-based isolation. More generally, any trusted software component may be referred to as a "trusted execution environment" (TEE).

[0078] Some processors support an instruction set architecture (ISA) (e.g., an ISA extension) to implement trust domains. For example, Trust Domain Extension ( Similarly, Advanced Micro Devices (AMD) has released Secure Encrypted Virtualization (SEV) extensions (SEV-SNP) with Secure Nested Paging (SNP) to deploy hardware-isolated VMs, referred to as "Trusted VMs" or "Secure VMs."

[0079] According to some examples, the hardware processor and its ISA can implement a management component (e.g., referred to as a trust domain manager or trusted security manager) that manages the isolation of trusted VMs from a VMM / hypervisor and / or other non-secure software (e.g., on a host platform). For these examples, the hardware processor and its ISA implement a trusted execution environment to enhance confidential computing by helping to protect trusted VMs from a wide range of software attacks and reducing the trusted computing base (TCB).

[0080] In some examples, the hardware processor and its ISA also support device I / O. For example, utilizing an ISA (e.g., TDX-IO) that supports Trust Domain Extensions (TDX) with device I / O (e.g., TDX-IO) TDX 2.0). For these examples, a hardware processor supporting device I / O and its ISA enable the assignment of a device's physical function (PF) and / or virtual function (VF) to a specific TD.

[0081] The I / O device can be an accelerator, and different types of accelerators can be used. For example, the first type of accelerator is an in-memory analytics accelerator (IAX). The second type of accelerator supports a set of transformation operations on the memory, such as a data streaming accelerator (DSA). For example, the accelerator is used to generate and test cyclic redundancy check (CRC) checksums or data integrity fields (DIF) to support storage and networking applications and / or for memory comparison and incremental generation / merging to support VM migration, VM fast check indication, and software management memory deduplication. The third type of accelerator supports security, authentication, and compression operations (e.g., encryption acceleration and compression operations), such as a Quick Assist Technology (QAT) accelerator. The fourth type of accelerator is a general-purpose graphics processor (GPGPU), which can also be referred to as a machine learning accelerator.

[0082] In some examples, to establish a trust relationship between a device and a trusted VM, certain architectures require that the trusted VM and / or a trusted security manager (e.g., a trusted execution environment (TEE) security manager (TSM)) establish a secure communication session between the device and the TSM (e.g., for a trust domain manager to allow a particular trusted VM to use the device or a subset of (one or more) functions of the device). For these examples, to establish a trust relationship between the device and the trusted VM, certain architectures require that the trusted VM and / or TSM authenticate the device (e.g., and collect device measurements) using various specifications, including but not limited to Distributed Management Task Force (DMTF), Secure Protocol and Data Model (SPDM) specifications (such as the SPDM specification, DSP0274, version 1.0.1, published by the DMTF's Platform Management Components Intercommunication (PMCI) working group in March 2021 (hereinafter referred to as the "SPDM specification")).

[0083] The TEE and / or TSM may also communicate with a device security manager (DSM) to manage the device's virtual function(s) using protocols and techniques described in the Peripheral Component Interconnect Special Interest Group (PCI-SIG) and / or the Trusted Device Interface Security Protocol (TDISP). An example of a TDISP implementation is the "TDX Connection Architecture," although the underlying principles of the present invention are not limited to any particular TDISP implementation.

[0084] According to some examples, the SPDM messaging protocol used in accordance with the SPDM specification defines a request-response messaging model between two endpoints to perform a message exchange, e.g., where each SPDM request message should be responded to with an SPDM response message. For these examples, a "measurement" of an endpoint (e.g., a device) describes the process of computing a cryptographic hash value for a piece of firmware / software or configuration data and tying the cryptographic hash value to the identity of the endpoint using a digital signature. This allows an authentication initiator to establish the identity and measurement of the firmware / software and / or configuration currently running on the endpoint.

[0085] In some examples, to help enforce the security policy of the trusted VM, the host processor uses a secure processor mode such as Secure-Arbitration Mode (SEAM) to implement a digitally signed but unencrypted security service module. For example, a trust domain manager (TDM) can be hosted in a reserved memory space identified by a SEAM-range register (SEAMRR). For this example, the processor can only allow software executing within the SEAM memory range to access the SEAM-memory range, while all other software accesses and direct-memory access (DMA) from the device to the SEAM-memory range are terminated. In some examples, the SEAM module does not have any memory access privileges to other protected memory areas in the compute / host platform, including system management mode (SMM) memory or protected memory (e.g., Software Guard Extensions (SGX).

[0086] The TDISP message protocol can be used by a TEE security manager (TSM) in a confidential computing environment (e.g., an environment implementing TDX-IO or SEV-IO). Some implementations include a secure startup service module (S3M) in the SoC for establishing secure communication sessions. In some examples, the secure startup service circuitry includes SPDM capabilities and stack / device attestation capabilities (e.g., to support TDX-IO, SEV-TIO, and other secure IO implementations).

[0087] Although described in the context of some of these security architectures, the underlying principles of the invention are not limited to any particular security architecture. For example, the embodiments described below can be used with any confidential / secure computing technology. For example, AMD Secure Encrypted Virtualization (e.g., SEV / SEV-ES / SEV-SNP) can use its specific components (e.g., Platform Security Processor (PSP)) to implement a TSM that includes two parts: (i) a manager that implements TEE isolation, and (ii) a PSP that handles communications with the Device Security Manager (DSM). For example, The Realm Management Extension (RME) can use its specific components (e.g. multiple One of the core Core) to implement TSM, for example, the entire TSM includes two components: (i) a trust domain manager that implements TEE isolation, and (ii) a trust domain manager that handles communication with the device security manager (DSM). core.

[0088] Recently, PCI-SIG released the PCIe Integrity & Data Encryption (IDE) standard, which provides a solution for end-to-end link encryption between PCIe root ports and devices. The IDE standard is implemented in a new IDE layer between the transaction layer and the data link layer. This new IDE layer uses encryption mechanisms to encrypt all data sent on both sides of each link. IDE supports different sets of "selective" IDE streams and "linked" IDE streams formed between different components. Each linked IDE stream is a secure channel within the corresponding link, and each selective IDE stream is an end-to-end secure channel between devices (not necessarily within the same link and may be across multiple links).

[0089] To ensure the security of the selective IDE stream for transaction layer packets (TLPs), the IDE standard requires that completions of non-posted requests must be transmitted on the same selective IDE stream as the one on which the device transmitted the request. The IDE "limited" stream feature was introduced to reduce the performance burden of IDE encryption. The limited stream feature allows a device to reduce the bandwidth and latency burden of encryption by encrypting only certain sensitive transactions rather than all transactions it sends to the root port. Apparatus and method for trusted access to extended memory across multiple potentially heterogeneous computing nodes

[0090] Memory capacity and bandwidth are becoming major bottlenecks in high-performance computing (HPC) and artificial intelligence / machine learning (AI / ML) systems. CXL Type 3 Multiple Logical Device (MLD) implementations with memory pooling are currently being used to expand memory capacity beyond classic host-attached memories such as DDR and HBM.

[0091] Supporting confidential computing on asymmetric types of storage (e.g., host-attached and CXL-attached storage) is a challenge. Host-attached storage is commoditized, and all confidential computing requirements for encryption are handled at the host. However, with CXL-attached storage, encryption and confidentiality schemes are inconsistent, and devices support different combinations of link encryption, memory encryption, and device attestation.

[0092] Embodiments of the present invention provide a solution: establishing trust boundaries across these asymmetric devices and enabling diverse CPUs, graphics processors, and accelerators (sometimes referred to as "XPUs") to be included in any CXL Type 3 memory pooling architecture. This is achieved through a common root of trust across multiple nodes. In particular, multiple host processors and accelerators in the ecosystem rely on a "Super Root of Trust" (SROT). In some implementations, the Super ROT is implemented using a microcontroller embedded in a CXL 2.0 switch that manages CXL Type 3 MLD devices (e.g., extended memory devices). It can also be configured to manage the allocation and deallocation of processor / accelerator workloads within a given extended memory pool of a processor / accelerator.

[0093] As used herein, the terms "binding" and "unbinding" refer to the process of allocating and deallocating pooled memory. In some embodiments, pooled memory is provided to processor or accelerator workloads (e.g., in response to software-generated requests) via a CXL Type 3 MLD interconnect. In some embodiments, a CXL switch is configured to couple heterogeneous devices (e.g., CPUs, graphics processors, accelerators) to pooled memory within the same trust boundary. In these embodiments, the CXL switch can enumerate support for Type 3 MLD devices.

[0094] Some implementations of the present invention rely on a host processor (e.g., a CPU core) to manage encryption and decryption requirements along the path to memory (referred to herein as host-based encryption (HBE)). In some embodiments, HBE provides both local integrity and encryption / decryption. Local integrity is provided using an access control table (ACT) and a corresponding instruction set that tracks ownership of each cache line entry during memory allocation.

[0095] In these embodiments, encryption / decryption is performed based on a KeyID corresponding to a given host physical address (HPA), which is used to encrypt and / or decrypt a payload in extended memory originating from a trusted execution environment (e.g., a virtual machine corresponding to a trust domain or other form of trusted software). In some specific implementations, the KeyID is a trust domain extension key (TDXKEY) or a multi-key total memory encryption (MKTME) key, although the underlying principles of the invention are not limited to any particular type of trusted execution environment (TEE).

[0096] These embodiments provide a Trusted Device Interface Security Protocol (TDISP) implementation within a CXL architecture (e.g., a CXL Type 3 architecture) (e.g., a TDX connection) by building a single trust boundary within a CXL switch across multiple nodes. Through the single trust boundary, the allocation (binding) or deallocation (unbinding) of secure and non-secure application data in a memory pool of an extended device memory is seamlessly managed. In some embodiments, the logical integrity of data stored in the extended device memory is achieved using a host-based encryption (HBE) scheme in the IO path (i.e., according to CXL.mem).

[0097] In some embodiments, multiple TDISP modules (e.g., TDX connection modules) are supported in a CXL type 3MLD device integrated with a host-based encryption scheme. These embodiments are scalable for XPU operating models as well as accelerator-enabled nodes such as Xe or any third-party GPGPU (e.g., Nvidia H100 GPGPU, AMD MI300 GPGPU, etc.). In some embodiments, hyper-ROT is utilized to manage scaling of nodes via lookup table (LUT) entries.

[0098] Figure 9 The diagram illustrates an example implementation including a host processor 920 having a local memory controller 922 for coupling the host processor to local memory 915 (e.g., HBM or DDR memory) and a CXL interface 924 for coupling the host processor to expansion memory 940 via a CXL Type 3 controller 930. As indicated, the CXL controller 930 may optionally support aspects of the PCIe Integrity & Data Encryption (IDE) standard (e.g., such as encrypted "linked" IDE streams). The physical address space visible to the host processor 920 may include, for example, on-die or host-attached memory 915 and CXL Type 3 expansion memory 940. By way of example and not limitation, if the on-die or host-attached memory 915 has a 2TB capacity and the expansion memory 940 has a 4TB capacity, then during the enumeration phase, the host processor 920 is informed of 6TB of available system memory.

[0099] refer to Figure 10 One embodiment of a host processor 920 includes multiple cores 1001 coupled to a home agent 1002, which provides the cores with access to host-attached memory 1098 and CXL type 3 MLD memory 1033 (via a CXL memory expansion device 1032). In these embodiments, the home agent 1002 is configured to distinguish between the on-die / host-attached memory subsystem 1090 and the expansion memory device 1032. Specifically, the illustrated home agent 1002 includes a source address decoder 1003, which uses a range register 1004 to perform address checking based on a host physical address (HPA) provided with the request. The source address decoder 1003 uses the range register 1004 to verify which memory path will accommodate the memory request: the host memory subsystem 1090, which provides access to the host-attached memory 1098, or the CXL controller 1030, which provides access to the CXL memory expansion device 1032 and the CXL type 3 MLD memory 1033.

[0100] In some embodiments, when the range register check indicates the host memory subsystem 1090, encryption / decryption circuitry 1094 of the host memory subsystem 1090 performs encryption / decryption based on the corresponding encryption key identified via key lookup circuitry 1092. When the range register check indicates the CXL type 3 expansion memory device 1032, host-based encryption 1051 managed by the host processor 920 is performed. In one embodiment, circuitry and instructions executed on the host processor 920 implement host-based encryption 1051, which tracks cache line ownership in an access control table (ACT) to maintain logical integrity. Details of one embodiment of host-based encryption 1051 are provided below.

[0101] Figure 11A The diagram illustrates an example embodiment including two processors 1122-1123, each including multiple cores, and a graphics processor 1124 coupled to a CXL switch 1110. This is merely one example implementation; embodiments of the present invention can be used with a variety of heterogeneous and homogeneous cores and processors. In the illustrated example, secure application 1191A and non-secure application 1192A execute on processor 1122; secure application 1191B and non-secure application 1192B execute on processor 1123; and secure application 1191C and non-secure application 1192C execute on processor 1124. Each processor 1122-1123 includes host-based encryption (HBE) circuitry / logic 1051A-1051B, respectively, for performing encryption and decryption operations as described herein on memory transactions on the CXL device memory path (e.g., using a corresponding KeyID to indicate the associated encryption key). GPU cryptographic circuitry 1051C may also perform HBE, or may perform different forms of encryption (e.g., SME / AES-XTS). In some embodiments, host security circuitry 1160A-1160C (such as a secure microcontroller) is coupled to or integrated with processors 1122-1124 and provides security services to implement host-based encryption (HBE), integrity and data encryption (IDE), and other host-side security functions described herein.

[0102] In some implementations, CXL switch 1110 includes a microcontroller 1140 to process requests from processors 1122-1124 and transmit responses via a Management Component Transport Protocol (MCTP) interconnect. The MCTP interconnect may include a system management bus and / or a PCIe link that supports MCTP (e.g., MCTP over PCIe). According to some embodiments, applications 1191A-1191C, 1192A-1192C executing or otherwise associated with processors 1122-1124 transmit memory requests (e.g., allocation / deallocation requests as described herein) to microcontroller 1140 via the MCTP interconnect 1180.

[0103] The MCTP payload of a memory allocation / deallocation request may include one or more of the following: an application ID indicating the corresponding application 1191A-1191C, 1192A-1192C generating the request; a host ID indicating the corresponding processor 1122-1124; a bit indicating whether the requesting application is a secure / trusted application (e.g., 1191A-1191C) or a non-secure / untrusted application (e.g., 1192A-1192C); the amount of memory requested (e.g., 2 GB, 4 GB, etc.); and a bit indicating whether the request is to bind / allocate memory or to unbind / deallocate memory.

[0104] The virtual-to-physical binding circuitry / logic 1050 includes a plurality of Multiple Logical Device (MLD) ports 1151 that provide access to a corresponding plurality of memory expansion devices D1-D9. Once portions of the memory expansion devices D1-D9 are mapped, memory access requests (e.g., load or store operations) originating from applications 1191A-1191C, 1192A-1192C can be directed toward the memory expansion devices D1-D9.

[0105] In some embodiments, in response to a request for expanded memory from a secure application 1191A-1191C or a non-secure application 1192A-1192C, the microcontroller 1140 allocates one or more logical devices (LDs) that are logical partitions of the memory expansion devices D1-D9. In the example shown, each device D1-D9 is divided into 16 logical devices LD0-LD15 that can be accessed via a corresponding MLD port 1151. The microcontroller 1140 can allocate a set of LDs to applications individually as needed. In some implementations, the microcontroller 1140 maintains a mapping of applications to logical devices within a super root of trust (SROT) 1142 that includes a lookup table, and the lookup table is updated after each allocation or deallocation of a logical device. Figure 11A The different shading patterns in FIG. 1 indicate a set of example assignments between applications 1191A-1191C, 1192A-1192C and logical devices (i.e., having the same pattern indicating that logical devices are assigned to corresponding applications). For example, LD0 of device D1 is assigned to non-secure application 1192B, and LD1 of device D1 and LD1 of device D2 are assigned to non-secure application 1192A.

[0106] In some embodiments, when a range check indication request issued by processors 1122-1124 is directed to one of CXL memory devices D1-D9, the memory access is encrypted using host-based encryption circuitry / logic 1051A-1051B according to the corresponding host security circuitry / controller 1160A-1160C. Encryption circuitry / logic 1051C may implement host-based encryption or another type of encryption (e.g., SME / AES-XTS). For example, when CXL memory expansion devices D1-D9 do not natively support IDE and / or memory encryption, host-based encryption (and / or other types of encryption, if used) may be performed.

[0107] In order to seamlessly scale up and down the architecture with multiple processors 1122-1124, some embodiments implement techniques that allow processors to interact within a single trust boundary. In these embodiments, the SROT 1142 of the microcontroller 1140 brings the various processors / nodes into a common trust boundary.

[0108] In one embodiment, to maintain the integrity of application data routed by each of processors 1122-1124 to CXL switch 1110 for allocation / deallocation in any expansion devices D1-D9, host-based encryption circuitry / logic 1051A-1051B and encryption circuitry 1051C of each respective processor tracks cache line ownership based on host physical addresses. As further described below, for example, HBE circuitry / logic 1051A-1051B and encryption circuitry 1051C may each cache portions of access control tables (ACTs) 1052A-1052C (e.g., generated by the corresponding processor's secure arbitration mode (SEAM)) to determine the characteristics of each application. If the application is trusted, the HBE may access corresponding keys (e.g., in some embodiments, TDXKEY and MKTME keys) to encrypt / decrypt memory accesses. For example, a processor in secure arbitration mode (SEAM) generates entries in ACTs 1052A-1052C while allocating host physical addresses for trusted workloads.

[0109] In some embodiments, each processor 1122-1124 has its own separate root of trust for establishing a secure flow within its respective host boundary. For example, the root of trust may be the host security circuitry 1160A-1160C of each respective processor 1122-1124. Each instance of the security circuitry 1160a-1160c may be a secure microcontroller, but various types of security circuitry may also be used (e.g., a converged security and management engine (CSME), an embedded security element (ESE), a secure processor (SP), or other dedicated security circuitry).

[0110] exist Figure 11A In the embodiment of the present invention, SROT 1142 managed by microcontroller 1140 establishes a trust boundary between processors 1122-1124, where the processors exchange messages with microcontroller 1140 via MCTP interconnect 1180 (e.g., SMBUS interface or PCIe interface). As further described below, the trust boundary is established via information provided in MCTP packets.

[0111] Figure 11B The figure shows an example of a lookup table 1147, which includes multiple entries (rows), each of which is associated with a specific memory allocation. The fields in each entry include: a device / logical device identifier (ID) for indicating a specific logical device of a specific device; an application ID for indicating the application to which the logical device is assigned; a host ID for indicating the associated node / processor; a trust bit for indicating whether the application is trusted or untrusted; an indication of the memory size requested by the application; a bind / unbind indication (indicating whether the memory is allocated / deallocated); and a mapping status, including a new / existing entry bit and a trusted / untrusted bit.

[0112] Figure 12 A method according to one embodiment is illustrated in FIG. The illustrated method can be implemented using the various architectural details described herein, but is not limited to these particular processor or system architectures.

[0113] At 1200, a corresponding application is classified based on whether it is trusted or untrusted. For example, an application may be trusted when the processor operates in Secure Arbitration Mode (SEAM) and generates trusted API calls (e.g., TDX API calls). In this environment, SEAMOPS Leaf0 (capability) can be used to append API call information to the MCTP format described herein. For example, an untrusted application operating in a virtualized execution environment may generate API calls (e.g., VMX API calls) initiated by the hypervisor of the virtualized environment that manages the application.

[0114] At 1201, in response to the request, a corresponding path to physical memory is determined based on the host physical address associated with the request. For example, a table or other data structure maintains a mapping of host physical addresses to different memory types, which may include, for example, CXL type 2 memory devices, CXL type 3 memory devices, and host-connected memory devices (e.g., system DRAM such as HBM or DDR). As described above, in some embodiments, a range check is performed by source address decoding logic 1003 in home agent 1002 (e.g., using range register 1004).

[0115] If the request is directed to host memory (e.g., system memory such as HBM or DDR DRAM), then at 1210, the memory request is transmitted to the corresponding memory controller of the processor. Note that this decision may have been made further upstream in the processing pipeline (e.g., before the start of the illustrated method). Therefore, the option of transmitting to the memory controller is optional at this point (as indicated by the dashed line).

[0116] If it is determined at 1202 that the request is directed to a CXL Type 2 device, security may be implemented using integrity and data encryption (IDE) at 1203, and a trust bit (e.g., a TEE bit) may be set to indicate whether the application is trusted or untrusted (based on the determination at 1200). If the request is directed to a CXL Type 3 device, security may be implemented using host-based encryption (HBE) at 1204, and the trust bit may be set to trusted or untrusted. Host-based encryption circuitry 1051A-1051B and encryption circuitry 1051C may determine the characteristics of the application via a cache access control table (ACT) located within the HBE boundary. If the application is trusted, the HBE may rely on corresponding keys (e.g., TDXKEY and MKTME keys) to encrypt / decrypt memory accesses. For example, a processor in secure arbitration mode (SEAM) may generate an entry in the ACT while allocating host physical addresses for trusted workloads.

[0117] At 1205, an MCTP payload is generated along with the data encrypted using host-based encryption. In some implementations, this payload information is appended by secure circuitry and transmitted via MCTP over SMBUS or PCIe. In one embodiment, each processor 1122-1124 uses a specific payload format (e.g., generated by SEAM OPS) to communicate with the super root of trust / LUT 1142 and establish a connection. As described above, the payload information may include one or more of the following: application ID, host ID, secure / non-secure indication, amount of memory requested, and indication of a bind / unbind request.

[0118] At 1206, microcontroller 1140 in CXL switch 1110 decodes the MCTP packet, which includes the attached payload. Microcontroller 1140 then accesses the appropriate entry in the lookup table based on the MCTP packet. At 1207, microcontroller 1140 allocates or de-allocates a logical device based on the payload.

[0119] Figure 13 The diagram illustrates one embodiment of the circuitry and logic between trusted / untrusted software 1390-1392, CXL device memory 1330, and HBM / DDR memory 1357. In this example, the software includes a VMX virtual execution environment 1390 running a guest OS, and two trusted virtual machines 1391-1392 running trusted operating systems, all managed by a hypervisor 1310. Hypervisor 1310 makes secure and non-secure API calls provided by dynamic memory range (DMR) logic 1312. By way of example and not limitation, trusted API calls may include TDX API calls in secure mode and VMX API calls in non-secure mode.

[0120] The API call is processed by secure arbitration mode (SEAM) logic 1312 operating on one or more processor cores 1312 (identified via a unique CPU ID). Source address decoding circuitry 1303 (e.g., shared by the processor cores 1312 in an area external to the core) decodes the address based on range information specified in one or more range registers 3004.

[0121] Source address decoding circuitry 1312 can access range register 3004 to determine whether the request is to be directed to processor system memory, such as HBM / DDR memory 1357 or CXL device memory 1330. If the request is to be directed to processor system memory, the MCTP packet format may not be required. For example, when the host physical address (HPA) assigned by security arbitration mode logic 1312 belongs to system memory 1357 and the corresponding trust bit indicates that the application data is to be encrypted or decrypted, the MTCP format is not required because the memory is within the processor's trust boundary. In some embodiments, the memory controller 1355 or the memory security engine in fabric 1350 manages encryption and decryption.

[0122] If the request is directed to CXL device memory 1330, the MCTP packet format is required. In these cases, the payload information is generated by the processor's HBE / IDE logic 1318 and sent via MCTP. This embodiment can be configured by the secure arbitration mode logic 1312, which programs the keys required for encryption / decryption. Per-cache line ownership information is obtained by implementing a local cache in the HBE / IDE logic 1318, which is coherent with system memory 1357 using ISA-based techniques (e.g., software-based coherency).

[0123] In one embodiment, PCIE / CXL control logic 1320 uses MCTP over PCIE to pass MCTP payloads to CXL switch 1210. Microcontroller 1240 in CXL switch 1210 manages the lookup table as described herein to establish a super root of trust and responsively allocates and deallocates logical devices of CXL device memory 1330 (e.g., as described with respect to Figure 11A described above).

[0124] In some embodiments, CXL switch platform enumeration operates as follows. During the enumeration process of multiple processors accessing a pool of CXL memory devices, the CXL switch will perform at least some of the enumeration operations. Specifically, referring again to Figure 11A The CXL switch configuration space 1155 (e.g., one or more configuration registers) stores data indicating the number of hosts connected to the CXL upstream port. The microcontroller 1140 of the CXL switch 1110 manages the detection of CXL Type 3 MLD devices D1-D9 and performs enumeration of these devices. The microcontroller 1140 also determines the memory size and logical partitioning of each device (e.g., logical devices 0-15 in each illustrated device) by accessing the configuration space 1155 during enumeration.

[0125] Additional details of allocation and deallocation by a CXL switch according to some embodiments are provided below. These details may be implemented based on the MCTP payload received by the microcontroller.

[0126] The following two bits indicate the possible states of an entry in the lookup table when it is first updated based on a received request: 00: Unmapped area 01: Non-secure mapping 10: Security Mapping 11: Undefined

[0127] In one embodiment, the microcontroller 1140 follows a set of rules when transitioning between certain states. For example, for any bound or unbound stream, a transition from the secure mapped state (10) to the non-secure mapped state (01) represents the unbinding of a slot for a secure application in a Type 3 MLD device and the binding of the same slot to a non-secure application. Therefore, one embodiment of the microcontroller 1140 performs this transition only after initiating a stream flush (e.g., flushing data from a secure application's slot before assigning it to a non-secure application). This prevents secure application data from being exposed to non-secure applications during reallocation.

[0128] Additionally, the transition from the non-secure mapping (01) to the secure mapping (10) represents the unbinding of the slot for the non-secure application in the Type 3 MLD device and the binding of the same slot to the secure application. In this case, the microcontroller similarly performs the transition only after initiating a stream flush to prevent the non-secure application data from being exposed to the secure application during the reallocation. This implementation addresses the threat model of a maliciously configured non-secure application attempting to gain secure mode privileges. The stream flush operation ensures that the secure application code will not be exposed.

[0129] In addition to the aforementioned set of rules, one embodiment of the microcontroller 1140 ensures that a processor's request for memory does not exceed the available device memory. In response to such a request, the microcontroller 1140 notifies the requesting processor of an error via an MCTP packet.

[0130] In one implementation, the aforementioned host-based encryption (HBE) circuitry / logic 1051A-1051B (and potentially 1051C) operates as follows. Devices attached to the processor are accessed over a CXL link that is not inherently encrypted. For example, CXL Type 3 expansion devices do not natively support SPDM sessions or other encryption protocols. In some embodiments, encryption is implemented using host-based encryption (HBE) as described herein. In addition to encryption, in some embodiments, HBE logical integrity is used to maintain data integrity by tracking ownership across cache line entries. These transactions are provided with a key for the host physical address targeted in the CXL Type 3 MLD device side of the memory. In some embodiments, HBE circuitry / logic 1051A-1051B (and potentially 1051C) supports the following features for Type 3 CXL devices: AES256XTS encryption (eg, using MKTME / MSE circuitry for encryption blocks). Multi-key support, such as available in the case of 1K keys Range-based key selection, such as maintaining a key for each 4KB page. • Logical integrity, achieved by tracking a trusted execution environment (eg, trust domain) ownership bit for each cache line entry. An access control table (ACT), such as an ACTRR stored in a protected area in DRAM, which indicates whether each host physical address maps to a trusted or untrusted domain. As described above, the HBE can maintain a local cache of access control table entries to select the key required to encrypt / decrypt data packets in the cxl.mem path.

[0131] In some embodiments, when a trust domain is created, SEAM logic 1312 calls the processor's security hardware to pick a random KEYID for the assigned trust domain page. The processor / core memory management unit (MMU) issues a secure EPT. The KEYID is selected based on this scheme. For example, the guest physical address (GPA) will indicate that this is a private page, so it goes into the secure EPT. The secure EPT then assigns a KEYID to the incoming host physical address (HPA).

[0132] The core 1312 transmits the request using this HPA[51:0] as [TDXKEYID, HPA[41:0]], and the key lookup table uses this KEYID for the HPA belonging to the created trust domain. In some implementations, the SEAM logic 1312 also ensures that no two private trust domain spaces share the same key (e.g., during programming, SEAM performs a conditional check to ensure that each KEYID is unique).

[0133] The source address decoding logic 3003 checks the range register(s) 3004 to determine if the incoming address belongs to a type 2 range or a type 3 range or system memory (e.g., HBM). Once this is determined, the following programming is performed:

[0134] HBM scope: MKTME implementation of MSE IP behavior.

[0135] Type 3 Scope: As described above, HBE is enabled for memory devices attached to a Type 3 scope. The HBE circuitry (e.g., 1318) uses the HPA flowed with the KEYID as part of the MSB bit to select a key slot. The trust domain now encrypts / decrypts data in the CXL Type 3 device (e.g., 1330) using the key assigned by the SEAM logic 1312 programming flow. Thus, a trusted private space is allocated in the CXL Type 3 device.

[0136] In some implementations, the HBE circuitry 1318 reuses the trust bit (i.e., "TEE bit") generated by the source address decoding (SAD) circuitry 3003 to determine whether a request is secure or non-secure. In a TDX implementation, the TDX detection register detects whether a particular transaction is for TDX (trusted) or non-TDX (non-trusted). Thus, TDX detection can be used to identify whether a private key or a shared key is required for trusted private memory access.

[0137] In some embodiments, an access control table (e.g., ACTRR) is implemented as a carve-out in system memory maintained by a specific SEAM instruction (ACT_en or ACT enable). The HBE circuitry 1318 maintains a local cache with ACT entries and performs a local cache lookup to identify the desired entry. The HBE circuitry 1318 determines whether an incoming HPA is associated with trusted or untrusted ownership based on a cache entry that may be maintained for each 4KB page. For example, each ACT entry may store an "owner bit" for each corresponding 4K memory range in memory, defining the "ownership" of the 4K page (e.g., TD / private data or non-secure "shared data"). Software coherency is achieved between the ACT in memory and the HBE circuitry local cache.

[0138] The cipher block uses the KEYID along with the trust bit indication to select a key for encryption of trusted private data stored in the CXL Type 3 device memory.

[0139] Merging the Super Root of Trust concept with the HBE circuitry implementation described herein allows multiple hosts to operate under a single root of trust for secure workloads in CXL Type 3MLD devices implemented using CXL 2.0 switches. Supporting encryption on the host can reduce the cost of CXL devices, enable the integration of third-party accelerator devices into heterogeneous implementations for confidential computing workloads (including AI / ML workloads), and extend the TDX connectivity architecture to the broader CXL market.

[0140] Figures 14A-14B The diagram illustrates a specific implementation that includes some of the security techniques described herein. A core 1101 is shown generating a request to a home agent 1102. Applications running on a host (e.g., core 1101) can initiate communications with either MKTME_KEYID or TDXKEYID based on the application's privilege state (e.g., from VMX or TD mode). The home agent first determines whether the request is for a Type 3 or Type 2 memory device. If the request is directed to a Type 3 device and is a TDX request, as determined at 1105, a lookup is performed in the TD key index table 1410 to determine the KeyID using the host physical address (HPA) of the request. At 1425, the KeyID and HPA are used for a key lookup to determine an encryption key, which is used to encrypt / decrypt 1430 to access a CXL Type 3 MLD device 1433 via a CXL 2.0 switch 1432.

[0141] If the request is a non-TDX request, the MKTME_KeyID and HPA are used to access the VMM key index table 1440 to locate the KeyID, which is then used to perform a lookup in the key lookup table 1445 to determine the encryption key associated with the HPA, which is then used to encrypt / decrypt 1430 to access the CXL Type 3 MLD device 1433 via the CXL 2.0 switch 1432.

[0142] Figure 14B The diagram illustrates the operation of a Type 2 CXL device. If the request is not a TDX request (e.g., not from a trusted domain), the request passes through host IO processing circuitry 1441 and bypasses encryption with MKTME circuitry 1445 to access Type 2 device 1454 via PCIe / CXL control circuitry 1451. If the request is a TDX request (e.g., from a trusted domain), the request passes through host IO processing circuitry 1446, which accesses secure EPT 1443 for memory translation. This request again bypasses encryption with MKTME circuitry 1445. However, PCIe / CXL control circuitry 1452 supports IDE link encryption for requests from a trusted domain targeting Type 2 devices. Therefore, access to Type 2 device 1454 is encrypted / decrypted via the IDE link.

[0143] Note that the terms "circuit" and "circuitry" are used interchangeably herein. As used herein, these terms, as well as the term "logic," are used alone or in any combination to refer to analog circuitry, digital circuitry, hardwired circuitry, programmable circuitry, processor circuitry, microcontroller circuitry, hardware logic circuitry, state machine circuitry, and / or any other type of physical hardware component. Embodiments may be used in many different types of systems. For example, in one embodiment, a communication device may be arranged to perform the various methods and techniques described herein. Of course, the scope of the present invention is not limited to communication devices, and instead, other embodiments may relate to other types of apparatus for processing instructions, or one or more machine-readable media including instructions that, in response to being executed on a computing device, cause the device to perform one or more of the methods and techniques described herein.

[0144] Example

[0145] The following are example implementations of different embodiments of the invention.

[0146] Example 1. An apparatus comprising: a plurality of cores of a host processor; a host processor memory subsystem for providing access to a host processor memory; a home agent for providing access to the host processor memory subsystem and an expansion memory subsystem by the plurality of cores, the home agent comprising a source address decoder for decoding memory requests generated from the plurality of cores to determine whether the memory requests are directed to the host processor memory subsystem or the expansion memory subsystem; the host processor memory subsystem for performing encryption and decryption of memory requests directed to the host processor memory subsystem based on a first key identified via key lookup circuitry; and host-based security circuitry for encrypting and decrypting memory requests directed to the expansion memory subsystem, the host-based security circuitry for performing encryption and decryption based on a second key stored in a cache maintained by the host-based security circuitry.

[0147] Example 2. The apparatus of Example 1, wherein the extended memory subsystem is to couple the plurality of cores to a CXL Type 3 Multi-Logic Device (MLD) memory.

[0148] Example 3. The apparatus of Example 1 or 2, wherein the expanded memory subsystem comprises: switching circuitry configurable to couple at least one memory expansion device to the plurality of cores; a microcontroller associated with the switching circuitry configured to securely partition the at least one memory expansion device into a plurality of logical devices and assign subsets of the logical devices to applications executing on the plurality of cores; and tracking circuitry storing data associated with the subsets of the logical devices assigned to the applications.

[0149] Example 4. The apparatus of Example 1, wherein the tracking circuitry comprises a lookup table to store a plurality of entries, each entry to store data associated with a subset of the subset of logic devices assigned to one of the applications.

[0150] Example 5. The apparatus of any of Examples 1-4, wherein each entry is configured to store one or more of: an application identifier (ID) indicating an application associated with a corresponding extended memory request; a host ID indicating a corresponding processor; a bit indicating whether the corresponding application is a trusted application or an untrusted application; an indication of an amount of requested extended memory; and one or more bits indicating whether the request is to bind / allocate or unbind / deallocate memory of an extended memory device.

[0151] Example 6. The apparatus of any of Examples 1-5, wherein the lookup table is used as a shared root of trust for the host processor and one or more additional processors coupled to the switching circuitry.

[0152] Example 7. The apparatus of any of Examples 1-6, wherein the host-based security comprises host-based cryptographic circuitry to encrypt and decrypt memory requests directed to the expansion memory subsystem based on a second key, wherein the cache comprises a cache of entries of an access control table (ACT), the ACT being stored in a protected area of ​​memory accessed via the host processor memory subsystem.

[0153] Example 8. The apparatus of any of Examples 1-7, wherein the ACT is to store keys associated with trusted applications executing on the plurality of cores.

[0154] Example 9. A method includes: receiving memory requests from a plurality of cores; decoding the memory requests generated from the plurality of cores to determine whether the memory requests are directed to a host processor memory subsystem or an expansion memory subsystem; performing, by the host processor memory subsystem, encryption and decryption for the memory requests directed to the host processor memory subsystem based on a first key identified via key lookup circuitry; and performing, by the host-based security circuitry, encryption and decryption for the memory requests directed to the expansion memory subsystem based on a second key stored in a cache maintained by the host-based security circuitry.

[0155] Example 10. The method of Example 9, wherein the extended memory subsystem is to couple the plurality of cores to a CXL Type 3 Multi-Logic Device (MLD) memory.

[0156] Example 11. The method of Examples 9-10, wherein the expanded memory subsystem includes switching circuitry configurable to couple the at least one memory expansion device to the plurality of cores, the method further comprising: securely partitioning the at least one memory expansion device into a plurality of logical devices; allocating subsets of the logical devices to applications executing on the plurality of cores; and tracking the allocation of the subsets of the logical devices by storing data associated with the subsets of the logical devices allocated to the applications.

[0157] Example 12. The method of any of Examples 9-11, wherein the data associated with the subset of logical devices is stored in entries of the lookup table, each entry storing data associated with one of the subsets of logical devices assigned to one of the applications.

[0158] Example 13. The method of any of Examples 9-12, wherein each entry is used to store one or more of: an application identifier (ID) indicating an application associated with the corresponding extended memory request; a host ID indicating the corresponding processor; a bit indicating whether the corresponding application is a trusted application or an untrusted application; an indication of an amount of requested extended memory; and one or more bits indicating whether the request is to bind / allocate or unbind / deallocate memory of the extended memory device.

[0159] Example 14. The method of any of Examples 9-13, further comprising: using the lookup table as a shared root of trust for the host processor and one or more additional processors coupled to the switching circuitry.

[0160] Example 15. The method of any of Examples 9-14, wherein the host-based security includes host-based cryptographic circuitry to encrypt and decrypt memory requests directed to the expansion memory subsystem based on the second key, wherein the cache includes a cache of entries of an access control table (ACT), the ACT stored in a protected area of ​​memory accessed via the host processor memory subsystem.

[0161] Example 16. The method of any of Examples 9-15, wherein the ACT is to store keys associated with trusted applications executing on the plurality of cores.

[0162] Example 17. A machine-readable medium having program code stored thereon, the program code, when executed by a machine, causing the machine to perform operations, the operations comprising: receiving memory requests from a plurality of cores; decoding the memory requests generated from the plurality of cores to determine whether the memory requests are directed to a host processor memory subsystem or an expansion memory subsystem; performing, by the host processor memory subsystem, encryption and decryption of the memory requests directed to the host processor memory subsystem based on a first key identified via key lookup circuitry; and performing, by the host-based security circuitry, encryption and decryption of the memory requests directed to the expansion memory subsystem based on a second key stored in a cache maintained by the host-based security circuitry.

[0163] Example 18. The machine-readable medium of Example 17, wherein the expansion memory subsystem is to couple the plurality of cores to a CXL Type 3 Multiple Logic Device (MLD) memory.

[0164] Example 19. The machine-readable medium of Example 17 or 18, wherein the expanded memory subsystem includes switching circuitry configurable to couple at least one memory expansion device to the plurality of cores, the machine-readable medium further comprising program code for causing a machine to: securely partition the at least one memory expansion device into a plurality of logical devices; assign subsets of the logical devices to applications executing on the plurality of cores; and track the assignment of the subsets of the logical devices by storing data associated with the subsets of the logical devices assigned to the applications.

[0165] Example 20. The machine-readable medium of any of Examples 17-19, wherein the data associated with the subset of logical devices is stored in entries of the lookup table, each entry storing data associated with one of the subset of logical devices assigned to one of the applications.

[0166] Example 21. The machine-readable medium of any of Examples 17-20, wherein each entry is to store one or more of: an application identifier (ID) indicating an application associated with the corresponding extended memory request; a host ID indicating the corresponding processor; a bit indicating whether the corresponding application is a trusted application or an untrusted application; an indication of an amount of requested extended memory; and one or more bits indicating whether the request is to bind / allocate or unbind / deallocate memory of an extended memory device.

[0167] Example 22. The machine-readable medium of any of Examples 17-21, further comprising program code that causes the machine to use the lookup table as a shared root of trust for the host processor and one or more additional processors coupled to the switching circuitry.

[0168] Example 23. The machine-readable medium of any of Examples 17-22, wherein the host-based security comprises host-based cryptographic circuitry to encrypt and decrypt memory requests directed to the expansion memory subsystem based on a second key, wherein the cache comprises a cache of entries of an access control table (ACT), the ACT stored in a protected area of ​​memory accessed via the host processor memory subsystem.

[0169] Example 24. The machine-readable medium of any of Examples 17-23, wherein the ACT is to store keys associated with trusted applications executing on the plurality of cores.

[0170] Embodiments may be implemented in code and may be stored on a non-transitory storage medium having stored thereon instructions that may be used to program a system to perform the instructions. Embodiments may also be implemented in data and may be stored on a non-transitory storage medium that, if used by at least one machine, causes the at least one machine to fabricate at least one integrated circuit to perform one or more operations. Further embodiments may be implemented in a computer-readable storage medium comprising information that, when fabricated in a SoC or other processor, configures the SoC or other processor to perform one or more operations. The storage medium may include, but is not limited to: any type of disk, including a floppy disk, an optical disk, a solid state drive (SSD), a compact disk read-only memory (CD-ROM), a compact disk rewritable (CD-RW), and a magneto-optical disk; a semiconductor device, such as a read-only memory (ROM), a random access memory (RAM) such as a dynamic random access memory (DRAM) and a static random access memory (SRAM), an erasable programmable read-only memory (EPROM), a flash memory, an electrically erasable programmable read-only memory (EEPROM); a magnetic or optical card; or any other type of medium suitable for storing electronic instructions.

[0171] While the present invention has been described with reference to a limited number of embodiments, numerous modifications and variations will occur to those skilled in the art. It is intended that the appended claims cover all such modifications and variations as fall within the true spirit and scope of the present invention. < / coherence> < / coherence>

Claims

1. A device comprising: Multiple cores of the host processor; a host processor memory subsystem for providing access to host processor memory; a home agent for providing access by the plurality of cores to the host processor memory subsystem and the expansion memory subsystem, the home agent comprising a source address decoder for decoding memory requests generated from the plurality of cores to determine whether the memory requests are directed to the host processor memory subsystem or the expansion memory subsystem; the host processor memory subsystem to perform encryption and decryption of memory requests directed to the host processor memory subsystem based on a first key identified via the key lookup circuitry; as well as Host-based secure circuitry is configured to encrypt and decrypt memory requests directed to the expansion memory subsystem, the host-based secure circuitry being configured to perform the encryption and decryption based on a second key stored in a cache maintained by the host-based secure circuitry.

2. The device according to claim 1, wherein The extended memory subsystem is configured to couple the plurality of cores to a CXL type 3 multi-logic device (MLD) memory.

3. The device according to claim 1 or 2, characterized in that The extended memory subsystem includes: switching circuitry configurable to couple at least one memory expansion device to the plurality of cores; a microcontroller associated with the switching circuitry for securely partitioning the at least one memory expansion device into a plurality of logical devices and assigning subsets of the logical devices to applications executing on the plurality of cores; and Tracking circuitry stores data associated with the subset of the logic devices assigned to the application.

4. The device according to claim 3, characterized in that The tracking circuitry includes a lookup table for storing a plurality of entries, each entry for storing data associated with a subset of the subsets of the logic devices assigned to one of the applications.

5. The device according to claim 4, characterized in that Each entry is used to store one or more of the following: an application identifier ID, the application ID indicating an application associated with a corresponding extended memory request; a host ID, the host ID indicating a corresponding processor; a bit for indicating whether the corresponding application is a trusted application or an untrusted application; an indication of the amount of requested extended memory; and one or more bits for indicating whether the request is to bind / allocate or unbind / deallocate memory of the expansion memory device.

6. The device according to claim 5, characterized in that The lookup table is used as a shared root of trust for the host processor and one or more additional processors coupled to the switching circuitry.

7. The device according to any one of claims 1 to 6, characterized in that The host-based security includes a host-based cryptographic circuit system for encrypting and decrypting memory requests directed to the expansion memory subsystem based on the second key, wherein the cache includes a cache of entries of an access control table ACT, the ACT being stored in a protected area of ​​memory accessed via the host processor memory subsystem.

8. The device according to claim 7, wherein The ACT is used to store keys associated with trusted applications executed on the plurality of cores.

9. A method comprising: receiving memory requests from a plurality of cores; decoding memory requests generated from the plurality of cores to determine whether the memory requests are directed to a host processor memory subsystem or an expansion memory subsystem; performing, by the host processor memory subsystem, encryption and decryption of memory requests directed to the host processor memory subsystem based on a first key identified via key lookup circuitry; as well as Encryption and decryption are performed, by host-based secure circuitry, on memory requests directed to the expansion memory subsystem based on a second key stored in a cache maintained by the host-based secure circuitry.

10. The method according to claim 9, wherein The extended memory subsystem is configured to couple the plurality of cores to a CXL type 3 multi-logic device (MLD) memory.

11. The method according to claim 9 or 10, wherein: The expansion memory subsystem includes switching circuitry configurable to couple at least one memory expansion device to the plurality of cores, the method further comprising: securely partitioning the at least one memory expansion device into a plurality of logical devices; allocating a subset of the logical devices to applications executing on the plurality of cores; and Allocation of the subset of the logical devices is tracked by storing data associated with the subset of the logical devices allocated to the application.

12. The method according to claim 11, wherein The data associated with the subset of the logical devices is stored in entries of a lookup table, each entry being used to store data associated with one of the subsets of the logical devices assigned to one of the applications.

13. The method according to claim 12, wherein: Each entry is used to store one or more of the following: an application identifier ID, the application ID indicating an application associated with a corresponding extended memory request; a host ID, the host ID indicating a corresponding processor; a bit for indicating whether the corresponding application is a trusted application or an untrusted application; an indication of the amount of requested extended memory; and one or more bits for indicating whether the request is to bind / allocate or unbind / deallocate memory of the expansion memory device.

14. The method of claim 13, further comprising: The lookup table is used as a shared root of trust for the host processor and one or more additional processors coupled to the switching circuitry.

15. The method according to any one of claims 9 to 14, characterized in that The host-based security includes a host-based cryptographic circuit system for encrypting and decrypting memory requests directed to the expansion memory subsystem based on the second key, wherein the cache includes a cache of entries of an access control table ACT, the ACT being stored in a protected area of ​​memory accessed via the host processor memory subsystem.

16. The method according to claim 15, wherein The ACT is used to store keys associated with trusted applications executed on the plurality of cores.

17. A machine-readable medium having program code stored thereon, the program code, when executed by a machine, causing the machine to perform operations comprising: receiving memory requests from a plurality of cores; decoding memory requests generated from the plurality of cores to determine whether the memory requests are directed to a host processor memory subsystem or an expansion memory subsystem; performing, by the host processor memory subsystem, encryption and decryption of memory requests directed to the host processor memory subsystem based on a first key identified via key lookup circuitry; as well as Encryption and decryption are performed, by host-based secure circuitry, on memory requests directed to the expansion memory subsystem based on a second key stored in a cache maintained by the host-based secure circuitry.

18. The machine-readable medium of claim 17, wherein: The extended memory subsystem is configured to couple the plurality of cores to a CXL type 3 multi-logic device (MLD) memory.

19. The machine-readable medium of claim 17 or 18, wherein: The expansion memory subsystem includes switching circuitry configurable to couple at least one memory expansion device to the plurality of cores, the machine-readable medium further including program code for causing a machine to: securely partitioning the at least one memory expansion device into a plurality of logical devices; allocating a subset of the logical devices to applications executing on the plurality of cores; as well as Allocation of the subset of the logical devices is tracked by storing data associated with the subset of the logical devices allocated to the application.

20. The machine-readable medium of claim 19, wherein: The data associated with the subset of the logical devices is stored in entries of a lookup table, each entry being used to store data associated with one of the subsets of the logical devices assigned to one of the applications.