Virtualized root of trust in distributed computing system

By virtualizing the root of trust (vRoT) on the management controller, the security risks and management complexity of ERoT chips in distributed computing systems are resolved, achieving more efficient security management and cost reduction.

CN120974493APending Publication Date: 2025-11-18NVIDIA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510624623.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-16
Filing Date
2025-05-15
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In distributed computing systems, physically deploying multiple ERoT chips increases security risks, management complexity, and costs, and the I/O and memory limitations of existing ERoT chips cannot meet the needs of large server systems.

Method used

By virtualizing the root of trust (vRoT) on the management controller, the high I/O and large memory resources of the management controller are utilized to achieve secure management of multiple application processors (APs), reduce the number of physical ERoT chips, and arbitrate I/O access and memory isolation through virtualization technology.

Benefits of technology

It reduces security risks and costs, improves system manageability and reliability, simplifies the distribution and enforcement of security policies, and reduces failure rates and bill of materials costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974493A_ABST
    Figure CN120974493A_ABST
Patent Text Reader

Abstract

The invention discloses a virtualized root of trust in a distributed computing system. A system includes a plurality of application processors (APs), a plurality of flash memory devices associated with the plurality of APs, and a plurality of multiplexers, each for selectively coupling one of the plurality of flash memory devices to one of the plurality of APs. The controller is operatively coupled to the plurality of multiplexers and provides a trusted execution environment to execute a virtual root of trust (vRoT) application for each respective AP of the plurality of APs. Each vRoT application accesses a corresponding one or more of the plurality of flash memory devices via a corresponding one or more of the plurality of multiplexers.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] At least one embodiment relates generally to distributed computing systems and more particularly, but not exclusively, to virtual root-of-trust (vRoT) in distributed computing systems. BACKGROUND

[0002] Some accelerated systems designed as distributed computing systems or platforms deploy many application processors (APs), such as modern graphics processing units (GPUs), central processing units (CPUs), and high-speed interconnects of GPUs and CPUs. For example, these accelerated systems support supercomputing for enterprise applications and computing functions related to artificial intelligence (AI).

[0003] These distributed computing systems tend to include multiple flash memories, often referred to as reprogrammable non-volatile memories, where each flash memory is used to store firmware and data for a corresponding AP of a plurality of APs. For example, known flash memory devices provide secure boot support and other configuration parameters for the operation of each AP. An independent root-of-trust (ERoT) is coupled with the flash memory devices to protect the flash memory devices and support secure operations related to each AP.

[0004] Such ERoT devices (also referred to as ERoT chips) can be physically deployed on servers and data center platforms as hardware security modules (HSMs), trusted platform modules (TPMs), or other such hardware modules. HSMs can provide a secure environment for cryptographic operations and key storage, while TPMs are often embedded in server hardware to securely store keys, digital certificates, and other sensitive data. These ERoT devices can ensure a secure boot process by attesting the integrity of server firmware, boot loaders, or operating systems. The ERoT devices can check the cryptographic signatures of these components and provide proofs of component integrity, for example, to ensure that no tampering has occurred. In data centers, ERoT devices can be implemented at a network level to secure network security tools, such as firewalls and intrusion detection and / or prevention systems.

[0005] High-availability data centers sometimes deploy ERoT devices to ensure continuous operation even when a device fails. However, the distributed use of many ERoT devices creates multiple points of failure and security risks when managing secure operations of multiple APs in a distributed computing system. Furthermore, the distribution and enforcement of security policies are more challenging when spread across multiple ERoT devices. BRIEF DESCRIPTION OF DRAWINGS

[0006] Various embodiments according to the present disclosure will be described with reference to the drawings, in which:

[0007] Figure 1 is a schematic block diagram of an example distributed computing system supporting a virtualized root-of-trust (vRoT) application of multiple APs from a trusted execution environment (TEE) according to some embodiments;

[0008] Figure 2 is a schematic block diagram of an example distributed computing system supporting a vRoT application of multiple APs from a TEE according to additional embodiments;

[0009] Figure 3 is a schematic block diagram of an example distributed computing system supporting a vRoT application from a management controller executing an unsecure kernel and a trusted operating system (OS) according to some embodiments;

[0010] Figure 4 is a schematic block diagram of an example distributed computing system supporting a vRoT application from a management controller executing an unsecure kernel, a trusted OS, and a secure kernel operation according to some embodiments;

[0011] Figure 5 is a schematic block diagram of an example distributed computing system according to various embodiments, the TEE availability of the system is different from Figure 4 ;

[0012] Figure 6 is a flowchart of an example method for performing a secure boot of an AP using a management controller according to some embodiments;

[0013] Figure 7 is a flowchart of an example method for performing a secure update of an AP through a coupled flash memory device according to some embodiments;

[0014] Figure 8 is a flowchart of an example method for performing a secure attestation of an AP according to some embodiments; and

[0015] Figure 9 is a flowchart of a method for operating a distributed computing system having the disclosed management controller according to at least one embodiment. DETAILED DESCRIPTION

[0016] Further discussion, in certain implementations of distributed computing systems, ERoT chips are physically distributed across multiple platforms of application processors (APs) in order to provide security for devices that do not meet the data center customer security or manageability requirements. On large platforms, this can involve deploying up to dozens of ERoT chips per platform, and the supply chain, un-auditable third party dependencies, bill of materials cost, board space cost, manufacturing defect, and component failure risks required by the ERoT all lead to security risks. For example, a single ERoT failure can result in the entire motherboard or system being returned. Furthermore, ERoT chips are typically low cost, but the risks associated with ERoT chips put billions of dollars of data center business at risk. Additionally, significant integration and customization work is typically required between the ERoT firmware and the components it is securing, leading to further security and manageability risks associated with such customization in addition to the associated costs.

[0017] Furthermore, because ERoT chips have limited IO, and both code space and static random access memory (SRAM) are limited, ERoT chips also cannot be used as platform root of trust for large server systems. ERoTs also have limited memory protection units (MPUs) for memory isolation functionality, which limits task isolation to between 5-8 regions, which requires security compromises in firmware design, and are also limited in processing power due to being built on smaller microcontrollers. These microcontrollers also lack advanced memory protection. Each physical ERoT is typically (but not always) associated with a single AP (e.g., GPU, CPU, etc.) and up to three or four flash memory devices, two of which are used for firmware, another for staging firmware updates, and a possible fourth as a minimum secure version. Multiplying the ERoT and flash memory requirements increases bill of materials cost, and the increased number of flash memory devices leads to increased failure rates.

[0018] Aspects and embodiments of the present disclosure address the above-mentioned deficiencies and other issues with using distributed ERoTs and flash memory by virtualizing ERoTs using a trusted execution environment of a management controller within a distributed computing system (e.g., a network platform or on-chip data center). In some embodiments, these virtual ERoTs (vERoTs) are active component (AC) RoTs (e.g., AC-RoTs). In various embodiments, a platform active root of trust (or PA-RoT) can also be implemented on an existing management chip (e.g., a baseboard management controller (BMC) capable of platform-wide security control management). In this way, each individual ERoT and PA-RoT can be virtualized in a central location of a management controller (or “BMC”) that includes a large number of IO controllers and a much larger memory footprint, while providing the necessary isolation to meet security requirements. The on-board CPU of such a management controller is an order of magnitude more powerful than an ERoT and has a full-featured memory management controller (MMU) and cache that can provide finer-grained isolation for better security. Access to IO can be arbitrated, time-sliced, and virtualized among vRoTs as needed, reducing the amount of IO required and improving the utilization of distributed computing.

[0019] For example, in some embodiments, a system includes a plurality of APs, a plurality of flash memory devices associated with the plurality of APs, and a plurality of multiplexers, each multiplexer to selectively couple one of the plurality of flash memory devices to one of the plurality of APs. A controller (e.g., a management controller or BMC discussed previously) can be operatively coupled to the plurality of multiplexers. The controller can be configured to provide a trusted execution environment (TEE) to execute a virtual root of trust (vRoT) application for each respective AP of the plurality of APs. In embodiments, each vRoT application accesses a corresponding one or more of the plurality of flash memory devices through a corresponding one or more of the plurality of multiplexers. In embodiments, an external processor includes a plurality of interface controllers, each vRoT application corresponding to one of the plurality of interface controllers, interacts with the plurality of multiplexers through the plurality of interface controllers, and the external processor includes control logic to control selection of inputs to the plurality of multiplexers, e.g., such that access to the flash memory devices is multiplexed between the vRoT application and the associated AP.

[0020] In other embodiments, a system includes one or more processor cores (e.g., processing devices) to execute an unsecure kernel, and a trusted operating system (OS) that provides a trusted execution environment (or TEE). A memory management unit (MMU) can be coupled to the one or more processor cores, input / output (IO) hardware can be coupled to the MMU, and a plurality of flash memory devices associated with a plurality of APs of a distributed computing system. In embodiments, the trusted OS executes a vRoT application for each respective AP of the plurality of APs and isolates the IO hardware using the MMU so that the trusted OS securely communicates with the plurality of flash memory devices while preventing intrusion by applications running on the unsecure kernel.

[0021] Accordingly, advantages of systems and methods implemented in accordance with some embodiments of the present disclosure include, but are not limited to, eliminating the need for dozens of ERoT chips and some associated flash memory devices, thereby avoiding the attendant security risks and costs, which have been discussed above. These advantages also include reducing supply chain risks, as the number of BMC (or other controller) chips required is only a fraction of the number of ERoT, and full control over source code and chip audits can be provided. BMCs can also be located on data center security control module (DC-SCM) cards, so if a BMC chip fails, only the module needs to be returned for repair, rather than the entire system. By eliminating ERoT between components (APs) and associated flash memory devices, integration-related issues due to flash monitoring functionality can be eliminated. Moreover, communication between vRoT and its corresponding AP can be simplified to a few well-defined messages. In some embodiments, since BMC (or other controller), vRoT, and PA-RoT functionality can be executed on a single DC-SCM card, staging flash can be consolidated into a single embedded multi-media card (eMMC) storage on the module, thereby reducing the failure rate and bill of materials cost due to flash memory devices. Other advantages will be apparent to those skilled in the art of distributed computing systems and platforms, such as in data centers, which will be discussed below.

[0022] Figure 1 is a schematic block diagram of an example distributed computing system 100 that supports a virtualized root-of-trust (vRoT) application for a plurality of APs from a trusted execution environment (TEE), in accordance with some embodiments. Figure 2is a schematic block diagram of an example distributed computing system 200 according to additional embodiments, which supports vRoT applications from multiple APs of a TEE. In embodiments, systems 100 and 200 include a management controller 102 that is operably coupled to a plurality of multiplexers 113, which are coupled to a plurality of APs 111 and a plurality of flash memory devices 115. In some embodiments, the plurality of APs 111 include, but are not limited to, GPUs, CPUs, data processing units (DPUs), and other computing devices, such as high-speed interconnects. In embodiments, the plurality of flash memory devices 115 are associated with (e.g., coupled to) individual APs of the plurality of APs 111. Further, each multiplexer can selectively couple one flash memory device of the plurality of flash memory devices 115 to one AP of the plurality of APs 111. In some embodiments, the management controller 102 is a baseboard management controller (BMC) or a controller designed specifically for control, security, and / or management of the system 100 or 200. In embodiments, the BMC is located on a DC-SCM card of the system 100 or 200.

[0023] In various embodiments, the management controller 102 includes one or more processor cores 104 (e.g., processing devices) configured to provide (e.g., execute) a trusted execution environment or TEE 105. In embodiments, the TEE 105 executes a vRoT application 106 for each respective AP of the plurality of APs 111, although a one-to-one correspondence is not required. In embodiments, each vRoT application 106 accesses a corresponding one or more flash memory devices of the plurality of flash memory devices 115 through a corresponding one or more multiplexers of the plurality of multiplexers 113. The management controller 102 can also include IO hardware 110 through which each vRoT application 106 running on the TEE 105 can communicate with each respective flash memory device 115. For example, the IO hardware 110 can include inter-integrated circuit (I2C), improved inter-integrated circuit (I3C), or peripheral component interconnect express (PCIe) circuitry, serial peripheral interface (SPI) circuitry, and the like.

[0024] As Figure 1As shown, according to some embodiments, a first vRoT application 106A can be coupled to a first multiplexer 113A (and / or a second multiplexer 113B), a second vRoT application 106B can be coupled to a second multiplexer 113B, and an nth vRoT application 106N can be coupled to an nth multiplexer 113N. In embodiments, the first multiplexer 113A is capable of selectively coupling the first vRoT application 106A and a first flash memory device 115A to the first AP 111A. In embodiments, the second multiplexer 113B is capable of selectively coupling the second vRoT application 106B (or the first vRoT application 106A, shown in dashed lines) and a second flash memory device 115B to the second AP 111A. In embodiments, the nth multiplexer 113N is capable of selectively coupling the nth vRoT application 106N and an nth flash memory device 115N to the nth AP 111N.

[0025] In some embodiments, each vRoT application 106 is capable of updating security data located in a flash memory device 115 with which the vRoT application 106 is coupled via a corresponding multiplexer. Each vRoT application 106 can also use the security data to perform at least one security operation on behalf of an AP associated with (e.g., coupled to) the flash memory device 115. In some embodiments, the security data includes firmware (FW) and / or configuration data, e.g., data used to enable secure boot and secure operation of the AP. In some embodiments, the security operation is secure boot of the AP, attestation of the AP, secure recovery of firmware or configuration data from the AP, installation of a debug token or debug firmware on the AP, and / or secure update of firmware of at least some of the plurality of flash memory devices 115 of the corresponding AP 111. In environments requiring strict security compliance, e.g., in the military, government, enterprise, or financial sectors, it can be necessary to measure the flash memory devices and record integrity checks for audit and compliance purposes. Thus, such updates and integrity checks can provide verifiable evidence that the integrity of the system 100 or 200 is being maintained.

[0026] In various embodiments, the management controller 102 also performs security-related updates to one or more vRoT applications 106. For example, the security-related updates can include distributing new or updated security policies to the vRoT applications 106 associated with the coupled flash memory devices 115. The security-related updates can also include enforcing new or updated security policies associated with particular APs that are selectively coupled to respective flash memory devices 115 through one or more multiplexers 113.

[0027] Further reference Figure 2 The system 200 can include a memory 260 for storing code or instructions to be executed by the one or more processor cores 104 as well as system and user data. In some embodiments, the memory 260 includes volatile memory and / or non-volatile memory to include computer storage. The memory 260 can also include dedicated memory devices for use by the management controller 102, such as flash memory or eMMC storage devices.

[0028] In some embodiments, the system 200 also includes a processor 202 that contains a plurality of interface controllers 220. In various embodiments, the external processor 202 is a system on a chip (SOC), such as a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a microcontroller, or a complex programmable logic device, among others. In embodiments, the IO hardware 210 and the external processor 202 communicate through a management component transport protocol (MCTP) PCIe interface, such as using the MCTP and PCIe protocols. Further, the IO hardware 210 can communicate with the memory 260 through any suitable memory type or interface protocol for memory.

[0029] In embodiments, each of the plurality of multiplexers 113 receives as input the output of one of the plurality of interface controllers 220 and one of the plurality of APs 111. In embodiments, each multiplexer 113 also receives as a control input (or select input) a multiplexer control signal (MUX ctrl.) from the external processor 202 that controls which multiplexer inputs are passed to the multiplexer output. In embodiments, the plurality of interface controllers 220 and the multiplexer control signal can be controlled by a respective vRoT application through the IO hardware. In the systems discussed below with reference to Figures 3 to 5 Any of the systems 300, 400, or 500 can have IO hardware that can communicate with and through the external processor 202 and / or the plurality of multiplexers 113 to ultimately access the plurality of flash memory devices 115.

[0030] Figure 3is a schematic block diagram of an example distributed computing system 300 according to some embodiments, which supports vRoT applications from a management controller on which a non-secure kernel and a trusted operating system (OS) are concurrently executing. For example, in embodiments, the system 300 includes a management controller 302 coupled to an external processor 202 (or directly to a plurality of multiplexers 113) and to a BMC firmware flash memory device 331, a BMC data flash memory device 333, and an eMMC storage device 335 (more on this later). In some embodiments, the management controller 302 is a BMC or a controller designed for control, security, and / or management of the system 300. In some embodiments, the BMC is located on a DC-SCM card of the system 300.

[0031] In some embodiments, the management controller 302 includes one or more processing cores 304 (e.g., processing devices) for executing a non-secure kernel 317 that provides a normal world and also executing a trusted OS 320. In some embodiments, the trusted OS 320 provides a secure world (e.g., TrustZone TM , which is a technology developed to create an isolated secure world within a processor to run trusted applications). The system 300 can provide a secure partition of operation between the normal world and the secure world. In some embodiments, for example, the trusted OS 320 is OPTEE, hafnium, or a Linux kernel, so the disclosed embodiments can be performed on existing systems. When employing the TrustZone TM technology, the one or more cores 304 can also provide an exception level 3 (or EL3) layer 305, which is the highest privilege level in an exception model, typically reserved for secure firmware, and supports the TrustZone TM .

[0032] In various embodiments, the non-secure kernel 317 executes unprivileged software 319, while the trusted OS 320 executes privileged software, such as a plurality of vRoT applications 306, a virtual SPI flash service 307 (which can support serial peripheral interface (SPI)-based flash memory devices with which the vRoT applications 306 are interacting), and a management component transport protocol (MCTP) bridge 308 for providing secure communication between MCTP-based interconnects.

[0033] In some embodiments, the management controller 302 includes an MMU 309 coupled to the one or more cores 304 and IO hardware 310 coupled between the MMU 309 and external devices, such as the external processor 202, the BMC firmware flash memory device 331, the BMC data flash memory device 333, and the eMMC storage device 335. In some embodiments, the BMC data flash memory device 333 is a non-volatile memory device coupled to the IO hardware 310 that is configured to store flash data for the vRoT application 106. In some embodiments, the BMC firmware flash memory device 331 is a non-volatile memory device coupled to the IO hardware 310 that is configured to store firmware for the vRoT application 106.

[0034] In various embodiments, the external processor 202 Figure 2 is coupled between the IO hardware 310 and the plurality of multiplexers 313. In some embodiments, the trusted OS 320 uses the MMU 309 to isolate the IO hardware 310 of the trusted OS 320, e.g., to securely communicate with the plurality of flash memory devices 115, while preventing intrusion by applications running on the non-secure kernel 317. In this way, the trusted OS 320 can arbitrate secure communications that are separate from the normal world of the non-secure kernel 317, despite operating on the same distributed computing system.

[0035] Figure 4 is a schematic block diagram of an example distributed computing system 400 that supports a vRoT application from a management controller on which a non-secure kernel, a trusted OS, and a secure kernel are executed, in accordance with some embodiments. For example, in embodiments, the system 400 includes a management controller 402 coupled to the external processor 202 (or directly to the plurality of multiplexers 113) and coupled to a BMC firmware flash memory device 431 and an eMMC storage device 435 (more on this later). In some embodiments, the management controller 402 is a BMC or a controller designed specifically for control, security, and / or management of the system 400. In some embodiments, the BMC is located on a DC-SCM card of the system 400.

[0036] In some embodiments, the management controller 402 includes one or more processor cores 404 (e.g., processing devices) on which a trusted hypervisor 405, such as Xen, KVM, Hyper-V, etc., is executed. In some embodiments, the unsecure kernel 417, the trusted OS 420, and the secure kernel 440 can run on the trusted hypervisor. For example, the unsecure kernel 417 can execute an open BMC virtual machine 412 through which unprivileged software is provided. Further, the trusted OS 420 can execute: a vRoT virtual machine 416 for running each of the plurality of vRoT applications 406; a virtual SPI flash service 407 capable of supporting SPI-based flash memory devices with which the vRoT applications 406 are interacting; and a MCTP bridge 408 for providing secure communication between MCTP-based interconnects.

[0037] In embodiments, the vRoT virtual machine 416 performs end-to-end encryption between the trusted OS 420 and each respective flash memory device of the plurality of flash memory devices 115. In some embodiments, the trusted OS 420 of 420 performs security-related updates to one or more vRoT applications 406, such as distributing new or updated security policies to the vRoT applications 406, or enforcing new or updated security policies associated with the corresponding AP 111.

[0038] In embodiments, the secure kernel 440 can execute a PA-RoT virtual machine 418 to run the PA-RoT applications 442, the attestation applications 444, and one or more additional platform security services 448. In embodiments, the PA-RoT 442 ensures that the system 400 is securely operating by establishing and managing a root of trust with the aid of hardware and firmware.

[0039] In various embodiments, the management controller 402 further includes IO hardware 410 coupled to the processing devices (e.g., one or more processor cores 404) and external devices (e.g., the external processor 202 (or directly coupled to the plurality of multiplexers 113), the BMC firmware flash memory device 431, and the eMMC storage device 435). In embodiments, a management controller bridge (e.g., the MCTP bridge 408) provides secure communication between the trusted hypervisor 405 and the IO hardware 410.

[0040] In some disclosed embodiments, the trusted hypervisor 405 also directly implements virtual fuses (vFuses), virtual crypto (vCrypto), and / or virtual system-on-a-chip (SOC) RoT (vSOC_RoT) to support the PA-RoT virtual machine 418 running on the secure kernel 440. Fuses in the PA-RoT virtual machine 418 context can refer to physical one-time programmable (OTP) storage cells used to store critical data that is protected from being modified, hence the use of the term “virtual” so that these storage cells can be logical and backed by secure cache. They are called fuses because once set (programmed), they cannot be changed; they are “blown” like an electrical fuse. Fuses can store cryptographic keys, device identity and authentication, and configuration settings. Cryptographic technology with reference to the secure kernel 440 encompasses algorithms and cryptographic processes used to protect data and ensure secure communication of the PA-RoT virtual machine 442. Thus, fuses and cryptographic mechanisms can work in concert to provide a strong security foundation. In embodiments, the vSOC_RoT can be or include Caliptra TM , which defines design criteria for silicon-internal RoT baselines. This standard meets the measurement root-of-trust (RTM) role. The open-source implementation of Caliptra TM improves transparency of the measurement mechanisms for RTM and anchored hardware attestation.

[0041] Figure 5 is a schematic block diagram of an example distributed computing system 500 according to various embodiments, the TEE availability of which differs from that in Figure 4 . For example, in embodiments, the system 500 includes a management controller 502 coupled to the external processor 202 (or directly to the plurality of multiplexers 113) as well as to the BMC firmware flash memory device 431 and the eMMC storage device 435. In some embodiments, the management controller 502 is a BMC or a controller designed for control, security, and / or management of the system 500. In embodiments, the BMC is located on a DC-SCM card of the system 500.

[0042] In some embodiments, the management controller 502 includes one or more processor cores 504 (e.g., processing devices) on which an untrusted hypervisor 505 executes. Thus, any programs executing on top of the untrusted hypervisor 505 can provide their own root of trust and / or trusted execution environment (TEE) because the hypervisor 505 is untrusted. Thus, in some embodiments, the non-secure kernel 417 executes an open BMC TEE 512 (or trusted VM) on which non-privileged software runs. Further, the secure kernel 420 can execute a vRoT trusted VM 516 on which the vRoT application 406, a virtual SPI flash service 407 (which can support a SPI-based flash memory device with which the vRoT application 406 is interacting), and an MCTP bridge 408 for facilitating secure communications across the MCTP-based interconnect run. Thus, the vRoT application 406 can be understood to be instantiated as one or more trusted virtual machines. Further, the secure kernel 440 can execute a PA-RoT TEE 518 (or PA-RoT trusted VM) on which the PA-RoT application 442, attestation application 444, and one or more additional secure services 448 run.

[0043] In embodiments, the open BMC TEE 512, vRoT trusted VM 516, and PA-RoT TEE 518 can communicate with each other through the untrusted hypervisor 505 using encrypted TEE inter-process communication (IPC). In embodiments, one or more cores 504 can include a trusted service manager (TSM) hardware 550 in order to provide a sufficient level of security for communications passing from the TEEs / trusted virtual machines through the IO hardware 510 to external devices. For example, the TSM hardware 550 of the trusted execution environment can implement confidential computing between the multiple APs 111 and the trusted virtual machines (e.g., vRoT trusted VM 516) using a device security interface protocol. In embodiments, the TSM hardware 550 executes firmware that is adapted to configure the processing devices to run each of the trusted virtual machines.

[0044] More specifically, in some embodiments, the management controller 502 further includes IO hardware 510 coupled to the processing devices (e.g., one or more processor cores 504) and external devices (e.g., the external processor 202 (or directly coupled to the multiple multiplexers 113), the BMC firmware flash memory device 431, and the eMMC storage device 435). In embodiments, a management controller bridge (e.g., the MCTP bridge 408) provides secure communications between the trusted hypervisor 405 and the IO hardware 410.

[0045] In embodiments, the PA-RoT TEE 518 (or PA-RoT trusted VM) uses a PCIe card 555 that supports a TEE device interface security protocol (TDISP) for the above fuse, crypto, and / or SOC_RoT, and can be protected by utilizing TDISP for end-to-end communication with the plurality of flash memory devices 115. See also Figure 2 , the external processor 202 can be assigned to the vRoT trusted VM 516 but is untrusted. In embodiments, the vRoT trusted VM 516 provides end-to-end encryption for the plurality of flash memory devices 115. Alternatively, if the external processor 202 supports TDISP, then the external processor 202 can be trusted and end-to-end encryption is performed between each vRoT 406 and the external processor 202.

[0046] Figure 6 is a flowchart of an example method 600 of performing AP secure boot using a management controller, in accordance with some embodiments. The method 600 can be performed by processing logic that can comprise hardware, software, firmware, or any combination thereof. For example, the method 600 can be performed by the systems 100, 200, 300, 400, and / or 500, or by particular components of each system, e.g., by the management controllers 102, 302, 402, and / or 502 in Figure 1 , Figure 3 , Figure 4 and Figure 5 . Although shown in a particular order or sequence, unless otherwise specified, the order or sequence can be modified. Thus, the shown embodiments should be understood only as an example, and the shown processes can be performed in a different order, and certain processes can be performed in parallel. Also, one or more processes can be omitted in various embodiments. Thus, not all processes are required in every embodiment. Other processes are also possible.

[0047] In operation 605, the processing logic receives a power-on signal (e.g., from the remote device 90) to power on the system and corresponding memory controller.

[0048] At operation 610, the processing logic requests that a TEE be loaded into the memory controller. In some embodiments, the processing logic (e.g., of the management controller) and the TEE run within the same silicon-based chip or component, thus as shown by the dashed box.

[0049] At operation 612, the AP can remain in a reset state, waiting for the processing logic to start. In some embodiments, this behavior of the AP is a policy choice. In some embodiments, the AP can be allowed to start without delay. In other embodiments, this policy can be configurable depending on the AP.

[0050] At operation 615, the processing logic loads the TEE into the memory controller for execution.

[0051] At operation 620, the processing logic causes the memory controller to continue booting.

[0052] At operation 625, the processing logic causes the TEE to prepare to boot the AP, and thus, it can be appreciated that the vRoT application of the TEE will now participate (from the TEE’s perspective) in the verification and secure boot of the AP in conjunction with the flash memory device, as previously described.

[0053] At operation 630, the processing logic causes the TEE to measure the flash memory device of the AP.

[0054] At operation 635, the processing logic verifies the measurements of the flash memory device of the AP. In embodiments, the measurement process involves computing a cryptographic hash of the firmware or software stored on the flash memory device. This hash can then be compared to a previously known good hash that represents a trusted state of the firmware. If the computed hash matches the trusted hash, it indicates that the firmware has not been altered or tampered with since the last attestation. The results of the measurement process can affect the boot process. For example, if a mismatch between the measured hash and the trusted hash is detected, the system can stop the boot process, enter a recovery mode, or take other predefined security measures. This would enforce a strict security policy aimed at protecting the system from running potentially harmful software.

[0055] At operation 640, assuming the measurements of the AP were successfully verified at operation 635, the processing logic causes the TEE to release the AP from the reset state and, at operation 642, receive an indication that the AP is under attack by a botnet.

[0056] At operation 645, the flash memory device retrieves the AP code.

[0057] At operation 650, the AP completes the secure boot and fully operates normally.

[0058] Figure 7 is a flowchart of an example method 700 for performing a secure update of an AP through a coupled flash memory device, in accordance with some embodiments. The method 700 can be performed by processing logic that comprises hardware, software, firmware, or any combination thereof. For example, the method 700 can be performed by the systems 100, 200, 300, 400, and / or 500, or by particular components of each system, such as by the TEE 110, 210, 310, 410, and / or 510, the memory controller 120, 220, 320, 420, and / or 520, the AP 130, 230, 330, 430, and / or 530, and / or the flash memory device 140, 240, 340, 440, and / or 540. Figure 1 、 Figure 3 、 Figure 4 and Figure 5The processes described herein can be performed by a management controller 102, 302, 402, and / or 502 in the system 100, 200, 300, 400, and / or 500. The processes performed by the management controller 102, 302, 402, and / or 502 are shown in a particular order or sequence, but unless otherwise specified, the order of the processes can be modified. Thus, the illustrated embodiments should be understood only as examples, the illustrated processes can be performed in a different order, and certain processes can be performed in parallel. Additionally, one or more processes can be omitted in various embodiments. Thus, not all processes are required in every embodiment. Other flows are possible.

[0059] In operation 705, the processing logic receives a power-on signal (e.g., from the remote device 90) to power on the system and corresponding memory controller.

[0060] In operation 710, the processing logic causes the TEE to update the AP's flash memory device. In some embodiments, the processing logic (e.g., of the management controller) and the TEE run within the same silicon-based chip or component, and thus are shown as dashed boxes. In some embodiments, the vRoT application executing within the TEE performs a secure update of the AP, e.g., by writing to the flash memory device associated with the AP. In some embodiments, AP behavior is a policy choice. In embodiments, the AP can run while its access to the flash memory device is temporarily denied. In other embodiments, the AP remains in a reset or idle state, enters a low-power or no-power state, e.g., while the flash memory device is being updated.

[0061] In operation 715, the processing logic causes the TEE to verify a flash image of the flash memory device. This verification ensures that the flash image (i.e., the binary data to be written to the flash memory) is authentic, unaltered, and can be safely installed. This image verification can involve several steps that overlap with the secure firmware update procedure, with a particular focus on ensuring the integrity and authenticity of the flash image.

[0062] In operation 720, the processing logic causes the TEE to write the flash image to the flash memory device, provided that the TEE successfully verified the flash memory device in operation 715.

[0063] In operation 725, the processing logic updates AP metadata associated with the programmed flash image at the flash memory device.

[0064] At operation 730, the processing logic receives an AP update completion message indicating that the flash image write on the flash memory device has successfully completed.

[0065] At operation 735, the processing logic sends the AP update completion message to the remote device 90.

[0066] Figure 8is a flowchart of an example method for performing a proof of security of an AP according to some embodiments. The method 800 can be performed by processing logic that can comprise hardware, software, firmware, or any combination thereof. For example, the method 800 can be performed by the systems 100, 200, 300, 400, and / or 500, or by a particular component of each system, such as by the management controller 102, 302, 402, and / or 502 in Figure 1 、 Figure 3 、 Figure 4 and Figure 5 . Although the flowcharts shown in the figures are shown in a particular order or sequence, unless otherwise specified, the order or sequence can be modified. Thus, the illustrated embodiments should be understood only as examples, the flowcharts shown can be performed in a different order, and certain steps can be performed in parallel. Additionally, one or more steps can be omitted in various embodiments. Thus, not all steps are required in all embodiments. Other steps are also possible.

[0067] In operation 805, the processing logic receives an AP attestation command from the remote device 90 to attest the integrity of firmware installed in the flash memory device.

[0068] In operation 810, the processing logic requests the TEE to perform an AP attestation of the flash memory device. In some embodiments, the processing logic (e.g., of the management controller) and the TEE run within the same silicon-based chip or component, thus as shown in the dashed box. In some embodiments, a vRoT application executing within the TEE performs the attestation of the AP, e.g., of the firmware stored in the flash memory device. In some embodiments, the AP behavior is a policy choice. In some embodiments, the AP can run while its access to the flash memory device is temporarily denied. In other embodiments, the AP remains reset or quiescent, entering a low-power or no-power state, e.g., while the flash memory device is being updated.

[0069] In operation 815, the processing logic causes the TEE to measure the flash memory device of the AP. In some embodiments, the measurement process includes reading data in the flash memory and computing a cryptographic hash of the firmware or software stored on the flash device. This hash value can then be compared to a previously known good hash value that represents a trusted state of the firmware.

[0070] In operation 820, the processing logic causes the TEE to sign the results of the measurement of the flash memory device of the AP. In embodiments, the signing process involves encrypting the measurement results using a private cryptographic key (e.g., a key specific to the vendor of the system that includes the management controller). The corresponding public key should already be trusted and securely stored in the remote device 90.

[0071] In operation 825, the processing logic receives the signed AP measurement results from the TEE.

[0072] In operation 830, the processing logic transmits the signed AP measurement results received from the TEE to the remote device 90. The remote device 90 can then decrypt the signed AP measurement results using the public key generated by the public key, thereby verifying or "attesting" the validity of the hash value, thereby generating the plaintext of the AP measurement results (e.g., hash value) of the flash memory device. The remote device 90 can then compare the computed hash value to a previously known good hash value, which represents a trusted state of the firmware. If the computed hash value matches the trusted hash value, it indicates that the firmware has not been altered or tampered with since the last verification.

[0073] Figure 9 FIG. 9 is a flow diagram of a method 900 for operating a distributed computing system having the disclosed management controller, in accordance with at least one embodiment. The method 900 can be performed by processing logic that can comprise hardware, software, firmware, or any combination thereof. For example, the method 900 can be performed by the system 100 or a particular component of the system 100 (see FIG. 1), e.g., a system comprising a plurality of APs, a plurality of flash memory devices, a plurality of multiplexers (each multiplexer to selectively couple one of the plurality of flash memory devices to one of the plurality of APs), and a controller coupled to the plurality of multiplexers. Although shown in a particular sequence or order, unless otherwise specified, the order of the processes can be modified. Thus, the illustrated embodiments should be understood only as examples, and the illustrated processes can be performed in a different order, and certain processes can be performed in parallel. Additionally, one or more processes can be omitted in various embodiments. Thus, not all processes are required in every embodiment. Other procedures can also be possible. Figure 1 ) performs, e.g., a system comprising a plurality of APs, a plurality of flash memory devices, a plurality of multiplexers (each multiplexer to selectively couple one of the plurality of flash memory devices to one of the plurality of APs), and a controller coupled to the plurality of multiplexers. Although shown in a particular sequence or order, unless otherwise specified, the order of the processes can be modified. Thus, the illustrated embodiments should be understood only as examples, and the illustrated processes can be performed in a different order, and certain processes can be performed in parallel. Additionally, one or more processes can be omitted in various embodiments. Thus, not all processes are required in every embodiment. Other procedures can also be possible.

[0074] In operation 910, the processing logic (e.g., of the controller) provides a trusted execution environment to execute a virtual root of trust (vRoT) application for each respective AP of the plurality of APs.

[0075] In operation 920, the processing logic accesses, by each vRoT application, a corresponding one or more flash memory devices of the plurality of flash memory devices via a corresponding one or more multiplexers of the plurality of multiplexers.

[0076] In extensions of the method 900, the processing logic also performs security-related updates to the first vRoT application, for example, including distributing a new or updated security policy to the first vRoT application associated with the first flash memory device, and / or enforcing a new or updated security policy associated with the first AP that is selectively coupled to the first flash memory device via a first multiplexer of the plurality of multiplexers. The method 900 can also include instantiating the vRoT application as one or more trusted virtual machines.

[0077] Other variations are within the spirit of the present disclosure. Thus, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrated embodiments thereof are shown in the drawings and have been described above in detail. It should, however, be understood that there is no intention to limit the disclosure to the specific form or forms disclosed, but on the contrary, the intention is to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of the disclosure, as defined in the appended claims.

[0078] The use of the terms "a" and "an" and "the" and similar referents in the context of describing the disclosed embodiments (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms "comprising," "having," "including," and "containing" are to be construed as open-ended terms (meaning "including, but not limited to,") unless otherwise noted. The term "connected" (and its derivatives) as used herein is intended to

[0079] Unless explicitly stated otherwise or apparent from context, a phrase such as "at least one of A, B, and C" or "at least one of A, B, or C" shall indicate that A is optionally present, with or without B and / or C. In other words, A, B, and C are each optional, independently of one another. For example, "at least one of A, B, and C" or "at least one of A, B, or C" shall mean, in the terms of sets, that A, B, and C can each be present, individually or in any combination, and that there is no presence of only the members of a set of A, B, and C. For example, in the case of the illustrative set of three members: A, B, and C, the

[0080] Unless otherwise indicated herein, or otherwise clearly contradicted by context, the operations of a process described herein can be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions, and are implemented as code (e.g., executable instructions, one or more computer programs or one or more applications) executing collectively on one or more processing units, by hardware or combinations thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, in the form of a computer program, comprising a plurality of instructions executable by one or more processors. In at least one embodiment, a computer-readable storage medium is a non-transitory computer-readable storage medium, excluding transitory propagating signals, but including non-transitory data storage circuits (e.g., buffers, cache and queues) within transceivers of transitory signals. In at least one embodiment, code (e.g., executable or source code) is stored on a set of one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) having executable instructions stored thereon which, when executed by one or more processors of a computer system (i.e., as a result of being executed), cause the computer system to perform operations described herein. In at least one embodiment, a set of non-transitory computer-readable storage media includes multiple non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media in the multiple non-transitory computer-readable storage media lack all of the code, with the multiple non-transitory computer-readable storage media collectively storing the entire code. In at least one embodiment, executable instructions are executed by different processors to cause different instructions to be executed.

[0081] Accordingly, in at least one embodiment, a computer system is configured to implement one or more services that individually or collectively perform operations of processes described herein, and such a computer system is configured with applicable hardware and / or software to enable performance of the operations. Moreover, a computer system implementing at least one embodiment of the present disclosure is a single device, and in another embodiment is a distributed computer system including multiple devices operating differently, with the distributed computer system performing operations described herein, and with a single device not performing all of the operations.

[0082] The use of any and all examples, or exemplary language (e.g., "such as") provided herein is intended merely to better illuminate embodiments of the disclosure and does not pose a limitation on the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.

[0083] All references cited herein, including publications, patent applications, and patents, are hereby incorporated by reference to the same extent as if each reference were individually and specifically incorporated by reference and was specifically stated to be incorporated by reference herein in its entirety.

[0084] In the description and claims, the terms "coupled" and "connected," along with derivatives thereof, can be used. It should be understood that these terms are not intended as synonyms for each other. Rather, in particular embodiments, "connected" or "coupled" can be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. "Coupled" can also mean that two or more elements are not in direct contact with each other, but yet still co-operate or interact with each other.

[0085] Unless specifically stated otherwise, it can be appreciated that throughout the specification terms such as "processing," "computing," "calculating," "determining," or the like, refer to the action and / or processes of a computer or computing system, or similar electronic

[0086] In a similar manner, the term "processor" can refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" can be a network device or a MACsec device. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities such as tasks, threads, and intelligent agents that perform work over time. Also, each process can refer to multiple processes to execute instructions sequentially or in parallel, continuously or intermittently. In at least one embodiment, the terms "system" and "method" can be used interchangeably herein as long as the system can embody one or more methods and the method can be considered as a system.

[0087] In this document, reference can be made to acquiring, obtaining, receiving analog or digital data or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. In at least one embodiment, the process of acquiring, obtaining, receiving, or inputting analog and digital data can be accomplished in various ways, such as by receiving data as a parameter of a function call or a call to an application programming interface. In at least one embodiment, the process of acquiring, obtaining, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In at least one embodiment, the process of acquiring, obtaining, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. In at least one embodiment, reference can also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be accomplished by transmitting data as an input or output parameter of a function call, a parameter of an application programming interface, or an interprocess communication mechanism.

[0088] While the description herein sets forth example embodiments of the described technology, other architectures can be used to implement the described functionality, and are intended to be within the scope of this disclosure. Moreover, while specific allocations of responsibilities have been defined above for the purpose of description, various functions and responsibilities can be distributed and divided among various parties in different ways, as circumstances can warrant.

[0089] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.

Claims

1. A system comprising: Multiple application processors (APs); Multiple flash memory devices associated with the plurality of APs; Multiple multiplexers, each multiplexer being used to selectively couple one of the multiple flash memory devices to one of the multiple APs; as well as A controller, operatively coupled to the plurality of multiplexers, wherein the controller is configured to provide a trusted execution environment for executing a Virtual Root of Trust (vRoT) application for each of the plurality of APs, wherein each vRoT application is configured to access one or more of the plurality of flash memory devices via one or more of the corresponding multiplexers.

2. The system of claim 1, wherein the first multiplexer of the plurality of multiplexers is configured to selectively couple a first AP of the plurality of APs or one of the controllers executing a first vRoT application to a first flash memory device of the plurality of flash memory devices, wherein the first vRoT application is configured to perform at least one of the following operations: Update the security data located in the first flash memory device; or This enables the use of the security data to perform at least one security operation on behalf of the first AP.

3. The system of claim 2, wherein the security data includes at least one of firmware or configuration data, and wherein the security operation is one of the following: secure boot of the first AP, authentication of the first AP, secure recovery of firmware or configuration data from the first AP, installation of a debug token or debug firmware on the first AP, or secure update of the firmware of at least some of the plurality of flash memory devices.

4. The system of claim 1, wherein the controller is further configured to perform security-related updates on the first vRoT application, the security-related updates comprising at least one of the following: Distribute a new or updated security policy to the first vRoT application associated with the first flash memory device; or Enforce the new or updated security policy associated with the first AP, which is selectively coupled to the first flash memory device via a first multiplexer of the plurality of multiplexers.

5. The system of claim 1, wherein the controller includes a baseboard management controller, and wherein the vRoT application is instantiated as one or more trusted virtual machines.

6. The system of claim 1, wherein the trusted execution environment includes a trusted operating system (OS) running on a processing device that also executes an insecure kernel, the system further comprising: Memory Management Unit (MMU); as well as Input / output I / O hardware coupled to the MMU; as well as An external processor is coupled between the I / O hardware and the plurality of multiplexers, wherein the trusted OS employs the MMU to isolate the I / O hardware so that the trusted OS can communicate securely with the plurality of flash memory devices while preventing intrusion by applications running on the insecure kernel.

7. The system of claim 6 further includes a non-volatile memory device coupled to the I / O hardware and used for storing flash data of the vRoT application.

8. The system of claim 6, wherein the external processor comprises a plurality of interface controllers, wherein each of the plurality of multiplexers is configured to receive the output of one of the plurality of interface controllers and one of the plurality of APs as input, and to receive a multiplexer control signal from the external processor as a control input, and wherein, The control signals of the multiple interface controllers and the multiplexer can be controlled by the corresponding vRoT application through the IO hardware.

9. The system according to claim 1, further comprising: The controller's processing device is used to execute a trusted management program, on which the following is performed: A trusted operating system that executes vRoT virtual machines to run each of the vRoT applications and manages the controller bridge; as well as A secure kernel that runs a Platform Active RoT (PA-RoT) virtual machine and one or more additional platform security services; Input / output (I / O) hardware coupled to the processing device, wherein the management controller bridge provides secure communication between the trusted management program and the I / O hardware; as well as An external processor coupled between the I / O hardware and the plurality of multiplexers.

10. The system of claim 9, wherein the vRoT virtual machine is a trusted virtual machine, and wherein the processing device includes the Trusted Execution Environment (TSM) hardware to implement confidential computation between the plurality of APs and the trusted virtual machine using a device security interface protocol, wherein, The TSM hardware executes firmware suitable for configuring the processing device to run the trusted virtual machine.

11. The system of claim 9, wherein the vRoT virtual machine is used to perform end-to-end encryption between the trusted operating system and each of the plurality of flash memory devices.

12. A processing apparatus, comprising: One or more processor cores are used to execute an insecure kernel and provide a trusted operating system (OS) that provides a trusted execution environment; A memory management unit (MMU) coupled to one or more processor cores; as well as Input / output (I / O) hardware coupled to the MMU and multiple flash memory devices associated with multiple application processors (APs) of the distributed computing system, wherein the trusted OS is used for: Execute the virtual root of trust (vRoT) application for each of the plurality of APs; as well as The MMU is used to isolate the I / O hardware so that the trusted OS can communicate securely with the plurality of flash memory devices, while preventing intrusion by applications running on the insecure kernel.

13. The processing apparatus of claim 12, wherein the first vRoT application for the first AP is configured to perform at least one of the following operations: Update security data in the first flash memory device coupled to the first AP; or This enables the use of the security data to perform at least one security operation on behalf of the first AP.

14. The processing apparatus of claim 13, wherein the security data includes at least one of firmware or configuration data, and the security operation is one of the following: secure boot of the first AP, authentication of the first AP, secure recovery of firmware or configuration data from the first AP, installation of a debug token or debug firmware on the first AP, or secure update of the firmware of at least some of the plurality of flash memory devices.

15. The processing apparatus of claim 13, wherein, in order to update the secure data in the first flash memory device, the trusted OS is configured to cause a multiplexer coupled between the first flash memory device and the first AP to select one or more processor cores executing the first vRoT application for output via a first interface controller coupled to the I / O hardware.

16. The processing apparatus of claim 13, wherein the trusted OS is further configured to perform security-related updates on the first vRoT application, the security-related updates comprising at least one of the following: Distribute the new or updated security policy to the first vRoT application; or Enforce the new or updated security policy associated with the first AP.

17. A method of operating a distributed computing system, the distributed computing system comprising a plurality of application processors (APs), a plurality of flash memory devices, a plurality of multiplexers, and a controller coupled to the plurality of multiplexers, each multiplexer being configured to selectively couple one flash memory device from the plurality of flash memory devices to one AP from the plurality of APs, wherein, The method includes: The controller provides a trusted execution environment to execute a virtual root of trust (vRoT) application for each of the plurality of APs; and Each vRoT application accesses one or more corresponding flash memory devices from the plurality of flash memory devices via one or more corresponding multiplexers from the plurality of multiplexers.

18. The method of claim 17, further comprising: The first multiplexer of the plurality of multiplexers selectively couples the first flash memory device of the plurality of flash memory devices to one of the first APs of the plurality of APs or to one of the controllers executing the first vRoT application; The first vRoT application updates the security data located in the first flash memory device; as well as This enables the first vRoT application to perform at least one security operation on behalf of the first AP using the security data.

19. The method of claim 18, wherein the security data comprises at least one firmware, and wherein the security operation is one of the following: secure boot of the first AP, authentication of the first AP, secure recovery of firmware or configuration data from the first AP, installation of a debug token or debug firmware on the first AP, or secure update of the firmware of at least some of the plurality of flash memory devices.

20. The method of claim 17, further comprising: The controller performs security-related updates on the first vRoT application, the security-related updates including at least one of the following: Distribute new or updated security policies to the first vRoT application associated with the first flash memory device; or Enforce the new or updated security policy associated with the first AP, which is selectively coupled to the first flash memory device via a first multiplexer of the plurality of multiplexers.

21. The method of claim 17, further comprising: Instantiate the vRoT application as one or more trusted virtual machines.