Methods, systems, and computer-readable media for protecting multiple computing nodes
By authenticating the hardware architecture and firmware of computing node components, generating firmware measurement certificates and implementing protection strategies, the problems of component counterfeiting and firmware version differences in computing node clusters are solved, improving security and availability.
Patent Information
- Application Number
- CN201980069596.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-01-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2039-01-08
AI Technical Summary
Counterfeit or malicious software in computing node components may become an attack vector, causing damage to the computing node. Existing technologies have difficulty in efficiently authenticating and managing firmware version differences of components in large-scale computing node clusters, leading to security and availability issues.
By authenticating the hardware architecture and firmware of multiple components of the computing node, generating firmware measurement certificates, and implementing protection strategies using the authentication database, the trust and security of the components are ensured.
It enables dynamic management of components in a computing node cluster, identifies and isolates potentially vulnerable or defective firmware versions, and improves the security and availability of computing nodes.
Smart Images

Figure CN112955888B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to compute nodes. Background Art
[0002] In the chassis or housing of a compute node, computer system, or mainframe, there may be hundreds of pluggable components, ranging from temperature sensors and power supplies to memory modules and processors. In a rack or cluster of compute nodes, there may be thousands of such components. However, each component may represent a security vulnerability, that is, a potential attack vector. If the component is counterfeit or contains malware that can harm the compute node, the component may be a potential attack vector. One potential way to compromise a component is to corrupt the firmware used to operate the component. Even simple components such as fans and sensors, if compromised, can cause damage to the compute node due to overheating or fire. Therefore, identifying compromised components may be useful to prevent component misuse. Summary of the Invention
[0003] According to one aspect of the present disclosure, a method for protecting multiple computing nodes is provided, the method comprising: authenticating a hardware architecture of each of multiple components of the multiple computing nodes; authenticating firmware of each of the multiple components, comprising: generating a firmware metric certificate for each layer of the firmware of each component, wherein the firmware metric certificate for each layer of the firmware is based on the firmware metric certificate for an authenticated previous layer of the firmware; and generating an authentication database, the authentication database comprising multiple authentication descriptions based on the authenticated hardware architecture and the authenticated firmware, wherein a policy for protecting a specified subset of the multiple computing nodes is implemented by using the authentication database.
[0004] According to another aspect of the present disclosure, a system for protecting multiple computing nodes is provided, comprising: a processor; and a memory component storing instructions that cause the processor to: authenticate the hardware architecture of each of multiple components of the multiple computing nodes; authenticate the firmware of each of the multiple components, comprising: generating a firmware metric certificate for each layer of the firmware of each component, wherein the firmware metric certificate for each layer of the firmware is based on the firmware metric certificate for a authenticated previous layer of the firmware; and generating an authentication database comprising multiple authentication descriptions based on the authenticated hardware architecture and the authenticated firmware, wherein a policy for protecting a specified subset of the multiple computing nodes is implemented by using the authentication database.
[0005] According to another aspect of the present disclosure, a non-transitory computer-readable medium is provided, which stores computer-executable instructions that, when executed, cause a computer to: authenticate the hardware architecture of each of a plurality of components of a plurality of computing nodes; authenticate the firmware of each of the plurality of components, including: generating a firmware metric certificate for each layer of the firmware of each component, wherein the firmware metric certificate for each layer of the firmware is based on the firmware metric certificate for a authenticated previous layer of the firmware; and generating an authentication database, wherein the authentication database includes a plurality of authentication descriptions based on the authenticated hardware architecture and the authenticated firmware, wherein a policy for protecting a specified subset of the plurality of computing nodes is implemented by using the authentication database. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The present disclosure will be understood from the following detailed description when read in conjunction with the accompanying drawings. According to standard practice in the industry, various features are not drawn to scale. In fact, the dimensions of the various features may be arbitrarily increased or decreased for clarity of discussion.
[0007] Some examples of the present application are described with respect to the following figures.
[0008] Figure 1 is an example compute node of a node group that can be protected.
[0009] Figure 2 is an example system for protecting a node group.
[0010] Figure 3 is a sample firmware metrics certificate used to secure a node group.
[0011] Figure 4 is an example system for protecting a node group.
[0012] Figure 5 is a process flow diagram of a method for generating a firmware metric certificate for securing a node group.
[0013] Figure 6 is a message flow diagram for authenticating components in a system for protecting a node group.
[0014] Figure 7 is an example system for protecting a node group.
[0015] Figure 8 is an example system for protecting a node group.
[0016] Figure 9 is a process flow diagram of a method for protecting a node group.
[0017] Figure 10An example system including a tangible, non-transitory computer-readable medium storing code for protecting a node group. DETAILED DESCRIPTION
[0018] Examples of authentication include Universal Serial Bus (USB) Type-C authentication, which enables computing nodes (i.e., hosts) to authenticate compliant USB components. USB Type-C authentication also forms the basis of a potential Peripheral Component Interconnect Express (PCIe) authentication mechanism that allows authentication of PCIe components. The authentication schemes in USB Type-C and PCIe can be extended to internal buses and other protocols and interconnects.
[0019] The purpose of component authentication is to establish trust in the component. The authentication mechanisms discussed above can verify that the component's hardware comes from a known and trusted manufacturer. However, verifying that a component comes from a known and trusted manufacturer does not necessarily mean that the firmware running within the component is correct and trustworthy. Correctness can mean that the correct firmware and the correct version of the firmware are installed in the component. Trustworthiness can also mean that the firmware can be trusted not to compromise the security of the component on which it is running.
[0020] Additionally, while certifying each component can be challenging, the challenge is even greater when considering that a compute node may include multiple components. Furthermore, compute nodes can be grouped into node groups, such as a chassis enclosure with two compute nodes, or a rack system where all compute nodes are located in a blade server rack. Node groups can also include node clusters, which can include hundreds or more compute nodes, each of which can include tens of thousands of components.
[0021] However, such scale may introduce additional challenges. For example, at the scale of a cluster of nodes, the overall rate at which components fail and are replaced can be such that the system configuration of many compute nodes (even those of the same type and from the same manufacturer) may differ from their original factory settings. As a result, attempting to use automated methods to certify components and verify the firmware of components in a cluster may be challenging. The different configurations may reduce the assumptions that can be made about the original factory settings at the scale of a cluster. Therefore, automated methods that might be more efficient in situations where such assumptions can be made at the scale of a cluster may not be useful. Instead, certification and verification may become customized, which can be cumbersome and costly.
[0022] Furthermore, in a node cluster, different compute nodes with the same hardware components may be running different versions of firmware at any given time. For example, as a practical consideration, component firmware updates can be performed in a phased manner. This ensures availability of the node cluster even if some components are unavailable due to the firmware update. Consequently, different groups of components may have different versions of firmware until all phases of the firmware update in the node cluster have been completed. Furthermore, components of the same type from the same manufacturer—whether initially installed or replaced—may have different factory-installed firmware versions. Furthermore, node clusters are complex computer systems, so higher-level data center and node cluster management systems may be tasked with creating a logical pool of resources based on the physical collection of components present in the node cluster. This can result in some parts of the node cluster being reconfigured, restarted, or taken offline more frequently than others. In some scenarios, restarting and taking a compute node offline may be an opportunity when certifying components on a compute node or updating them with new firmware versions. Therefore, if different compute nodes in a node cluster are restarted or taken offline at different times, the components on those compute nodes may end up with different firmware versions, some of which may be compromised.
[0023] Therefore, in examples of the present disclosure, a node cluster-scale component authentication and verification system can be provided to dynamically manage such differences at scale by providing the ability to identify, report, and manage components on a node group according to specified policies. In this way, the component authentication and verification system can identify components within a node group that may be running vulnerable or defective firmware versions.
[0024] Figure 1An example compute node 100 is provided for a group of nodes that can be secured. The node group can be authenticated by ensuring that the components of each example compute node 100 in the group are trustworthy. The example compute node 100 includes multiple components that can be authenticated along with the firmware installed on each component. The firmware can be computer instructions that operate the various components of the compute node 100. In an example, some components can be installed by the manufacturer of the compute node 100. Additionally, some components can be field-replaceable, meaning that after purchasing the compute node 100, the components can be replaced by installing them on the compute node 100 while the compute node 100 is powered off. Furthermore, some components can be hot-swappable. Hot-swappable means that the components are physically connected to the interconnects of the compute node 100 while the compute node 100 is powered on. The compute node 100 can include components with a range of functions, including components with little or no processing power, such as sensors 102, fans 104, and power supplies 106, as well as components with complex processing power, such as general-purpose processors 108. Additional components of the compute node 100 may include, for example, a Gen-Z component 110, a Gen-Z switch 112, a USB component 114, a baseboard management controller (BMC) 116, BMC software 118, a plurality of network interface controllers (NICs) 120, a memory 122, and a serial peripheral interconnect (SPI) switch 124. The SPI switch 124 may be in a bus used to access read-only memory (ROM). The SPI switch 124 may enable the BMC 116 to check whether the general-purpose processor 108 is loading the correct firmware. For example, before the basic input and output system (BIOS) 126 can be loaded by the processor 108, the SPI switch 124 may enable the BMC 116 to read the BIOS 126. If the BIOS 126 is correct, the BMC 116 may open a switch to allow the processor 108 to load the BIOS 126. The SPI switch 124 may also enable the BMC 116 to restore or update the BIOS 126. Further, components may include BIOS 126, trusted platform module 128, PCIe component 130, serial attached SCSI (SAS) / SATA controller 132, SATA component 134, and SAS component 136. These components may also be connected within compute node 100 through a series of interconnects including Inter-Integrated Circuit (I2C), PCIe, USB, Double Data Rate (DDR), High Bandwidth Memory (HBM), and Gen-Z.
[0025] The example computing node 100 may be a network switch, a router, a network server, etc. As such, the components of the example computing node 100 may vary to include fewer components or additional components. For example, the computing node 100 may include multiple processors 108 or a platform controller hub. Alternatively, the computing node 100 may include one or more processors 108 but not a BMC 116. Thus, authentication and verification may be performed out-of-band via the BMC 116. Alternatively, authentication and verification may be performed in-band via the processor 108.
[0026] In an example of the present disclosure, one or more of the components of the system 100 may include a firmware metric certificate 146 that can be used to authenticate that the firmware of the component of the compute node is correct and trustworthy. The firmware metric certificate 146 can provide a trustworthy metric of the firmware running on the component. The metric is the value of a binary image of the firmware loaded into memory for execution. Therefore, the example authentication and verification system can use the firmware metric certificate 146 to determine whether the associated component is running firmware according to the manufacturer's specifications or whether the firmware may be compromised. The example compute node 100 can communicate with other compute nodes in the same node group or on different networks across Gen-Z or IP networks. For example, the authentication and verification of components can be managed from outside the compute node 100 on a rack, node cluster, or fabric management system. Authentication managed from outside the compute node 100 can be performed across the Gen-Z interconnect 140 or IP networks 142, 144.
[0027] The hardware and firmware of the component can be authenticated in response to a request from an authentication initiator (also referred to herein as an initiator). The initiator can authenticate a computer component (e.g., USB component 114) by executing a series of calls to the component, which are also referred to herein as responders. Examples of initiators can include software or firmware executing on the computing node 100. Example initiators can include an operating system executing on the processor 108 and firmware of the BMC 116.
[0028] Figure 2 2 is an example system 200 for securing a node group. Example system 200 may include a node group 202, a node group manager 204, an authentication database 206, and a known metric server 208. Node group 202 may be a chassis enclosure, a rack, a node cluster, or other grouping of computing nodes. Node group 202 may include a plurality of nodes 210, which may be computing nodes such as switches, routers, blade servers, etc. Node 210 may include a controller 212 and one or more components 214. Components 214 may include, for example, Figure 5Various computing node components, such as the components described above, may be included. Components may include firmware 216 and one or more firmware metric certificates (FMCs) 218. Firmware 216 may be instructions installed on component 214 for operating component 214. FMCs 218 may be similar to digital certificates, which are electronic documents that can be distributed by a certification authority. Digital certificates can ensure the trustworthiness of component 214, compute node 210, and the like. Firmware metric certificates 218 enable component certification manager 220 to ensure the trustworthiness of firmware 216 by providing accurate measurements of firmware 216 loaded into memory for operating component 214. In an example, the measurements in firmware metric certificates 218 may be compared with measurements in known metric servers 208. Known metric servers 208 may include measurements of firmware provided by the component manufacturer, indicating measurements of a binary image of the firmware installed on component 214 during manufacturing or during a legitimate update from the manufacturer. Therefore, such a comparison may be useful for determining whether firmware 216 is trustworthy or compromised.
[0029] The controller 212 can be a BMC or a processor, such as the BMC 116 and the processor 108. The example node 210 can include multiple controllers 212, including a combination of BMCs and processors. The controller 212 includes a component authentication manager 220, which can be firmware that performs authentication of the components 214 on all nodes 212 in the group when the controller 212 is powered on, exits a low-power state, or as needed. The component authentication manager 220 can perform authentication of the hardware and firmware 216 of the components 214. Therefore, the controller 212 can represent an initiator, and the component can represent a responder. In an example, a controller 212 can authenticate itself and another controller 212. For example, the BMC 116 can authenticate itself, another BMC 116, and the processor 108.
[0030] The component authentication manager 220 may also provide an API that can be used to perform authentication operations on individual components 214 or groups of components. Further, when authentication is performed, the component authentication manager 220 may store details of the authentication in the authentication database 206. Details about the authentication may include, for example, when the authentication was performed. The details in the example node 210 may specify one or more interconnects between the processor 108 and the BMC 116: PCIe, low pin count (LPC), inter-integrated circuit (I2C), through which the authentication is performed. Thus, the API can be used to query the authentication database 206 for details about the authentication.
[0031] The node group manager 204 can verify the authentication of components 214 in the node group 202 by issuing calls to the component authentication manager 220 API. The node group manager 204 can issue these calls in scripts and computer programs that automatically enforce security policies for the node group 202 by checking, for example, whether the authentication of the components is current. If the authentication is not current according to the security policy for the node group 202, the node group manager 204 can perform authentication on one or more components 214. In another example, the node group manager 204 can use the component authentication manager 220 API to check whether the authentication on a particular component 214 is performed through the same interconnect that the controller 212 currently uses to execute the firmware 216 of the component 214. If the authenticated interconnect is different from the current interconnect, the component 214 may be compromised. Therefore, in such an example, the node group manager 204 can take additional steps to isolate the potentially compromised component 214 according to the predetermined security policy.
[0032] Node group 202, node group manager 204, authentication database 206, and known metric server 208 may communicate via a network 222. Network 222 may be a computer communication network or collection of networks, including a local area network, a wide area network, the Internet, and the like.
[0033] Figure 3 2 is an example firmware metric certificate 300 used in a system for protecting a node group. The firmware metric certificate 300 is an example of a firmware metric certificate 218 that can be used to protect the firmware 216 of each component 214. Return to Reference Figure 3Firmware metric certificate 300 includes issuer 302 and attributes 304. Issuer 302 can be the name of the issuer. This name can be associated with a public key signed by the certificate issued to the issuer, with the root being self-signed. The root certificate authority key can be known. There may be no higher authority than the root certificate authority, which represents a trust anchor known to the authenticated initiator. The trust anchor can be represented in a local data store that identifies trustworthy certificate authorities. For example, when the identity certificate is used in a web browser, the web browser provider can configure the trust anchor into the web browser before releasing it for general use. Similarly, the initiator can have a trust anchor, i.e., one or more known root certificate authorities trusted by the initiator's manufacturer. A signature can be applied to the entire certificate, but the signature is a separate structure (not shown). In this example, issuer 302 can represent a core root of trust or a specific layer of firmware 216. Attributes 304 can include component ID 306, cumulative hash 308, metric 310, alias ID 312, and random number 314. Similar to the subject of the identity certificate, alias ID 312 may be a public key identifying the owner of the firmware metric certificate 300. Alias ID 312 may be authenticated by the initiator against a trust anchor. To limit authentication of the firmware metric certificate 300 to one layer of firmware 216, alias ID 312 may be made unavailable to higher layers of firmware 216. Component ID 306 identifies a public-private key pair assigned to component 214. In an example, component ID 306 may be used to sign the first firmware metric certificate in a hierarchy (i.e., layer L0 of firmware 216). The firmware metric certificates 300 for multiple layers of firmware 216 may form a hierarchy, where each firmware metric certificate 300 is issued using an alias ID in the firmware metric certificate 300 of the previous layer of firmware 216. Component ID 306 may be used to sign the first firmware metric certificate 300 in the hierarchy, and subsequent firmware metric certificates 300 may be signed using alias ID 312 in the previous layer of firmware 216. In an example, component ID 306 may not be accessible outside of the core root of trust.
[0034] The cumulative hash 308 can be a cryptographic hash representation of all layers of the firmware 216 up to the protected layer of the firmware 216. An example equation for calculating the cumulative hash 308 on layers 0 through n of the firmware 216 is shown in Equation 1. In Equation 1, H_ represents a cumulative hash function, and H represents a hash function approved by the National Institute of Standards and Technology (NIST). Additionally, in Equation 1, the symbol "||" represents the concatenation of fields or functions.
[0035] H_(L0)=H(0||H(L0))
[0036] H_(Ln )=H(H_(L n-1 )||H(L n ))
[0037] Equation 1
[0038] As previously described, firmware metric certificate 300 can be generated by non-updatable, trusted hardware or firmware code, such as a core root of trust, which runs during the first phase of initialization component 214. The core root of trust can measure the next layer or layers of firmware 216 by obtaining their cryptographic hashes. Each measurement can be provided to cumulative hash 308 and included in measurement 310. In an example, alias ID 312 can be generated by the core root of trust to authenticate firmware metric certificate 300. More specifically, alias ID 312 can be generated based on cumulative hash 308 or measurement 310. Alias ID 312 can be used to sign the next layer. Therefore, once initialization is complete, alias ID 312 can be used for the next layer of firmware 216 within the responder and used to digitally sign the firmware metric certificate 300 of the subsequent next layer of firmware 216. For example, alias ID ID1 can be used for layer L1. Layer L1 can measure layer L2, which can generate alias ID ID2, and publish a firmware metric certificate 300 signed with alias ID ID1. The signature certifies the cumulative hash 308 and metric 310 of layer L2, as well as alias ID 2. Alias ID ID2 can then be made available to layer L2.
[0039] Figure 44 is an example system 400 for protecting a group of nodes. System 400 includes multiple components, an initiator 402 in communication with a responder 404 via one or more interconnects 406 and / or a network 408. In an example, a controller such as controller 212 can initiate an authentication process for a component such as component 214. Thus, controller 212 can represent initiator 402, and the authenticated component 214 can represent responder 404. For example, initiator 402 can be a general-purpose computer processor that performs in-band authentication of component 214 in a compute node 210 without a BMC. Alternatively, initiator 402 can be a BMC that performs out-of-band authentication of other components 214 in a compute node or a general-purpose computer processor using controller 218 such as a processor and a BMC. Although the authentication performed in the examples of the present disclosure can be performed on various components 214, for simplicity, the authentication is described in the context of authentication of responder 404 as a network interface controller (NIC). Thus, the BMC initiator and the NIC responder can be connected to the PCIe interconnect and communicate via the PCIe interconnect. Additionally, the NIC responder can be connected to the network 408. The BMC initiator can authenticate all components 214 in the computing node 210, or perform authentication on the NIC in particular in response to a call from the node group manager 204. In order to determine whether the responder 404 is trustworthy, the initiator 402 can use the authentication service 428. In an example, the authentication service 428 can be a computer application that uses a digital certificate under a public key infrastructure (PKI) to determine the trustworthiness of the responder 404. If a hacker or other malicious user has control, the responder 404 may not be trustworthy. If the responder 404 is a counterfeit hardware component, or if the firmware 216 on the responder 404 is counterfeit, the malicious user can control the responder 404.
[0040] Initiator 402 and responder 404 can reside on the same computing node and therefore communicate via an interconnect 406 that passes messages between initiator 402 and responder 404 based on a specific protocol. Interconnect 406 can include one or more interconnects such as USB, PCIe, Gen-Z, etc. Initiator 402 and responder 404 can include a protocol engine 426 that can ensure that messages between initiator 402 and responder 404 are provided in a format that conforms to the protocol of the relevant interconnect 406. In some embodiments, initiator 402 can use multiple protocol engines 426 to handle interconnections with different types of components such as baseboard management controllers (BMCs) and general-purpose computer processors. Initiator 402 and responder 404 can also reside on different computing nodes 210. In this case, a network component on the computing node can provide a connection to network 408, which can include an Internet Protocol network such as a local area network, a wide area network, and the Internet.
[0041] To determine whether responder 404 is trustworthy, initiator 402 can authenticate responder 404 by verifying the public-private key pair 416 of responder's component identification (ID) certificate 410 to confirm that responder 404's hardware is authentic, i.e., not counterfeit. Component ID certificate 410 can be a public key certificate that verifies the identity of the manufacturer of responder 404. Proving the identity of the responder's manufacturer can ensure that responder 404's hardware is trustworthy. Any entity, such as initiator 402, wishing to authenticate responder 404 can read component ID certificate 410. In an example, initiator 402 can authenticate responder 404 by confirming the public key in public-private key pair 416 and determining whether responder 404 possesses the private key corresponding to the public key. If responder 404 possesses the private key, initiator 402 can determine that the responder's hardware is trustworthy. The public key can be confirmed by verifying that component ID certificate 410 or a certificate chain is signed by a trusted party. Once the public key can be trusted, the initiator 402 can challenge the responder 404 to prove possession of the corresponding private key. The challenge may involve having the responder 404 sign a random number 438 using the private key. The random number 438 can be a relatively large random number, for example, 256 bits that are used only once. The initiator 402 can also use the public key to apply an algorithm to the random number 438 and use the resulting value to determine whether the random number 438 signed by the responder 404 was signed using the corresponding private key. If so, the initiator 402 can determine that the responder's hardware is trustworthy.
[0042] In addition, responder 404 includes firmware 216, which can be a computer application that performs the operations of responder 404. For example, a NIC responder can operate a physical network such as an Ethernet, a wireless network, or a radio network. A NIC responder can also send and receive data packets from one computing node 210 to another computing node. Another example responder 404 can be a disk controller. Disk controller responder 404 can read data from a hard drive and write data to a hard drive according to a storage device protocol such as Serial Advanced Technology Attachment (SATA). Firmware 216 can include one or more layers, each of which represents a computer application that is executed in a specified order. Therefore, the operations of the responder are performed by executing the layers of firmware 216 in this order. In an example, determining whether responder 404 is trustworthy can also involve determining whether firmware 216 is trustworthy. In such an example, initiator 402 can determine whether firmware 216 is trustworthy by verifying one or more firmware metric certificates 414 of firmware 216.
[0043] As previously described, firmware metric certificate 414 may be an attribute certificate, a digital document that describes attributes associated with the issuer and the holder. In an example, the attributes described by firmware metric certificate 414 may be the measurements of a binary image of firmware 216 loaded into memory for execution. Attribute certificates may be associated with public key certificates, such as component ID certificate 410. In this manner, firmware metric certificate 414 may extend the certificate chain used to authenticate responder 404. Thus, component ID certificate 410 may describe the identity of the manufacturer of responder 404, while firmware metric certificate 414 may describe the properties of firmware 216 used to operate responder 404. Similar to component ID certificate 410, firmware metric certificate 414 may include public key 432. Public key 432 may be part of a public-private key pair of mathematically related keys used in asymmetric encryption schemes. In an example, an algorithm may be used to derive the public-private key pair for each firmware metric certificate 414, the output of which depends on the properties of the firmware or hardware logic that generated certificate 414, such that changes in the properties will result in a different public-private key pair. Additionally, firmware metric certificate 414 may include a cumulative hash 418, which may represent measurements of firmware 216 when loaded into computer memory (not shown) for execution. In an example, cumulative hash 418 may be compared to a binary image of expected firmware 422 stored on known metric server 434. In an example, initiator 402 may cache the measurements from known metric server 434 for comparison with firmware metric certificate 414. Expected firmware 422 may be a binary image of firmware installed on responder 404 during manufacturing or during a legitimate update from the manufacturer. Therefore, if cumulative hash 418 does not match the cumulative hash of expected firmware 422, firmware 216 may not be trustworthy. Therefore, initiator 402 may refuse to use responder 404. In an example, firmware metric certificate 414 may be issued by a verified component (i.e., responder 404) during the initialization (power-up) process and possibly at other times. Therefore, firmware metric certificate 414 may reside in and be retrieved from responder 404. Alternatively, the firmware metric certificate 414 may be temporarily stored in memory for caching after being retrieved from the responder 404. Since the firmware metric certificate 414 is signed, it can be cached safely. Any tampering will invalidate the signature and, therefore, the firmware metric certificate 414.
[0044] Each layer of firmware 216 can be associated with one of firmware metric certificates 414. In some examples, each firmware metric certificate 414 can include a unique public key 432. In such an example, a chain of firmware metric certificates 414 can be created between different layers of firmware 216. In other words, layer n of firmware 216 can attest to the public key of the next layer (layer n+1). Layer n+1 then uses the private key associated with the attested public key to sign the firmware metric certificate 414 of layer n+2. In other examples, a single private-public key pair can be used for all layers of firmware 216 on responder 404. In such an example, the different layers of firmware 216 can be linked together by updating the cumulative hash 418 of the firmware metric certificate 414 of each layer of firmware 216. Therefore, to verify the link between two layers of firmware 216, initiator 402 can compare the cumulative hash 418 of the firmware metric certificate 414 of each layer.
[0045] To ensure its trustworthiness, firmware metric certificate 414 may be generated by core root of trust 424. Core root of trust 424 may comprise non-updatable hardware or firmware installed by the original manufacturer of responder 404, which may be trusted to create firmware metric certificate 414 representing a binary image of actual measurements of firmware 216. In an example, firmware 216 may comprise multiple layers. Each layer may represent a portion of computer instructions for operating responder 404. Each layer may be executed in a prescribed order. Because firmware 216 may comprise multiple layers, each layer may be vulnerable to compromise by a malicious user. Therefore, firmware metric certificate 414 may be generated for each layer. To ensure the trustworthiness of the generated firmware metric certificate 414, firmware metric certificate 414 may be generated for each layer using a previously authenticated layer. In an example responder 404 having multiple layers of firmware 216, core root of trust 424 may generate a first firmware metric certificate 414 representing the first layer of firmware 216. Subsequently, before executing the second layer of firmware 216, the first layer can generate a firmware measurement certificate 414 for the second layer, thereby ensuring that the cumulative hash 418 of the second layer accurately represents the measured binary image of the second layer. Alternatively, the core root of trust 424 can generate a single firmware measurement certificate 414 that can be used to authenticate all layers of firmware 216.
[0046] Alternatively, a single firmware metric certificate 414 may be used to authenticate multiple layers of firmware 414. Thus, a single firmware metric certificate 414 may include metrics 436. Metrics 436 may represent a hash of a binary image of each layer of firmware 216. Thus, there may be one metric 436 for each layer of firmware 216 up to the layer of firmware 216 represented by firmware metric certificate 414. For example, if firmware 216 includes layers L0, L1, and L2, metrics 436 of firmware metric certificate 414 may include three hashes: a hash of each of the binary images of layers L0, L1, and L2.
[0047] Additionally, the firmware metric certificate 414 may include a random number 438. The random number 438 may ensure the freshness of the measurement 436 and ensure execution of the core root of trust 424. The random number 438 may be provided by the initiator 402 to the responder 404 during the challenge-response protocol for authentication. Alternatively, the initiator 402 may write the random number 438 to a specific memory location or register in the responder 404. Since the firmware metric certificate 414 is generated at power-up or after a reset, the random number 438 may be stored in a persistent location, such as in the responder 404. It should be noted that for the first authentication of the responder 404, there may not be a random number 438 available for the firmware metric certificate 414. However, after the first authentication, the initiator 402 may provide the random number 438 that can be written to a persistent storage device in the responder 404.
[0048] Responder 404 may include a signature service 430 and a protocol engine 426. Signature service 430 may provide secure storage. Examples of signature services include trusted platform modules and field programmable gate arrays. A trusted platform module may be a secure coprocessor that operates in response to a prescribed command set that can be used to securely store data, including the operating state of a computing platform such as a compute node. A field programmable gate array (FPGA) may be an integrated circuit that can be programmed using a hardware description language to execute specific instructions. In this way, an FPGA is similar to a processor. However, in contrast, a processor may be pre-programmed with a complex instruction set.
[0049] Figure 5is a process flow diagram of a method 500 for generating a firmware metric certificate for protecting a node group. At block 502, a component of a compute node having firmware may be powered on or begin reinitialization. The component may be a responder, such as responder 404. An example responder 404 may include any component of the compute node 100, such as the processor 108, the Gen-Z component 110, etc. For the purposes of this discussion, one of the PCIe components 130 of the compute node 100 is used as an example component. The PCIe component 130 may be powered on during a power cycle of the compute node 100. Alternatively, the PCIe component 130 may be powered on when the compute node 100 exits a low-power or reset state. A random number, such as random number 438, may also be read from a persistent storage device. In an example, the random number 438 may be provided by the initiator 402. For example, the random number 438 may be provided from a previous authentication request (e.g., authentication request 314). In such an example, the responder 404 may store the random number 438 in a persistent storage device within the responder 404. In another example, the initiator 402 can write the random number 438 to a register in the responder 404. For example, a PCIe device such as the PCIe component 130 can expose a register that can be read and written via the PCIe bus. The value stored in such a register can make the random number 438 persistent across resets or power cycles.
[0050] Additionally, the first layer L0 of the firmware 412 of the PCIe component 130 may be loaded into the memory. As previously described, the PCIe component 130 may be a Figure 4 An example of a responder 404 is described. The PCIe component 130 may be powered on by an operating system of the computing node 100 at block 502. The method 500 may be further performed by a core root of trust, such as the core root of trust 424, of the PCIe component 130.
[0051] At block 504, a hash may be generated for the next layer of firmware 412 of the PCIe component 130. For example, after power-up, layer L0, which may represent the immutable core root of trust 424, may measure layer L1 of the firmware 412. The hash may be a NIST-approved hash of a binary image of layer L1.
[0052] At block 506, a firmware metric certificate, such as firmware metric certificate 300, may be generated for the next layer of firmware 412. For example, core root of trust 424 may generate firmware metric certificate 300 for layer L1. Firmware metric certificate 300 may include issuer 302, component ID 306, alias ID 312, and cumulative hash 308 or metrics 310. Issuer 302 and component ID 306 may be the component ID of PCIe component 130. The component ID of PCIe component 130 may be considered layer L0, which may be considered to issue firmware metric certificate 300. Alias ID 312 may be a public key identifying the owner of firmware metric certificate 300. The owner of firmware metric certificate 300 may be the next layer, i.e., layer L1 after power-up. Cumulative hash 308 or metrics 310 may be populated based on the hash generated at block 504. Cumulative hash 308 may be determined based on Equation 1. Alternatively, metrics 310 may be populated with the generated hash. Additionally, the keyCertSign bit of the firmware metric certificate 300 may be cleared to prevent malicious users from creating fake certificate authorities.
[0053] In an example, firmware 412 may include one or more layers. Thus, blocks 504 through 506 may be repeated for each subsequent layer of firmware 412. However, blocks 504 through 506 may be performed by the current layer of firmware 412 rather than by the core root of trust 424. Thus, layer L0 may generate a firmware metric certificate 300 for layer L1. Layer L1 may generate a firmware metric certificate for layer L2, and so on. If firmware 412 includes one layer, method 500 may proceed to block 508.
[0054] At block 508 , layers of firmware 412 associated with the generated firmware metric certificate 300 may be executed. Executing firmware 412 may involve operating PCIe component 130 .
[0055] At block 510, an initiator, such as initiator 402, may authenticate the firmware metric certificate 300 generated at block 506. Figure 3 Authentication is performed in the manner described. If authentication fails, method 500 may end. However, if firmware metric certificate 300 is authenticated, method 500 for executing firmware 412 may continue. Further, in some scenarios, the same firmware metric certificate 300 may be authenticated multiple times. If compute node 100 has not yet been power cycled but requests authentication of firmware 412, multiple authentications may be performed. Additionally, the random number provided by initiator 402 during authentication may be written to PCIe component 130.
[0056] In an example, additional firmware metric certificates 300 may be added after the firmware 412 is executed. If the operating system loads the additional firmware 412 into the PCIe component 130 shortly after the operating system boots, additional firmware metric certificates 300 may be generated. If the operating system updates the firmware 412, additional firmware metric certificates 300 may also be generated. In this case, the method may proceed to block 504. In the case where a single firmware metric certificate 300 is used for all layers of the firmware 412, the firmware metric certificate 300 may be updated rather than adding a new firmware metric certificate 300.
[0057] In some examples, multiple firmware certificates 300 may be generated in blocks 504 through 510, one for each layer of firmware 412. In such an example, the component ID 306 of each firmware metric certificate 300 may be the component ID of the PCIe component 130. The issuer 302 of such a firmware metric certificate 300 may be the previous layer of firmware 412. The alias ID 312 may be the public key identifying the current layer of firmware 412. The random number 438 may be a random number stored in the PCIe component 130 by the initiator 402. In examples with multiple firmware metric certificates 300 for a component (such as PCIe component 130), a cumulative hash 308 may be populated. As described with respect to Equation 1, the cumulative hash 308 may be the concatenation of the cumulative hash of the previous layer and the NIST-approved hash function of the binary image of the current layer. The NIST-approved hash function of the binary image of the current layer may be additionally appended to the metric 310.
[0058] In some examples, a single firmware metric certificate 300 representing all layers of firmware 412 can be generated. In such examples, instead of generating a new firmware metric certificate 300 for each layer, a new firmware metric certificate 300 can be issued with an updated signature. This new firmware metric certificate 300 can replace the previously issued firmware metric certificate 300. This new firmware metric certificate 300 can be generated by the current layer of firmware 412. In the new firmware metric certificate 300, component ID 306 and random number 314 may not be changed. However, issuer 302 may be the current layer of firmware 412. Alias ID 312 may be a public key identifying the next layer of firmware 412. In an example with a single firmware metric certificate 300 for PCIe component 130, cumulative hash 308 and metrics 310 may be populated. As described with respect to Equation 1, cumulative hash 308 may be a concatenation of the cumulative hash of the previous layer and a NIST-approved hash function of the binary image of the current layer. Thus, the metrics 310 from the previously issued firmware metric certificate 300 may be supplemented with new metrics 310 for the next layer of firmware 412. The metrics 310 may be a NIST-approved hash function of a binary image of the next layer of firmware 412.
[0059] It should be understood that Figure 5 The process flow diagram is not intended to indicate that method 500 includes Figure 5 Further, any number of additional blocks may be included within method 500, depending on the details of the specific implementation. In addition, it should be understood that Figure 5 The process flow diagram is not intended to indicate that the method 500 must be performed in every case. Figure 5 For example, block 504 may be rearranged to occur before block 502.
[0060] Figure 6 6 is a message flow diagram 600 for authenticating components in a system for protecting a node group. Message flow diagram 600 may represent a message flow between an authentication initiator 602 and a responder 604. Initiator 602 may represent a component such as initiator 402 and may include an authentication service 606 and a protocol engine 608. Protocol engine 608 may translate messages between initiator 602 and responder 604 based on the interconnections between initiator 602 and responder 604. Responder 604 may represent a component with firmware, such as responder 404. Message 610 represents a request from initiator 602 for a responder's certificate chain (or multiple certificate chains). Message 612 represents a certificate chain sent by responder 604 to initiator 602 in response to the request. In response to receiving the responder's certificate chain, initiator 602 may verify the one or more certificate chains and select a public key to be authenticated by responder 604. The public key may be selected from the leaf certificates of a valid certificate chain.
[0061] Message 614 may represent an authentication request from initiator 602 to responder 604. The authentication request may consist of a large random number and the selected public key to be authenticated. Typically, the public key to be authenticated is identified because responder 604 may have multiple public-private key pairs for different purposes.
[0062] Upon receiving the authentication request, the responder's protocol engine can extract the random number and the identity of the public key to be authenticated from the authentication request. Alternatively, the responder's signing service can sign the concatenation of the random number and the internally generated random obfuscation using the private key corresponding to the identified public key. The purpose of the obfuscation is to prevent chosen-plaintext attacks, so the obfuscation should be unpredictable to the initiator. Message 616 can represent the responder's response to the authentication request containing the obfuscation and the signature of the concatenation of the random number and the obfuscation.
[0063] Upon receiving the response to the authentication request, the initiator 602 can verify that the random number and the obfuscation have been signed by the private key corresponding to the public key in the leaf certificate. If the verification is successful, the responder 604 has been authenticated.
[0064] Figure 7 7 is an example system 700 for protecting a node group. The example system 700 includes a node group manager 702, a baseboard management controller (BMC) 704, a processor 706, and a known metric server 708. The BMC 704 and the processor 706 may represent information about Figure 2 Specific examples of controllers 212 are described. Additionally, a BMC 704 and a processor 706 may be implemented on one or more nodes 210 within a node group (not shown). Similar to system 200, an example node group manager 702 may protect one or more node groups by authenticating each node 210. Authentication of each node 210 may include authenticating the trustworthiness of the hardware and firmware of the installed BMC 704, processor 706, and components 712. In an example, components 712 may be connected to either or both of the BMC 704 and processor 706 via one or more interconnects including PCIe, USB, I2C, etc. Additionally, the node group manager 702 may be configured to report inventory status of each node via an IP network. The inventory status may indicate various details regarding the authentication of the inventory of components 712 for a node 210. The inventory status may include identifiers of the authenticated BMC 704, processor 706, or component 712, the time of authentication, the interconnect(s) over which authentication was performed, and the like. For example, being able to determine the inventory status of parts 712 can enable detection and monitoring of service engineer activities. In this way, the possibility of erroneous or malicious insertions during service activities can be reduced. Service activities refer to the act of installing new parts 712 or updating firmware.
[0065] BMC 704 and processor 706 may include a protocol engine 714, an authentication service 716, and a component authentication manager 718. The authentication service 714 and protocol engine 716 may represent Figure 6 Examples of authentication services 606 and protocol engines 608 are described. Return to Reference Figure 7 , the component authentication manager 718 can use the authentication service 714 and the protocol engine 716 to perform authentication during initialization and at the request of the node group manager 702.
[0066] During initialization, the component authentication manager 718 can direct the authentication service 714 to authenticate each component 712 that is reachable on each interconnect or structure connected to the corresponding BMC 704 and processor 706. Some components 712 can be accessed via more than one interconnect. For example, the BMC 704 can access the component 712 via both PCIe and I2C. In such an example, a point-to-point I2C connection can be used to distribute confidential material, such as encryption or authentication keys. In this way, the I2C connection can be used instead of using a shared PCIe structure. The shared PCIe structure can have a relatively higher bandwidth than the I2C interconnect and can be used for data transmission. However, the PCIe structure may be more vulnerable to snooping than the I2C. Therefore, by authenticating the component 712 via both its PCIe and I2C connections, the BMC 704 can ensure that the component 712 is securely connected. The keys distributed via the I2C connection can be used to encrypt and decrypt data stored within the compute node. Alternatively, the encryption key may be used for authenticated and encrypted data transfer of data stored in unencrypted computing nodes.
[0067] By authenticating components 712 on available interconnects and structures, component authentication manager 718 can generate an authenticated list of all active components 712 connected to the corresponding BMC 704 or processor 706. More specifically, component authentication manager 718 can include a data store with a list of authenticated components, their certificate chains, and a status of whether each certificate chain has been authenticated. Generating and maintaining this authenticated list over time can provide an inventory tracking record that can be verified against a known metric server 708 to ensure the trustworthiness of components 712 and associated firmware. Component authentication manager 718 can also build and maintain a connection topology for components 712. Alternatively, BMC 704 can be connected to processor 706 via an interconnect or structure. In this way, BMC 704 and processor 706 can authenticate each other.
[0068] In an example, the BMC 704 may serve as the local authority for all communications with the node group manager 702. Figure 7A single BMC 704 is shown connected to a node group manager 702, but in the example, the node group manager 702 can maintain connections to multiple BMCs. Similarly, although the system 700 includes only one node group manager 702, in the example, a hierarchical, federated, or other clustered authentication and verification manager group can exist. The component authentication manager 718 in the BMC 704 can obtain an authenticated list from the component authentication manager 718 in the processor 706. In the example, the authenticated list can be stored in an attribute certificate containing a list of component identifiers that have been authenticated by the controller, as well as metadata such as an interconnect identifier, a timestamp, etc. Such an attribute certificate can also list authentication failures (e.g., identifiers for components 712 that failed authentication, are unresponsive, or have unresponsive connections). Once the processor 706 has been authenticated to the BMC 704, an attribute certificate can be issued by the processor 706. Because the attribute certificate can be a signed assertion, it can be stored, copied, and transmitted over a network without the risk of a malicious user adding or deleting information about which components 712 have been authenticated.
[0069] Authentication failures may be detected by electrical connections to unresponsive components 712, loss of authentication of previously authenticated components 712, and changes to the connection topology of components 712. Therefore, component authentication manager 718 and node group manager 702 may store previous connection topologies, inventory reports, and authentication statuses for comparison.
[0070] The node group manager 702 can authenticate the BMC 704 by using the BMC's component authentication manager 718 to obtain an attribute certificate that includes an authenticated list of components 712 authenticated by the BMC 704. Additionally, the node group manager 702 can use the BMC's component authentication manager 718 to obtain an authenticated list of processors 706. In an example, the list of components for a node, such as node 210, can be included in a single signed assertion or certificate. In such an example, the certificate or assertion can reference other certificates and assertions.
[0071] The node group manager 702 can also use the component authentication manager 718 to obtain all certificate chains of all components 712 and can request the component authentication manager 718 to authenticate any certificate chain of any connected component 712 or controller. In the case of a component connected to a processor 706, the component authentication manager 718 of the BMC can forward the authentication request to the component authentication manager 718 of the processor.
[0072] In an example, the component authentication manager 718 may provide an application programming interface (API). In such an example, the API may include operations for listing components, obtaining certificate chains, querying status, authenticating components, and other related operations. The list components operation may return one or more attribute certificates. Each attribute certificate may list the authenticated components 712 associated with a particular component authentication manager 718, as well as related metadata such as interconnect identifiers and any authentication failures that may have been detected. In an example, the component authentication managers 718 may be logically configured in a hierarchical structure. Therefore, certain component authentication managers 718 may represent different levels of the architecture in a tree-like structure. Thus, the component authentication managers 718 at the bottom of the tree-like structure may be referred to as leaves. In an example, a component authentication manager 718 at a leaf in the hierarchy may return a single attribute certificate. For component managers 718 higher in the hierarchy, multiple attribute certificates may be returned. In an example, variations of the list components operation may also be included. Such variations may provide information about the topology of the components 712, as well as authentication information such as authentication timestamps.
[0073] The get certificate chain operation may provide a certificate chain retrieved from component 712. Thus, a call to the get certificate chain operation may specify the component 712 for which to retrieve the certificate chain.
[0074] The query status operation can provide the authentication status of a specified component 712. If the component 712 has been authenticated, the query status operation can provide the authenticated public-private key pair 416. Additionally, for each public-private key pair 416, the query status operation can provide the number of times the public-private key pair 416 has been authenticated since the component 712 was last initialized and the timestamp of the last authentication. Furthermore, for authenticated components 712, the identity of the authenticated initiator can be provided. For components 712 that fail authentication, an error code can be provided.
[0075] The authentication operation may specify an identifier of the component 712 being authenticated, a public key from the public-private key pair 416 of the component 712, and a random number. Thus, the component authentication manager 718 may use the given random number to authenticate the identified component 712 as possessing the given public key. However, in some cases, the authentication operation may be denied if the component 712 cannot be authenticated. For example, if the processor 706 is running a workload such as an application or an operating system, the component authentication manager 718 of the processor 706 may not be able to authenticate the DDR component. Instead, such authentication may be limited to the period when the processor 706 is in a specific operating mode or running specific firmware: for example, during a power-on self-test.
[0076] Using firmware metric certificates 218 to represent the authenticated manifest can prevent tampering because the manifest can be compared over time. For example, each time a node is restarted, the component authentication manager 718 can compare the authenticated manifest with the actual manifest. Once the node group manager 702 has access to the firmware metric certificates 218 for the authenticated component 712 and the certificate chain corresponding to the authenticated component 712, the component authentication manager 718 can verify the correctness of any firmware metric contained in the authenticated certificate chain. The component authentication manager 718 can verify correctness by comparing the metric or cumulative hash with the value contained in the known metric server 708.
[0077] The manner in which authentication of components 712 is accomplished may vary depending on whether the system 700 is being initialized or is running a workload. After a reset or power cycle, as part of the initialization process, the component authentication manager 718 may authenticate all components 712 on any bus or structure connected to the BMC 704 described above. The results of the authentication may be stored in the authentication database 206 in the form of an attribute certificate. Alternatively, the BMC 704 may boot the processor 706 using firmware or a reduced-function operating system so that the processor 706 may authenticate all components 712 connected to any bus or structure connected to the processor 706, such as DDR, Gen-Z, PCIe, and USB connections. Alternatively, the BMC 704 may use information about Figure 6 7. The system 700 may authenticate the processor 706 using the authentication mode described above. In an example, the system 700 may include a security coprocessor (not shown) such as a Trusted Platform Module (TPM). In such an example, the processor 706 may authenticate keys and certificates that have been issued to the TPM during manufacturing. The TPM may be attached to the same board as the processor 706. Alternatively, the TPM may be integrated within the processor package. In an example, the BMC 704 may obtain an attribute certificate (not shown) from the processor's component authentication manager 718 that lists an authenticated list of components 712 connected to the processor 706. The attribute certificate may be made available to the node group manager 702 where it may be archived for future comparison and inventory tracking. In this way, all components 712 within the system 700 may be authenticated during power cycling and initialization.
[0078] However, the system 700 may not be power cycled or reinitialized for months or years. Therefore, during runtime, the BMC 704 can run a component authentication manager 718 to authenticate all connected components 712. Additionally, the BMC 704 can authenticate any new component 712 that is the subject of a hot plug or hot insertion event, such as a USB device. In an example, the operating system and device drivers can be enhanced to support component authentication of the processor 706 and connected components. In such an example, the operating system can run an authentication service and can also run a component authentication manager 718 to report the authentication status of any component 712 that is authenticated. The operating system can maintain the presence, insertion, and removal information of all connected components 712. Further, the operating system can periodically authenticate components 712 in response to an insertion event and during runtime. When exiting from a low-power state, the operating system can also authenticate the processor 706 and connected components 712.
[0079] Alternatively, a secure processor and trusted execution environment (TEE) different from the processor 706 can be used to run the component authentication manager 718 and authentication services. In another example, a reduced-function hypervisor can be used to isolate component authentication from other security-critical functions from the main operating system. In another example, a system management interrupt (SMI) can be used to run the component authentication manager 718 and authentication services on the processor 706. In such an example, the BMC 704 can generate an SMI to trigger the execution of a handler that runs the component authentication manager 718 and authentication services. In this way, the BMC 704 can trigger authentication of one or more components 712 connected to the processor 706. In some scenarios, the component 712 may experience a deep low power event. In such a scenario, runtime authentication can be used to detect any component replacement that occurs during such an event.
[0080] Figure 8 800 is an example system for protecting a node group. Other arrangements of authentication are possible and may be more useful for different architectures. For example, system 800 includes a cluster authentication and verification manager 802, a rack authentication and verification manager 804, a processor 806, and a known good metric 808. Cluster authentication and verification manager 802 and rack authentication and verification manager 804 may be similar to those described with respect to Figure 7 Node Group Manager 702 as described above, but operates at a different hierarchical level to address manageability and scalability requirements. Figure 8 , the processor 806 may include a protocol engine 814, an authentication service 816, and a component authentication manager 818. The authentication service 814 and the protocol engine 816 may represent Figure 6 Examples of authentication services 606 and protocol engines 608 are described. Return to Reference Figure 8 , the component authentication manager 818 may use the authentication service 814 and the protocol engine 816 to perform authentication during initialization and at the request of the cluster authentication and verification manager 802 .
[0081] The system 800 may represent a node without a BMC, where a single component authentication manager 818 runs on the processor 806 and is connected to the rack authentication and validation manager 804. The system 800 may perform "in-band" authentication, meaning that authentication may occur over the same interconnect used by applications and services running on the processor 806. In an example, the rack authentication and validation manager 804 may be connected to multiple processors, for example, all processors in a rack. The rack authentication and validation manager 804 may perform the same functions as a cluster manager, such as validating firmware, but only for those components 812 in the rack. Additionally, as discussed with respect to Figure 7 As described, the rack certification and validation manager 804 can provide the same API to the cluster certification and validation manager 802 that the component certification manager 718 provides to the node group manager 702. In this way, the cluster certification and validation manager 802 can obtain a cluster-wide view of the certified inventory and enforce re-certification of any component 812 through the rack certification and validation manager 804. In another example, the system 800 can include a dedicated network interface for the management plane, wherein access to the component certification manager 818 can be restricted to the dedicated network interface.
[0082] In another example, some components 812 may not be associated with a node. Instead, such components 812 may be associated with a rack or another component such as a power distribution unit, a fan, or a sensor. In such an example, the same mechanisms and principles described above may be used to authenticate the components 812. Thus, the component authentication manager 818 and the authentication service 814 may be embedded in the rack authentication and verification manager 804 and perform authentication using any protocol used to control the rack power distribution unit, fan, or sensor.
[0083] In an example, known good metrics 808 may include a manifest describing the expected metrics of component 812. This manifest may be created by the manufacturer of a platform such as node 210. More specifically, during the final stages of manufacturing of node 210, for example, during testing, component 812 may be powered on and initialized. Furthermore, a firmware metric certificate may be generated for the component. Thus, cluster certification and validation manager 802 and rack certification and validation manager 804 may determine the component identity and firmware metrics of component 812. The manufacturer or system integrator may then issue a set of "manifest" certificates specifying which components and firmware are in node 210, i.e., the "known good state" of the entire system at the time of shipment. Therefore, when node 210 arrives at a customer site, the customer can compare the manifest with the information reported by cluster certification and validation manager 802 or rack certification and validation manager 804 to ensure that node 210 has not been tampered with at the time of shipment. In one example, the manifest may be a platform certificate, which is currently under development by the Trusted Computing Group.
[0084] Regarding the authentication of node groups, a hierarchy of services can be used to manage the authentication and verification of components at the cluster level. Therefore, Table 1 provides an example hierarchy that lists the services and responsibilities for each node hierarchy level in the hierarchy. The lowest level in the hierarchy is a single node, while the highest level can be a cluster. A node can deliver compute, storage, or networking (switching) services. Most components can be contained within a node. An enclosure can contain multiple nodes and can have one or more power supplies. A rack can have multiple enclosures and power supplies or UPS (uninterruptible power supplies). A cluster can contain multiple racks.
[0085]
[0086]
[0087]
[0088] Table 1
[0089] Depending on the size and configuration of the computer hardware architecture, some elements of the hierarchy may be omitted. For example, if there are a small number of racks, the cluster service may interact directly with the enclosure authentication and verification service. It should be noted that Table 1 is only one possible hierarchy of nodes and node groups. In other examples, other hierarchical organizations may be used. For example, a cluster may be included within the hierarchy level of a data center. Furthermore, for example, a data center may be included within the hierarchy level of an information technology department or enterprise.
[0090] Figure 99 is a process flow diagram of a method 900 for securing a node group. The method 900 may be performed by a node group manager, such as the node group manager 204. At block 902, the node group manager 904 may authenticate the hardware architecture of a component 214 of the node group 202. The node group 202 may be described as a node hierarchy level. The node hierarchy level may include one or more compute nodes 210. Authenticating the hardware architecture may include verifying that the hardware of the component 214 is from a known and trusted manufacturer. Verifying that the hardware is from a known and trusted manufacturer may include validating a certificate chain for the component 214.
[0091] At block 904, node group manager 204 may authenticate firmware 216 for each component 214 of node group 202. Authenticating firmware 216 may include comparing metrics in firmware metric certificate 218 for component 214 to metrics in known metric server 208.
[0092] At block 906, the node group manager 204 may generate a certification database 206 based on the certified hardware architecture and the certified firmware 216. The certification database 206 may include a description of the certifications, such as when each component 214 was certified.
[0093] At block 908, the node group manager 204 may protect the node group 210 by using the authentication database 206. In an example, a policy for protecting a specified node group 202 may be implemented by using the authentication database 206. For example, the authentication database 206 may describe when a component 214 of the node group 202 was last authenticated. Further, an example policy for protecting the node group 202 may specify that the component 214 be authenticated at least once a month. Thus, the node group manager 204 may execute a policy script each month that checks when the components 214 of the node group 202 were last authenticated. If any component 214 has not been authenticated within the last month, the node group manager 204 may automatically authenticate the component 214 that has not yet been authenticated according to the policy.
[0094] It should be understood that Figure 9 The process flow diagram is not intended to indicate that method 900 includes Figure 9 Further, any number of additional blocks may be included within method 900, depending on the details of the specific implementation. In addition, it should be understood that Figure 9 The process flow diagram is not intended to indicate that the method 900 must be performed in every case. Figure 9 For example, block 904 may be rearranged to occur before block 902.
[0095] Figure 10is an example system 1000 that includes a tangible, non-transitory computer-readable medium 1002 storing code for protecting a group of nodes. The tangible, non-transitory computer-readable medium is generally indicated by the reference numeral 1002. The tangible, non-transitory computer-readable medium 1002 may correspond to any typical computer memory that stores computer-implemented instructions such as program code. For example, the tangible, non-transitory computer-readable medium 1002 may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, disk and optical disk include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Optical disks, where magnetic disks typically reproduce data magnetically, while optical disks reproduce data optically with lasers.
[0096] Processor 1004 may access tangible, non-transitory, computer-readable medium 1002 via computer bus 1006. Processor 1004 may be a central processing unit (CPU) to execute an operating system in system 1000. Area 1008 of tangible, non-transitory, computer-readable medium 1002 stores computer-executable instructions for authenticating the hardware architecture of each of a plurality of components of a computing node. Area 1010 of tangible, non-transitory, computer-readable medium stores computer-executable instructions for authenticating the firmware of each component. Area 1012 of tangible, non-transitory, computer-readable medium stores computer-executable instructions for generating an authentication database comprising a plurality of authentication descriptions based on the authenticated hardware architecture and the authenticated firmware, wherein a policy for protecting a specified subset of the computing nodes is implemented using the authentication database. Area 1014 of tangible, non-transitory, computer-readable medium stores computer-executable instructions for protecting a group of nodes using the authentication database.
[0097] Although shown as consecutive blocks, the software components may be stored in any order or configuration. For example, if the tangible, non-transitory computer-readable medium 1002 is a hard drive, the software components may be stored in non-consecutive or even overlapping sectors.
[0098] For the purpose of explanation, the foregoing description uses specific terms to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that specific details are not required to practice the systems and methods described herein. The foregoing description of specific examples is provided for the purpose of illustration and description. The description is not intended to be exhaustive or to limit the present disclosure to the precise form described. Obviously, in view of the above teachings, many modifications and variations are possible. These examples are shown and described to best explain the principles and practical applications of the present disclosure, thereby enabling other persons skilled in the art to best utilize the present disclosure and various examples with various modifications suitable for the intended specific use. The scope of the present disclosure is intended to be defined by the claims and their equivalents.
Claims
1. A method for protecting a plurality of computing nodes, the method comprising: authenticating a hardware architecture of each of a plurality of components of the plurality of computing nodes; Authenticating the firmware of each of the plurality of components includes: generating a firmware metric certificate for each layer of the firmware for each component, wherein the firmware metric certificate for each layer of the firmware is based on the firmware metric certificate for a certified previous layer of the firmware; and generating a certification database comprising a plurality of certification descriptions based on the certified hardware architecture and the certified firmware, wherein a policy for securing a specified subset of the plurality of computing nodes is implemented using the certification database, wherein the certification database indicates when the plurality of components were last certified, wherein the policy specifies a certification period for the plurality of components, and wherein the plurality of components are automatically certified based on the policy.
2. The method of claim 1, further comprising implementing the policy by using (i) operation of an application program interface for authenticating a group of computing nodes and (ii) operating the authentication database.
3. The method according to claim 1, wherein One of the plurality of computing nodes includes a controller connected to the plurality of components, and wherein the controller authenticates the hardware architecture of the plurality of components and authenticates the firmware of the plurality of components, wherein the controller compares the authenticated hardware architecture and the authenticated firmware to one or more previous authentications to detect changes to the hardware architecture and the firmware.
4. The method according to claim 1, wherein One of the plurality of computing nodes includes a baseboard management controller and a computer processor, and wherein the baseboard management controller authenticates the hardware architecture of the computer processor and authenticates the firmware of the computer processor.
5. The method of claim 1, wherein: A core root of trust generates the firmware measurement certificate for the first layer of the firmware; the firmware metric certificate of the first layer of the firmware specifies a trusted certificate authority; and Authentication of the firmware is based on the firmware metric certificate.
6. The method of claim 5, wherein: the firmware of the plurality of components operating the plurality of components; The firmware metric certificate includes an attribute certificate; and The firmware metric certificate includes: a cumulative hash of the layer of firmware, wherein the cumulative hash comprises a concatenation of: a hash of the layer of firmware; and a hash of each of the one or more lower layers of the firmware; and Random number.
7. The method according to claim 6, wherein: The firmware is authenticated using a trusted data store including a binary image of the firmware and a certificate chain including a hardware digital certificate and the firmware metric certificate.
8. The method of claim 1, wherein: The plurality of computing nodes are associated with a node hierarchy level, and wherein an authentication and verification manager for the node hierarchy level performs the authentication of the hardware architecture and the authentication of the firmware.
9. A system for protecting a plurality of computing nodes, comprising: processor; as well as a memory component storing instructions that cause the processor to: authenticating a hardware architecture of each of a plurality of components of the plurality of computing nodes; Authenticating the firmware of each of the plurality of components includes: generating a firmware metric certificate for each layer of the firmware for each component, wherein the firmware metric certificate for each layer of the firmware is based on the firmware metric certificate for a certified previous layer of the firmware; and generating a certification database comprising a plurality of certification descriptions based on the certified hardware architecture and the certified firmware, wherein a policy for securing a specified subset of the plurality of computing nodes is implemented using the certification database, wherein the certification database indicates when the plurality of components were last certified, wherein the policy specifies a certification period for the plurality of components, and wherein the plurality of components are automatically certified based on the policy.
10. The system of claim 9, wherein: The instructions cause the processor to implement the policy by (i) operating an application program interface for authenticating a group of computing nodes and (ii) operating the authentication database.
11. The system of claim 9, wherein: One of the plurality of computing nodes includes a controller connected to the plurality of components, and wherein the controller authenticates the hardware architecture of the plurality of components and authenticates the firmware of the plurality of components, wherein the controller compares the authenticated hardware architecture and the authenticated firmware to one or more previous authentications to detect changes to the hardware architecture and the firmware.
12. The system of claim 9, wherein: One of the plurality of computing nodes includes a baseboard management controller and a computer processor, and wherein the baseboard management controller authenticates the hardware architecture of the computer processor and authenticates the firmware of the computer processor.
13. The system of claim 9, wherein: A core root of trust generates the firmware measurement certificate for the first layer of the firmware; the firmware metric certificate of the first layer of the firmware specifies a trusted certificate authority; and Authentication of the firmware is based on the firmware metric certificate.
14. The system of claim 13, wherein: the firmware of the plurality of components operating the plurality of components; The firmware metric certificate includes an attribute certificate; and The firmware metric certificate includes: a cumulative hash of the layer of firmware, wherein the cumulative hash comprises a concatenation of: a hash of the layer of firmware; and a hash of each of the one or more lower layers of the firmware; and Random number.
15. The system of claim 14, wherein: The firmware is authenticated using a trusted data store including a binary image of the firmware and a certificate chain including a hardware digital certificate and the firmware metric certificate.
16. A non-transitory computer-readable medium storing computer-executable instructions that, when executed, cause a computer to: authenticating a hardware architecture of each of a plurality of components of a plurality of computing nodes; Authenticating the firmware of each of the plurality of components includes: generating a firmware metric certificate for each layer of the firmware for each component, wherein the firmware metric certificate for each layer of the firmware is based on the firmware metric certificate for a certified previous layer of the firmware; and generating a certification database comprising a plurality of certification descriptions based on the certified hardware architecture and the certified firmware, wherein a policy for securing a specified subset of the plurality of computing nodes is implemented using the certification database, wherein the certification database indicates when the plurality of components were last certified, wherein the policy specifies a certification period for the plurality of components, and wherein the plurality of components are automatically certified based on the policy.
17. The non-transitory computer readable medium of claim 16, wherein: The computer-executable instructions cause the computer to implement the policy by (i) operating an application program interface for authenticating a group of computing nodes and (ii) operating the authentication database.
18. The non-transitory computer readable medium of claim 16, wherein: One of the plurality of computing nodes includes a controller connected to the plurality of components, and wherein the controller authenticates the hardware architecture of the plurality of components and authenticates the firmware of the plurality of components, wherein the controller compares the authenticated hardware architecture and the authenticated firmware to one or more previous authentications to detect changes to the hardware architecture and the firmware.
19. The non-transitory computer readable medium of claim 16, wherein: One of the plurality of computing nodes includes a baseboard management controller and a computer processor, and wherein the baseboard management controller authenticates the hardware architecture of the computer processor and authenticates the firmware of the computer processor.
20. The non-transitory computer-readable medium of claim 16, wherein: A core root of trust generates the firmware measurement certificate for the first layer of the firmware; the firmware metric certificate of the first layer of the firmware specifies a trusted certificate authority; and Authentication of the firmware is based on the firmware metric certificate.
Citation Information
Patent Citations
Firmware verification method and device
CN108345805A
Program authentication on environment
US20060200859A1
Methods and systems for controlling access to computing resources based on known security vulnerabilities
US9923918B2