Data processing unit integration

By using a trusted firmware interface to abstract signaling between the DPU and the server, unified deployment of DPU operating systems across vendors and OEMs is achieved, solving the integration complexity between different DPUs and servers and improving the ease of management and maintenance.

CN118369648BActive Publication Date: 2025-10-28VMWARE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380014958.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-01-24
Filing Date
2023-01-18
Publication Date
2025-10-28
Estimated Expiration
2043-01-18

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve unified DPU operating system integration across DPUs and servers from different manufacturers, resulting in the need to deploy different software images for each type of DPU and server, increasing the complexity of management and maintenance.

Method used

By using a trusted firmware interface to abstract low-level signaling and high-level protocol signaling between the DPU and the server, a vendor-neutral DPU operating system is provided, which can uniformly deploy a single operating system image on different types of DPUs and servers, achieving cross-vendor and OEM compatibility.

Benefits of technology

It enables unified management and operating system integration across different DPUs and servers, simplifying the deployment and maintenance process and improving the system's flexibility and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118369648B_ABST
    Figure CN118369648B_ABST
Patent Text Reader

Abstract

This disclosure describes a combined DPU and server solution with integrated data processing unit (DPU) and operating system (OS). The DPU OS executes on the DPU or other computing device, wherein the DPU OS utilizes secure calls provided by a trusted firmware component of the DPU, which can be invoked by the DPU OS component to abstract DPU vendor-specific and server vendor-specific integration details. One of the secure calls made on the DPU is identified for communication with its associated server computing device. In an example where one of the secure calls is invoked, the invoked secure call is translated into an architecture-specific call or request specific to the server computing device, and the call is executed, which may include sending signals to the server computing device in a format interpretable by the server computing device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] Claiming the benefit of 35U.SC119(a) through (d) of U.S. Application No. 17 / 582,055, entitled “BARE-METAL HYPERVISOR INTEGRATION WITH NETWORKARCHITECTURE”, filed January 24, 2022, which is incorporated herein by reference in its entirety for all purposes. Background Technology

[0003] Data centers and other computing infrastructure employ various types of physical hardware, such as central processing units (CPUs), graphics processing units (GPUs), network interface cards (NICs), data processing units (DPUs), memory, and the like. Data centers use physical hardware to host customer workloads. Some customer workloads contain computing resources virtualized by a hypervisor to provide a large number of virtual machines (VMs), which provide virtualized computing software and hardware.

[0004] In some server designs and scenarios, it is beneficial to offload server input / output (I / O) functions, as well as acceleration and management tasks, to a Data Processing Unit (DPU). A DPU is a network-oriented microserver that combines general-purpose computing with high-performance networking and storage I / O, and is packaged within a server adapter form factor via PCIe, CCIX, CXL, GenZ, NVLINK, CAPI, AXI, or other buses and I / O structures, becoming a physical part of the computer server. Other recognized industry names for DPUs include SmartNIC or Intelligent Processing Unit (IPU), but other types of "intelligent" I / O adapters for general-purpose computing, such as compute storage devices (CSDs) and GPUs, can also exist.

[0005] The DPU operating system, also known as the I / O hypervisor, is a multiplexed I / O accelerator that provides additional I / O services to servers containing DPU devices by implementing additional I / O devices (such as PCIe functionality) or by preprocessing and post-processing (or filtering) the data passing through the I / O devices. A good example of this processing by the DPU is a firewall application. The DPU OS also separates I / O management from the server hypervisor that performs workload management, thereby separating server (workload) tenant permissions from the management permissions of the underlying networking / storage infrastructure.

[0006] Therefore, the DPU operating system provides enhanced access and control over basic hardware resources and allows the inherent values ​​of I / O virtualization present in the hypervisor to support bare-metal workloads (i.e., workloads running on servers without supported hypervisors, such as Linux). TM (Containers). However, DPUs come from different manufacturers. This combination with multiple server manufacturers results in many vendor-specific aspects integrating everything together to form a solution. Attached Figure Description

[0007] Many aspects of this disclosure can be better understood by referring to the following figures. The components in the figures are not necessarily to scale, but are intended to clearly illustrate the principles of this disclosure. Furthermore, in the figures, the same reference numerals are used throughout several views to represent corresponding parts.

[0008] Figure 1 This is a diagram illustrating an example of a combined DPU and server solution.

[0009] Figures 2A to 2B It is an architecture diagram illustrating a DPU with a DPU operating system running on it and a server with a hypervisor or other operating system (e.g., a bare metal operating system) running on it.

[0010] Figure 3 It is a sequence diagram illustrating the boot process, in which an application programming interface is used to control the watchdog operation implemented by the underlying hardware.

[0011] Figure 4 This is an explanation of the reason. Figure 1 The flowchart illustrates the functional implementation of components in a solution environment such as a DPU or other computing device. Detailed Implementation

[0012] This disclosure relates to the integration of a DPU operating system in a combined server and DPU solution, wherein the DPU may comprise different types, versions, and original equipment manufacturer (OEM) devices. The DPU operating system is currently being deployed and optimized for execution on DPUs and similar computing devices. The DPU may comprise 64-bit based... The implementation scheme of the architecture, and / or other similar computing devices that perform operations similar to those of traditional servers.

[0013] There is a desire for a DPU operating system capable of supporting multiple types of DPUs, regardless of brand, model, version, manufacturer, or similar. Furthermore, there is a desire to abstract various aspects through a vendor-neutral interface to support a single binary version of the DPU operating system. In particular, various network service providers want to adopt the DPU operating system in diverse solutions comprised of DPUs and other types of computing devices from different vendors, as well as servers from different original equipment manufacturers. However, the way the DPU communicates with the server backplane management controller (BMC) or server firmware may vary from device to device. To handle different types of DPUs, the DPU operating system will need to be deployed across numerous different software images for different types of devices. However, having a single DPU operating system and / or operating system (OS) image with as similar a code path as possible across different types of computing devices is attractive.

[0014] The different aspects for each DPU and each OEM include low-level signaling from the DPU OS, such as fault, service status, and diagnostic status sent to the platform (e.g., BMC and / or server firmware). Other aspects that differ for each DPU and each OEM include high-level protocol signaling from the DPU OS to the BMC and platform (e.g., BMC and server firmware) signaling to the DPU OS (e.g., power off, non-maskable interrupt (NMI), etc.).

[0015] Therefore, various embodiments for DPU OS integration in combined servers and DPU or similar solutions are described. In some embodiments, a system is described comprising a server computing device and a network computing device (e.g., a DPU or the like) communicatively coupled to each other, wherein each includes at least one hardware processor. Program instructions are stored in memory and, when executed by the at least one hardware processor of the network computing device, instruct the network computing device or other desired device to execute the DPU OS on the network computing device, wherein the DPU OS utilizes secure calls implemented by DPU vendor-specific trusted firmware (e.g., for example, ...). (Architecturally secure firmware calls). The trusted firmware of the network computing device can further identify one of the secure calls made by the DPU OS on the network computing device to communicate with the server computing device, and when invoked, translates said secure call into an architecture-specific call (e.g., vendor-specific or OEM-specific call) for the server computing device. The network computing device can execute the call, which may include sending signals to the server computing device in a format that the server computing device can interpret.

[0016] In some embodiments, the network computing device is a data processing unit that implements a Reduced Instruction Set Computer (RISC) architecture within one or more embedded processors. However, in other embodiments, instead of a data processing unit, a Field Programmable Gate Array (FPGA) card or other similar computing device capable of running a DPU operating system or other similar operating systems may be used.

[0017] In various embodiments, the server computing device may include a backplane management controller. Therefore, secure calls made by the DPU OS to the DPU trusted firmware can be transformed in a vendor-specific proprietary manner to invoke or communicate with the server computing device's backplane management controller. In various embodiments, secure calls made by the DPU OS can be transformed in a vendor-specific proprietary manner to invoke or communicate with the server firmware components of the server device. In various embodiments, secure calls made by the DPU OS to the DPU trusted firmware abstract away functionality implemented entirely by the DPU hardware or firmware components. In various embodiments, calls made by the DPU OS to the DPU trusted firmware abstract away the integration of this DPU OS with the combined server and DPU solution, thereby abstracting away differences between various DPUs or between various servers. This firmware abstraction mechanism includes not only calls made by the DPU OS, but also how the DPU firmware abstracts notifications (e.g., power-down requests, crash requests, etc.) in common ways and delivers them to the DPU OS.

[0018] In some instances, signals sent to the server computing device may include fatal peripheral component fast interconnect (PCIe) error interruptions that notify the server computing device's operating system of a fatal error in the DPU OS or DPU components, causing impairment of the functionality presented by the DPU to the server computing device. In additional instances, signals sent to the server computing device may be power-on self-test codes or DPU operating system service states sent to the server computing device's backplane management controller. In other instances, power-down signals sent by the server computing device (e.g., OS, BMC, or other components therein) may be abstracted and communicated as Advanced Configuration and Power Interface (ACPI) power button device events delivered to the DPU OS. In still other instances, signals sent to the server computing device may include backplane management controller crash requests.

[0019] For reference Figure 1This illustrates an example of a networked environment 100. The networked environment 100 may include computing environments 103, client devices 106, and various computing systems 109 that communicate with each other via a network 112. The network 112 may include, for example, the Internet, an intranet, an extranet, a wide area network (WAN), a local area network (LAN), a wired network, a wireless network, other suitable networks, or any combination of two or more such networks.

[0020] Network 112 of the networking environment 100 may include satellite networks, wired networks, Ethernet, telephone networks, and other types of networks. Computing system 109 may include devices mounted on racks 115a…115n (collectively referred to as “rack 115”), which may form a group of computers in a server cluster, aggregated computing system, or data center or other similar facility. In some instances, computing system 109 may include a high-availability computing system comprising a group of computing devices that act as a single system and provide continuous uptime. The devices in computing system 109 may include any number of physical machines, virtual machines, virtual devices, and associated software (e.g., operating systems, drivers, hypervisors, DPU OS, scripts, applications, etc.).

[0021] The computing system 109 and the various hardware and software components contained therein may include the infrastructure of a networking environment 100 that provides one or more computing services 118. The computing services 118 may include, for example, network-based API services that can be invoked via network-based application programming interfaces (APIs).

[0022] Computing environment 103 may comprise an enterprise computing environment comprising hundreds or even thousands of physical machines, virtual machines, virtual devices, and other software implemented in devices stored in rack 115, geographically distributed, and interconnected via network 112. Thus, in some instances, computing environment 103 may be referred to as a distributed computing environment. It should be understood that any virtual machine or virtual device is implemented using at least one physical device (e.g., a server or other computing device).

[0023] The devices in rack 115 may include various physical computing resources. These physical computing resources may include, for example, physical computing hardware such as memory and storage devices, servers 121a…121n (collectively referred to as “server 121”), switches 124a…124n, GPUs 130a…130n, DPUs 133a…133n (collectively referred to as “DPU 133”), central processing units (CPUs), power supplies, etc. For example, devices such as server 121, switch 124, GPU 130, DPU 133, and the like may have dimensions suitable for rapid installation in slots 136a…136n (collectively referred to as “slot 136”) on rack 115.

[0024] In each instance, server 121 may include physical hardware and software for creating and managing virtualized infrastructure, cloud computing environments, on-premises deployment environments, and / or serverless computing environments. Furthermore, in some instances, physical computing resources may be used to provide virtual computing resources (such as virtual machines or other software) as computing services 118. In each instance, virtual machines may provide virtual desktops or other virtualized computing infrastructure.

[0025] Each server 121 in network environment 100 may therefore contain one or more virtual machines (VMs) running thereon. Referring to a representative DPU 133, DPU 133 may include an accelerator 139 that offloads tasks (e.g., tasks managing distributed and virtualized applications) from the CPU of server 121. As will be understood, accelerator 139 may perform networking and storage tasks more efficiently than the CPU of server 121. In some implementations, DPU 133 includes a CPU and memory 142, such that the operation of accelerator 139 can be configured by developers and / or administrators (e.g., through programming and execution of specific applications or other processes).

[0026] In some instances, DPU OS 145 is an operating system that can be installed on one or more DPUs 133, and hypervisor 202 can be installed on server 121 to support a virtual machine execution space, in which one or more virtual machines can be instantiated and executed simultaneously with networking and storage virtualization offloaded to the DPU and DPU OS. DPU OS 145 may include... ESXio TM A hypervisor, or in some instances, a hypervisor-like program. Similarly, server hypervisor 202 may include... ESXi TM A hypervisor or similar hypervisor. In these embodiments, the DPU OS (I / O hypervisor) and the server OS (compute hypervisor) work together without any trust or isolation boundaries, and the DPU OS 145 is effectively managed by the server hypervisor.

[0027] In some instances, the DPU OS145 is installed on one or more DPU 133s, with any customer-selected operating system or hypervisor, including Linux. TM Windows TM Hyper-V TM Xen TM ESXi TMThese are installed on server 121. In these embodiments, DPU 133 and DPU OS 144 offload I / O functions to the server OS without requiring specific integration with the server's virtualization environment (if any). Therefore, a clear isolation and trust boundary exists between server 121 and DPU OS 145, which is particularly advantageous in cloud service provider (CSP) environments where the server OS and underlying infrastructure (including networking, storage, and DPU 133) are managed by different organizational units, legal entities, and the like, when impacting some available functionality. In some embodiments, DPU OS 145 is managed separately from the server hypervisor. This is also referred to as support for bare-metal computing, in which even non-virtualized server workloads can utilize virtualized networking and storage provided by DPU 133.

[0028] It should be understood that the computing system 109 is scalable, which means that the computing system 109 in the networked environment 100 can be dynamically added or removed to include or remove servers 121, switches 124, GPUs 130, DPUs 133, power supplies and other components without downtime or otherwise impairing the performance of the computing services 118 provided by the computing system 109.

[0029] Referring now to computing environment 103, computing environment 103 may include, for example, server 121 or any other system providing computing power. Alternatively, computing environment 103 may include, for example, one or more computing devices arranged in one or more server groups, computer groups, computing clusters, or other arrangements. Computing environment 103 may include grid computing resources or any other distributed computing arrangement. Computing devices may be located in a single installation or may be distributed across many different geographical locations. In some instances, computing environment 103 may include or operate as one or more virtualized computer examples. Although shown separately from computing system 109, it should be understood that in some instances, computing environment 103 may be included as all or part of computing system 109.

[0030] For convenience, computing environment 103 is referred to herein in the singular. Although computing environment 103 is referred to in the singular, it should be understood that multiple computing environments 103 may be employed in the various arrangements described above. When computing environment 103 communicates with computing system 109 and client device 106 via network 112, it is sometimes remote communication, and in some instances, computing environment 103 may be described as remote computing environment 103. Furthermore, in various instances, computing environment 103 may be implemented in server 121 of rack 115 and may manage the operation of virtualized or cloud computing environments through interaction with computing service 118.

[0031] Computing environment 103 may include data storage device 150, which in some instances may include one or more databases. Data storage device 150 may include the memory of computing environment 103, mass storage resources of computing environment 103, or any other storage resources on which computing environment 103 may store data. In some instances, data storage device 150 may include the memory of server 121. Data storage device 150 may include one or more relational databases, such as Structured Query Language databases, non-SQL databases, or other relational or non-relational databases. For example, data stored in data storage device 150 may be associated with the operation of various services or functional entities described below. Components executing on computing environment 103 may include, for example, virtualization service 153, network service 156, and other applications, services, processes, systems, engines, or functionalities not discussed in detail herein.

[0032] Ultimately, the various physical and virtual components of computing system 109 can handle workloads 180a…180n. Workload 180 may refer to the amount of processing or routing that has been instructed to be handled or routed at a given time by server 121, switch 124, GPU 130, DPU 133, or other physical or virtual components. Workload 180 may be associated with the execution of virtual machines, public cloud services, private cloud services, hybrid cloud services, virtualization services, device management services, containers, or other software running on server 121 (and therefore within computing environment 103).

[0033] Referring back to the representative DPU 133a, the DPU 133a (or other computing device) may include a hardware-implemented watchdog 159. The hardware-implemented watchdog 159 may comprise a watchdog configured in physical circuitry, an application-specific integrated circuit (ASIC), or a computing system to send a reset signal after a predetermined amount of time without receiving a refresh signal. For example, a timer increments downwards until the predetermined amount of time has expired. If no refresh signal is received before the predetermined amount of time has expired, the hardware-implemented watchdog 159 sends a reset signal. It is understood that the reset signal may guide the device into a safe operating mode, performing a system reset, recycling, or rebooting of the device, or similar operations. The hardware-implemented watchdog 159 can be contrasted with a software-implemented watchdog that requires software to refresh and / or send a reset signal, as is understood, the software-implemented watchdog requires the use of a CPU.

[0034] Firmware 148 may further include a runtime watchdog service 162. It may be desirable to have a single image of an operating system (e.g., DPUOS 145) that can be installed on and operate on the device, regardless of the device's type, model, manufacturer, specifications, etc. For example, the same image of DPU OS 145 that is expected to perform on a particular model of DPU 133 manufactured by DeltaCo, a general example of a first OEM, may also be expected to perform on different models of DPU 133 manufactured by BetaCo, a general example of a second OEM. It should be understood that server 121, DPU 133, and the like may have different models, manufacturers, specifications, etc., and existing systems require a consistent type of device.

[0035] Furthermore, in order to perform the boot operations associated with DPU OS145, it may be desirable for the hardware-implemented watchdog 159 to be able to remain idle for a long period of time without sending a reset signal. In other words, it is undesirable for the hardware-implemented watchdog 159 to send a retransmission signal when DPU OS145 is booted or otherwise brought online. Therefore, it may be desirable for the hardware-implemented watchdog 159 to be able to remain idle for a predetermined amount of time (e.g., approximately five minutes, as an example only) without sending a reset signal. For example, The Basic System Architecture (BSA) compliant watchdog has a 48-bit watchdog offset register (WOR), which is sufficient to allow the hardware-implemented watchdog 159 to be idle for approximately five minutes. It is further desirable that the hardware-implemented watchdog 159 be able to perform a "bite" operation that causes a system reset.

[0036] If the hardware-implemented watchdog 159 is unable to idle for the scheduled time and / or perform a biting operation, then appropriate watchdog functionality can be paravirtualized via a runtime watchdog service 162 implemented in the DPU firmware. In other words, the DPU 133 can be configured to handle longer idle times and perform other operations as needed to boot the DPU OS 145.

[0037] In some embodiments, the runtime watchdog service 162 may use the same unit as a general-purpose timer (e.g., driven by CNTFRQ_EL0) and may have the same constraints as a BSA general-purpose watchdog. While implementations utilizing only a security timer are possible, other implementations include using and refreshing a hardware-implemented watchdog 159 to prevent system resets, for example, during the boot of the DPU OS 145. Through the operation of the runtime watchdog service 162, the device will be able to recover from situations where all processing cores crash due to programming errors or external events and are unable to handle exceptions.

[0038] Now for reference Figure 2A and2B The display may include Figure 1 An example of the architecture diagram 200 of the components of the networked environment 100. For example, architecture diagram 200 includes a DPU OS 145a installed and executed on a DPU 133 (or a similar computing device) and a hypervisor 202 installed and executed on a server 121. In some instances, client device 106 may execute server management interface 209 to bootstrap the execution of virtualization services.

[0039] The hypervisor 202 manages the DPU 133 through interaction with the DPU OS145a, which is a standard PCIe peripheral. As may be understood, the backplane management controller 212 of the server 121 can be the source of platform management. In some embodiments, communication with the backplane management controller 212 can be performed by providing a dedicated control and status channel between the DPU 133 and the backplane management controller 212. For example, a high-bandwidth network connection can be provided for communication using a communication protocol such as Redfish, which employs a RESTful interface to manage networking, storage, servers, and converged infrastructure. Alternatively, a low-bandwidth control channel interface (e.g., a Network Controller Sideband Interface (NC-SI)) can be provided for communication.

[0040] The communication between the DPU OS 145 and the backplane management controller 212 and / or the server 121 is now described. In some embodiments, the DPU OS 145 communicates with the backplane management controller 212 via the Redfish communication protocol, but other suitable communication protocols may be used. In some embodiments, the DPU OS 145 uses the backplane management controller 212 of the server 121 for the lifecycle operation of the DPU 133. For this purpose, the DPU OS 145 may communicate with the backplane management controller 212 of the server 121 for hypervisor-related preset purposes and configurations (e.g., imaging), hypervisor-related control operations (e.g., rebooting and power-off handling), collecting troubleshooting information, etc.

[0041] In some embodiments, the backplane management controller 212 of server 121 may provide a LAN-based channel to allow management from hypervisor 202. The management LAN may be encrypted or otherwise protected. For example, in some embodiments, communication between hypervisor 202 and backplane management controller 212 of server 121 must be authenticated. Furthermore, backplane management controller 212 may be resistant to external attacks such as denial-of-service (DoS) attacks. Components interacting with hypervisor 202 (e.g., Redfish) may be able to interact with hypervisor 202 using only cryptography validated by the Federal Information Processing Standard (FIPS). Server 121 may use approved (or unapproved) cryptography, but may not require hypervisor 202 to use unapproved cryptography. Management protocols may be implemented using Redfish or similar communication protocols.

[0042] In some embodiments, to facilitate discovery operations, the backplane management controller 212 of server 121 may provide the following information for the supported DPU 133: slot information (e.g., the physical location of DPU 133 within server 121); bus identification information (e.g., the programming location of the functional portion of DPU 133, i.e., segment, bus, device, and function information of a PCIe-like system), which can be used to perform Hardware Compatibility List (HCL) verification and to take DPU 133 offline via OS-triggered Delayed Procedure Call (DPC); and UUID information (e.g., a vendor-unique identifier for consistent identification of DPU 133 in DPU OS 145a, hypervisor 202, backplane management controller 212, and data center / clustering management).

[0043] The functions performed by server 121 as described herein may be performed by a server hypervisor or a bare-metal workload, or otherwise. Hypervisor 202 may be a proprietary hypervisor as described above. Thus, functionality is provided by avoiding a trust boundary between the server hypervisor and DPU 133 (or other I / O software on DPU 133).

[0044] However, in the case of bare-metal computing, such as Figure 2B As shown and described above, a trust boundary exists between the DPU 133 and the server operating system, which can be useful in methods where the server 121 and workload 180 are managed by different entities (e.g., in a cloud computing environment). The streamlined and disparate integration between the DPU 133 and server operating system management allows support for arbitrary server operating system environments, such as Linux. TM OS, Windows TM OS or including KVMTM HyperV TM Xen TM or The solution is a virtualized environment.

[0045] Now for reference Figure 3 The following illustrates non-limiting examples of sequence diagrams according to various embodiments. Various stages of the sequence diagrams can be executed during a boot process, which may include procedures for loading operating system components into random access memory or other desired memory. Initially, the DPU 133 or other device may include firmware 148 with UEFI or BIOS firmware supervising the boot operation. Thus, the sequence diagrams can be executed in a boot loading environment (e.g., a UEFI boot loading environment) by an application running in that boot loading environment.

[0046] First, at box 303, during the power-up phase (e.g., immediately following the physical power-up of the device, such as DPU 133 or server 121), the UEFI bootloader 199 on the device can initiate an EFI infrastructure (e.g., execution software) that allows the execution of EFI-compatible executables. The EFI infrastructure can allow, for example, the execution of an application in the first-stage bootloader 193, such as to boot the DPU OS. This interface RUNTIME_WATCHDOG_PROTOCOL provides a simple way not only to query the available facilities and / or specifications of any hardware-implemented watchdog 159 on the device, but also to activate the hardware-implemented watchdog 159 in the first-stage bootloader 193 and put the hardware-implemented watchdog 159 on standby after a call to ExitBootServices() (e.g., UEFI to operating system switch). Furthermore, the application programming interface can be used not only to activate the hardware-implemented watchdog 159, but also to handle any required periodic updates to the hardware-implemented watchdog 159, removing such engineering requirements from the first-stage bootloader 193.

[0047] In box 306, the UEFI system may, for example, install a runtime watchdog protocol during the power-up phase. The runtime watchdog protocol may include an application programming interface (API), which, as will be described, can be invoked to initialize runtime watchdog services for a watchdog 159 that oversees the hardware implementation. In some embodiments, the runtime watchdog protocol (e.g., the API) is installed by storing a driver in a directory where the boot UEFI bootloader environment installs drivers during the power-up phase of the boot process.

[0048] The runtime watchdog protocol may include an application programming interface, wherein API calls cause the execution of at least one of the following: enabling a hardware-implemented watchdog; disabling a hardware-implemented watchdog; accessing a type of hardware-implemented watchdog; accessing the physical memory address of a hardware-implemented watchdog; identifying the minimum countdown period that the hardware-implemented watchdog can be configured for; and identifying the maximum countdown period that the hardware-implemented watchdog can be configured for.

[0049] Subsequently, the process proceeds to the operating system loading phase. There, at box 309, the UEFI system executes a boot manager configured to handle and supervise the boot process. At box 312, the boot manager initiates an operating system bootloader containing executable code that initializes and starts the operating system. At box 315, the first-stage bootloader 193 initializes a runtime watchdog service 162. Initializing the runtime watchdog service 162 may include invoking runtime watchdog protocol functions using input parameters. Furthermore, initializing the runtime watchdog service 162 may include enabling a hardware-implemented watchdog 159.

[0050] Subsequently, in blocks 318 and 321, the first-stage bootloader 193 may, for example, set the runtime watchdog refresh timer by calling the RUNTIME_WATCHDOG_SET function of the runtime watchdog protocol (“RUNTIME_WATCHDOG_PROTOCOL”). For example, if the watchdog refresh timer is successfully set on the hardware-implemented watchdog 159, then in block 324, the UEFI system may respond by returning a success signal (“EFI_SUCCESS”) to the first-stage bootloader 193.

[0051] In box 327, the first-stage bootloader 193 may load the DPU OS145 components for executing the DPU OS145. In other words, the first-stage bootloader 193 may load or store the operating system components in random access memory or other memory. In box 330, the first-stage bootloader 193 may construct boot information data that may contain tables, data objects, or other data sets. In box 333, the first-stage bootloader 193 may construct runtime watchdog entries for tables, databases, or other suitable memory locations.

[0052] Subsequently, the process proceeds to the operating system switchover phase. In box 336, the ExitBootServices() function is called after a set of predetermined boot operations have completed. Next, in box 339, the UEFI bootloader 199 (e.g., under the instruction of the first-stage bootloader 193) may perform a final watchdog refresh to prevent the hardware-implemented watchdog 159 from failing during the switchover from the UEFI system to the operating system. In box 342, the UEFI system's runtime completes, and the UEFI system will no longer refresh the watchdog. Therefore, in box 345, the UEFI system sends an EFI success signal to the first-stage bootloader 193, which then switches the operation of the hardware-implemented watchdog 159 to the operating system kernel in box 348. The process can then proceed to completion.

[0053] In particular, a UEFI (or similar) protocol for keeping the hardware-implemented watchdog 159 ready is described. In some embodiments, the hardware-implemented watchdog 159 is not enabled by default and will be activated by the UEFI bootloader 199. When activated, the UEFI bootloader 199 may be responsible for refreshing the hardware-implemented watchdog 159 until ExitBootServices() is called (e.g., in the case where UEFI is switched to the operating system). In hardware-allowed examples, if the boot process aborts and execution is passed back to the UEFI Boot Device Selection (BDS), then the hardware-implemented watchdog 159 may be deactivated by the UEFI bootloader 199.

[0054] In some embodiments, when ExitBootServices() is called, the hardware-implemented watchdog 159 can be put on standby. The UEFI bootloader 199 can perform a final watchdog refresh to ensure that the operating system is not relinquished control at the end of the refresh cycle. In some embodiments, the operating system may then be responsible for the refresh. In the event of an operating system stop or crash, the operating system may be responsible for refreshing the watchdog, if necessary, to avoid a hard reset (e.g., physical power-on of the device). The bootloader (e.g., the DPU OS bootloader) can use RUNTIME_WATCHDOG_SET to set a watchdog period long enough to cover the boot of the DPU OS 145 or other software (e.g., 5 minutes). In the future, the bootloader may alternatively choose a shorter period, as the protocol definition includes automatic refresh for use in UEFI environments.

[0055] Continue to Figure 4 This is a flowchart illustrating an example of the operation of a part of the networked environment 100. Figure 4The flowchart can be viewed as an instance of elements depicting a method implemented by DPU OS145, which executes in DPU 133 or other computing devices, according to one or more instances. Functional separations or divisions as discussed herein are presented for illustrative purposes only.

[0056] As described above, it is ideal to provide a combined server and DPU solution with server 121 and DPU 133, where those computing devices are manufactured and sold by different vendors, different OEMs, etc. For example, a solution employing only a specific type of DPU 133 and a specific type of server 121 can be limiting. By providing the ability to use different types of server 121, DPU 133, and the like, the way DPU 133 or other devices communicate with the backplane management controller 212 (or server firmware) of server 121 can differ. However, it is ideal to have a single DPU OS 145 operating system image that can be installed and executed on network devices of different types and architectures (e.g., DPU 133 or the like) while maintaining the same code path as much as possible.

[0057] In different types of network devices, different areas are adapted, including low-level signaling (e.g., fault, service status, diagnostic status) from DPU OS145 to the platform (e.g., backplane management controller 212 and / or server firmware); high-level protocol signaling from DPU OS145 to the backplane management controller 212; and platform signaling (e.g., backplane management controller 212 and server firmware) to DPU OS145 (e.g., shutdown operation, NML and the like).

[0058] Therefore, the various embodiments described herein relate to abstracting low-level signaling using a trusted firmware interface with a set of security surveillance calls that abstract and hide DPU hardware-specific and OEM-specific calls and operations. Thus, the DPU 133 firmware abstraction interface (e.g., security surveillance calls and other calls) can be provided by the DPU firmware (e.g., trusted firmware, UEFI, and ACPI) to facilitate DPU OS145 or other I / O hypervisors. Therefore, firmware abstraction can use the same DPU OS145 image on different combinations of server 121 and DPU 133 vendors without requiring specific approaches for each combination (e.g., Dell). TM and NVIDIA TM Dell TM and Pensando TM Dell TM and Intel TM HPE TM and Dell TM LenovoTM and Dell TM The unique architecture of the DPU OS145 (etc.). In some implementations, there are no multiple types of DPU 133. In other words, in some solutions, the brand, model and / or manufacturer of DPU 133 may be the same, while the brand, model and / or manufacturer of server 121 and / or DPU 133 may be different.

[0059] Beginning with box 403, DPU 133 may execute DPU OS145 or other types of operating systems. DPU OS145 may execute on DPU 133 to provide enhanced access and control to the underlying hardware resources used for virtualization-related services or other network-related services. DPU OS145 may invoke multiple security calls (e.g., SMCs) that abstract certain DPU or SmartNIC mechanisms and integrations within a larger composite server and DPU solution. These SMCs are used by the DPU OS145 kernel and user components (e.g., applications, services, engines, scripts, and the like) executing on it. Therefore, DPU OS145 components or other software executing on a first DPU 133 with a first vendor-specific architecture (requiring vendor-specific calls) will execute as expected on the first DPU 133, and the same application executing on a second DPU 133 with a second vendor-specific architecture different from the first vendor-specific architecture (also requiring vendor-specific calls) will execute as expected on the second DPU 133, despite the architectural differences. The SMC, including the DPU firmware interface for DPU OS145, can be implemented as part of the trusted firmware on DPU 133, a low-level component critical to booting DPU OS145, and performs a higher level of privilege than DPU OS145 itself.

[0060] Next, in box 409, DPU 133 can recognize one of the secure calls made on DPU 133 to communicate with server 121. Communication with server 121 may include communication with server 121's backplane management controller 212, server 121's firmware, server 121's operating system, and the like. It should be understood that server 121 may be one of many servers 121 in the networking environment 100, where each server 121 may be one of many different types and models, and manufactured and sold by different suppliers, different OEMs, etc.

[0061] Therefore, in box 412, in the example where one of the secure calls is invoked, the trusted firmware on DPU 133 can transform the invoked secure call into an architecture-specific call (e.g., vendor-specific or OEM-specific call) for server 121. Thus, high-level signaling, low-level signaling, and the like can be abstracted into sample secure calls that can be identical across the image of DPU OS 145 deployed on different types and models of network devices.

[0062] In some embodiments, targeting the use of Alternatively, a RISC architecture implementation can use a trusted firmware interface with a set of Secure Monitoring Calls (SMCs) to abstract low-level signaling between the DPU 133 and the server 121. This abstraction of low-level signals hides hardware-specific and OEM-specific operations of the DPU 133.

[0063] In one instance, a SignalFatalError call may be provided to signal fatal PCIe errors on each endpoint visible to the OS or hypervisor running on server 121. In some implementations, such as (e.g., Enhanced Downstream Port Containment (eDPC) implementations, AER implementations), server 121 may contain or ignore the error as needed. Unlike the passive status reporting of the SignalServiceStatus call, which is primarily intended for platform firmware and boot integration, SignalFatalError can cause an erroneous interruption of the operating system delivered to server 121. This call may be used by DPU OS 145, for example, as part of system crash and / or emergency handling to notify server 121. Furthermore, the call may also be used by DPU 133 firmware 148 in the event of an unexpected CPU reset (e.g., thermal, power-related, or watchdog-related reset) to signal server 121 before operation of DPU 133 ceases.

[0064] In another instance, a PostUpdate call may be provided to report the Power-On Self-Test (POST) code or DPU operating system service status to the backplane management controller 212 of server 121. The POST code or DPU operating system service status may be intended to be delivered to the backplane management controller 212 of server 121 (e.g., pushed to the backplane management controller 212 or polled by the backplane management controller 212). In some embodiments, it may not be required that the backplane management controller 212 see the POST code transition or change. However, the backplane management controller 212 may always be able to recognize the last known POST code.

[0065] In addition, brief informational status updates can be passively provided by multiple components using the PostUpdate OEM service call for debugging and OEM server-specific purposes (e.g., boot progress across different layers of firmware, boot logic, and DPU OS145 software). Interfaces can be implemented based on vendor-specific functionality, such as SMBUS traffic, registers visible from server 121 in the PCIe configuration space, custom interfaces, RAM-based logging, etc. `postcode_t` can be a 32-bit value, with 4 bits containing an entity number used to divide the value space into network service provider and vendor-specific ranges.

[0066] In other instances, the SignalServiceStatus call can be provided to perform platform-specific actions based on the DPU OS145 status. Some DPU OS145-based solutions may require additional DPU 133 or Server 121 OEM-specific steps, such as those triggered by a full DPU OS145 boot, entry into a crash handler, etc. For example, certain server-visible PCIe configuration space registers used to expose functionality to the DPU 133 may take special values, or the firmware may optionally report special status codes. The SignalServiceStatus OEM call can be used for such actions. For instance, the SignalServiceStatus call can be used primarily to passively report the DPU OS145's operational status to the server platform firmware as part of boot integration.

[0067] In some embodiments, high-level protocol signaling can be executed from the DPU OS145 to the backplane management controller 212, which may rely on a dedicated TCP / IP connection existing between the DPU OS145 and the backplane management controller 212 via an NC-SI connection. In some embodiments, the high-level protocol is the Redfish communication protocol. While the actual details of the command set may differ, the operations are common and differences can be abstracted in software using, for example, a BMCAL layer (e.g., updating the status of DPUOS provisioned operations, reporting logs, reporting high-level service status, etc.). For some solutions that report complete operations, communication can be performed via Redfish while using the low-level signaling described above to report error states.

[0068] In other embodiments, platform signaling can be provided from the server 121 platform to the DPU OS 145. In some instances, the power status notification path (via) can be abstracted for a regular ACPI power button notification. Although the actual mechanism used by the backplane management controller 212 is OEM-specific for the server 121 and vendor-specific for the DPU 133 (e.g., NC-SI, GPIO, SMBUS write, etc.), the power-down request can be delivered as an ACPI power button device (PNP0C0C) event.

[0069] Furthermore, the crash request (“NMI”) of the backplane management controller 212 can be abstracted. While the actual mechanisms used by the backplane management controller 212 are OEM-specific for the server 121 and vendor-specific for the DPU 133 (e.g., NC-SI, GPIO, SMBUS writes, etc.), the crash dump request can be treated as a RAS event, for example, as a custom unmasked SError interrupt (SEI) for EL2, with further guidance for the current set of Cortex-A72-based DPUs 133. For this purpose, the SEI can be signaled regardless of the abnormal masking state (PSTATE.A), since the NMI SEI may be unmasked.

[0070] Finally, in box 415, DPU 133 can execute the call transformed in box 412. Executing the call may involve sending signals from DPU 133 to server 121 in a format that can be interpreted by server 121, for example. For example, the actual call sent from DPU 133 to server 121 may be in a vendor-specific or OEM-specific format, while the secure call identified in box 409 is a secure call from the DPU OS to the DPU trusted firmware.

[0071] Therefore, in some instances, for example, a security call to be made in box 409 can be transformed into a call specific to the architecture of the first server 121a by: identifying the architecture of the first server 121a; identifying the function to be called (e.g., vendor-specific or OEM-specific function) based on the architecture of the first server 121a; and performing the function in response to the security call being invoked.

[0072] Conversely, if the same call is made to a second server of a different brand or manufacturer, in some instances, it can be transformed, for example, into an architecture-specific call to the second server 121b by the following secure call, which will be made in box 409: identifying the architecture of the second server 121b; identifying the function to be called (e.g., vendor-specific or OEM-specific function) based on the architecture of the second server 121b; and executing the function in response to the secure call being invoked. The process can then proceed to completion.

[0073] The above text is about Figure 4 The various operations described can be executed by a computing device by executing program instructions. These program instructions may be portions of firmware 148 stored in non-volatile memory (e.g., memory 142 of the DPU 133 or other computing device). While many of the examples described herein relate to a DPU OS 145 executing on the DPU 133, it should be understood that the DPU 133 may instead be a similar kind of device on which an I / O OS or hypervisor can run, such as a computing storage device or a GPU.

[0074] The memory device stores both data and several components executable by a processor. The memory may also store data storage device 150, firmware 148, and other data. Many software components are stored in the memory and are executable by a processor. In this respect, the term "executable program" means a program file in a form ultimately executable by a processor. Examples of executable programs may be, for example, a compiler that can be translated into machine code in a format that can be loaded into the random access portion of one or more memory devices and executed by a processor; code expressed in a format (e.g., object code) that can be loaded into the random access portion of one or more memory devices and executed by a processor; or code that can be interpreted by another executable program to generate instructions in the random access portion of the memory device to be executed by a processor. Executable programs may be stored in any part or component of a memory device, including, for example, RAM, ROM, hard disk drives, solid-state drives, USB flash drives, memory cards, optical discs (e.g., optical discs (CDs) or digital versatile optical discs (DVDs)), floppy disks, magnetic tapes, or other memory components.

[0075] Memory may include both volatile and non-volatile memory, as well as data storage components. Additionally, a processor may represent multiple processors and / or multiple processor cores, and one or more memory devices may represent multiple memories operating in parallel processing circuitry. Memory devices may also represent various types of storage devices, such as combinations of RAM, mass storage devices, flash memory, or hard disk storage devices. In this case, the local interface may be a suitable network facilitating communication between any two of the multiple processors or between any processor and any memory device. The local interface may include additional systems designed to coordinate this communication (e.g., including load balancing). Processors may be electrically powered or have some other available configuration.

[0076] Client device 106 can be used to access a user interface generated to configure computing environment 103 or otherwise interact with computing environment 103. These client devices 106 may include a display on which the user interface generated by a client application for providing a virtual desktop session (or other session) can be presented. In some instances, the user interface may be generated using user interface data provided by computing environment 103. Client device 106 may also include one or more input / output devices, which may include, for example, a capacitive touchscreen or other types of touch input devices, a fingerprint reader, or a keyboard.

[0077] While the various services and applications described herein may be embodied as software or code executed by the general-purpose hardware discussed above, alternatively, they may also be embodied as dedicated hardware or a combination of software / general-purpose hardware and dedicated hardware. If embodied as dedicated hardware, each service and application may be implemented as a circuit or state machine employing any or a combination of various technologies. These technologies may include discrete logic circuits with logic gates for implementing various logical functions when one or more data signals are applied, application-specific integrated circuits (ASICs) with appropriate logic gates, field-programmable gate arrays (FPGAs), or other components.

[0078] Sequence diagrams and flowcharts illustrate examples of the functionality and operation of partial implementations of the components described herein. If represented as software, each box may represent a module, section, or portion of code, which may contain program instructions for implementing a specified logical function. The program instructions may comprise source code of human-readable statements written in a programming language, or may be represented in the form of machine code containing numerical instructions recognizable by a suitable execution system, such as a processor in a computer system or other system. Machine code can be derived from source code. If represented as hardware, each box may represent a circuit or multiple interconnected circuits for implementing a specified logical function.

[0079] Although sequence diagrams and flowcharts illustrate a specific execution order, it should be understood that the execution order may differ from the one depicted. For example, the execution order of two or more boxes may be perturbed relative to the order shown. Additionally, two or more boxes shown consecutively may execute simultaneously or partially simultaneously. Furthermore, in some instances, one or more boxes shown in the diagram may be skipped or omitted.

[0080] Furthermore, any logic or application program containing software or code described herein can be embodied in any non-transitory computer-readable medium for use by or in conjunction with an instruction execution system (e.g., a processor in a computer system or other system). In this sense, logic may include, for example, statements containing program code, instructions, and declarations that can be extracted from a computer-readable medium and executed by an instruction execution system. In the context of this disclosure, "computer-readable medium" can be any medium that can contain, store, or maintain the logic or application program described herein for use by or in conjunction with an instruction execution system.

[0081] Computer-readable media can comprise many physical media, such as any of magnetic, optical, or semiconductor media. More specific examples of suitable computer-readable media include solid-state drives or flash memory. Furthermore, any logic or application described herein can be implemented and structured in various ways. For example, one or more applications may be implemented as modules or components of a single application. Additionally, one or more applications described herein may execute in shared or separate computing devices or combinations thereof. For example, multiple applications described herein may execute on the same computing device or on multiple computing devices.

[0082] It should be emphasized that the examples described above are merely possible examples of embodiments presented to clearly understand the principles of this disclosure. Many variations and modifications can be made to the above embodiments without substantially departing from the spirit and principles of this disclosure. All such modifications and variations are intended to be included herein within the scope of this disclosure.

Claims

1. A system for integrating a DPU operating system in a combined server and data processing unit (DPU) solution, comprising: Server computing devices and data processing units are communicatively coupled to each other, and each includes at least one hardware processor. and Program instructions, stored in memory, when executed by the at least one hardware processor of the DPU, guide the DPU to: A DPU operating system is executed on the DPU, wherein the DPU operating system runs multiple security calls implemented by a DPU trusted firmware component, the multiple security calls being implemented by the DPU trusted firmware to be invoked by components of the DPU operating system; Identify a call to one of the secure calls made on the DPU to communicate with the server computing device or to perform another function associated with the DPU; and In one example of invoking one of the secure calls, the secure call is transformed into a call or request specific to the architecture of the server computing device and the call is executed, wherein the executed call includes sending a signal to the server computing device in a format that can be interpreted by the server computing device; in The DPU implementation is a Reduced Instruction Set Computer (RISC) architecture capable of executing the DPU operating system; The server computing device includes a baseboard management controller (BMC); and One of the security calls made to the DPU is a Security Monitoring Call (SMC) that transforms into a request to the backplane management controller of the server computing device, the security monitoring call being a call compatible with the RISC architecture.

2. The system of claim 1, wherein the signal sent to the server computing device is a signal that instructs the operating system of the server computing device to notify of a fatal peripheral component fast interconnect PCIe error interruption at an endpoint visible to the operating system of the server computing device.

3. The system of claim 1, wherein the signal sent to the server computing device is a power-on self-test code sent to the backplane management controller of the server computing device.

4. The system of claim 1, wherein the signal sent to the server computing device is a DPU operating system status signal sent to the backplane management controller of the server computing device.

5. The system of claim 1, wherein the signal sent to the server computing device is a DPU stop request communicated as an Advanced Configuration and Power Interface (ACPI) power button device event.

6. The system of claim 1, wherein the signal sent to the server computing device is a DPU crash request from the backplane management controller.

7. The system of claim 1, further comprising a second server computing device, wherein: The server computing device is a first server computing device; Transforming one of the secure calls from the DPU operating system into an architecture-specific call for the first server computing device by: identifying the architecture of the first server computing device; and identifying the function to be called based on the architecture of the first server computing device. and in response to the security call being invoked to perform the function; and The at least one hardware processor of the DPU is further guided to identify the architecture of the second server computing device; identify the function to be invoked based on the architecture of the second server computing device; and execute the function in response to the security call being invoked.

8. A method for DPU integration in a combined server and data processing unit (DPU) solution, comprising: A networking environment is provided that includes server computing devices and DPUs that are communicatively coupled to each other, wherein the server computing devices contain the DPUs, and the server computing devices and the DPUs have different types and manufacturers; A DPU operating system is executed on the DPU, wherein the DPU operating system runs multiple secure calls implemented by the DPU trusted firmware and invoked by components of the DPU operating system; Identify one of the secure calls made on the DPU to communicate with the server computing device containing the DPU; In the example where one of the security calls is invoked, the security call is transformed into a call or request specific to the architecture of the server containing the DPU; and Executing the call, wherein the executed call includes sending a signal to the server computing device in a format that can be interpreted by the server computing device; in The DPU implementation is a Reduced Instruction Set Computer (RISC) architecture capable of executing the DPU operating system; The server computing device includes a baseboard management controller (BMC); and One of the security calls made to the DPU is a Security Monitoring Call (SMC) that transforms into a request to the backplane management controller of the server computing device, the security monitoring call being a call compatible with the RISC architecture.

9. The method of claim 8, wherein the signal sent to the server computing device is to guide the operating system of the server computing device to detect a fatal peripheral component fast interconnect PCIe error interruption on an endpoint visible to the operating system of the server computing device, the fatal error being caused by a failure of one of the DPUs or the DPU operating system of the one of the DPUs.

10. The method of claim 8, wherein the signal sent to the server computing device is a power-on self-test code or DPU operating system service status sent to the backplane management controller of the server computing device.

11. The method of claim 8, wherein the signal sent to the server computing device is a DPU stop request communicated as an Advanced Configuration and Power Interface (ACPI) power button device event.

12. The method of claim 8, wherein the signal sent from the server computing device to the DPU is a crash request requested by the backplane management controller.

13. The method according to claim 8, wherein: The server computing device is a first server computing device; The network environment further includes a second server computing device; Transforming one of the secure calls from the DPU operating system into an architecture-specific call for the first server computing device by: identifying the architecture of the first server computing device; and identifying the function to be called based on the architecture of the first server computing device. and in response to the security call being invoked to perform the function; and At least one hardware processor of the DPU is further guided to identify the architecture of the second server computing device; identify the function to be invoked based on the architecture of the second server computing device; and execute the function in response to the security call being invoked.

14. A non-transitory computer-readable medium having stored thereon program instructions executable by a data processing unit (DPU) having at least one hardware processor, the program instructions, when executed by the at least one hardware processor, instructing the DPU to: A DPU operating system is executed on a network interface card, wherein the DPU operating system runs multiple security calls implemented by the DPU trusted firmware and invoked by the processes of the DPU operating system; Identify one of the secure calls made on the DPU to communicate with the server computing device; and In one example of invoking one of the secure calls, the secure call is transformed into a call or request specific to the architecture of the server computing device and the call is executed, wherein the executed call includes sending a signal to the server computing device in a format that can be interpreted by the server computing device; in The DPU implementation is a Reduced Instruction Set Computer (RISC) architecture capable of executing the DPU operating system; The server computing device includes a baseboard management controller (BMC); and One of the security calls made to the DPU is a Security Monitoring Call (SMC) that transforms into a request to the backplane management controller of the server computing device, the security monitoring call being a call compatible with the RISC architecture.

15. The non-transitory computer-readable medium of claim 14, wherein the signal sent to the server computing device is a signal that instructs the operating system of the server computing device to notify of a fatal peripheral component fast interconnect PCIe error interruption at an endpoint visible to the operating system of the server computing device.

16. The non-transitory computer-readable medium of claim 14, wherein the signal sent to the server computing device is a power-on self-test code or DPU operating system service status sent to the backplane management controller of the server computing device.

17. The non-transitory computer-readable medium according to claim 14, wherein: The signal sent to the server computing device is a DPU stop request communicated as an event of the Advanced Configuration and Power Interface (ACPI) power button device; or The signal sent to the server computing device is a backplane management controller crash request.

Citation Information

Patent Citations

  • Data processing unit for compute nodes and storage nodes

    CN110915173A

  • Process-based virtualization system for executing a secure application process

    US20210232693A1