A network interface device as a computing platform
The network interface device addresses latency and resource constraints in edge computing by offering flexible deployment and failover capabilities, enhancing efficiency and reliability in edge and data center environments.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- INTEL CORP
- Filing Date
- 2023-11-03
- Publication Date
- 2026-06-03
AI Technical Summary
Edge computing systems face challenges with latency and resource constraints due to the physical distance of data centers from client devices, especially for workloads requiring fast data processing and proximity to data sources, and there is a need for flexible and efficient deployment of workloads in edge environments.
A network interface device (NID) that operates as a standalone computing platform or companion to a server, providing AI inference, video analytics, and access to accelerators, processors, and storage resources, with modes for standalone operation or offloading from a host, and supports failover and seamless migration of applications.
Enables reduced latency and efficient resource utilization by deploying compute resources closer to data sources, adapting to varying demands, and ensuring continuous operation through failover and migration, suitable for edge and data center environments.
Smart Images

Figure 00000023_0000 
Figure 00000024_0000 
Figure 00000025_0000
Abstract
Description
PRIORITY CLAIM
[0001] This application claims priority pursuant to 35 USC § 365(c) of US application No. 18 / 370,621, filed on September 20, 2023, which in turn claims priority and benefit of Indian provisional patent application No. 202341046012, filed with the Indian Patent Office on July 8, 2023, the entire contents of which are incorporated by reference in their entirety. STATE OF THE ART
[0002] Edge computing clusters and data center clusters encompass client use cases such as smart cities, augmented reality (AR) / virtual reality (VR), assisted / autonomous vehicles, retail, proximity-based services, and other applications with a wide variety of workload behaviors and requirements. However, using a data center that is physically miles away from the client device can introduce latency in completing a work request. Furthermore, with a moving client device, such as a car or other high-speed vehicle, the client device can quickly enter and then quickly exit a computing cluster region.Devices that are in motion pose a challenge to the need for even faster deployment of work by the cluster to provide data to the client device in a timely manner.
[0003] Some workloads rely on machine learning (ML) or deep learning (DL) inferences tailored to specific groups of users, while others utilize hardware-accelerated execution to complete within tight timeframes. Some workloads can be delivered as Function-as-a-Service (FaaS) operations. Many workloads require proximity to specific data consumed during their operation, due to both speed and data movement challenges, while others may have unique requirements. For example, a wearable device might only be able to access an AR service from a limited set of service providers that employ a specific algorithm for high accuracy.As another example, a driver with driver assistance systems can request a specific navigation service from a particular service provider.
[0004] Edge computing aims to place compute and data storage resources physically closer to data sources and recipients to reduce data processing and access latency and network bandwidth usage. Edge cloud architectures utilize network interface devices, such as Intel® Infrastructure Processing Units (IPUs), to manage the infrastructure and enable central processing units (CPUs), graphics processing units (GPUs), and other processors (e.g., xPUs) to perform core application-level functions. Far-edge can include large, distributed devices (e.g., multiple edge devices), but these devices may be limited by power consumption and space constraints. The data center edge can include devices with higher computing power and resource sharing among multiple tenants.On-site edge can combine devices with both far-field and data center edge features. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 represents an exemplary system. Fig. Figure 2 represents an exemplary system. Fig. 3A-3C represent exemplary systems. Fig. 4A and Fig. 4B represent exemplary application models. Fig. Section 5 presents exemplary application models. Fig. Figure 6 represents an exemplary process. Fig. Figure 7 represents an example network interface. Fig. Figure 8 represents an exemplary system. DETAILED DESCRIPTION
[0005] Edge cloud architectures can leverage network interface devices, such as Intel® Infrastructure Processing Units (IPUs), to manage the infrastructure and enable central processing units (CPUs), graphics processing units (GPUs), and other processors (e.g., xPUs) to perform core application-level functions. To provide devices capable of deployment in far-edge, data center-edge, or on-premises environments, some examples offer a network interface device capable of operating as a standalone computing platform system or as a companion to a server platform. Various examples provide a network interface device that can operate in one or more modes: standalone system-on-module mode, standalone system mode, or companion mode.Various examples can be used by telecommunications providers, communication service providers (CoSPs), and cloud service providers (CSPs).
[0006] The standalone system-on-module mode can provide computing capabilities offloaded to a network interface device, such as artificial intelligence (AI) inference and video analytics (e.g., image recognition, incident detection, alert generation, decision-making, issuing commands to an autonomous vehicle, or other tasks). The standalone system-on-module mode can provide AI inference and video analytics, and based on increasing computing demands (e.g., the number of video streams received from one or more cameras and processed by a video analytics application), the network interface device can access connected accelerators to meet variations in demand and provide modularity and flexibility. The standalone system mode allows a network interface device to operate independently of a host server and can access accelerators, processors (e.g., CPUs, GPUs, etc.).XPUs) over a fabric and non-volatile memory express (NVMe) over a fabric or other means to dynamically create virtual systems with disaggregated resources to run heterogeneous workloads and adapt to changing resource requirements over time. Companion mode can provide a network interface device to offload operations from a host, including, in the event of host failure, running an application previously executed by the host.
[0007] Some examples provide a network interface device as a standalone system without a connected host, and the network interface device can run workloads and access resources such as accelerators (e.g., field-programmable gate arrays (FPGAs)), processors, memory, network interface, and others. A network interface device as a standalone system without a connected host can be used in edge or non-edge deployments (e.g., data center or cloud). Some examples provide a network interface device that includes: (1) a media and AI accelerator circuit arrangement to which a host server processor can access as a virtual function (VF) or physical function (PF) (e.g.,Single Root I / O Virtualization - SR-IOV and Sharing Specification or Intel® Scalable I / O Virtualization - SIOV)) to provide access to a network interface and AI processing; (2) an AI accelerator circuit arrangement in the network interface device to perform inline network analytics; and / or (3) processors to run applications (e.g., virtual machines, containers, microservices, or other processes) in the network interface device by accessing an AI circuit arrangement or other resources.
[0008] Example deployments can be used to access the edge, such as an edge compute layer where multiple edge nodes can be connected. Connectivity between a first layer of edges and the access edge can utilize one or more of the following: wireless connectivity, which can be in standard wireless form or more specialized protocols, such as 802.11p (Cellular V2X) or LoRa; wired connectivity, such as Ethernet; or LTE or 5G wireless, which can be deployed in private spectrum or commercial telecommunications spectrum (telco spectrum).
[0009] Exemplary deployments can be used at the network edge, which can be implemented in base stations or in some type of distribution point at the roadside. The network edge may be subject to strict security, space, heat, and power consumption constraints.
[0010] It should be noted that, although examples relating to the edge or edge computing are described here, examples can apply to any environment, such as data centers, servers within a rack, or other systems.
[0011] Fig. Figure 1 represents an exemplary environment. For example, in the case of an intelligent sensor or gateway 100, a network interface device in standalone system-on-module mode 102 can be used to provide software and service access to computing, media, and AI accelerator devices. Services targeting an intelligent gateway can include video analytics (e.g., cameras mounted on a pole or an Internet of Things (IoT) gateway) or industrial control loops.
[0012] In an intelligent edge system 110, for example, a network interface device in standalone system mode 112 can be deployed to provide software and services with access to compute, media, and AI accelerator devices, as well as storage resources. Similarly, in a data center 120 (e.g., at the edge), a private and / or public cloud environment 130, a network interface device in standalone system mode 132 can be deployed to provide software and services with access to compute, media, and AI accelerator devices, as well as storage resources. A chassis height for standalone or companion mode deployments can be 1U or 2U (depending on available space). Data centers utilizing a network interface device can be hosted on-premises, at a telecommunications operator's central office, in government offices, or in a private cloud (the last layer of the edge).
[0013] Fig. Figure 2 represents an exemplary environment. For example, in a smart sensor or gateway 200, a network interface device in standalone system-on-module mode 202 can be deployed to provide software and services with access to hardware resources including the CPU 204, GPU 206, and / or GPU 208. In a smart edge system 220, for example, a network interface device in standalone system mode 222 can be deployed to provide software and services with access to hardware resources including the GPU 224, GPU 226, and / or GPU 228. For example, in a data center 230 (e.g., edge), a private, and / or a public cloud environment 240, a network interface device in standalone system mode 232 can be deployed to provide software and services with access to hardware resources including the CPU 234 and / or GPU 236.Examples of software and services may include one or more of the following: video analytics, sensor processing (e.g., light detection and distance measurement (LIDAR)), sensor processing, image recognition, content delivery network (CDN), video analytics, or others.
[0014] Fig. 3A represents an exemplary system. The Host 300 can include processors, memory devices, device interfaces, and other circuit arrangements, such as those relating to one or more of the Fig. 3B, Fig. 7 and / or 8 described. Host 300 processors can run software, such as applications (e.g., microservices, virtual machines (VMs), microVMs, containers, processes, threads, or other virtualized execution environments), operating systems (OS), and device drivers. An OS or device driver can configure a network interface device or a packet processing device 310 to use one or more control planes to communicate with the software-defined networking (SDN) controller 350 over a network to configure the operation of the one or more control planes.
[0015] The packet processing device 310 can include several compute complexes, such as an acceleration compute complex (ACC) 320 and a management compute complex (MCC) 330, as well as a packet processing circuit arrangement 340 and network interface technologies for communicating with other devices over a network. The ACC 320 can be implemented as one or more of the following: a microprocessor, a processor, an accelerator, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or a circuit arrangement that, at least with respect to Fig. 3B, Fig. 7 and / or 8. Likewise, the MCC 330 can be implemented as one or more of the following: a microprocessor, a processor, an accelerator, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or a circuit arrangement that, at least with respect to Fig. 3B, Fig. 7 and / or 8 are described. In some examples, the ACC 320 and the MCC 330 can be implemented as separate cores in one CPU, as different cores in different CPUs, as different processors in the same integrated circuit, or as different processors in different integrated circuits.
[0016] The packet processing device 310 can be implemented as one or more of the following: a microprocessor, a processor, an accelerator, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or a circuit arrangement that, at least with respect to Fig. 3B, Fig. 7 and / or 8 are described. The packet processing pipeline circuit 340 can process packets according to the instructions or configuration of one or more control levels executed by multiple compute complexes. In some examples, the ACC 320 and the MCC 330 can execute corresponding control levels 322 and 332.
[0017] The SDN controller 350 can update or reconfigure software running on the ACC 320 (e.g., control plane 322 and / or control plane 332) using the contents of packets received by the packet processing device 310. In some examples, the ACC 320 can run a control plane operating system (OS) (e.g., Linux) and / or a control plane application 322 (e.g., user-space or kernel modules) that are used by the SDN controller 350 to configure the operation of the packet processing pipeline 340. The control plane application 322 can include generic flow tables (GFT), ESXi, NSX and Kubernetes control plane software, application software for managing crypto configurations, a runtime daemon for protocol-independent packet processors (P4), a target-specific daemon, CSI (Container Storage Interface) agents or RDMA (Remote Direct Memory Access) configuration agents.
[0018] In some examples, the SDN Controller 350 can communicate with the ACC 320 using a remote procedure call (RPC), such as Google Remote Procedure Call (gRPC), or another service, and the ACC 320 can convert the request into a destination-specific protocol buffer (Protobuf) request to the MCC 330. gRPC is a remote procedure call solution based on data packets sent between a client and a server. Although gRPC is just one example, other communication schemes can be used, such as, but not limited to, Java Remote Method Invocation, Modula-3, RPyC, Distributed Ruby, Erlang, Elixir, Action Message Format, Remote Function Call, Open Network Computing RPC, JSON-RPC, and so on.
[0019] In some examples, the SDN controller 350 can provide packet processing rules for execution by the ACC 320. The ACC 320 can, for example, program table rules (such as header field matching and corresponding action) that are applied by the packet processing pipeline circuit 340 based on policy changes and changes in VMs, containers, microservices, applications, or other processes. The ACC 320 can be configured to provide a network policy as flow cache rules in a table to configure the operation of the packet processing pipeline 340. For example, the ACC-executing control plane application 322 can configure rule tables, applied by the packet processing pipeline circuit 340, with rules to define a traffic destination based on packet type and content. The ACC 320 can provide table rules (such as...)Match action) in memory that the packet processing pipeline circuitry 340 can access, based on a policy change and changes in VMs.
[0020] A flow can be a sequence of packets transmitted between two endpoints, generally representing a single session using a protocol. Accordingly, a flow can be identified by matching a set of defined tuples, and for routing purposes, a flow is identified by the two tuples that identify the endpoints, such as the source and destination addresses. For content-based services (e.g., load balancers, firewalls, intrusion detection systems, etc.), flows can be identified with finer granularity by using N-tuples (e.g., source address, destination address, IP protocol, transport layer source port, and destination port). A packet in a flow is expected to have the same set of tuples in its packet header. A packet flow to be controlled can be identified by a combination of tuples (e.g.,Ethernet type field, source and / or destination IP address, source and / or destination User Datagram Protocol (UDP) ports, source / destination TCP ports or any other header field) and a unique source and destination queue pair (QP - Queue Pair) number or identifier.
[0021] For example, the ACC 320 can run a virtual switch, such as vSwitch or Open vSwitch (OVS), Stratum, or Vector Packet Processing (VPP), which provides communication between virtual machines running on the host 300 or with other devices connected to a network. For example, the ACC 320 can configure the packet processing pipeline circuitry 340 to specify which VM should receive traffic and what type of traffic a VM can transmit. For example, the packet processing pipeline circuitry 340 can run a virtual switch, such as vSwitch or Open vSwitch, which provides communication between virtual machines running on the host 300 and the packet processing device 310.
[0022] The MCC 330 can run a host management control plane, a global resource manager, and perform hardware register configuration. The control plane 332, run by the MCC 330, can provision and configure the packet processing circuitry 340. For example, a virtual machine (VM) running on the host 300 can use the packet processing device 310 to receive or transmit packet traffic. The MCC 330 can execute startup, performance, management, and manageability software (SW) or firmware (FW) code to start and initialize the packet processing device 310, manage device power consumption, provide connectivity to a baseboard management controller (BMC), and perform other operations.
[0023] One or both control planes of the ACC 320 and the MCC 330 can define the contents of the traffic routing table and the network topology applied by the packet processing circuit arrangement 340 to select a path for a packet in a network to the next hop or to a device connected to the destination network. For example, a VM running on the host 300 can use the packet processing device 310 to receive or transmit packet traffic.
[0024] The ACC 320 can execute control plane drivers to communicate with the MCC 330. At a minimum, to provide a configuration and deployment interface between control planes 322 and 332, the communication interface 325 can enable control plane-to-control plane communication. Control plane 332 can perform a gatekeeper operation to configure shared resources.For example, the ACC control plane 322 can communicate with the control plane 332 via the communication interface 325 to perform one or more of the following: determining hardware capabilities, accessing the data plane configuration, reserving hardware resources and configuration, communicating between ACC and MCC through interrupts or queries, subscribing to receive hardware events, performing indirect hardware register read / write operations for debuggability, configuring flash and physical layer (PHY) interfaces, or performing system provisioning for different uses of a network interface device, such as: storage nodes, tenant hosting nodes, microservice backends, compute nodes, or others.
[0025] The communication interface 325 can be used by a negotiation protocol and configuration protocol that are executed between the ACC control plane 322 and the MCC control plane 332. The communication interface 325 can include a general-purpose mailbox for various operations performed by the packet processing circuit arrangement 340. Examples of packet processing circuit arrangement 340 operations include NVMe (Non-Volatile Memory Express) read or write output, NVMe-oF™ (Non-Volatile Memory Express over Fabrics) read or write output, LCE (Lookaside Crypto Engine) (e.g., compression or decompression), and ATE (Address Translation Engine) (e.g.,Input / Output Memory Management Unit (IOMMU) for providing a translation from virtual to physical address, encryption or decryption, configuration as a storage node, configuration as a tenant hosting node, configuration as a compute node, provision of several different types of services between different Peripheral Component Interconnect Express (PCIe) endpoints, or others.
[0026] Communications interface 325 can contain one or more mailboxes, which can be accessed as registers or memory addresses. For communications from control plane 322 to control plane 332, communications can be written to one or more mailboxes by control plane drivers 324. Communications written to mailboxes can include descriptors containing message opcodes, message errors, message parameters, and other information. Communications written to mailboxes can contain messages in a defined format that transmit data.
[0027] The 325 communication interface can provide communication based on write or read operations to specific memory addresses (e.g., dynamic random access memory (DRAM)), registers, and other mailboxes to which data is written and from, in order to forward commands and data. To provide secure communication between the 322 and 332 control planes, registers and memory addresses (and memory address translations) can be available for communication only to or from the 322 and 332 control planes, or to cloud service provider (CSP) software running on the ACC 320, and device provider software, embedded software, or firmware running on the MCC 330.The communication interface 325 can support communication between several different computing complexes, such as from the host 300 to the MCC 330, from the host 300 to the ACC 320, from the MCC 330 to the ACC 320, from the baseboard management controller (BMC) to the MCC 330, from the BMC to the ACC 320 or from the BMC to the host 300.
[0028] The packet processing circuitry 340 can be implemented using one or more of the following: application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), software-executing processors, or another circuitry. The control plane(s) 322 and / or 332 can configure the packet processing pipeline circuitry 340 or other processors to perform operations related to NVMe, NVMe-oF read or write operations, the lookaside crypto engine (LCE), the address translation engine (ATE), the local area network (LAN), compression / decompression, encryption / decryption, or other accelerated operations.
[0029] Various message formats can be used to configure the ACC 320 or the MCC 330. In some examples, a P4 program can be compiled and provided to the MCC 330 to configure the packet processing circuit arrangement 340. The following is a JSON configuration file that can be transferred from the ACC 320 to the MCC 330 to enable the capabilities of the packet processing circuit arrangement 340 and / or another circuit arrangement in the packet processing device 310.In particular, the file can be used to specify a number of send queues, a number of receive queues, a number of supported traffic classes (TC), a number of available interrupt vectors, a number of available virtual ports and the types of ports, a size of allocated memory, supported parser profiles, exactly matching table profiles, packet mirroring profiles, among others.
[0030] As described in this document, the Network Interface Device (NID) 310 can operate in standalone mode or companion mode. The NID 310 may include an interface or registers that allow configuration of an operating mode 342 of the NID 310. For example, the host 300, an SDN controller 350, or another entity can configure an operating mode of the NID 310.
[0031] In standalone mode, the NID 310 can control one or more processes using the ACC 320 and a different circuit arrangement (e.g., one related to Fig. (circuit arrangement described in 3B). For example, in standalone mode, the NID 310 can run one or more applications and does not use the CPU of the host 300. In companion mode, the host 300 can offload a workload to the NID 310 for various reasons (e.g., the host 300's power consumption is reached or exceeded, the temperature limit is reached or exceeded, insufficient capacity, excessive load) and can transfer the workload (e.g., a process) to the NID 310 for execution.
[0032] The NID 310 can operate in a standalone mode in the event of a host 300 failure. The host 300 can expose its heartbeat and operational monitoring to the NID 310 to monitor the host 300's state (e.g., processors and / or memory). For example, the NID 310 can monitor specific memory areas or access to certain types of registers (e.g., model-specific registers - MSR) that store the host 300's state. If the NID 310 identifies that the host 300 is not operating or functioning (e.g., excessive temperature, insufficient power, or excessive load), the NID 310 can transition to failover mode and take over applications previously running on the host 300. The NID 310 can migrate or instantiate the applications running on the host 300 to run on cores or processors of the ACC 320.In some cases, because the cores of the ACC 320 and the host 300 may be of different technologies (e.g., ARM in the ACC 320 versus x86 in the host 300), the NID 310 can utilize a table of applications to be instantiated and pointers to the ARM binaries to be started on the ACC 320. The NID 310 can copy the process state from memory or a specific memory area of the host 300 to resume the operation of an application that was previously running on the host 300. The NID 310 can coordinate with a management controller (e.g., BMC) to put the host 300 (e.g., processors and / or memory) into a low-power state (e.g., C6 state), thus making it the primary system.
[0033] In some examples, the NID 310 can operate in a standalone mode based on detecting that the host interface 344 (e.g., CXL, PCIe, UPI) to the host 300 is not operational or accessible.
[0034] In some examples, based on reducing the power to a CPU of the host 300 and with the NID 310 operating in standalone mode to run one or more applications previously executed by the host 300, the discrete devices connected to the CPU as endpoints (and the corresponding CXL connectivity) can be mapped to the NID 310 and made accessible to processors or accelerators of the NID 310. Conversely, based on reducing the power to a processor or accelerator of the NID 310 and with the host 300 operating in standalone mode to run one or more applications previously executed by the NID 310, the discrete devices connected to the NID 310 as endpoints (and the corresponding CXL connectivity) can be mapped to the host 300 and made accessible to processors or accelerators of the host 300.
[0035] In some examples, memory on host 300 can be shared by applications running on host 300 and applications running on NID 310. Execution pointers of specific threads can pause or begin using interfaces. In such cases, the state might not be migrated to NID 310, but the final application state can be marked, and a pointer can be changed from the source to the target system via interfaces.
[0036] An Orchestrator or SDN Controller 350 can trigger a seamless migration of the entire active state of work jobs with a stateful failover from Host 300 (or another NID) to NID 310. Specifically, example failover scenarios using consistent Compute Express Link (CXL) storage interfaces are as follows: Scenario (1), migrating the application and application state from NID 1 to NID 2, with the source and destination NIDs connected via CXL interfaces. Scenario (2), migrating the application and application state from Host 300 to NID 310, which are connected via CXL. Scenario (3), migrating the application and application state from NID 310 to Host 300, with NID 310 as the source and Host 300 as the destination.
[0037] In some examples, the instruction semantics of a processor or accelerator in host 300, which is running an application, may be compatible with the instruction semantics of a processor or accelerator in NID 310, which is selected to run the migrated application. In such cases, the application can be migrated from host 300 to NID 310 or instantiated on NID 310.
[0038] In some examples, the instruction semantics of a processor or accelerator in the Host 300 running an application may differ from the instruction semantics (e.g., Instruction Set Architecture - ISA) of a processor or accelerator in the NID 310 selected to run the migrated application. In such cases, an executable binary or kernel that can run on a selected processor or accelerator of the NID 310 can be retrieved from, connected to, or transferred to the NID 310 and executed on the NID 310. In such cases, the NID 310 can translate a binary file associated with the migrated application from the host 300 into a format (e.g., ISA or kernel) that can be run on a selected processor or accelerator of the NID 310.In some cases, a selected processor or accelerator of the NID 310 can perform processor emulation to run the migrated application by translating processor instructions and operating system calls while an application is running.
[0039] Migrating an application running on the NID 310 to run it on a target processor or accelerator in the Host 300 may involve operations similar to those used to migrate an application from the Host 300 to the NID 310, in order to provide an application that can be run by a selected processor in the Host 300.
[0040] Example failover operations might look like this. In (1), the application runs on the source host system 300. In (2), the orchestrator triggers a failover from the source system 300 to the target system (e.g., NID 310). In (2A), the orchestrator can attest to the source system 300. In (2B), the orchestrator attests to the target system (e.g., NID 310). In (2C), attestation results are compared to determine if the trusted environments are approximately the same. If the trusted environments are approximately the same, the triggered failover proceeds. If the trusted environments are not approximately the same, another target can be identified, and operations 2A-2C are repeated for that target. In (3), the memory contents are copied from local memory to shared memory (e.g., via the CXL interface). In (4) a corresponding application is produced on the target system.In step (5), storage content is copied from the shared storage to the target system. Content can be authenticated by providing metadata describing the environment from which the content originated, as well as metadata from an environment that processed the data. The Coalition for Content Provenance and Authenticity (C2PA) version 1.0 is an exemplary standard for content exchange that can be extended to include the attestation context.
[0041] In (6), the execution state on the source system (e.g., host 300) is suspended (standby) and switched to the target system. The source system's performance state can be placed into a deep sleep if needed, or it can continue processing other applications. In (7), returning from standby or sleep triggers a layered attestation environment or a root of trust to collect attestation measurements of the newly established execution state. If cryptographic keys are derived using a Device Identity Composition Engine (DICE), layer keys can be regenerated before an attestation report. Note that failover operations can occur from NID 310 to another NID, or from NID 310 to host 300 or another host system.After migrating an application to the NID 310 and detecting that the host 300 is running, the application can be migrated back to run on the host 300.
[0042] For example, a processor core can be chosen as the primary attestation environment, which executes MCHECK, querying other cores to verify that they possess an attestation key. The primary key can attest to the existence of other keys, allowing a verifier to perform revocation and authorization checks. If ARM cores follow a similar approach, a second-level function is required that enables a PASID context for a Level 2 primary key / leading attestation, which either verifies Level 1 attestations or forwards Level 1 attestations to a remote verifier (outside the host 300). The Level 2 primary key can then use its attestation key to sign secondary Level 2 keys (which become the primary Level 1 entities).
[0043] The NID 310 can provide an endpoint for devices connected to the storage complex (e.g., Compute Express Link (CXL) connected devices) of the Host 300. As described in this document, the NID 310 can be connected to a CXL fabric, which provides access to discrete devices (e.g., GPUs, accelerators, processors, or storage devices). In companion mode, the NID 310 may not have access to the CXL fabric, but in standalone mode, the NID 310 can access the CXL fabric, thus enabling access to discrete devices.
[0044] Fig. Figure 3B presents an exemplary network interface device system. Various examples of a packet handling device or a network interface device 301 can be components of the system. Fig. Use 3B. In some examples, a packet processing device or network interface device may refer to one or more of the following: a network interface controller (NIC), an RDMA (Remote Direct Memory Access)-enabled NIC, SmartNIC, router, switch, forwarding element, infrastructure processing unit (IPU), data processing unit (DPU), or edge processing unit (EPU). An edge processing unit (EPU) may include a network interface device that utilizes processors and accelerators (e.g., digital signal processors - DSPs), signal processors, or wireless-specific accelerators for virtualized radio access networks (vRANs), cryptographic operations, compression / decompression, and so on. The network subsystem 360 may be communicatively coupled to the compute complex 380.Device interface 362 can provide an interface for communicating with a host. Various examples of device interface 362 can utilize protocols based on Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), or others, as well as a virtual device interface, such as virtual device interfaces.
[0045] Interfaces 364 can initiate and terminate at least offload RDMA (Remote Direct Memory Access) operations, NVMe (Non-Volatile Memory Express) read or write operations, and LAN operations. Packet processing pipeline 366 can perform packet processing (e.g., packet headers and / or packet payloads) based on a configuration and support Quality of Service (QoS) and telemetry reporting. Inline processor 368 can perform offloaded encryption or decryption of packet communications (e.g., Internet Protocol Security (IPSec) or others). Traffic structuring 370 can schedule the transmission of communications. Network interface 372 can provide an interface to at least one Ethernet network through Media Access Control (MAC) and serializer / deserializer (Serdes) operations.
[0046] The 382 cores can be configured to perform infrastructure operations, such as storage initiator, transport layer security (TLS) proxy, virtual switch (e.g., vSwitch), or other operations. Memory 384 can store applications and data to be executed or processed. The 386 swap circuitry can perform at least cryptographic and compression operations for the host or for use by the 380 compute complex. The 386 swap circuitry can include one or more graphics processing units (GPUs) that can access memory 384. The 388 management complex can perform secure boot, lifecycle management, and management of the 360 network subsystem and / or the 380 compute complex.
[0047] Fig. 3C presents an exemplary system. The Media / AI Tile 390 can be integrated into a network interface device. The Media / AI 390 can be enabled or disabled by configuration from an orchestrator or administrator, and when enabled, it can be accessed through a network interface device. The Media / AI 390 can utilize a root port (RP) to read from or write to a storage subsystem (e.g., Double Data Rate (DDR)-compatible storage and cache). A bridge from the Media / AI 390 to the ARM subsystem can be provided by APB. Storage traffic from the Media / AI Tile 390 can be transferred to the ARM subsystem using an ARM-compatible Advanced Peripheral Bus (APB).An input / output memory management unit (IOMMU) can provide memory address translation between devices.
[0048] Fig. 4A and Fig. Figure 4B represents exemplary configurations. Depending on the deployment model, a network interface device can operate as a standalone computing platform or as a counterpart to a computing platform. Fig. 4A represents several independent model configurations. Applications running on the NID 400's compute complex 406 (e.g., the processors) can utilize ASICs 402, a GPU 404, and memory 408 to perform operations.
[0049] Applications running on the NID 410's computing complex 416 (e.g., the processors) can utilize ASICs 412, a GPU 414, a memory 418, and a discrete GPU 420 connected to the NID 410 via a device interface (e.g., PCIe or CXL) to perform operations.
[0050] Applications running on the NID 410's compute complex 436 (e.g., the processors) can utilize ASICs 432, a GPU 434, a memory 438, and a discrete GPU 440, as well as discrete NVMe storage drives connected to the NID 410 via a device interface (e.g., PCIe or CXL), to perform operations.
[0051] It should be noted that in one or more standalone modes, the host system connected to an NID may remain connected to the NID but be shut down or in power-reduction mode. Discrete devices connected to a host may be accessible to the NID in some examples. For instance, host system 409, which is communicatively coupled to NID 400, may provide access to discrete devices (such as accelerators, data storage, memory, GPU, or others) to NID 400. Similarly, host system 419, which is communicatively coupled to NID 410, may provide access to discrete devices (such as accelerators, data storage, memory, GPU, or others) to NID 410. Similarly, the host system 439, which is communicatively coupled with the NID 430, can access discrete devices (e.g.Provide accelerators, (data) storage, (working) memory, GPU or other) to the NID 430.
[0052] When a host core or CPU reduces its power or is turned off, discrete devices exposed as endpoints for the CPU (and the corresponding CXL functionalities) can be mapped to the NID and made accessible to cores or processors of the NID.
[0053] Fig. 4B represents an exemplary companion model configuration. The host server 460 can run one or more applications that utilize resources of the NID 450 (e.g., ASICs 452, GPU 454, compute unit 456 and / or memory 458).
[0054] Fig. Figure 5 presents an example of a service migration. Scenario 502 involves a standalone deployment of a service on an NID infrastructure. Scenario 504 involves a Service 1 that utilizes CPU computational capacity to perform calculations. Scenario 506 involves a Service 1 that significantly reduces computational demands, and the system may attempt to achieve power savings (e.g., to reduce battery consumption). Accordingly, the NID on Edge 1 migrates Service 1 to its cores, switches to standalone mode, and the CPU enters a lower-power state. In Scenario 506, Service 1, running in standalone mode on the NID, can access memory and GPUs connected to the CPU via PCIe or another interface. Devices connected via PCIe can act as a root port for the NID, either instead of or in addition to being exposed to the host.
[0055] At 508, the computational demands of Service 1 increase, but Edge 1 cannot provide the power to resume Service 1 on the CPU. The NIDs on Edge 1 and Edge 2 communicate to migrate Service 1 to the CPU on Edge 2. The NID on Edge 1 remains in standalone mode, and the NID on Edge 2 releases media and AI processors as VFs on the CPU on Edge 2 to perform operations for Service 1.
[0056] Fig. Section 6 presents an example process. The process can be performed by an orchestrator. At 602, a determination of the operating mode of a network interface device can be made. For example, the network interface device can operate in companion mode or in standalone mode. Based on the network interface device operating in standalone mode, the process can proceed to 604. Based on the network interface device operating in companion mode, the process can proceed to 610.
[0057] With option 604, in a standalone mode, an application can be assigned to run exclusively on the network interface device. For example, the application can run on processors of the network interface device and utilize a circuit arrangement within the network interface device to perform media processing and inference operations. In some examples, the network interface device can be selected to run an application that has been migrated from a server or was previously running on a server. Application state data can be migrated or copied from the server to the network interface device. In some examples, an orchestrator or the network interface device can detect a server failure and initiate a migration of the application and its associated state to the network interface device.
[0058] With version 610, in companion mode, an application can be assigned to run on a server and utilize hardware resources of the network interface device. For example, the server and the network interface device can be communicatively coupled using a network, a storage interface (such as CXL), or a device interface (such as PCIe). In some examples, the network interface device can be selected to run an application that was migrated from a server or was previously run by a server. Various application and state migration technologies are described in this document.
[0059] Fig. Figure 7 illustrates an exemplary network interface device or packet processing device. In some examples, the network interface device circuitry can be used to run applications or provide hardware resources in standalone or companion mode, as described in this document. In some examples, the packet processing device 700 can be implemented as a network interface controller, a network interface card, a host fabric interface (HFI), or a host bus adapter (HBA), and such examples can be interchangeable. The packet processing device 700 can be coupled to one or more servers using a bus, PCIe, CXL, or dual data rate (DDR).The packet processing device 700 can be implemented as part of a system-on-chip (SoC) that includes one or more processors, or it can be contained on a multi-chip package that also includes one or more processors.
[0060] Some examples of the Packet Processing Device 700 are part of, or utilized by, an Infrastructure Processing Unit (IPU) or a Data Processing Unit (DPU). An xPU can refer to at least one IPU, DPU, GPU, GPGPU, or other processing units (e.g., accelerator devices). An IPU or DPU may include a network interface with one or more programmable or fixed-function processors to offload operations that could have been performed by a CPU. The IPU or DPU may include one or more storage devices. In some examples, the IPU or DPU may perform virtual switch operations, manage storage transactions (e.g., compression, encryption, virtualization), and manage operations performed on other IPUs, DPUs, servers, or devices.
[0061] The network interface 700 can include a transceiver 702, processors 704, a transmit queue 706, a receive queue 708, a (working) memory 710, a host interface 712, and a DMA engine 772. The transceiver 702 can be capable of receiving and transmitting packets in accordance with applicable protocols, such as Ethernet as described in IEEE 802.3, although other protocols may be used. The transceiver 702 can receive packets from a network via a network medium (not shown) and transmit them to that network. The transceiver 702 can include a PHY circuitry 714 and a media access control (MAC) circuitry 716.The PHY circuit arrangement 714 can include an encoding and decoding circuit arrangement (not shown) configured to encode and decode data packets according to applicable physical layer specifications or standards. The MAC circuit arrangement 716 can be configured to group data to be transmitted into packets containing destination and source addresses along with network control information and error detection hash values.
[0062] The System-on-Chip (SoC) 750 and the 704 processors can be any combination of the following: processor, core, graphics processing unit (GPU), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or any other programmable hardware device that enables programming of the 700 network interface. For example, a "smart network interface" can provide packet processing capabilities in the network interface using 704 processors.
[0063] The 704 processors can include one or more packet processing pipelines, which can be configured to perform a matching action on received packets to identify packet processing rules and next jumps using information stored in ternary content-addressable memory tables (TCAM tables) or, in some embodiments, exact-match tables. For example, matching action tables or circuit arrangements can be used, where a hash of a portion of a packet is used as an index to locate an entry. The packet processing pipelines can perform one or more of the following: packet parsing, exact-match action (e.g.,A small exact match (SEM) engine or large exact match (LEM) engine, a wildcard match action (WCM), a longest prefix match block (LPM), a hash block (e.g., receive-side scaling - RSS), a packet modifier, or a traffic manager (e.g., send rate metering or shaping). Packet processing pipelines can, for example, implement an access control list (ACL) or packet loss due to queue overflow.
[0064] The configuration of the operation of the 704 processors, including their data plane, can be programmed based on one or more of the following: Protocol-independent Packet Processors (P4), Software for Open Networking in the Cloud (SONiC), Broadcom® Network Programming Language (NPL), NVIDIA® CUDA®, NVIDIA® DOCA™, Infrastructure Programmer Development Kit (IPDK), among others.
[0065] A packet allocator 724 can distribute received packets for processing by multiple CPUs or cores using time slot allocation or RSS, as described in this document. If the packet allocator 724 uses RSS, it can calculate a hash or perform another determination based on the contents of a received packet to determine which CPU or core should process it.
[0066] An interrupt 722 interface can perform interrupt moderation, where it waits for multiple packets to arrive or for a timeout to expire before generating an interrupt for the host system to process one or more received packets. Receive segment coalescing (RSC) can be performed by network interface 700, which combines portions of incoming packets into segments of a single packet. Network interface 700 then provides this coalesced packet to an application.
[0067] A direct memory access (DMA) engine 772 can copy a packet header, packet payload, and / or descriptor directly from host memory to the network interface, or vice versa, instead of copying the packet to an intermediate buffer in the host and then using another copy operation from the intermediate buffer to the destination buffer.
[0068] The (working) memory 710 can be any type of volatile or non-volatile storage device and can store any queue or instructions used to program the network interface 700. The transmit queue 706 can contain data or references to data for transmission through the network interface. The receive queue 708 can contain data or references to data received by the network interface from a network. Descriptor queues 720 can contain descriptors that reference data or packets in the transmit queue 706 or the receive queue 708. The host interface 712 can provide an interface to the host device (not shown).For example, the host interface 712 can be compatible with PCI, PCI Express, PCI-x, Serial ATA and / or a USB-compatible interface (although other connection standards may be used).
[0069] Fig. Figure 8 represents a system. In some examples, the circuit arrangement of the network interface device can be used to run applications or provide hardware resources in standalone or companion mode, as described in this document. The System 800 includes a Processor 810, which provides processing, operations management, and instruction execution for the System 800. The Processor 810 can include any type of microprocessor, central processing unit (CPU), graphics processing unit (GPU), XPU, processing core, or other processing hardware to provide processing for the System 800, or a combination of processors. An XPU can include one or more of the following: a CPU, a graphics processing unit (GPU), a general-purpose GPU (GPGPU), and / or other processing units (for example, accelerators or programmable or fixed-function FPGAs).The Processor 810 controls the overall operation of the System 800 and may be or include one or more programmable general-purpose or special-purpose microprocessors, digital signal processors (DSPs), programmable controllers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or the like, or a combination of such devices.
[0070] In one example, the System 800 includes an 812 interface coupled to the 810 processor. This interface can represent a higher-speed or high-throughput interface for system components requiring higher-bandwidth connections, such as the 820 memory subsystem, the 840 graphics interface components, or the 842 accelerators. The 812 interface represents an interface circuit that can be a standalone component or integrated on a processor chip. Where present, the 840 graphics interface provides an interface with graphics components to provide a visual display to a System 800 user. In one example, the 840 graphics interface can drive a display that provides output to a user. In another example, the display can be a touchscreen display.In one example, the 840 graphics interface generates a display based on data stored in the 830 memory, or based on operations performed by the 810 processor, or both.
[0071] The Accelerators 842 can be a programmable or fixed-function swap engine that an 810 processor can access or use. For example, an Accelerator 842 can provide data compression (DC) capability, cryptographic services such as public key encryption (PKE), ciphers, hash / authentication capabilities, decryption, or other capabilities or services. In some cases, the Accelerators 842 can be integrated into a CPU socket (e.g., a connector on a motherboard or a circuit board containing a CPU, forming an electrical interface with the CPU).The Accelerator 842 can include, for example, a single- or multi-core processor, a graphics processing unit (GPU), a single- or multi-level cache of a logic execution unit (LEU), functional units that can be used to independently execute programs or threads, application-specific integrated circuits (ASICs), neural network processors (NNPs), programmable control logic, and programmable processing elements such as field-programmable gate arrays (FPGAs). The Accelerator 842 can provide multiple neural networks, CPUs, processor cores, general-purpose graphics units, or graphics processing units that can be deployed for use by artificial intelligence (AI) or machine learning (ML) models.The AI model can, for example, use or include one or a combination of the following: a reinforcement learning scheme, a Q learning scheme, deep Q learning or asynchronous advantage actor-critic (A3C), a combinatorial neural network, a recurrent combinatorial neural network, or another AI or ML model. Multiple neural networks, processor cores, or graphics processing units can be made available to AI or ML models to perform learning and / or inference operations.
[0072] The (working) memory subsystem 820 represents the main memory of the System 800 and provides storage for code to be executed by the Processor 810 or data values to be used when executing a routine. The memory subsystem 820 can include one or more memory devices 830, such as read-only memory (ROM), flash memory, one or more types of random-access memory (RAM), such as DRAM, or other memory devices, or a combination of such devices. Memory 830 stores and hosts, among other things, the Operating System (OS) 832 to provide a software platform for executing instructions in the System 800. In addition, applications 834 can be executed on the OS 832 software platform from memory 830. Applications 834 represent programs that have their own operating logic for performing one or more functions.The processes 836 represent agents or routines that provide auxiliary functions to the OS 832, one or more applications 834, or a combination thereof. The OS 832, the applications 834, and the processes 836 provide software logic to provide functions to the System 800. In one example, the memory subsystem 820 includes a memory controller 822, which is a memory controller for generating and issuing instructions to the memory 830. It is understood that the memory controller 822 can be a physical part of the processor 810 or a physical part of the interface 812. The memory controller 822 can, for example, be an integrated memory controller that is integrated on a circuit with the processor 810.
[0073] Applications 834 and / or processes 836 can instead or additionally reference a virtual machine (VM), container, microservice, processor, or other software. Several examples described in this document can run a microservice-based application, with each microservice running in its own process and communicating using protocols (such as the Application Programming Interface (API), Hypertext Transfer Protocol (HTTP) Resource API, messaging service, Remote Procedure Calls (RPC), or Google RPC (gRPC)). Microservices can communicate with each other using a service mesh and can run in one or more data centers or edge networks. Microservices can be deployed independently using centralized management of these services.The management system can be written in various programming languages and use different data storage technologies. A microservice can be characterized by one or more of the following: polyglot programming (e.g., code written in multiple languages to capture additional functionality and efficiency not available in a single language), deployment in a lightweight container or virtual machine, and decentralized continuous delivery of microservices.
[0074] In some examples, the operating system can be Linux®, Windows® Server or Personal Computer, FreeBSD®, Android®, macOS®, iOS®, VMware vSphere, openSUSE, RHEL, CentOS, Debian, Ubuntu, or any other operating system. The operating system and driver can run on a processor sold or developed by, among others, Intel®, ARM®, AMD®, Qualcomm®, IBM®, NVIDIA®, Broadcom®, and Texas Instruments®.
[0075] Although not explicitly illustrated, it is understood that System 800 may include one or more buses or bus systems between devices, such as a memory bus, a graphics bus, interface buses, or others. Buses or signal lines may couple components communicatively or electrically, or couple components both communicatively and electrically. Buses may include physical communication lines, point-to-point connections, bridges, adapters, controllers, or other circuit arrangements, or a combination thereof. For example, buses may include a system bus and / or a Peripheral Component Interconnect (PCI) bus and / or a HyperTransport or Industry Standard Architecture (ISA) bus and / or a Small Computer System Interface (SCI) bus and / or a Universal Serial Bus (USB) or a bus conforming to Institute of Electrical and Electronics Engineers (IEEE) Standard 1394 (FireWire).
[0076] In one example, the System 800 includes an 814 interface, which can be coupled to the 812 interface. In another example, the 814 interface represents an interface circuit that can include standalone components and integrated circuit arrangements. In yet another example, multiple user interface components or peripheral components, or both, are coupled to the 814 interface. The 850 network interface provides the System 800 with the ability to communicate with remote devices (such as servers or other computing devices) over one or more networks. The 850 network interface can include an Ethernet adapter, wireless connectivity components, cellular connectivity components, USB (Universal Serial Bus), or other interfaces based on wired or wireless standards, or proprietary interfaces.The Network Interface 850 can transmit data to a device located in the same data center or rack, or to the same remote device, which may involve sending data stored in memory. The Network Interface 850 can receive data from a remote device, which may involve storing received data in memory. In some examples, the packet handling device or the Network Interface 850 device may reference one or more of the following: a Network Interface Controller (NIC), an RDMA (Remote Direct Memory Access)-enabled NIC, a SmartNIC, a router, a switch, a forwarding element, an Infrastructure Processing Unit (IPU), or a Data Processing Unit (DPU). An example IPU or DPU is shown with reference to... Fig. 7 described.
[0077] In one example, the System 800 includes one or more input / output (I / O) interfaces 860. The I / O interface 860 can include one or more interface components through which a user interacts with the System 800. The peripheral interface 870 can include any hardware interface not previously specifically mentioned. Peripherals generally refer to devices that establish a dependent connection to the System 800.
[0078] In one example, the System 800 includes a (data) storage subsystem 880 for storing data in a non-volatile manner. In another example, in certain system implementations, at least some components of the (data) storage 880 may overlap with components of the (working) storage subsystem 820. The storage subsystem 880 includes one or more storage devices 884, which may be or include any conventional medium for the non-volatile storage of large amounts of data, such as one or more magnetic, solid-state, or optical disks, or a combination thereof. The storage 884 contains code or instructions and data 886 in a persistent state (e.g., the value is retained even if power is interrupted to the System 800).The (data) memory 884 can generally be considered a "memory," although the (working) memory 830 is typically the execution or working memory for providing instructions to the processor 810. While memory 884 is non-volatile, memory 830 may contain volatile memory (e.g., the value or state of the data is indeterminate if power to system 800 is interrupted). In one example, the memory subsystem 880 includes a controller 882 for forming an interface with memory 884. In another example, the controller 882 is a physical part of interface 814 or processor 810, or it may include circuitry or logic in both processor 810 and interface 814.
[0079] A volatile memory is a memory device whose state (and therefore the data stored in it) is indeterminate if the power supply to the device is interrupted. A non-volatile memory device (NVM device) is a memory device whose state is definite, even if the power supply to the device is interrupted.
[0080] In one example, System 800 can be implemented using interconnected computing sleds with processors, (working) memory, (data) storage, network interfaces, and other components. High-speed interconnects can be used, such as Ethernet (IEEE 802.11).3), Remote Direct Memory Access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), quick UDP Internet Connections (QUIC), RDMA over Converged Ethernet (RoCE), Peripheral Component Interconnect express (PCIe), Intel QuickPath Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), Omni-Path, Compute Express Link (CXL), HyperTransport, a high-speed fabric, NVLink, Advanced Microcontroller Bus Architecture (AMBA) interconnect, OpenCAPI, Gen-Z, Infinity Fabric (IF), a cache coherent interconnect for accelerators (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G and variations thereof. Data can be copied or stored in virtualized storage nodes or accessed using a protocol such as NVMe over Fabrics (NVMe-oF) or NVMe (e.g.,can operate an NVMe (Non-Volatile Memory Express) device in a manner consistent with the Non-Volatile Memory Express (NVMe) specification, Revision 1.3c, published on May 24, 2018 (“NVMe Specification”) or derivatives or variations thereof).
[0081] Communications between devices can take place using a network that provides die-to-die communication; chip-to-chip communication; circuit board-to-circuit board communication and / or enclosure-to-enclosure communication.
[0082] In one example, the System 800 can be implemented using interconnected computing sleds with processors, (working) memory, (data) storage, network interfaces, and other components. High-speed interconnects can be used, such as PCIe, Ethernet, or optical interconnects (or a combination thereof).
[0083] Examples in this document can be implemented in various types of computing and networking equipment, such as switches, routers, racks, and blade servers, like those used in a data center and / or server farm environment. Servers used in data centers and server farms include arrayed server configurations, such as rack-based servers or blade servers. These servers are interconnected via various network facilities, such as partition sets of servers in local area networks (LANs) with appropriate switching and routing equipment between the LANs to form a private intranet. For example, cloud hosting facilities typically utilize large data centers with a multitude of servers. A blade comprises a separate computing platform configured to perform server-like functions, i.e., a "server on a card."Accordingly, each blade includes components common to conventional servers, including a printed main circuit board (mainboard) that provides internal wiring (for example, buses) for coupling suitable integrated circuits (ICs) and other components mounted on the board.
[0084] Various examples can be implemented using hardware elements, software elements, or a combination of both. In some examples, hardware elements might include devices, components, processors, microprocessors, circuits, circuit elements (for example, transistors, resistors, capacitors, inductors, and so on), integrated circuits, ASICs, PLDs, DSPs, FPGAs, memory units, logic gates, registers, semiconductor devices, chips, microchips, chipsets, and so on. In some examples, software elements might include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, APIs, instruction sets, computational codes, computer codes, code segments, computer code segments, words, values, symbols, or any combination thereof.The decision of whether to implement an example using hardware and / or software elements can vary depending on any number of factors, such as desired computing power, power consumption, thermal tolerances, processing cycle budget, input data rates, output data rates, memory resources, data bus speeds, and other design or performance constraints as desired for a particular implementation. A processor can be one or more combinations of a hardware state machine, digital control logic, a central processing unit, or any hardware, firmware, and / or software elements.
[0085] Some examples may be implemented using one or more manufacturing items or at least one computer-readable medium. A computer-readable medium may include a non-transitory storage medium for storing logic. In some examples, the non-transitory storage medium may include one or more types of computer-readable storage media capable of storing electronic data, including volatile or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, and so on.In some examples, the logic may include various software elements, such as software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, procedures, software interfaces, APIs, instruction sets, data processing code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof.
[0086] According to some examples, a computer-readable medium can include a non-transitory storage medium for storing or maintaining instructions which, when executed by a machine, computing device, or computing system, cause the machine, computing device, or computing system to perform procedures and / or operations according to the examples described. The instructions can include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, and the like. The instructions can be implemented according to a predefined computer language, manner, or syntax to instruct a machine, computing device, or computing system to perform a specific function.The instructions can be implemented using any suitable higher-level, lower-level, object-oriented, visual, compiled and / or interpreted programming language.
[0087] One or more aspects of at least one example can be implemented by representative instructions stored on at least one machine-readable medium representing different logics within the processor. When read by a machine, computing device, or system, these instructions cause the machine, computing device, or system to produce logic for executing the techniques described in this document. Such representations, known as "IP kernels," can be stored on a tangible, machine-readable medium and supplied to various customers or manufacturing facilities for loading into the manufacturing machines that actually produce the logic or the processor.
[0088] The phrase "(exactly) one example" or "(any) example" does not necessarily always refer to the same example or embodiment. Any aspect described in this document may be combined with any other aspect or similar aspect described in this document, regardless of whether the aspects are described with respect to the same figure or element. The splitting, omission, or inclusion of block functions shown in the accompanying figures does not imply that the hardware components, circuits, software, and / or elements for implementing these functions must necessarily be split, omitted, or included in embodiments.
[0089] Some examples can be described using the terms "coupled" and "connected," along with their respective derivatives. These terms are not necessarily synonymous. For example, descriptions using the terms "connected" and / or "coupled" may indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" can also mean that two or more elements are not in direct contact with each other but nevertheless work together or interact.
[0090] The terms "first," "second," and the like do not denote order, quantity, or importance in this document, but are used to distinguish one element from another. The term "one" in this document does not denote a quantity limitation, but rather the presence of at least one of the elements mentioned. The term "set" as used in this document with respect to a signal denotes a signal state in which the signal is active and which can be achieved by applying any logical level, either logical 0 or logical 1, to the signal. The terms "following" or "after" can refer to immediate continuation or to continuation following one or more other events. In alternative embodiments, other sequences of operations may also be performed. Furthermore, depending on the application, additional operations may be added or removed.Any combination of changes may be used, and average experts familiar with the doctrine of this revelation would understand the many variations, modifications, and alternative embodiments thereof.
[0091] Disjunctive language, such as the expression "at least one of X, Y, or Z," is understood, unless specifically stated otherwise, in the context in which it is generally used to indicate that an element, concept, etc., can be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Therefore, such disjunctive language is not intended to, and does not imply, that any particular embodiment requires at least one X, at least one Y, or at least one Z for each. Furthermore, conjunctive language, such as the expression "at least one of X, Y, and Z," should also be understood, unless explicitly stated otherwise, as X, Y, Z, or any combination thereof, including "X, Y, and / or Z."
[0092] Illustrative examples of the devices, systems, and methods disclosed in this document are provided below. An embodiment of the devices, systems, and methods may include one or more of any, or any combination of, the examples described below.
[0093] Example 1 includes one or more examples and includes a device comprising: a network interface device comprising: a network interface, a direct memory access (DMA) circuitry, a host interface, memory, one or more processors, and a circuitry designed to: based on an operating configuration specifying standalone operation, cause the network interface device to operate in standalone mode to run one or more applications; and based on an operating configuration specifying companion operation, cause the network interface device to operate in companion mode to provide access to one or more hardware resources that the network interface device can access to at least one host system.
[0094] Example 2 includes one or more examples and includes a second circuit arrangement for performing video analysis operations, wherein the second circuit arrangement is enabled or disabled based on a request to enable or disable the second circuit arrangement.
[0095] Example 3 includes one or more examples, wherein the circuit arrangement is designed to: monitor the operation of a host system to determine whether one or more processors of the host system executing a specific process are operational, and based on a determination that one or more processors of the host system are not operational, access the memory of the host system to perform a failover operation to cause the specific process to run on one or more processors of the network interface device and to copy state data associated with the specific process from the memory of the host system and / or from non-operational processors into the memory of the network interface device.
[0096] Example 4 includes one or more examples, wherein the circuit arrangement is designed to perform a binary translation of the specific process or processor emulation for execution on the one or more processors of the network interface device.
[0097] Example 5 includes one or more examples, wherein the circuit arrangement is designed to: cause the specific process to be executed on the one or more processors and copy state data of the specific process from the memory of the host system to the memory of the network interface device, based on the detection of the operation of at least one of the one or more processors of the host system, causing the execution of the specific process on the at least one of the one or more processors of the host system.
[0098] Example 6 includes one or more examples, wherein the circuit arrangement is designed to: after causing the specific process to be executed on the one or more processors and copying state data of the specific process from the memory of the host system to the memory of the network interface device, based on the detection of the operation of at least one processor of a second network interface device, causing the execution of the specific process on the at least one processor of the second network interface device.
[0099] Example 7 includes one or more examples, where the specific process includes one or more of the following: a virtual machine, a container, or a microservice.
[0100] Example 8 includes one or more examples, wherein the network interface device comprises a circuit arrangement for selecting a binary file for the specific process that is compatible with the one or more processors of the network interface device.
[0101] Example 9 includes one or more examples, wherein the network interface device is designed to: access, via a secure connection, the operational state of the at least one host system and the state data from one or more of the following: memory, registers or cache of the at least one host system.
[0102] Example 10 includes one or more examples and includes a non-transitory, computer-readable medium comprising instructions stored thereon which, when executed by one or more processors, cause the one or more processors to: configure a network interface device to operate as a standalone computing platform to run one or more applications, or to operate as a companion to provide access to one or more hardware resources accessible by the network interface device to at least one host system, wherein the network interface device is configured to: detect whether a host interface is present, and, based on a lack of access to the host interface, operate as the standalone computing platform, and wherein the network interface device is a network interface,includes a Direct Drive Access (DMA) circuit arrangement and a host interface.
[0103] Example 11 includes one or more examples, wherein the network interface device comprises a circuit arrangement for performing video analysis operations, and wherein the circuit arrangement is enabled or disabled based on a request to enable or disable the circuit arrangement.
[0104] Example 12 includes one or more examples and includes instructions stored therein which, when executed by one or more processors, cause the one or more processors to do the following: configure the network interface device to monitor the operation of a host system to determine whether one or more processors of the host system running a specific process are operational, and based on a determination that the one or more processors of the host system are not operational, access the memory of the host system to perform a failover operation to cause the specific process to run on the one or more processors of the network interface device and to copy state data associated with the specific process from the memory of the host system to the memory of the network interface device.
[0105] Example 13 includes one or more examples and includes instructions stored therein which, when executed by one or more processors, cause the one or more processors to: select an executable binary associated with the specific process for execution on the one or more processors of the network interface device.
[0106] Example 14 includes one or more examples and includes instructions stored therein which, when executed by one or more processors, cause the one or more processors to do the following: after causing the specific process to run on the one or more processors and copying state data of the specific process from the memory of the host system to the memory of the network interface device, based on the detection of the operation of at least one of the one or more processors of the host system, cause the specific process to run on the at least one of the one or more processors of the host system.
[0107] Example 15 includes one or more examples and includes instructions stored therein which, when executed by one or more processors, cause the one or more processors to do the following: after causing the specific process to run on the one or more processors and copying state data of the specific process from the memory of the host system to the memory of the network interface device, based on the detection of the operation of at least one processor of a second network interface device, cause the specific process to run on the at least one processor of the second network interface device.
[0108] Example 16 includes one or more examples and includes a method that includes an orchestrator that performs the following: causing an application to run on one or more of a server or network interface device, wherein the network interface device comprises a network interface, a direct storage access (DMA) circuit arrangement, and a host interface, and wherein the network interface device operates in stand-alone mode to run one or more applications, or operates as a companion to provide access to one or more hardware resources accessible through the network interface device to at least one host system.
[0109] Example 17 includes one or more examples and involves monitoring the operation of a host system to determine whether one or more processors of the host system executing a specific process are operational, and, based on a determination that one or more processors of the host system are not operational, accessing the memory of the host system to perform a failover operation to cause the specific process to run on one or more processors of the network interface device and to copy state data associated with the specific process from the memory of the host system to the memory of the network interface device.
[0110] Example 18 includes one or more examples and involves selecting an executable binary file associated with the specific process for execution on the one or more processors of the network interface device.
[0111] Example 19 includes one or more examples and includes, after causing the specific process to be executed on the one or more processors and copying state data of the specific process from the memory of the host system to the memory of the network interface device, based on the detection of the operation of at least one of the one or more processors of the host system, causing the execution of the specific process on the at least one of the one or more processors of the host system.
[0112] Example 20 includes one or more examples and involves accessing, via a secure connection, the operational state of the at least one host system and the state data from one or more of the following: memory, registers or cache of the at least one host system. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 18 / 370,621
[0001] IN 202341046012
[0001]
Claims
[1] Institution, comprising the following: a network interface device, comprising the following: a network interface a direct memory access (DMA) circuit arrangement, a host interface, a storage one or more processors and a circuit arrangement designed to do the following: based on an operating configuration that specifies standalone operation, causing the network interface device to operate in standalone mode to run one or more applications, and Based on an operating configuration that specifies companion operation, cause the network interface device to operate in companion mode to provide at least one host system with access to one or more hardware resources that the network interface device can access. [2] Device according to claim 1, comprising a second circuit arrangement for performing video analysis operations and wherein the second circuit arrangement is activated or deactivated based on a request to activate or deactivate the second circuit arrangement. [3] Device according to claim 1, wherein the circuit arrangement is designed as follows: Monitoring the operation of a host system to determine whether one or more of the host system's processors running a specific process are operational, and Based on a determination that one or more processors of the host system are not operational, accessing the memory of the host system to perform a failover operation to cause the specific process to run on one or more processors of the network interface device and to copy state data associated with the specific process from the memory of the host system and / or from non-operational processors into the memory of the network interface device. [4] Device according to claim 3, wherein the circuit arrangement is designed to perform a binary translation of the specific process or the specific processor emulation for execution on the one or more processors of the network interface device. [5] Device according to claim 3, wherein the circuit arrangement is designed as follows: after causing the specific process to be executed on the one or more processors and copying state data of the specific process from the memory of the host system to the memory of the network interface device, based on the detection of the operation of at least one of the one or more processors of the host system, causing the execution of the specific process on the at least one of the one or more processors of the host system. [6] Device according to claim 3, wherein the circuit arrangement is designed as follows: after causing the specific process to be executed on one or more processors and copying state data of the specific process from the memory of the host system to the memory of the network interface device, based on the detection of the operation of at least one processor of a second network interface device, causing the execution of the specific process on the at least one processor of the second network interface device. [7] Device according to claim 3, wherein the specific process comprises one or more of the following: a virtual machine, a container or a microservice. [8] Device according to claim 3, wherein the network interface device comprises a circuit arrangement for selecting a binary file for the specific process that is compatible with the one or more processors of the network interface device. [9] Device according to claim 3, wherein the network interface device is designed to: access, via a secure connection, the operating state of the at least one host system and the state data from one or more of the following: memory, registers or cache of the at least one host system. [10] Non-transitory computer-readable medium comprising instructions stored on it which, when executed by one or more processors, cause the one or more processors to: Configuring a network interface device to operate as a standalone computing platform to run one or more applications, or as a companion device to provide at least one host system with access to one or more hardware resources that the network interface device can access, wherein The network interface device is designed to: detect whether a host interface is present, and, based on a lack of access to the host interface, operate as the standalone computing platform, and The network interface device comprises the following: a network interface, a direct storage access (DMA) circuit arrangement, and a host interface. [11] Computer-readable medium according to claim 10, wherein the network interface device comprises a circuit arrangement for performing video analysis operations and wherein the circuit arrangement is activated or deactivated based on a request to activate or deactivate the circuit arrangement. [12] Computer-readable medium according to claim 10, comprising instructions stored thereon which, when executed by one or more processors, cause the one or more processors to do the following: Configuring the network interface device to: Monitoring the operation of a host system to determine whether one or more of the host system's processors running a specific process are operational, and Based on a determination that one or more processors of the host system are not operational, accessing the memory of the host system to perform a failover operation to cause the specific process to run on one or more processors of the network interface device and to copy state data associated with the specific process from the memory of the host system to the memory of the network interface device. [13] Computer-readable medium according to claim 12, comprising instructions stored thereon which, when executed by one or more processors, cause the one or more processors to do the following: Selecting an executable binary file associated with the specific process for execution on one or more processors of the network interface device. [14] Computer-readable medium according to claim 12, comprising instructions stored thereon which, when executed by one or more processors, cause the one or more processors to do the following: after causing the specific process to be executed on the one or more processors and copying state data of the specific process from the memory of the host system to the memory of the network interface device, based on the detection of the operation of at least one of the one or more processors of the host system, causing the execution of the specific process on the at least one of the one or more processors of the host system. [15] Computer-readable medium according to claim 12, comprising instructions stored thereon which, when executed by one or more processors, cause the one or more processors to do the following: after causing the specific process to be executed on one or more processors and copying state data of the specific process from the memory of the host system to the memory of the network interface device, based on the detection of the operation of at least one processor of a second network interface device, causing the execution of the specific process on the at least one processor of the second network interface device. [16] Procedures, comprising the following: an orchestrator who performs the following: Causing the execution of an application on one or more from a server or network interface device, wherein the network interface device comprises a network interface, a direct storage access (DMA) circuit arrangement and a host interface, and wherein the network interface device operates in a standalone operation to execute one or more applications, or operates as a companion to provide at least one host system with access to one or more hardware resources accessible through the network interface device. [17] The method of claim 16, comprising the following: Monitoring the operation of a host system to determine whether one or more of the host system's processors running a specific process are operational, and Based on a determination that one or more processors of the host system are not operational, accessing the memory of the host system to perform a failover operation to cause the specific process to run on one or more processors of the network interface device and to copy state data associated with the specific process from the memory of the host system to the memory of the network interface device. [18] The method of claim 17, comprising the following: Selecting an executable binary file associated with the specific process for execution on one or more processors of the network interface device. [19] The method of claim 17, comprising the following: after causing the specific process to be executed on the one or more processors and copying state data of the specific process from the memory of the host system to the memory of the network interface device, based on the detection of the operation of at least one of the one or more processors of the host system, causing the execution of the specific process on the at least one of the one or more processors of the host system. [20] The method of claim 17, comprising the following: Access, via a secure connection, to the operational state of the at least one host system and the state data from one or more of the following: memory, registers or cache of the at least one host system.