AI with high availability via a programmable network interface device

Programmable network interface devices manage replicas and facilitate load balancing to ensure high availability of vector databases, addressing node unavailability issues in AI inference pipelines.

DE102025146931A1Pending Publication Date: 2026-06-18INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
INTEL CORP
Filing Date
2025-11-13
Publication Date
2026-06-18

AI Technical Summary

Technical Problem

Maintaining the availability of vector database nodes in AI inference pipelines is challenging due to node unavailability, which disrupts access to data required for RAG setups.

Method used

Utilizing programmable network interface devices such as IPUs, DPUs, EPUs, and smart NICs to manage replicas, provide a unified front end, track heartbeats, and facilitate load balancing and recovery to ensure high availability of sharded vector databases.

Benefits of technology

Enhances the high availability of vector databases by dynamically replicating the state of programmable network interface devices, ensuring immediate activation in case of planned or unplanned shutdowns, thereby maintaining data accessibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The techniques described herein address the challenges encountered when using host-executed software to manage vector databases by providing vector database accelerator and shard management offload logic implemented within hardware and software running on device processors and programmable data layers of a programmable network interface device. In one embodiment, a programmable network interface device includes an infrastructure management circuitry designed to enable data access for a neural network inference engine with a distributed data model through dynamic management of a node associated with the neural network inference engine, where the node contains a database shard of a vector database.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE REVELATION

[0001] Infrastructure use for artificial intelligence (AI) has expanded from model provisioning and inference to AI data management, organization, and retrieval. For example, vector databases are used to store and organize AI embeddings in memory for similarity search queries. Vector databases are typically designed to keep all data in memory because large language models (LLMs) require real-time similarity searches against embeddings in RAG (Retrieval Augmented Generation) setups. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] The embodiments described here are illustrated by way of example and without limitation in the figures of the accompanying drawings, in which the same reference numerals indicate similar elements, and in which the following applies: Fig. Figure 1 is a block diagram illustrating a computer system configured to implement one or more aspects of the embodiments described herein; Fig. Figure 2 is a block diagram of a system containing selected components of a data center; Fig. Figure 3 is a block diagram of a part of a data center according to one or more examples in the present patent specification; Fig. 4A-4C illustrate programmable forwarding elements and adaptive routing; Fig. Figures 5A-5B show an example of network interface devices; Fig. Figure 6 is a block diagram illustrating a programmable network interface and a data processing unit; Fig. Figure 7 is a block diagram illustrating an IP core development system; Fig. Figures 8A-8B illustrate a system in which a vector database is distributed across multiple shards for load balancing and high availability; Fig. Figures 9A-9B illustrate systems in which programmable network interface devices provide vector database acceleration and network monitoring, according to embodiments; Fig. Figure 10 illustrates a programmable network interface device that can be configured as a vector database accelerator; Fig. Figure 11 illustrates a system designed to enable flow replication between programmable network interface devices according to one embodiment; and Fig. Figure 12 illustrates a method for enabling AI applications with high availability via programmable network interface devices according to embodiments described herein. DETAILED DESCRIPTION

[0003] RAG is a technique in AI inference that combines the strengths of information retrieval and language generation models to improve answer accuracy and relevance. RAG enhances AI model inference capability by enabling the retrieval of relevant documents or knowledge sources and then using this content to generate inferential results. The data to be referenced is transformed into AI model embeds, which are numerical representations in the form of large vectors. These embeds are then stored in a vector database to facilitate document retrieval. The vector database can be searched via a query to retrieve the matching database record that most closely matches the query. The vector database can be divided into shards and distributed across multiple database nodes.A vector database is a data storage and retrieval system used to manage vector data associated with AI model embeddings. The vectors are mathematical representations of data in a high-dimensional space, where the dimensions correspond to a feature of the data, and the feature is an individual, measurable property or characteristic of the data.

[0004] One of the emerging challenges in vector database implementations is maintaining the availability of the vector database nodes containing information used by the inference pipeline of the associated AI model. If a node is unavailable, the model lacks access to the data used to operate the RAG setup. This paper describes a mechanism to maintain high availability for sharded vector databases and / or virtual machines via programmable network interface devices, such as infrastructure processing units (IPUs), data processing units (DPUs), embedded or edge processing units (EPUs), and / or smart network interface controllers (smart NICs).

[0005] Programmable network interface devices are clearly positioned as a distinct fault domain with tight access to the network. Such devices can be designed to manage replicas, provide a unified front end, track heartbeats, provide load balancing, mitigate node failures, and manage recovery and migration as needed. In one embodiment, high availability is enhanced by enabling the state of a programmable network interface device to be dynamically and / or in real time replicated to a standby device designed to become active immediately in the event of planned or unplanned shutdown of the active device.

[0006] Numerous specific details are set forth in the following description to provide a more comprehensive understanding. However, those skilled in the art will recognize that the embodiments described herein can be implemented in practice without one or more of these specific details. In other cases, well-known features have been omitted to avoid obscuring the details of the present embodiments.

[0007] Fig. Figure 1 is a block diagram illustrating a computing system 100 configured to implement one or more aspects of the embodiments described herein. The computing system 100 includes a processing subsystem 101 with one or more processors 102 and a system memory 104, which communicate via a link path that may include a memory hub 105. The memory hub 105 may be a separate component within a chipset component or may be integrated into the one or more processors 102. The memory hub 105 is coupled to an I / O subsystem 111 via a communication link 106. The I / O subsystem 111 includes an I / O hub 107, which may allow the computing system 100 to receive input from one or more input devices 108.Furthermore, the I / O hub 107 can enable a display controller, which may be contained in one or more processors 102, to provide outputs to one or more display devices 110A. In one embodiment, the one or more display devices 110A coupled to the I / O hub 107 may include a local, internal, or embedded display device.

[0008] The processing subsystem 101 includes, for example, one or more parallel processors 112, which are coupled to the memory hub 105 via a communication link 113, such as a bus or a fabric. The communication link 113 can be one of any number of standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or it can be a vendor-specific communication interface or a communication fabric. The one or more parallel processors 112 can form a computationally focused parallel or vector processing system, which may include a large number of processing cores and / or processing clusters, such as a MIC processor (MIC: Many Integrated Core).For example, the one or more parallel processors 112 form a graphics processing subsystem that can output pixels to one or more display devices 110A coupled via the I / O hub 107. The one or more parallel processors 112 can also include a display controller and a display interface (not shown) to enable a direct connection to one or more display devices 110B.

[0009] Within the I / O subsystem 111, a system storage unit 114 can connect to the I / O hub 107 to provide a storage mechanism for the computing system 100. An I / O switch 116 can be used to provide an interface mechanism to enable connections between the I / O hub 107 and other components, such as a network adapter 118 and / or a wireless network adapter 119, which may be integrated into the platform, and various other devices that can be added via one or more add-in devices 120. The one or more add-in devices 120 may, for example, also include one or more external graphics processing devices, graphics cards, and / or compute accelerators. The network adapter 118 may be an Ethernet adapter or another wired network adapter.The wireless network adapter 119 can include a Wi-Fi and / or Bluetooth and / or Near Field Communication (NFC) and / or other network device that has one or more wireless radios.

[0010] The Computing System 100 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video recording devices, and the like, which may also be connected to the I / O Hub 107. Communication paths connecting the various components in Fig. 1. Connecting them together can be implemented using any suitable protocols, such as PCI (Peripheral Component Interconnect)-based protocols (e.g., PCI Express) or any other bus or point-to-point communication interfaces and / or protocol(s), such as the NVLink high-speed interconnect, Compute Express Link™ (CXL™) (e.g., CXL.mem), Infinity Fabric (IF), Ethernet (IEEE 802.3), Remote Direct Memory Access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Quick UDP Internet Connections (QUIC), RDMA over Converged Ethernet (RoCE), Ultra Ethernet Transport (UET), Intel QuickPath Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), OmniPath, HyperTransport, Advanced Microcontroller Bus Architecture (AMBA) interconnect, OpenCAPI, Gen-Z, Cache Coherent Interconnect for Accelerators (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G and variations thereof, or wired or wireless interconnect protocols known in the art. In some examples, data can be copied or stored in virtualized storage nodes using a protocol such as Non-Volatile Memory Express (NVMe)-over-Fabrics (NVMe-oF) or NVMe.In one embodiment, time-aware communication protocols are supported, including time-aware RDMA, time-aware NVME, and time-aware NVME-oF, where an accurate time and rate of data consumption is used to control the transfer of data.

[0011] The one or more parallel processors 112 can comprise a circuit arrangement optimized for graphics and video processing, including, for example, a video output circuit arrangement, and form a graphics processing unit (GPU). Alternatively or additionally, the one or more parallel processors 112 can comprise a circuit arrangement optimized for general-purpose processing, while retaining the underlying computing architecture, which is described in more detail herein. Components of the computing system 100 can be integrated with one or more other system elements on a single integrated circuit. For example, the one or more parallel processors 112, the memory hub 105, the one or more processors 102, and the I / O hub 107 can be integrated in an integrated system-on-a-chip (SoC) circuit.Alternatively, the components of the computer system 100 can be integrated into a single package to form a system-in-package (SIP) configuration. In one embodiment, at least some of the components of the computer system 100 can be integrated into a multi-chip module (MCM), which can be interconnected with other multi-chip modules to form a modular computer system.

[0012] In some configurations, the Computing System 100 includes, in addition to the one or more processors 102 and the one or more parallel processors 112, one or more accelerator devices 130 coupled to the memory hub 105. The one or more accelerator devices 130 are designed to perform domain-specific acceleration of workloads to handle computationally intensive or high-throughput tasks. The one or more accelerator devices 130 can reduce the load placed on the one or more processors 102 and / or the one or more parallel processors 112 of the Computing System 100.The one or more accelerator devices 130 may, but are not limited to, include intelligent network interface cards, data processing units, cryptographic accelerators, storage accelerators, artificial intelligence (AI) accelerators, neural processing units (NPUs), storage accelerators and / or video transcoding accelerators.

[0013] It is understood that the computing system 100 shown herein is for illustrative purposes only and that variations and modifications are possible. The connection topology, including the number and arrangement of bridges, the number of processors 102, and the number of parallel processors 112, can be modified as desired. For example, the system memory 104 can be directly connected to the one or more processors 102 instead of via a bridge, while other devices communicate with the system memory 104 via the memory hub 105 and the one or more processors 102. In other alternative topologies, the one or more parallel processors 112 are connected to the I / O hub 107 or directly to one of the one or more processors 102 instead of to the memory hub 105. In other embodiments, the I / O hub 107 and the memory hub 105 can be integrated on a single chip.It is also possible that two or more sets of processor(s) 102 are connected via multiple sockets that can couple with two or more instances of one of the multiple parallel processors 112.

[0014] Some of the special components shown herein are optional and need not be included in all implementations of the Computing System 100. For example, any number of add-in cards or peripherals may be supported, or some components may be omitted. Furthermore, some architectures may use different terminology for components, similar to that used in Fig. 1 are illustrated.

[0015] Fig. Figure 2 is a block diagram of a System 200 containing selected data center components. The components of the depicted data center could be located at a cloud service provider (CSP) or another data center, such as a traditional enterprise data center, a private cloud, or a public cloud providing services like Infrastructure as a Service (IaaS), Platform as a Service (PaaS), or Software as a Service (SaaS). The System 200 includes a number of workload clusters, including but not limited to Workload Cluster 218A and Workload Cluster 218B. The workload clusters can consist of single servers, blade servers, rackmount servers, or any other suitable server topology.

[0016] The System 200 can contain workload clusters 218A-218B. Workload clusters 218A-218B can occupy a Rack 248, housing multiple servers (e.g., Server 246). The Rack 248 and the servers of the Workload Clusters 218A-218B can conform to the rack unit standard (“U” standard), where a rack unit corresponds to a 19-inch wide rack frame and a full-size industry-standard rack can accommodate 42 units (42U) of equipment. A device unit (1U) (e.g., a 1U server) can be 1.75 inches high and approximately 36 inches deep. In various configurations, compute resources such as processors, memory, storage, accelerators, and switches can fit into multiples of rack units within a Rack 248.

[0017] A Server 246 can host a standalone operating system configured to provide server functions, or the servers can be virtualized. A virtualized server can be controlled by a Virtual Machine Manager (VMM), hypervisor, and / or orchestrator and host one or more virtual machines, virtual servers, or virtual devices. Workload clusters 218A-218B can reside in a single data center or in different geographic data centers. Depending on contractual agreements, some servers may be dedicated to specific enterprise customers or tenants, while other servers may be shared.

[0018] The various devices in a data center can be interconnected via a switching fabric 270, which may include one or more high-speed routing and / or switching devices. The switching fabric 270 can provide north-south traffic 202 (e.g., traffic to and from the wide area network (WAN), such as the internet) and east-west traffic 204 (e.g., traffic throughout the data center). Historically, north-south traffic 202 constituted the majority of network traffic, but as web services have become more complex and distributed, the volume of east-west traffic 204 has increased. In many data centers, east-west traffic 204 now represents the largest share of traffic. Furthermore, the volume of traffic can increase further with the increasing performance of a server 246.For example, a Server 246 can offer multiple processor slots, with each slot accommodating a processor with four to eight cores and sufficient memory for those cores. This allows a server to host a number of virtual machines (VMs), which can generate traffic.

[0019] To handle the large traffic volume in a data center, a high-performance implementation of the Switching Fabric 270 can be deployed. The implementation of the Switching Fabric 270 shown is an example of a flat network where a Server 246 can have a direct connection to a Top-of-Rack Switch (ToR Switch 220A-220B) (e.g., a "star" configuration). The ToR Switch 220A can connect to a Workload Cluster 218A, while the ToR Switch 220B can connect to the Workload Cluster 218B. A ToR Switch 220A-220B can be coupled to a Core Switch 260. This two-tier flat network architecture is shown only as an illustrative example, and other architectures can also be used, such as...Three-tiered star or leaf-spine topologies (also called "fat tree" topologies) based on the "Clos" architecture, hub-and-spoke topologies, mesh topologies, ring topologies or 3D mesh topologies, to name just one example.

[0020] The Switching Fabric 270 can be deployed through any suitable interconnect using any suitable interconnect protocol. For example, a Server 246 can contain a Fabric interface (FI), a Network Interface Card (NIC), or another host interface. The host interface itself can be connected to one or more processors via an interconnect or bus such as PCI, PCIe, or similar, and in some cases, this interconnect bus can be considered part of the Switching Fabric 270. The Switching Fabric can also utilize physical PCIe interconnects to implement more advanced protocols such as Compute Express Link (CXL).

[0021] Interconnect technology can be provided through a single interconnect or a hybrid interconnect. For example, PCIe might handle on-chip communication, 1Gb or 10Gb copper Ethernet provides relatively short connections to a 220A-220B toroidal switch, and optical cabling enables relatively long connections to the 260 core switch. Interconnect technologies include Ultra Path Interconnect (UPI), Fibre Channel, Ethernet, Fibre Channel over Ethernet (FCoE), InfiniBand, PCIe, NVLink, and fiber optic cables, to name a few. Some are better suited for certain applications or functions than others, and selecting the appropriate fabric for a given application is a task for the average professional.

[0022] In one embodiment, the switching elements of the Switching Fabric 270 are configured to implement switching techniques to improve network performance in high-utilization scenarios. Examples of advanced switching techniques include adaptive routing, adaptive fault resolution, and adaptive and / or telemetry-based congestion control.

[0023] Adaptive Routing allows a ToR switch 220A-220B and / or a core switch 260 to select the outbound port to which traffic is directed based on the load of the selected port, provided unrestricted port selection is enabled. An adaptive routing table can configure the forwarding tables of the switches in the Switching Fabric 270 to select between multiple ports between switches when multiple connections exist between a given set of switches in an adaptive routing group. Adaptive fault correction (e.g., self-healing) allows the automatic selection of an alternative port if the port selected by the forwarding table is in a failed or inactive state, enabling rapid recovery in the event of a switch-to-switch port failure.A notification can be sent to neighboring switches when adaptive routing or adaptive troubleshooting becomes active on a particular switch. With adaptive congestion control, a switch is configured to send a notification to neighboring switches when the congestion of a port on that switch exceeds a configured threshold. This can then prompt those neighboring switches to adaptively switch to uncongested ports on that switch or to switches connected via an alternative route to the destination.

[0024] Telemetry-based congestion control utilizes dynamic and / or real-time monitoring of telemetry from network devices, such as switches within the Switching Fabric 270, to detect when congestion begins to impact the Switching Fabric 270's performance and to proactively adjust the switching tables within the network devices to prevent or mitigate the impending congestion. A ToR Switch 220A-220B and / or Core Switch 260 can implement an integrated telemetry-based congestion control algorithm or provide an application programming interface (API) through which a programmable telemetry-based congestion control algorithm can be implemented.A continuous feedback loop can be implemented in which the telemetry-based congestion control system continuously monitors the network and dynamically adjusts the traffic flow based on the ongoing telemetry data.

[0025] In a data center network, a network flow (e.g., flow) represents the directed, ordered sequence of packets transmitted from a source node to a destination node, typically through a series of network devices such as routers, switches, and load balancers. A flow can be identified by a unique 5-tuple (source IP, source port, destination IP, destination port, and protocol). Flows can be managed using network flow tables, which optimize traffic paths and ensure quality of service (QoS). The telemetry-based congestion control system can learn and adapt by adjusting to changing network conditions and improving its congestion control strategies based on historical data and trends.

[0026] It should be noted, however, that while high-end fabrics are provided here for illustrative purposes, Switching Fabric 270 can generally encompass any suitable interconnect or bus for the specific application, including legacy interconnects used to implement local area networks (LANs), synchronous optical networks (SONETs), asynchronous transfer-mode (ATM) networks, wireless networks such as Wi-Fi and Bluetooth, 4G wireless, 5G wireless, digital subscriber line connections (DSL), Multimedia Over Coax Alliance (MoCA) interconnects, or similar wired or wireless networks. It is also expressly assumed that new network technologies will emerge in the future that complement or replace some of those listed here, and that such future network topologies and technologies may be, or become, part of Switching Fabric 270.

[0027] Fig. Figure 3 is a block diagram of a part of a data center 300 according to one or more examples in the present patent specification. The depicted part of the data center 300 is not intended to include all components of a data center. The depicted part can be duplicated multiple times within the data center 300 and / or the data center 300 can include parts beyond those depicted, depending on the capacity and functionality to be provided by the data center 300. The data center 300 can, in various embodiments, include components of the data center of the system 200. Fig. 2 or another data center.

[0028] The Data Center 300 comprises a set of logic elements that form a multitude of nodes, where a node can be a physical server, a group of servers, or other hardware. A server can also host one or more virtual machines, depending on the application. A Fabric 370 is provided to connect various aspects of the Data Center 300. The Fabric 370 can be provided using any suitable interconnect technology, including but not limited to InfiniBand, Ethernet, PCIe, or CXL. The Data Center 300's Fabric 370 can be a version of the System 200's Switching Fabric 270. Fig. 2 and / or include elements thereof. The Fabric 370 of the Data Center 300 can connect elements of the data center, including server nodes (e.g., storage server node 304, heterogeneous compute server node 306, CPU server node 308, storage server node 310), accelerators 330, gateways 340A-340B to other fabrics, fabric architectures or interconnect technologies, and an Orchestrator 360.

[0029] The server nodes of data center 300 can include, but are not limited to, a storage server node 304, a heterogeneous compute server node 306, a CPU server node 308, and a storage server node 310. The heterogeneous compute server node 306 and a CPU server node 308 can perform independent operations for different tenants or cooperative operations for a single tenant. The heterogeneous compute server node 306 and a CPU server node 308 can also host virtual machines that provide virtual server functionality to the data center's tenants.

[0030] The server nodes can be connected to the Fabric 370 via a Fabric 372 interface. The specific type of Fabric 372 interface used depends, at least in part, on the technology or protocol used to implement the Fabric 370. For example, if the Fabric 370 is an Ethernet fabric, the Fabric 372 interface can be an Ethernet network interface controller. If the Fabric 370 is a PCIe-based fabric, the Fabric interfaces can be PCIe-based interconnects. If Fabric 370 is an InfiniBand fabric, the Fabric interface 372 of the heterogeneous compute server node 306 and a CPU server node 308 can be a Host Channel Adapter (HCA), while the Fabric interface 372 of the storage server node 304 and the storage server node 310 can be a Target Channel Adapter (TCA).The TCA functionality can be an implementation-specific subset of the HCA functionality. The various fabric interfaces can be implemented as IP (Intellectual Property) blocks that can be inserted as a modular unit into an integrated circuit, just like other circuits within the Data Center 300.

[0031] The heterogeneous compute server node 306 includes multiple CPU sockets that can accommodate a CPU 319, which can be, but is not limited to, a multi-core Intel® Xeon™ processor. For example, the CPU 319 could also be a data center-class ARM® multi-core CPU, such as an NVIDIA® Grace™ CPU. The heterogeneous compute server node 306 includes storage devices 318 for storing runtime data and storage devices 316 for persistent data storage on non-volatile storage devices. The heterogeneous compute server node 306 can perform heterogeneous processing through the presence of GPUs (e.g., GPU 317), which can be used for high-performance computing (HPC), media servers, cloud gaming servers, and / or machine learning compute.In one configuration, the GPUs and CPUs of the heterogeneous compute server node 306 can be interconnected via interconnect technologies such as PCIe, CXL or NVLink.

[0032] The CPU server node 308 contains a variety of CPUs (e.g., CPU 319), memory (e.g., memory devices 318), and storage (storage devices 316) to run applications and other program code that provide server functions, such as web servers or other types of functions that clients of the CPU server node 308 can access remotely. The CPU server node 308 can also run program code that provides services or microservices enabling complex enterprise functions.Fabric 370 is provided with sufficient throughput to allow concurrent access to CPU server node 308 by a large number of clients, while simultaneously maintaining sufficient throughput for use by heterogeneous compute server node 306 and enabling the use of storage server node 304 and storage server node 310 by heterogeneous compute server node 306 and CPU server node 308. Furthermore, in a configuration where CPU server node 308 primarily relies on distributed services provided by storage server node 304 and storage server node 310, the memory and storage capacity of CPU server node 308 may not be sufficient for all operations it may require.Instead, a large pool of high-speed or specialized storage can be dynamically provisioned among a number of nodes, giving nodes access to a large resource pool without those resources remaining unused when not needed. This type of distributed architecture is made possible by the high speeds and low latency offered by the Fabric 370 architecture of modern data centers and can be advantageous because it avoids the need for excessive resource provisioning to server nodes.

[0033] Storage server node 304 can include storage node 305 with memory technologies suitable for storing data used during the execution of program code by heterogeneous compute server node 306 and CPU server node 308. Storage node 305 can include volatile memory modules, such as DRAM modules, and / or non-volatile memory technologies capable of operating at speeds similar to DRAM, providing sufficient throughput and latency performance metrics to be used as a layer of system memory at runtime. Storage server node 304 can be connected to heterogeneous compute server node 306 and / or CPU server node 308 via technologies such as CXL.mem, which enable memory access from a host to a device.In such a configuration, CPU 319 of heterogeneous compute server node 306, CPU server node 308, can connect to storage server node 304 and access storage node 305 of storage server node 304 in a similar way to how CPU 319 of heterogeneous compute server node 306 can access the device memory of a GPU within heterogeneous compute server node 306. For example, storage server node 304 can grant remote DMA (Relative Direct Access) access to storage node 305, allowing CPU server node 308, for instance, to access the memory resources on storage server node 304 via Fabric 370 using DMA operations, much like the CPU would access its own onboard memory.

[0034] Storage server node 304 can be used by heterogeneous compute server node 306 and CPU server node 308 to extend the runtime memory available for memory-intensive activities such as training machine learning models. A multi-tiered storage system can be enabled, allowing model data to be swapped in and out of storage devices 318 of heterogeneous compute server node 306 to the memory of storage server node 304, which offers higher performance and / or lower latency than local storage (e.g., storage devices 316). During workload setup, the entire workload dataset can be loaded into one or more of the storage nodes 305 of storage server node 304 and the storage devices 318 of heterogeneous compute server node 306, as needed during the execution of the heterogeneous workload.

[0035] Storage server node 310 provides storage functionality for heterogeneous compute server node 306, CPU server node 308, and potentially storage server node 304. Storage server node 310 can provide a networked bundle of disks (JBOD), program flash memory (PFM), a redundant array of independent disks (RAID), a redundant array of independent nodes (RAIN), network-attached storage (NAS), or other non-volatile storage solutions. In one configuration, storage server node 310 can be coupled with heterogeneous compute server node 306, CPU server node 308, and / or storage server node 304, such as via NVMe-oF, enabling the implementation of the NVMe protocol over Fabric 370. In such configurations, the Fabric interfaces 372 of these servers can be intelligent interfaces containing hardware to accelerate NVMe-oF operations.

[0036] The Accelerators 330 within Data Center 300 can provide various accelerated functions, including hardware or coprocessor acceleration for functions such as packet processing, encryption, decryption, compression, decompression, network security, or other accelerated functions within the data center. In some examples, the Accelerators 330 may include deep learning accelerators, such as neural processing units (NPUs), which can offload matrix multiplication operations or other neural network operations from the heterogeneous compute server node 306 or the CPU server node 308. In some configurations, the Accelerators 330 may reside in a dedicated accelerator server or be distributed across the various server nodes of Data Center 300.For example, an NPU can be directly connected to one or more CPU cores within the heterogeneous compute server node 306 or the CPU server node 308. In some configurations, the Accelerators 330 can include or be contained within intelligent network controllers, infrastructure processing units (IPUs), or data processing units that combine network control functions with accelerator, processor, or coprocessor functions. The Accelerators 330 can also include edge processing units (EPUs) to perform real-time inference operations at the network edge.

[0037] In a configuration, the Data Center 300 can include gateways 340A-340B connecting Fabric 370 to other fabrics, fabric architectures, or interconnect technologies. For example, if Fabric 370 is an InfiniBand fabric, gateways 340A-340B can be gateways to an Ethernet fabric. If Fabric 370 is an Ethernet fabric, gateways 340A-340B can contain routers to route data to other parts of Data Center 300 or to a larger network, such as the internet. For example, a first gateway 340A can connect to another network or subnet within Data Center 300, while a second gateway 340B can be a router to the internet.

[0038] Orchestrator 360 manages the provisioning, configuration, and operation of network resources within Data Center 300. Orchestrator 360 can consist of hardware or software running on a dedicated orchestration server. It can also be included in software running, for example, on CPU Server Node 308, which configures the Software Defined Networking (SDN) functionality of components within Data Center 300. In various configurations, Orchestrator 360 can enable the automated provisioning and configuration of Data Center 300 components by managing network resource allocation and template-based deployment.Template-based deployment is a method for provisioning and managing IT resources using predefined templates. These templates can be based on standards required by government agencies, service providers, financial institutions, standards bodies, or customers. Service Level Agreements (SLAs) or Service Level Obligations (SLOs) can also be defined within the template. Orchestrator 360 can perform, but is not limited to, functions such as load balancing and traffic engineering, network segmentation, security automation, real-time telemetry monitoring, and adaptive switching management, including telemetry-based adaptive switching.In some configurations, Orchestrator 360 can also provide multi-tenant and virtualization support by enabling the management of virtual networks, including the creation and deletion of virtual LANs (VLANs) and virtual private networks (VPNs), and tenant isolation for multi-tenant data centers.

[0039] Fig. 4A-4C illustrates programmable forwarding elements and adaptive routing. Fig. 4A illustrates a forwarding element that includes a control plane and a programmable data plane. Fig. Figure 4B illustrates a network with switching devices configured for adaptive routing and telemetry-based congestion control. Fig. Figure 4C illustrates an InfiniBand switch with multiport IB interfaces.

[0040] Fig. Figure 4A shows a forwarding element 400 that can be configured to forward data messages within a network based on a user-provided program. In some embodiments, the program includes instructions for forwarding data messages as well as for performing other processes such as firewalling, denial-of-service protection, and load balancing operations. The forwarding element 400 can be any type of forwarding device, including but not limited to a switch, router, or bridge. The forwarding element 400 can forward data messages associated with various technologies, such as but not limited to Ethernet, Ultra Ethernet, InfiniBand, or NVLink.

[0041] In various network configurations, the forwarding element is used as a non-edge forwarding element within the network to forward data messages from a source device to a destination device. In network configurations, the forwarding element 400 is used as an edge forwarding element at the network edge to connect to computing devices (e.g., standalone or host computers) that serve as sources and destinations of the data messages. As a non-edge forwarding element, the forwarding element 400 forwards data messages between forwarding elements within the network, for example, via an intermediary network fabric. As an edge forwarding element, the forwarding element 400 forwards data messages to and from edge computing devices, to other edge forwarding elements, and / or to non-edge forwarding elements.

[0042] The forwarding element 400 contains circuit arrangements for implementing a data plane 402, which performs the forwarding operations of the forwarding element 400 to forward data messages received by the forwarding element to other devices. The forwarding element 400 also includes circuit arrangements for implementing a control plane 404, which configures the data plane circuitry. Furthermore, the forwarding element 400 includes physical ports 406, which receive data messages from and send data messages to devices outside the forwarding element 400. The data plane 402 includes ports 408, which receive data messages from the physical ports 406 for processing. The data messages are processed and forwarded to another port on the data plane 402, which is connected to another physical port of the forwarding element 400.In addition to the physical ports of the forwarding element 400, some of the ports 408 on data plane 402 may also be associated with other modules of data plane 402.

[0043] The data plane contains programmable packet processor circuits that provide several programmable message processing stages. These stages can be configured to perform the forwarding operations of the forwarding element 400 data plane to process and forward data messages to their destinations. These message processing stages perform these forwarding operations by processing data tuples (such as message headers) associated with the data messages received by the data plane 402 to determine how the messages should be forwarded. The message processing stages include match action units (MAUs) that attempt to match data tuples (such as header vectors) from messages with tables that specify which actions should be performed on the data tuples.In some embodiments, the table sets are populated by the control plane 404 and are not known when the data plane is configured to execute a program provided by a network user. The programmable message-handling circuits are grouped into multiple message-handling pipelines. The message-handling pipelines can be input or output pipelines before or after the traffic management stage of the routing element, which directs the messages from the input pipelines to the output pipelines.

[0044] The hardware specifications of data plane 402 depend on the communication protocol implemented via the forwarding element 400. Ethernet switches use application-specific integrated circuits (ASICs) designed to process Ethernet frames and the TCP / IP protocol stack. These ASICs are optimized for a wide range of traffic types, including unicast, multicast, and broadcast. Ethernet switch ASICs are typically designed for a balance of cost, power consumption, and performance, although high-end Ethernet switches may also support advanced features such as deep packet inspection and enhanced quality of service (QoS). InfiniBand switches use specialized ASICs designed for extremely low latency and high throughput.These ASICs enable features such as optimized processing of the InfiniBand protocol and offer support for RDMA and other features that require precise timing and high-speed data processing, although high-end Ethernet switches can support RoCE (RDMA over Converged Ethernet), which offers similar advantages to InfiniBand but has higher latency compared to native InfiniBand RDMA.

[0045] The 400 forwarding element can also be configured as an NVLink switch (e.g., NVSwitch), which connects multiple GPUs using the NVLink connection protocol. When configured as an NVLink switch, the 400 forwarding element can provide GPU servers with higher GPU-to-GPU bandwidth compared to GPU servers connected via InfiniBand. An NVLink switch can reduce network traffic hotspots that can occur when interconnected GPU-equipped servers perform tasks such as distributed neural network training.

[0046] When data plane 402, in coordination with a program running on data plane 402 (e.g., a program written in the P4 language), performs message or packet forwarding operations on incoming data, control plane 404 generally determines how messages or packets are forwarded. The behavior of a program running on data plane 402 is partly determined by control plane 404, which populates match-action tables with specific forwarding rules. The forwarding rules used by the program running on data plane 402 are independent of the data plane program itself. In a configuration, the control plane can be coupled to an administration port 410, which allows the administrator to configure the forwarding element 400.The data connection established via the management port 410 is separate from the data connections for the input and output data ports. In one configuration, the management ports 410 can be connected to a management level 405, which facilitates administrative access to the device, enables analysis of the device's condition and health, and allows for device reconfiguration. The management level 405 can be part of the control level 404 or communicate directly with it. In one implementation, the administrator does not have direct access to the components of the control level 404. Instead, information is collected by the management level 405, and changes to the control level 404 are made by the management level 405.

[0047] Fig. Figure 4B shows a Network 420 with Switches 432A-432E supporting adaptive routing and telemetry-based congestion control. The Network 420 can be implemented using a variety of the communication protocols described here. In one embodiment, the Network 420 is implemented using the InfiniBand protocol. In another embodiment, the Network 420 is an Ethernet, Converged Ethernet, or Ultra-Ethernet network. The Network 420 can incorporate aspects of Fabric 370. Fig. 3 included. Switches 432A-432E may be an implementation of the forwarding element 400 of Fig. Network 420 handles packet-based communication for multiple nodes (e.g., node 424, node 446), including a source node 422 and a destination node 442, of a data transfer to be carried out over network 420. The packets of a flow are routed through network 420 along a route that traverses the switches (switches 432A-432E) and links (links 426A-426B, 427A-427B, 428, 429A-429B, 430A-430B) of network 420. In an InfiniBand application, the switches and links belong to a specific InfiniBand subnet, which is managed by a subnet manager (SM) that may be located in one of the switches (e.g., switch 432D). Source node 422 and destination node 442 are the source and destination nodes for an example data flow. Depending on the configuration of network 420, packets can flow from any node to any other node via one or more paths.

[0048] The 432A-432E switches include a 402 data plane, a 404 control plane, a 405 management plane, and 406 physical ports, as described in the forwarding element 400 of Fig. 4A. A 404 control plane processor can be used to implement adaptive routing techniques to adjust a route between source node 422 and destination node 442 based on the current state of the network. During network operation, the route from source node 422 to destination node 442 may become unsuitable or impaired in its ability to transmit packets due to various events, such as congestion, link failures, or head-of-line blocking. Should such a scenario occur, switches 432A-432E can be configured to dynamically adjust the route of packets flowing over a compromised path.

[0049] An adaptive routing (AR) event can be detected by any switch along a compromised route, for example, when the switch attempts to output packets to a specific output port. For instance, data from source node 422 to destination node 442 might traverse links across network switches. An AR event might be detected by switch 432D for link 429B, for example, in response to congestion or a link failure associated with link 429B. Upon detection, switch 432D, as the detecting switch, generates an adaptive routing notification (ARN) with an identifier that distinguishes an ARN packet from other packet types. In various implementations, the ARN includes parameters such as an identifier for the detecting switch, the type of AR event, the source and destination addresses of the flow that triggered the AR event, and / or other appropriate parameters.The detecting switch sends the ARN backward along the route to the preceding switches. The ARN can contain a request to the notified switches to change their route to avoid passing over the detected switch. A notified switch can then check if its routes can be modified to bypass the detected switch. Otherwise, the switch forwards the ARN to the preceding switch along the route. In this scenario, switch 432B is unable to bypass switch 432D and forwards the ARN to switch 432A. Switch 432A can determine that the route to destination node 442 should be adjusted via link 427A to switch 432C. Switch 432C can reach Switch 432E via Link 429A, allowing packets from source node 422 to reach destination node 442, while bypassing the AR event with respect to Link 429B.

[0050] In various configurations, the 420 network can also adapt to congestion scenarios via programmable data planes within switches 432A-432E. These data planes are capable of executing data plane programs to implement network-internal congestion control algorithms (CCAs) for TCP over Ethernet-based fabrics. Using in-band network telemetry (INT), the programmable data planes in switches 432A-432E can detect when a port or link along a route is congested and proactively attempt to reroute packets via alternative paths. For example, switch 432A can balance traffic to destination node 442 between links 427A and 427B based on the degree of congestion on the routes to those links.

[0051] Fig. 4C shows an InfiniBand switch 450, which is an implementation of the forwarding element 400 of Fig. The InfiniBand Switch 450 can handle 4A. It includes a programmable data plane and is configurable to perform adaptive routing and telemetry-based congestion control, as described herein. The InfiniBand Switch 450 comprises the multiport IB interfaces 460A-460D and the core switch logic 480. The multiport IB interfaces 460A-460D include multiple ports. In one embodiment, a single physical interface (IB-PHY 453) with input and output buffers is associated with one port. In another embodiment, the ports have separate physical interfaces. The ports can be coupled, for example, to an HCA 452, a TCA 461, or another InfiniBand Switch 432. The multiport IB interfaces 460A-460D can include a crossbar switch 454 configured to selectively couple input and output port buffers to the local memory 456.The Crossbar Switch 454 is a non-blocking crossbar switch that enables direct switching with low latency and fixed or variable packet size.

[0052] Local memory 456 contains several queues, including an outer receive queue 462, an outer send queue 463, an inner receive queue 464, and an inner send queue 465. The outer queues are used for data received at a specific Multiport IB interface and intended to be sent back over the same Multiport IB interface. The inner queues are used for data that is forwarded over a different Multiport IB interface than the one through which the data was received. Other types of queue configurations can be implemented in local memory 456. For example, different queues can exist to support multiple traffic classes, either based on individual ports, shared ports, or a combination thereof.The 460A-460D multiport IB interfaces include a 455 power management circuit that can adjust the power state of the circuits within the respective multiport IB interface. Additionally, power management logic performing similar operations can be implemented as part of the core switch logic.

[0053] The multiport IB interfaces 460A-460D include the packet processing and switching logic 458, which is generally used to perform aspects of packet processing and / or switching operations that are executed locally across multiple ports, rather than via the IB switch as a whole. Depending on the implementation, the packet processing and switching logic 458 may be configured to perform a subset of the operations of the packet processing and switching logic 478 within the core switching logic 480, or it may be configured with the full functionality of the packet processing and switching logic 478 within the core switch logic 480. The processing functionality of the packet processing and switching logic 458 can vary depending on the complexity of the operations and / or the speed at which the operations are to be performed.The packet processing and switching logic of version 458 can, for example, include processors ranging from microcontrollers to multi-core processors. Different types or architectures of multi-core processors can also be used. Additionally, some of the packet processing operations can be implemented using embedded hardware logic.

[0054] The core switch logic 480 includes a crossbar 482, a memory 470, a subnet management agent (SMA 476), and packet processing and switching logic 478. The crossbar 482 is a non-blocking, low-latency crossbar that connects the multiport IB interfaces 460A-460D and establishes a connection to the memory 470. The memory 470 includes receive queues 472 and transmit queues 474. In one embodiment, packets to be routed between the multiport IB interfaces 460A-460D can be received by the crossbar 482, stored in one of the receive queues 472, processed by the packet processing and switching logic 478, and stored in a transmit queue 474 for transmission to the outgoing multiport IB interface.In implementations that do not use the multiport IB interfaces 460A-460D, the core switch logic 480 and the crossbar 482 directly transfer packets between I / O buffers with receive queues 472 and send queues 474 within the memory 470.

[0055] The Packet Processing and Switching Logic 478 incorporates programmable functions and can execute data-plane programs across a variety of multi-core processor types and architectures. The Packet Processing and Switching Logic 478 represents the applicable circuit layouts and logic for implementing switching and packet processing operations that may extend beyond those performed at the ports themselves. The processing elements of the Packet Processing and Switching Logic 478 execute software and / or firmware commands configured to implement packet processing and switching operations. This software and / or firmware can be stored in non-volatile memory on the switch itself. Alternatively, the software can be downloaded or updated over a network in conjunction with the initialization of the InfiniBand Switch 450.

[0056] The SMA 476 is configurable for managing, monitoring, and controlling the functions of the InfiniBand 450 switch. The SMA 476 also acts as an agent of the subnet manager (SM) for the subnet associated with the InfiniBand 450 switch and communicates with it. The SM is the instance that discovers devices within the subnet and performs periodic checks of the subnet to detect changes in the subnet topology. An SMA within a subnet can be designated as the primary SMA for that subnet and act as the SM. Other SMAs within the subnet then communicate with this SMA. Alternatively, the SMA 476 can cooperate with other SMAs in the subnet and function as a distributed SM. In some embodiments, the SMA 476 includes or runs on self-contained circuitry and logic, such as a microcontroller, a single-core processor, or a multi-core processor.In other embodiments, the SMA 476 is implemented via software and / or firmware instructions executed on a processor core or other processing element that is part of a processor or other processing element used to implement the packet processing and switching logic 478.

[0057] The embodiments are not specifically limited to implementations with multiport IB interfaces 460A-460D. In one embodiment, the ports are connected to their own receive and transmit buffers, with the crossbar 482 configured to connect these buffers to the receive queues 472 and the transmit queues 474 in memory 470. Packet processing and switching are then primarily performed by the packet processing and switching logic 478 of the core switch logic 480.

[0058] Fig. Figures 5A-5B show an example of network interface devices. Fig. 5A shows a network interface device 500 that can be configured as a Smart Ethernet device. Fig. Figure 5B shows a 550 network interface device that can be configured as an InfiniBand channel adapter.

[0059] As in Fig. As shown in Figure 5A, the network interface device 500 can, in one configuration, include a transmit receiver 502, a transmit queue 507, a receive queue 508, a memory 510, a bus interface 512, and a DMA engine 526. The network interface device 500 can also include a SoC / SIP 545, which contains processors 505 for implementing the intelligent network interface device functionality, as well as accelerators 506 for various accelerated functions such as NVMe-oF or RDMA. The specific characteristics of the network interface device 500 depend on the protocol implemented through the network interface device 500.

[0060] In various configurations, the Network Interface Device 500 can be configured to connect to networks, including but not limited to Ethernet, including Ultra Ethernet. However, the Network Interface Device 500 can also be configured as an InfiniBand or NVLink interface by modifying various components. For example, the Transceiver 502 can receive and transmit packets in accordance with the InfiniBand, Ethernet, or NVLink protocols. Other protocols can also be used. The Transceiver 502 can receive and transmit packets to and from a network over a network medium. The Transceiver 502 can include a PHY circuit assembly 514 and a Media Access Control (MAC) circuit assembly 516.The PHY circuit arrangement 514 can include an encoding and decoding circuit arrangement for encoding and decoding data packets according to valid physical layer specifications or standards. The MAC circuit arrangement 516 can be configured to assemble data to be transmitted into packets containing destination and source addresses along with network control information and error-detection hash values.

[0061] The SoC / SIP 545 can include processors that can be any combination of the following: CPU processor, core, graphics processing unit (GPU), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or other programmable hardware device that enables programming of the Network Interface Device 500. For example, a smart network interface can provide packet processing capabilities in the network interface using 505 processors. The configuration of the operation of the 505 processors, including the programmable data plane processors, can be programmed using Protocol-Independent Packet Processors (P4), C, Python, Broadcom Network Programming Language (NPL), x86- or ARM-compatible executable binaries, or other executable binaries.

[0062] A packet assigner 524 can distribute received packets for processing by multiple CPUs or cores using time slot allocation. An interrupt coalescing circuit 522 can perform interrupt moderation, where it waits for multiple packets to arrive or for a timeout to expire before generating an interrupt for the host system to process one or more received packets. Receive segment coalescing (RSC) can be performed by the network interface device 500, where portions of incoming packets are combined into segments of a packet. The network interface device 500 can then deliver this coalesced packet to an application.A DMA engine 526 can copy a packet header, packet payload, and / or a descriptor directly from host memory to the network interface, or vice versa, instead of copying the packet to an intermediate buffer on the host and then using a different copy operation from the intermediate buffer to the destination buffer. Memory 510 can be any type of volatile or non-volatile storage device and can store any queue or instructions used to program the network interface device 500. The send queue 507 can contain data or references to data for transmission through the network interface. The receive queue 508 can contain data or references to data received by the network interface from a network. The descriptor queues 520 can contain descriptors that reference data or packets in the send queue 507 or receive queue 508.The 512 bus interface can provide an interface with the host device. For example, the 512 bus interface can be PCI Express compatible, although other connection standards can also be used.

[0063] As in Fig. As shown in Figure 5B, a network interface device 550 can be configured as an implementation of the network interface device 500 to implement an InfiniBand HCA. The network interface device 550 includes network ports 552A-552B, memory 554A-554B, a PCIe interface 558, and an integrated circuit 556 containing hardware, firmware, and / or software for implementing, managing, and / or controlling the HCA functionality. In one implementation, the integrated circuit includes a hardware transport engine 560, an RDMA engine 562, congestion control logic 563, virtual endpoint logic 564, offload engines 566, QoS logic 568, GSA / SMA logic 569, and a management interface. Different implementations of the network interface device 550 may include additional components or omit some components.A network interface device 550 configured as a TCA contains an implementation-specific subset of the functionality of an HCA. The integrated circuit 556 contains programmable and fixed-function hardware for implementing the described functionality.

[0064] While the illustrated implementation of the Network Interface Device 550 uses a PCIe 558 interface, other implementations may use different interfaces. For example, the Network Interface Device 550 may use an Open Compute Project (OCP) mezzanine connector. Furthermore, the PCIe 558 interface can also be configured with a multi-host solution, allowing multiple compute or storage hosts to connect to the Network Interface Device 550. The PCIe 558 interface can also support a technology that enables direct PCIe access to multiple CPU sockets, eliminating the need for network traffic to traverse the inter-processor bus of a multi-socket server motherboard when the server includes the Network Interface Device 550.

[0065] The 550 network interface device implements endpoint elements of the InfiniBand architecture, which is based on queue pairs and RDMA. InfiniBand offloads traffic control from software by using execution queues (e.g., work queues) that are initiated by a software client and managed in hardware. Communication endpoints include a queue pair (QP) with a send queue and a receive queue. A QP is a memory-based abstraction where communication occurs between memory-to-memory transfers between applications or between applications and devices. Communication with the QPs takes place over virtual lanes of network ports 552A-552B, which allow multiple independent data flows to use the same link, with separate buffering and flow control for each flow.

[0066] Communication occurs via channel I / O, where a virtual channel directly connects two applications located in separate address spaces. The Hardware Transport Engine 560 contains hardware logic for performing transport-layer operations across the QP for an endpoint. The RDMA Engine 562 leverages the Hardware Transport Engine 560 to perform RDMA operations between endpoints. Implementing RDMA operations in hardware, the RDMA Engine 562 allows an application to read and write the memory of a remote system without OS kernel intervention or unnecessary data copying by permitting one endpoint of a communication channel to directly place information into the memory of another endpoint. The Virtual Endpoint Logic 564 manages the operation of a virtual endpoint for channel I / O, that is, a virtual instance of a QP used by an application.The virtual endpoint logic 564 assigns the QPs to the virtual address space of an application connected to a virtual endpoint.

[0067] Congestion control logic 563 performs operations to reduce the occurrence of congestion on a channel. In various implementations, congestion control logic 563 can perform flow control over a channel to limit congestion at the destination of a data transfer. Congestion control logic 563 can perform flow control at the link level to manage source congestion on virtual connections of network ports 552A-552B. In some implementations, congestion control logic can perform operations to limit congestion at intermediate points (e.g., IB switches) along a channel.

[0068] The Offload Engines 566 enable the offloading of network tasks that would otherwise be performed in software on the Network Interface Device 550. The Offload Engines 566 can support the offloading of operations, including, but not limited to, offloading receive-side scaling from a device driver or stateless network operations, such as for TCP implementations over InfiniBand, like stateless TCP / UDP / IP offloading or VXLAN offloading. The Offload Engines 566 can also offload operations of an interrupt coalescing circuit 522 of the Network Interface Device 500. Fig. Implement 5A. The Offload Engines 566 can also be configured to support offloading NVMe-oF or other storage acceleration operations from a CPU.

[0069] The QoS logic 568 can perform QoS operations, including the QoS functionality built into InfiniBand's basic service provisioning mechanism. The QoS logic 568 can also implement enhanced InfiniBand QoS, such as fine-grained end-to-end QoS. The QoS logic 568 can implement queuing services and management to prioritize flows and ensure service levels or bandwidth according to flow priority. For example, the QoS logic 568 can configure virtual lane arbitration for virtual lanes of network ports 552A-552B according to flow priority. The QoS logic 568 can also work in conjunction with the congestion control logic 563.

[0070] The GSA / SMA logic 569 implements General Services Agent (GSA) operations for managing the network interface device 550 and the InfiniBand structure, as well as for performing subnet management agent operations. GSA operations include device-specific management tasks such as querying device attributes, configuring device settings, and controlling device behavior. The GSA / SMA logic 569 can also implement SMA operations, including a subset of those available from the SMA 476 of the InfiniBand switch 450. Fig. 4C can be executed. For example, the GSA / SMA logic 569 can handle management requests from the subnet manager, including requests to reset the device, update the firmware, or change configuration parameters.

[0071] The 570 management interface provides support for a hardware interface to perform out-of-band management of the 550 network interface device, such as a connection to a Board Management Controller (BMC) or a hardware debug interface.

[0072] Fig. Figure 6 is a block diagram illustrating a programmable network interface 600 and a data processing unit. The programmable network interface 600 is a programmable network engine that can be used to accelerate network-based computing tasks within a distributed environment. The programmable network interface 600 can be coupled to a host system via a host interface 670. The programmable network interface 600 can be used to accelerate network or storage operations for CPUs or GPUs of the host system. The host system could, for example, be a node of a distributed learning system used to perform distributed training, as in Fig. Figure 6 shows that the host system can also be a data center node within a data center.

[0073] In one embodiment, access to remote storage containing model data can be accelerated by the programmable network interface 600. For example, the programmable network interface 600 can be configured to present remote storage devices to the host system as local storage devices. The programmable network interface 600 can also accelerate RDMA operations performed between GPUs of the host system and GPUs of remote systems. In one embodiment, the programmable network interface 600 can enable storage functionality, such as, but not limited to, NVMe-oF.The programmable network interface 600 can also accelerate encryption, data integrity, compression and other remote storage operations for the host system, enabling remote storage to approach the latencies of storage devices directly connected to the host system.

[0074] The Programmable Network Interface 600 can also perform resource allocation and management for the host system. Storage security operations can be delegated to the Programmable Network Interface 600 and performed in accordance with the allocation and management of remote storage resources. Network-based operations for managing access to remote storage, which would otherwise be performed by a processor in the host system, can instead be performed by the Programmable Network Interface 600.

[0075] In one embodiment, network and / or data security operations can be delegated from the host system to the programmable network interface 600. Data center security policies for a data center node can be handled by the programmable network interface 600 instead of the host system's processors. For example, the programmable network interface 600 can detect and contain an attempted network-based attack (e.g., DDoS) on the host system, thereby preventing the attack from impacting the host system's availability.

[0076] The programmable network interface 600 can include a system-on-a-chip (SoC / SIP 620) that runs an operating system across multiple processor cores 622. The processor cores 622 can include general-purpose processor (e.g., CPU) cores. In one embodiment, the processor cores 622 can also include one or more GPU cores. The SoC / SIP 620 can execute instructions stored in a storage device 640. A storage device 650 can store local operating system data. The storage device 650 and storage device 640 can also be used to cache remote data for the host system. Network ports 660A-660B provide connectivity to a network or fabric and enable network access for the SoC / SIP 620 and, via the host interface 670, for the host system.In one configuration, a first network port 660A can be connected to a first forwarding element, while a second network port 660B can be connected to a second forwarding element. Alternatively, both network ports 660A-660B can be connected to a single forwarding element via a link aggregation protocol (LAG). The programmable network interface 600 can also include an I / O interface 675, such as a USB interface. The I / O interface 675 can be used to connect external devices to the programmable network interface 600 or as a debug interface.

[0077] The programmable network interface 600 also includes a management interface 630, which allows software on the host device to manage and configure the programmable network interface 600 and / or the SoC / SIP 620. In one embodiment, the programmable network interface 600 can also include one or more accelerators or GPUs 645 to accept the offloading of parallel computing tasks from the SoC / SIP 620, the host system, or remote systems connected via network ports 660A-660B. For example, the programmable network interface 600 can be configured with a graphics processor and participate in general-purpose or graphics computing operations in a data center environment.

[0078] One or more aspects can be implemented by a representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, the machine-readable medium may contain instructions representing different logics within the processor. When read by a machine, these instructions can cause the machine to construct the logic to execute the techniques described herein. Such representations, known as "IP kernels," are reusable logic units for an integrated circuit that can be stored on a tangible, machine-readable medium as a hardware model describing the structure of the integrated circuit.The hardware model can be supplied to various customers or manufacturing facilities, which load the hardware model into manufacturing machines that produce the integrated circuit. The integrated circuit can be manufactured such that the circuit performs operations described in connection with any of the embodiments described herein.

[0079] Fig. Figure 7 is a block diagram illustrating an IP core development system 700. The IP core development system 700 can be used to fabricate an integrated circuit to perform the fabric and data center component operations described here. The IP core development system 700 can be used to create modular, reusable designs that can be integrated into a larger design or used to construct an entire integrated circuit (e.g., an integrated SOC circuit). A design facility 730 can generate a software simulation 710 of an IP core design in a higher-level programming language (e.g., C / C++). The software simulation 710 can be used to design, test, and verify the behavior of the IP core using a simulation model 712. The simulation model 712 can include functional, behavioral, and / or timing simulations.A register transfer level (RTL) design (715) can then be generated or synthesized from the simulation model 712. The RTL design 715 is an abstraction of the integrated circuit's behavior, modeling the flow of digital signals between hardware registers, including the associated logic executed using the modeled digital signals. In addition to an RTL design 715, lower-level designs at the logic or transistor level can also be generated, designed, or synthesized. Therefore, the specific details of the initial design and simulation can vary.

[0080] The RTL design 715 or an equivalent can further be synthesized by the design facility into a hardware model 720, which may be in a hardware description language (HDL) or another representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. The IP core design can be stored for delivery to a manufacturing facility 765 using non-volatile memory 740 (e.g., hard disk, flash memory, or any non-volatile storage medium). The manufacturing facility 765 can be a third-party manufacturing facility. Alternatively, the IP core design can be transmitted via a wired connection 750 or a wireless connection 760 (e.g., over the Internet). The manufacturing facility 765 can then manufacture an integrated circuit based at least partially on the IP core design.The manufactured integrated circuit can be configured to perform operations according to at least one embodiment described herein. High availability of distributed AI software

[0081] AI deployments continue to grow and scale as models become larger with an increasing number of parameters. With the growing volume of real-time data from applications like smart cars and smart sensors, the amount of data that needs to be accommodated in RAG and RAG-like setups will increase significantly. This surge in data volume places considerable pressure on the infrastructure to manage these large datasets, including sharding and replication.

[0082] Fig. Figures 8A-8B illustrate a System 800 in which a vector database is distributed across multiple shards for load balancing and high availability. Fig. Figure 8A illustrates vector database sharding for the System 800. Fig. Figure 8B illustrates the high availability of the vector database via shard replication.

[0083] As in Fig. As shown in Figure 8A, the System 800 includes an application 802, local application access interfaces 804A-804C, and database shards 806A-806C of a vector database. The application 802 can be an AI application that uses an LLM configured for RAG and / or a front-end API for a remotely accessible LLM that utilizes RAG. The local application access interfaces 804A-804C provide access to respective database shards 806A-806C that are geographically distributed (e.g., London, Paris, Rome). The local application access interfaces 804A-804C can also provide access to a geographically adjacent database shard. The database shards 806A-806C include respective replication units 807 that can replicate between shards.The respective replication units 807 can contain query nodes that can be searched by the application 802 through the local application access interfaces 804A-804C in response to a RAG query during an inference operation.

[0084] As in Fig. As shown in Figure 8B, the respective local application access interfaces 804A-804C can include a proxy 808 that enables a search query 809 against the vector database. The respective replication units 807 of the vector database contain query nodes 810A-810D, which are searched in response to the search query 809. If the search query 809 attempts to search a segment of the database (e.g., query node 810A) that is unavailable, the proxy 808 can be configured to search another accessible query node (e.g., query node 810B).

[0085] Techniques known in engineering manage the data and data infrastructure of the system. Fig. The System 800 shown in Figures 8A-8B is implemented in software, where a software process tracks the location of replicas and manages replication and sharding, including ensuring that the replica spans rack-level, data center-level, or geographic separation depending on use-case-specific configurations and constraints. However, there are several challenges in implementing vector database management in the host software. As data volumes continue to increase across model sizes and embeddings, challenges arise in tracking the presence and availability of replicas across racks, data centers, and geographic areas using software running on one of the nodes. For example, a host processor failure on a node can cause the tracking software running on that node to fail.Furthermore, the tracking software triggers interrupts on the host CPU to regularly check and confirm that replicas are alive and active. The host software will also be responsible for monitoring network traffic and latency to access a specific replica, potentially located in a different geographic area, in order to direct a query to the nearest replica. The host software is also responsible for synchronizing replicas, which may occur across geographical areas, and managing the recovery process from a replica in response to a node failure. Additionally, the host software is tasked with predicting node failure and taking corrective action. This task is challenging when performed on the node itself, which is capable of failing.

[0086] The techniques described herein can help overcome the challenges mentioned above, which arise when using host-executed software to manage vector databases, by providing vector database accelerator and shard management offload logic implemented within hardware and through software running on device processors and programmable data planes of a programmable network interface device (e.g., IPU, DPU, EPU, Smart NIC), as described herein. Programmable network interface devices are uniquely positioned as distinct failover domains that can run independently of a node's host processors.A node's programmable network interface device (PND) can continue to function through a software or hardware failure of the host processor on that node and can remain operational while the host environment is reset or restored. PNDs are tightly integrated with the network and can be configured to manage replicas, provide a unified front end, track heartbeats, minimize load balancing, mitigate node failures, and manage recovery and migration based on circumstances.

[0087] Fig. Figures 9A-9B illustrate systems in which programmable network interface devices provide vector database acceleration and network monitoring, according to embodiments. Fig. Figure 9A illustrates a System 900 in which network-linked nodes include IPUs designed to accelerate vector database access and provide high availability for vector database shards. Fig. Figure 9B illustrates a System 950 in which INT (In-Band Network Telemetry) is used to track network health and enable failure prediction.

[0088] As in Fig. As shown in Figure 9A, a System 900 can include head nodes 902A-902F, which are coupled to a Network 908. The System 900 can be analogous to the part of a Data Center 300 that is shown in Figure 9A. Fig. Figure 3 shows that the 908 network can be a Fabric 370 as shown in Fig. 3. For example, the network 908 can include Ethernet and / or InfiniBand interconnects. The head nodes 902A-902F can include elements of the heterogeneous compute server node 306. Fig. 3 and include a GPU array 904 and an IPU 906. The use of an IPU 906 is exemplary, and the techniques described here are not limited to a particular implementation of a programmable network interface device. The System 900 can also be implemented with a DPU, an EPU, or other analogous network interface devices with programmable data planes and general-purpose processors that can be configured to run network infrastructure management software. In various embodiments, the IPU 906 includes elements and functionality within the head nodes 902A-902F, including, but not limited to, the network interface device 500. Fig. 5A, the network interface device 550 of Fig. 5B and / or the programmable network interface 600 from Fig. 6.

[0089] System 900 implements mechanisms via the IPU 906 of the head nodes 902A-902F to manage high availability and data replication of distributed AI software, such as a shared vector database. The IPU 906 within the head nodes 902A-902F is positioned with close access to the network 908, accelerators (e.g., GPU array 904), and host cores within the head nodes. In one configuration, the head nodes 902A-902F contain multiple host cores, with multiple host cores connected to the IPU 906 within a single head node. The head nodes 902A-

[0090] The functionality of the IPU 906 within the head node 902E is shown in detail. The illustrated functionality can be implemented, in whole or in part, by the IPU 906 within any or all of the head nodes 902A-902F, depending on the functionality assigned to those nodes. The IPU 906 can provide a unified front-end service for an inference engine by orchestrating critical operational tasks. The IPU 906 is configurable to manage replicas, track heartbeats, predict host node failures, and manage recovery and migration via a P4 programmable data plane circuit arrangement. The IPU 906 manages database shards, replicas of AI models, and vector database replicas across the entire infrastructure, enabling high availability and fault tolerance.The IPU 906 monitors the health of system components by tracking heartbeats and telemetry, enabling it to detect and respond to anomalies. The IPU 906 implements logic for load balancing capabilities, optimizing the distribution of inference requests, preventing bottlenecks, and facilitating efficient resource utilization. In the event of a node failure, the IPU 906 proactively mitigates the impact of this failure by initiating failover procedures to maintain uninterrupted service. The IPU 906 also oversees recovery and migration processes to allow AI services to adapt to infrastructure changes without compromising performance or data integrity. This orchestration by the IPU 906 enables robust, scalable AI deployments.For failed nodes after migration, the failed node can be examined, and based on the results, the node can be brought back online, repurposed for lower priority tasks, or scheduled for maintenance or replacement.

[0091] Vector database acceleration and high availability logic is implemented via hardware and / or firmware, a programmable data plane, and / or software running on processor cores within the IPU 906. In one embodiment, the logic includes a replica and redundancy execution tracker 910, a failure predictor 912, a per-GPU domain execution tracker 914, heartbeat logic 916, migration logic 918, recovery logic 920, and error correction logic 922.

[0092] The Replica and Redundancy Execution Tracker 910 tracks the location of replicas throughout a deployment with the goal of maintaining adequate redundancy for these replicas at the rack, data center, or geographic level. The Replica and Redundancy Execution Tracker 910 can track a replica ID / address, a sibling replica ID / address, a replica network node, the replica heartbeat result, the last downtime and recovery time, and the current probability of failure for the replica.

[0093] The Failure Predictor 912 tracks a set of telemetry metrics to maintain a running failure predictor. The Failure Predictor 912 tracks network and device telemetry, including, but not limited to, a number of correctable faults that are corrected by a node's memory controller, temperature metrics for a node, or other hardware metrics that may indicate an impending device failure. The Failure Predictor 912 can also track network telemetry to determine whether a potential node failure may occur due to the failure of the node's IPU 906. When the Failure Predictor value exceeds a threshold, policy-based corrective, mitigation, or migration actions can be taken. These policy actions may include creating an additional replica to ensure probabilistic coverage or migrating the node's data to another domain.

[0094] The Pro-GPU Domain Execution Tracker 914 tracks the load on GPUs and other relevant accelerator devices (e.g., NPUs, AI accelerators, etc.) in a head node. These metrics are a relevant factor to consider when determining which shards to use for a given task. For example, inference requests can be routed to different accelerator devices on a node based on the current load distribution across the accelerator devices. In one embodiment, the IPU 906 is designed to route inference requests to remote accelerators or remote accelerator nodes. The Pro-GPU Domain Execution Tracker 914 enables the IPU 906 to route inference requests to remote accelerators or remote accelerator nodes based on the load distribution across the remote accelerators or remote accelerator nodes.

[0095] The heartbeat logic 916 within an IPU associated with a database shard exchanges heartbeats with other IPUs connected to replicas of a given shard. The heartbeat logic 916 can generate a heartbeat message or a ping and transmit this message to the relevant IPUs, monitoring the return status of these messages. The heartbeat logic 916 within the various IPUs can exchange specific heartbeat status data. This heartbeat exchange periodically confirms that the other nodes are powered on and reachable within reasonable delays. In one embodiment, the heartbeat logic 916 is implemented via a programmable data plane of the IPU 906.

[0096] Migration logic 918 is activated when a node failure is imminent, based on a calculated probability, or in the case of an actual failure condition where the node's data is still accessible. Migration logic 918 can enable zero-downtime migration by initiating a copy operation to a new node while tracking changes that occur during the copy. Once the copy is complete, a transaction can be transferred to the new node.

[0097] Recovery logic 920 is activated to perform a node recovery. If a node is reset, recovers from a failure, or if a new node comes online, the previous state for the node can be restored from the node's IPU. The node is restored using state information stored on the node's IPU, along with updates from the other IPUs hosting shards for the task being processed.

[0098] Error correction logic 922 is activated when the IPU detects errors. These errors can be corrected via application-level programmed logic on the IPU. For example, checksums can be stored, exchanged, and propagated at the shard level to account for and correct network or other errors.

[0099] As in Fig. As shown in Figure 9B, a System 950 can enable in-band network telemetry (INT) by embedding relevant element telemetry data into packets as these packets traverse the System 950 from a first endpoint H1 (e.g., Host 952A) to a second endpoint H2 (e.g., Host 952B) through several forwarding elements 954A-954C (FE1-FE3). INT can be enabled via programmable data planes, which include, for example, P4 programmable packet processors that can define user-defined in-band telemetry behaviors and can be used to embed telemetry information (e.g., latency, queue depth) into live packets as these packets travel across the network.

[0100] Endpoint H1 can contain a segment of a vector database, and endpoint H2 can contain a replica of that segment. Forwarding elements 954A-954C can include, among other things, a switch, a switch chip, a router, or a bridge. The structure of System 950 is exemplary and is not limited to any specific embodiments described herein. In one embodiment, System 950 is a subsystem of System 900. Fig. 9A.

[0101] During operation, endpoint H1 can transmit a data packet 953A, addressed to the second endpoint H2, to FE1 954A. FE1 954A is an INT source that inserts an INT header and adds metadata relating to its own element ID (e.g., switch ID) and a forwarding delay before forwarding a data packet 953B to FE2 954B, which is an INT transit element. FE2 954B appends its own INT metadata and forwards a data packet 953C to FE3 954C. FE3 954C is an INT sink that appends its own INT metadata and then creates a copy of data packet 953D with all the collected INT metadata and forwards this data packet 953D to a failure predictor 912 for analysis. As the INT sink, FE3 954C also removes the INT header and metadata to restore the original instance of data packet 953A. FE3 954C then forwards data packet 953A to the second endpoint H2 (host 952B).

[0102] The data packet 953D, which is forwarded to the failure predictor 912, contains an INT metadata stack 955, which includes the forward element IDs and the latency for jumps along the path between endpoints. The INT metadata stack 955 contains information that can be used by the failure predictor 912 to monitor the health and performance of the network within the system 950 in real time.

[0103] Fig. Figure 10 illustrates a programmable network interface device 1000 that can be configured as a vector database accelerator. The programmable network interface device 1000 incorporates functionality from the network interface device 500. Fig. 5A and / or the network interface device 550 from Fig. 5B and is configured in various embodiments as an Ethernet device and / or an InfiniBand device. The programmable network interface device 1000 can be an implementation of the IPU 906 from Fig. It should be 9A.

[0104] The programmable network interface device 1000 includes a network subsystem 1010, which enables network interface functionality, and a computing complex 1030, which enables program execution. The network subsystem 1010 includes a circuit arrangement of a host interface SerDes 1011 (serializer / deserializer) that is configurable for coupling with a host interconnect (e.g., PCIe). The network subsystem 1010 also includes a circuit arrangement of the network interface SerDes 1028 and a network media access control (MAC) circuit arrangement that is configurable for coupling with a physical interface of a network.

[0105] The host interface SerDes 1011 is coupled to a circuit arrangement to provide virtual functions (VFs 1012) and physical functions (PFs 1013) for single-root I / O virtualization (SR-IOV) and scalable I / O virtualization (SIOV). Multiple instances of PFs 1013 enable the programmable network interface device 1000 to be used concurrently by multiple host processors and / or multiple physical hosts, which can virtualize the programmable network interface device 1000 via the VFs 1012 associated with the respective instances of PFs 1013.

[0106] The programmable network interface device 1000 includes an RDMA circuit arrangement 1014 for accelerating RDMA operations and an NVMe circuit arrangement 1016 for providing an NVMe device interface for NVMe-oF devices. The LAN circuit arrangement 1018 accelerates local network functionality and is coupled with a packet processing pipeline 1020, which in one embodiment is a P4 programmable pipeline. The inline cryptography circuit arrangement 1022 enables line-rate packet encryption and decryption, for example, for IPsec (Internet Protocol Security) protocols and / or VPN functionality, and the traffic shaper circuit arrangement 1024 enables traffic shaping via transmission planning.

[0107] The computing complex 1030 includes a processor core array 1032, which can execute infrastructure software directly on the programmable network interface device 1000, thus enabling such functionality to be offloaded from host processors. The processor core array 1032 is coupled with a system cache 1033, which is secured by multiple memory channels. A lookaside cryptography and compression engine 1036 provides cryptography and compression acceleration functionality for the processor core array 1032 and for host processors. Additionally, the management complex circuit arrangement 1038 includes a dedicated management processor or microcontroller that provides secure boot and lifecycle management functionality and enables remote management of the programmable network interface device 1000.

[0108] In one embodiment, the management processor within the management complex circuit arrangement 1038 can configure the execution of the vector database accelerator logic 1040 across one or more processor cores of the processor core array 1032. The vector database accelerator logic 1040 is configurable to execute program code to provide at least part of the functionality of the system 900. Fig. 9A, of system 950 from Fig. 9B to provide. Other parts of the functionality of System 900 from Fig. 9A and / or system 950 from Fig. 9B are implemented in various embodiments via the management complex circuit arrangement 1038 and / or the packet processing pipeline 1020. Access to vector data of the vector databases can be accelerated via the RDMA circuit arrangement 1014 and / or the NVMe circuit arrangement 1016. In one embodiment, the packet processing pipeline 1020 and the vector database accelerator logic 1040 are configurable to implement flow replication between instances of the programmable network interface devices 1000, as shown in Fig. 11 shown.

[0109] Fig. Figure 11 illustrates a system 1100 designed to enable flow replication between programmable network interface devices according to one embodiment. In one embodiment, the system includes an active IPU 1000A and a standby IPU 1000B, which are implementations of the programmable network interface device 1000. Fig. 10 or other programmable network interface devices described herein. The logic of System 1100 facilitates availability for a vector database by safeguarding against device failure of a programmable network device that provides network access, data access, and / or node replication management.

[0110] In one embodiment, the System 1100 can access software and data from the System 900. Fig. 9A replicates from an active IPU 1000A to a standby IPU 1000B. The System 1100 can also perform flow replication from the active IPU 1000A to a standby IPU 1000B. Flow replication enables planned switchovers, for example, in the case of a firmware update on the active IPU 1000A. Flow replication also enables unplanned switchovers in the event of a hardware failure. In such scenarios, a switchover occurs in which the standby IPU 1000B becomes the active device with minimal packet loss, while limiting the potential for link loss. Instead of using a dedicated out-of-band link between the active IPU 1000A and the standby IPU 1000B, flow replication events are transmitted in-band as network packets between the active IPU 1000A and the standby IPU 1000B.

[0111] In one embodiment, the active IPU 1000A and the standby IPU 1000B each include a flexible pipeline receive circuit arrangement (RX 1110) and a flexible pipeline receive transmit circuit arrangement (TX 1120). The RX circuit arrangement 1110 includes direction and lookup logic 1111 of an elastic network interface (ENI), overlay routing logic 1112, and ACL policy lookup logic 1114. The RX circuit arrangement 1110 also includes connection tracing and TCP state machine logic 1116 and high availability logic 1118, managing a flow table 1125A of the active IPU 1000A. The high availability logic 1118 is designed to replicate updates to the flow table 1125A of the active IPU 1000A to the standby flow table 1125B of the standby IPU 1000B.In response to an update of flow table 1125A, the high-availability logic 1118 can transmit an update data packet 1102 via an underlay routing circuit arrangement 1121 of the TX circuit arrangement 1120. The update data packet 1102 is processed by the RX circuit arrangement 1110 of the standby IPU 1000B and recognized as an HA packet. The standby IPU 1000B uses data from the update data packet 1102 to update the standby flow table 1125B. The HA logic of the standby IPU 1000B can terminate an acknowledgment data packet 1104 back to the active IPU 1000A to acknowledge the update.

[0112] In one embodiment, a control plane configuration API is provided to enable scope and role definition for active and standby devices. This APU allows the active IPU 1000A to declare and configure itself as active and to specify a standby IPU 1000B to which migration and flow data are replicated. When a standby IPU 1000B is defined for an active IPU 1000A, the state, configuration data, and flow table 1125A of the active IPU 1000A are applied to the standby IPU 1000B, causing the standby IPU 1000B to become a replica of the active IPU 1000A. Configuration and flow update events occurring on the active IPU 1000A (e.g., flow add, flow delete) are translated into network messages and transmitted internally to the standby IPU 1000B. In one embodiment, the messages are transmitted using a reserved Layer 4 UDP port (e.g.,(Transport layer). In one embodiment, the messages include a header or identifier indicating that the message contains flow events that are replicated by the active IPU 1000A. In another embodiment, flow events are transmitted between the active IPU 1000A and the standby IPU 1000B over a predetermined VLAN. In such an embodiment, the messages have a VLAN tag of the predetermined VLAN.

[0113] In one embodiment, aspects of System 1100 can be incorporated into software for open networking in the cloud (SONiC) disaggregated APIs for SONiC hosts (DASH) implementation to enable high availability for cloud applications. In one embodiment, the active IPU 1000A is located in a first host, the standby IPU 1000B is located in a second host, and network traffic between the active IPU 1000A and the standby IPU 1000B traverses at least one forwarding element. For example, System 1100 can be located within System 950. Fig. 9B, wherein the active IPU 1000A is located in host 952A and the standby IPU 1000B is located in host 952B. In such an embodiment, host 952B can be associated with replicas of one or more database shards that are stored on, associated with, or managed by host 952B. In one embodiment, the active IPU 1000A and the standby IPU 1000B are located in the same host, with the standby IPU 1000B functioning as a redundant network interconnect and / or a redundant infrastructure management resource.

[0114] Fig. Figure 12 illustrates a method 1200 for enabling AI applications with high availability via programmable network interface devices according to embodiments described herein. In one embodiment, the method 1200 can be implemented via logic of the system 900. Fig. 9A and the 950 system from Fig.9B will be implemented. Method 1200 involves providing access to a physical function of a programmable network interface device via a host interface, wherein the physical function enables access to an infrastructure management circuit arrangement within the programmable network interface (1202). The programmable network interface device includes a network interface and a packet processing circuit arrangement coupled to the network interface.

[0115] The infrastructure management circuit arrangement can monitor the status of a first node associated with a database shard of a vector database (1204). The infrastructure management circuit arrangement can additionally monitor the status of individual replicas of multiple replicas of the first node and / or a second node via heartbeat data and / or telemetry data. The heartbeat data indicates the network availability of the respective replicas, and the telemetry data includes in-band network telemetry generated via forwarding elements between the network interface and the node. The telemetry data can also include hardware metrics associated with the node, including hardware metrics indicating a probability of failure associated with the node. The infrastructure management circuit arrangement can determine a probability of failure associated with the first node (1206).The infrastructure management circuit arrangement can configure a query for data from the first node to be served via a replica of the first node in response to a determination that the probability of failure exceeds a threshold (1208).

[0116] The techniques described herein address the above challenges that arise when using host-executed software to manage vector databases by providing vector database accelerator and shard management offload logic implemented within hardware and through software running on device processors and programmable data layers of a programmable network interface device.

[0117] One embodiment provides a device comprising a network interface, a packet processing circuit arrangement coupled to the network interface, a host interface coupled to the packet processing circuit arrangement and the network interface, wherein the host interface includes a circuit arrangement to enable access to a physical function, and an infrastructure management circuit arrangement coupled to the network interface and the host interface and accessible via the physical function.The infrastructure management circuit arrangement is designed to enable data access for a neural network inference engine with a distributed data model that includes multiple distributed data shards, via dynamic and / or real-time management of a node associated with the neural network inference engine, where the node contains a database shard of a vector database.

[0118] In one embodiment, the infrastructure management circuitry monitors the node's status using heartbeat or telemetry data, calculates a failure probability for the node based on this status, and, if the failure probability exceeds a threshold, configures data queries to be served by a data replica. Additionally, in cases where a failure probability is determined to be high, the circuitry can initiate a live data migration to generate this data replica. The heartbeat data provides insight into the node's network availability, while telemetry data includes in-band network telemetry from forwarding elements between the network interface and the node, as well as hardware metrics associated with the node.In one embodiment, the hardware metrics refer to an accelerator device that processes inference requests from the neural network inference engine.

[0119] In one embodiment, the infrastructure management circuit arrangement is designed to identify a remote network interface device to be configured as a standby network interface device and to replicate a network flow configuration associated with the packet processing circuit arrangement to the standby network interface device. The infrastructure management circuit arrangement is configurable to receive an event involving an adaptation to a network flow configuration associated with the packet processing circuit arrangement, translate the event into a network message containing the adaptation, and transmit the network message to the standby network interface device.

[0120] One embodiment provides a method that includes providing access to a physical function of a programmable network interface device via a host interface. The programmable network interface device comprises a network interface and a packet processing circuit arrangement coupled to the network interface. The physical function enables access to an infrastructure management circuit arrangement coupled to the network interface and the host interface.The procedure additionally includes monitoring, via the infrastructure management circuit arrangement, the status of a first node associated with a database shard of a vector database, determining a failure probability associated with the first node, and configuring a query for data from the first node to be served via a replica of the first node, in response to a determination that the failure probability exceeds a threshold.

[0121] In one embodiment, the method further comprises receiving a request from a neural network inference engine for access to data associated with a second node, selecting one of several replicas of the second node, and transmitting a request for the data associated with the second node to the selected replica of the second node. The method may further include monitoring, via the infrastructure management circuit arrangement, the status of each of the second node's multiple replicas using heartbeat data and / or telemetry data. The heartbeat data may indicate the network availability of the second node and / or its respective replicas, and the telemetry data includes in-band network telemetry generated via forwarding elements between the network interface and the second node and / or its respective replicas.The telemetry data may additionally or alternatively include hardware metrics associated with the second node and / or its respective replicas. In one embodiment, any aspect of the method described above may be implemented by a system that includes means for performing the operations of the method. Additionally, a non-volatile, machine-readable medium may store instructions to cause a processor to perform aspects of the method described above.

[0122] One embodiment provides a system comprising a storage device, a first host processor coupled to the storage device, and a second host processor coupled to the storage device. The system further includes a first network interface device comprising a first network interface, a first packet processing circuitry coupled to the first network interface, and a host interface comprising a first circuitry configured to provide access to a first physical function and a second circuitry configured to provide access to a second physical function. The first host processor is coupled to the first network interface device via the first physical function, and the second host processor is coupled to the first network interface device via the second physical function.The system additionally includes a third circuit arrangement designed to identify a second network interface device to be configured as a standby network interface device, and to replicate a network flow configuration associated with the first packet processing circuit arrangement to the second network interface device.

[0123] In one embodiment, the third circuit arrangement is configured to receive an event involving an adaptation to a network flow configuration associated with the first packet processing circuit arrangement, to translate the event into a network message containing the adaptation, and to transmit the network message to the second network interface device. The second network interface device is configurable to receive the network message, determine that the network message contains the adaptation, and apply the adaptation to a flow table associated with a second packet processing circuit arrangement within the second network interface device.In one embodiment, the second network interface device determines that the network message includes the adaptation based on an identifier associated with the network message and / or a port associated with the network message. The first packet processing circuit arrangement and the second packet processing circuit arrangement include programmable packet processor circuits.

[0124] References in the specification to "an embodiment," "an illustrative embodiment," etc., indicate that the described implementation may have a specific feature, structure, or characteristic. However, each embodiment may or may not contain this specific feature, structure, or characteristic. Furthermore, such expressions do not necessarily refer to the same embodiment. Moreover, if a particular feature, structure, or characteristic is described in connection with an embodiment, it is also assumed that it is within the knowledge of a person skilled in the art to achieve such a feature, structure, or characteristic in connection with other embodiments, whether explicitly described or not.Additionally, it is understood that entries in a list of the form "at least one of A, B and C" can mean: (A); (B); (C); (A and B); (A and C); (B and C); or (A, B and C). Likewise, elements listed in the form of "at least one of A, B or C" can mean: (A); (B); (C); (A and B); (A and C); (B and C); or (A, B and C).

[0125] The disclosed embodiments may in some cases be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions that are carried by or stored on a volatile or non-volatile machine-readable (e.g., computer-readable) storage medium and that can be read and executed by one or more processors. A machine-readable storage medium may be any storage device, mechanism, or other physical structure for storing or transmitting information in a machine-readable form (e.g., volatile or non-volatile memory, a media disc, or other media device).This machine-readable storage medium can contain instructions which, when executed, cause one or more processors to perform the operations described here.

[0126] In the drawings, some structural or process features may be shown in specific arrangements and / or sequences. However, it should be recognized that such specific arrangements and / or sequences may not be necessary. In some embodiments, such features may be arranged in a different manner and / or sequence than shown in the illustrative figures. Additionally, the inclusion of a structural or process feature in a specific figure does not mean that such a feature is required in all embodiments, and in some embodiments it may not be included or may be combined with other features.

[0127] The foregoing description and drawings are to be considered illustrative and non-limiting. Those skilled in the art will understand that various modifications and alterations can be made to the embodiments described herein without deviating from the overall concept and scope of protection of the features as set forth in the attached claims.

Claims

[1] Device comprising: a network interface; a packet processing circuit arrangement coupled with the network interface; a host interface coupled to the packet processing circuitry and the network interface, wherein the host interface includes a circuitry to enable access to a physical function; and An infrastructure management circuit arrangement coupled to the network interface and the host interface and accessible via the physical function, wherein the infrastructure management circuit arrangement is designed to enable data access for a neural network inference engine with a distributed data model via real-time management of a node associated with the neural network inference engine, wherein the node contains a database shard of a vector database. [2] Device according to claim 1, wherein the infrastructure management circuit arrangement is designed to: Monitoring the status of the node via heartbeat data and / or telemetry data; Determining a failure probability associated with the node, based on its status; and In response to a determination that the probability of failure exceeds a threshold, configure a query for data associated with the node to be served via a data replica. [3] Device according to claim 2, wherein the infrastructure management circuit arrangement is designed to initiate a live migration of data associated with the node in order to generate the data replica in response to a determination that the probability of failure exceeds the threshold. [4] Device according to claim 2 or 3, wherein the heartbeat data indicates network availability of the node. [5] Device according to one of claims 2-4, wherein the telemetry data includes in-band network telemetry generated via forwarding elements between the network interface and the node. [6] Device according to one of claims 2-5, wherein the telemetry data includes hardware metrics associated with the node. [7] Device according to claim 6, wherein the hardware metrics include metrics associated with an accelerator device of the node, and the infrastructure management circuit arrangement is designed to: Monitoring the metrics associated with an accelerator device of the node, wherein the accelerator device performs an operation in response to an inference request received from the neural network inference engine; and Determine whether the inference request should be routed to the accelerator device, based on the metrics. [8] Device according to claim 7, wherein the metrics are configurable to include a load associated with the accelerator device. [9] Device according to any one of claims 1-8, wherein the infrastructure management circuit arrangement is designed to: Identifying a remote network interface device to be configured as a standby network interface device; and Replicating a network flow configuration associated with the packet processing circuitry arrangement to the standby network interface device. [10] Device according to claim 9, wherein the infrastructure management circuit arrangement is designed to: Receiving an event that involves an adaptation to a network flow configuration associated with the packet processing circuitry; Translating the event into a network message, including adaptation; and Transmitting the network message to the standby network interface device. [11] Procedure, encompassing: Providing access to a physical function of a programmable network interface device via a host interface, wherein the programmable network interface device includes a network interface and a packet processing circuit arrangement coupled to the network interface, and wherein the physical function provides access to an infrastructure management circuit arrangement coupled to the network interface and the host interface; Monitor, via the infrastructure management circuit arrangement, the status of a first node associated with a database shard of a vector database; Determining a failure probability associated with the first node; and Configure a query for data from the first node, to be served via a replica of the first node, in response to a determination that the probability of failure exceeds a threshold. [12] The method of claim 11, comprising: Receiving a request from a neural network inference engine to access data associated with a second node; Selecting one of several replicas of the second node; and Transferring a request for the data associated with the second node to a selected copy of the second node. [13] The method of claim 12, comprising: Monitor, via the infrastructure management circuit arrangement, the status of each replica of the multiple replicas of the second node via heartbeat data and / or telemetry data. [14] Method according to claim 13, wherein the heartbeat data indicate a network availability of the respective replicates. [15] Method according to claim 13 or 14, wherein the telemetry data includes in-band network telemetry generated via forwarding elements between the network interface and the second node, and / or hardware metrics associated with the second node. [16] Non-volatile machine-readable medium containing instructions stored thereon, wherein the instructions cause one or more processors of a programmable network interface device to perform operations of a method according to any one of claims 11-15. [17] Equipment comprising means for carrying out a method according to any one of claims 11-15. [18] System, encompassing: a storage device; a first host processor coupled with the storage device; a second host processor coupled with the storage device; and a first network interface device, which includes the following: a first network interface; a first packet processing circuit arrangement coupled with the first network interface; a host interface comprising a first circuit arrangement designed to provide access to a first physical function and a second circuit arrangement designed to provide access to a second physical function, wherein the first host processor is coupled to the first network interface device via the first physical function, and wherein the second host processor is coupled to the first network interface device via the second physical function; and a third circuit arrangement designed to: Identifying a second network interface device for configuration as a standby network interface device; and Replacing a network flow configuration associated with the first packet processing circuitry to the second network interface device. [19] System according to claim 18, wherein the third circuit arrangement is designed to: Receiving an event that involves adapting to a network flow configuration associated with the first packet processing circuit arrangement; Translating the event into a network message, including adaptation; and Transmitting the network message to the second network interface device. [20] System according to claim 19, wherein the second network interface device is configurable to: Receiving the network message; Determine that the network message includes the adaptation; and Applying the adaptation to a flow table associated with a second packet processing circuit arrangement within the second network interface device. [21] System according to claim 20, wherein the second network interface device is to determine that the network message includes the adaptation based on an identifier associated with the network message and / or a port associated with the network message. [22] System according to claim 20 or 21, wherein the second network interface device is configurable to send an acknowledgment to the first network interface after applying the adaptation to the flow table associated with the second packet processing circuit arrangement. [23] System according to one of claims 20-22, wherein the first packet processing circuit arrangement and the second packet processing circuit arrangement comprise programmable packet processor circuits.