Prompt adaptation for artificial intelligence models in programmable network interface devices
Programmable network interface devices with accelerators and connectivity offload infrastructure tasks, reducing server load and enhancing security in virtualized environments by managing infrastructure functions.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- INTEL CORP
- Filing Date
- 2025-11-04
- Publication Date
- 2026-06-11
AI Technical Summary
In highly virtualized environments, significant server resources are dedicated to processing tasks beyond user applications, including hypervisors, container engines, networking, and storage functions, which can be offloaded using programmable network interface devices (PNDs) to reduce infrastructure load and provide an additional layer of security.
Implementing programmable network interface devices (PNDs) with accelerators and network connectivity to offload infrastructure tasks, utilizing infrastructure processing units (IPUs), data processing units (DPUs), and edge processing units (EPUs) to manage infrastructure functions and reduce server overhead.
PNDs effectively reduce server load and enhance security by offloading infrastructure tasks, optimizing resource utilization and improving network performance in data centers.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND OF THE REVELATION
[0001] In highly virtualized environments, significant amounts of server resources are dedicated to processing tasks beyond user applications. These processing tasks can include hypervisors, container engines, networking and storage functions, security, and large volumes of network traffic. To handle these diverse processing tasks, programmable network interface devices (PNDs) with accelerators and network connectivity have been introduced. These PNDs are referred to as infrastructure processing units (IPUs), data processing units (DPUs), edge processing units (EPUs), programmable network devices, and so on. The PNDs can accelerate and manage infrastructure functions using dedicated and programmable cores deployed within the devices.Programmable network interface devices (PNDs) can reduce infrastructure load and provide an additional layer of security by acting as the host's control point for running infrastructure applications. Using PNDs allows the overhead associated with executing infrastructure tasks to be offloaded from a server device. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] The embodiments described herein are illustrated by way of example and without limitation in the figures of the accompanying drawings, in which the same references denote similar elements and in which: Fig. a block diagram showing a computer system configured to implement one or more aspects of the embodiments described herein; Fig. a block diagram of a system that contains selected components of a data center; Fig. a block diagram of a part of a data center according to one or more examples of the present description; Fig. Show programmable forwarding elements and adaptive routing. Fig. represent an example of network interface devices; Fig. a block diagram showing a programmable network interface and a data processing unit; Fig. a block diagram that shows an IP core development system; Fig. a block diagram that shows an example of a computing environment for providing prompt adaptation for artificial intelligence (AI) models in programmable network interface devices according to the implementations described herein; Fig. a block diagram showing an example of a programmable network interface device for providing prompt customization for AI models in programmable network interface devices according to the implementations described herein; Fig. a block diagram representing a data center environment that includes an AI deployment supporting prompt customization for AI models according to the implementations described herein; Fig. a block diagram showing a prompt augmentation microservice according to the implementations described herein. Fig. a flowchart that shows an embodiment of a method for providing prompt adaptation for AI models in a programmable network interface device; and Fig. a flowchart that shows an embodiment of a method for providing a prompt augmentation tracking table to support prompt adaptation for AI models in a programmable network interface device. DETAILED DESCRIPTION
[0003] The following description includes numerous specific details to facilitate better understanding. However, it is obvious to a person skilled in the art that the embodiments described herein can also be implemented without one or more of these specific details. In other cases, known features have not been described so as not to obscure the details of the present embodiments.
[0004] Fig. Figure 1 is a block diagram showing a computer system 100 configured to implement one or more aspects of the embodiments described herein. In the implementations described herein, the exemplary computer system 100 can consist of Fig. Components for implementing prompt adaptation for artificial intelligence (AI) models in programmable network interface devices as discussed below with reference to Fig. The computing system 100 includes a processing subsystem 101 with one or more processors 102 and a system memory 104, which communicate via a connection path that may include a memory hub 105. The memory hub 105 may be a separate component within a chipset component or it may be integrated into the one or more processors 102. The memory hub 105 is coupled to an I / O subsystem 111 via a communication link 106. The I / O subsystem 111 includes an I / O hub 107, which allows the computing system 100 to receive inputs from one or more input devices 108. The I / O hub 107 may additionally allow a display controller, which may be included in the one or more processors 102, to provide outputs for one or more display devices 110A.In one embodiment, the one or more display devices 110A coupled to the I / O hub 107 may include a local, internal or embedded display device.
[0005] The processing subsystem 101, for example, contains one or more parallel processors 112 connected to the memory hub 105 via a communication link 113, such as a bus or fabric. The communication link 113 can be any number of standards-based communication link technologies or protocols, such as PCI Express, or a vendor-specific communication interface or structure. The one or more parallel processors 112 can form a computationally focused parallel or vector processing system that may have a large number of processing cores and / or processing clusters, such as a multi-core processor (MIC processor).For example, the one or more parallel processors 112 form a graphics processing subsystem that can output pixels to one or more display devices 110A coupled via the I / O hub 107. The one or more parallel processors 112 can also include a (not shown) display controller and a (not shown) display interface to enable direct connection to one or more display devices 110B.
[0006] Within the I / O subsystem 111, a system storage unit 114 can be connected to the I / O hub 107 to provide a storage mechanism for the computing system 100. An I / O switch 116 can be used to provide an interface mechanism to enable connections between the I / O hub 107 and other components, such as a network adapter 118 and / or a wireless network adapter 119, which may be integrated into the platform, and various other devices that can be added via one or more add-in devices 120. The one or more add-in devices 120 may, for example, also include one or more external graphics processing units, graphics cards, and / or compute accelerators. The network adapter 118 may be an Ethernet adapter or another wired network adapter.The wireless network adapter 119 can contain one or more wireless radio devices from a Wi-Fi, Bluetooth, Near Field Communication (NFC) or other networking device.
[0007] The Computing System 100 may include other components not explicitly shown, such as USB or other connectors, optical storage drives, video recording devices, and the like, which may also be connected to the I / O Hub 107. Communication paths connecting the various components in Fig. Connecting them together can be implemented using any suitable protocols, such as PCI (Peripheral Component Interconnect)-based protocols (e.g., PCI Express) or any other bus or point-to-point communication interfaces and / or protocol(s), such as the NVLink high-speed connection, Compute Express Link™ (CXL™) (e.g., CXL.mem), Infinity Fabric (IF), Ethernet (IEEE 802.3), Remote Direct Memory Access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Quick UDP Internet Connections (QUIC), RDMA over Converged Ethernet (RoCE), Ultra Ethernet Transport (UET), Ultra Accelerator Link (UALink), Intel QuickPath Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), Omnipath, HyperTransport, Advanced Microcontroller Bus Architecture (AMBA) Interconnect, OpenCAPI, Gen-Z, Cache Coherent Interconnect for Accelerators (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G and variations thereof, or wired or wireless connection protocols known in the art. In some examples, data can be copied or stored on virtualized storage nodes using a protocol such as Non-Volatile Memory Express (NVMe) over Fabrics (NVMe-oF) or NVMe.In one embodiment, time-dependent communication protocols are supported, including time-dependent RDMA, time-dependent NVME, and time-dependent NVME-oF, where an accurate time and rate of data consumption is used to control data transmission.
[0008] The one or more parallel processors 112 can include a circuit arrangement optimized for graphics and video processing, including, for example, a video output circuit arrangement, and they can form a graphics processing unit (GPU). Alternatively or additionally, the one or more parallel processors 112 can include a circuit arrangement optimized for general-purpose processing, while retaining the underlying computing architecture, which is described in more detail herein. Components of the computing system 100 can be integrated with one or more other system elements on a single integrated circuit. For example, the one or more parallel processors 112, the memory hub 105, the one or more processors 102, and the I / O hub 107 can be integrated on a single system-on-a-chip (SoC) circuit.Alternatively, the components of the computing system 100 can be integrated in a single package to form a system-in-package (SIP) configuration. In one embodiment, at least some of the components of the computing system 100 can be integrated in a multi-chip module (MCM), which can be interconnected with other multi-chip modules to form a modular computing system. An MCM or SIP configuration can contain multiple integrated circuits, chiplets, dielets, tiles, or other circuit forms. The term chip can refer to a packaged die, while the die can refer to a single, unpackaged instance of a chip design. However, the terms chip and die are often used interchangeably in the prior art. When the term chiplet is used herein, it refers to an at least partially packaged integrated circuit that can be integrated with other circuits in an MCM or SIP configuration.
[0009] In some configurations, the Computing System 100 includes, in addition to the one or more processors 102 and the one or more parallel processors 112, one or more accelerator devices 130 coupled to the memory hub 105. The one or more accelerator devices 130 are designed to perform domain-specific acceleration of workloads to handle computationally intensive or high-throughput tasks. The one or more accelerator devices 130 can reduce the load placed on the one or more processors 102 and / or the one or more parallel processors 112 of the Computing System 100.The accelerator device(s) 130 may include, but are not limited to, intelligent network interface cards, data processing units, cryptographic accelerators, memory accelerators, artificial intelligence (AI) accelerators, compression units, neural processing units (NPUs), memory accelerators and / or video transcoding accelerators.
[0010] It is understood that the computing system 100 shown herein is for illustrative purposes only and that variations and modifications are possible. The connection topology, including the number and arrangement of the bridges, the number of processors 102, and the number of parallel processors 112, can be changed as desired. For example, the system memory 104 can be connected directly to the one or more processors 102 instead of via a bridge, while other devices communicate with the system memory 104 via the memory hub 105 and the one or more processors 102. In other alternative topologies, the one or more parallel processors 112 are connected to the I / O hub 107 or directly to one of the one or more processors 102 instead of to the memory hub 105. In other embodiments, the I / O hub 107 and the memory hub 105 can be integrated on a single chip.It is also possible for two or more sets of processors 102 to be mounted across multiple sockets, which may be coupled with two or more instances of the one or more parallel processors 112.
[0011] Some of the special components shown herein are optional and need not be included in all implementations of the Computing System 100. For example, any number of expansion cards or peripherals may be supported, or some components may be omitted. Furthermore, some architectures may use different terminology for components similar to those described in Fig. be shown.
[0012] Fig. This is a block diagram of a System 200 containing selected components of a data center. In the implementations described herein, the exemplary System 200 can be derived from Fig. Components for implementing prompt adaptation for AI models in programmable network interface devices as discussed below with reference to Fig. Included are the components of the depicted data center. These components may be located, for example, at a cloud service provider (CSP) or another data center, such as a traditional enterprise data center, a private cloud, or a public cloud providing services like Infrastructure as a Service (IaaS), Platform as a Service (PaaS), or Software as a Service (SaaS). System 200 includes a number of workload clusters, including but not limited to workload clusters 218A and 218B. The workload clusters may consist of single servers, blade servers, rackmount servers, or any other suitable server topology.
[0013] The System 200 can contain workload clusters 218A-218B. Workload clusters 218A-218B can contain a Rack 248, which houses multiple servers (e.g., Server 246). The Rack 248 and the servers of the Workload Clusters 218A-218B can conform to the rack unit standard (“U” standard), where a rack unit corresponds to a 19-inch wide rack frame and a full-size industry-standard rack can accommodate 42 units (42U) of equipment. A unit of equipment (1U) (e.g., a 1U server) can be 1.75 inches high and approximately 36 inches deep. In various configurations, compute resources such as processors, memory, storage, accelerators, and switches can fit into multiples of rack units within a Rack 248.
[0014] Server 246 can contain a standalone operating system configured to provide server functions, or the servers can be virtualized. A virtualized server can be controlled by a Virtual Machine Manager (VMM), hypervisor, and / or orchestrator and host one or more virtual machines, virtual servers, or virtual devices. Workload clusters 218A-218B can be located in a single data center or in different geographic data centers. Depending on contractual agreements, some servers may be dedicated to specific enterprise customers or tenants, while other servers may be shared.
[0015] The various devices in a data center can be interconnected via a Switching Fabric 270, which may include one or more high-speed routing and / or switching devices. The Switching Fabric 270 can provide north-south traffic 202 (e.g., traffic to and from a wide area network (WAN), such as the internet) and east-west traffic 204 (e.g., traffic throughout the data center). Historically, north-south traffic 202 constituted the majority of network traffic, but as web services have become more complex and distributed, the volume of east-west traffic 204 has increased. In many data centers, east-west traffic 204 now represents the largest share of traffic. Furthermore, traffic volumes may increase further with the growing capabilities of the server 246.For example, the Server 246 can offer multiple processor slots, each capable of accommodating a processor with four to eight cores and sufficient RAM for those cores. This allows the Server 246 to host a number of virtual machines (VMs), which can then act as a source of traffic.
[0016] To handle the high traffic volume in a data center, a high-performance implementation of Switching Fabric 270 can be deployed. The Switching Fabric 270 implementation shown is an example of a flat network where Server 246 can have a direct connection to a top-of-rack switch (ToR switch 220A-220B) (e.g., a "star" configuration). ToR switch 220A can connect to workload cluster 218A, while ToR switch 220B can connect to workload cluster 218B. ToR switch 220A-220B can be coupled to a core switch 260. This two-tier flat network architecture is shown as an illustrative example, and other architectures can also be used, such as...Three-tiered star or leaf rib topologies (also called "Fat Tree" topologies) based on the "Clos" architecture, hub-and-spoke topologies, mesh topologies, ring topologies or 3D mesh topologies, to name just a few examples.
[0017] The Switching Fabric 270 can be deployed through any suitable connection using any suitable connection protocol. For example, the Server 246 can contain a Fabric interface (FI), a Network Interface Card (NIC), or another host interface. The host interface itself can be connected to one or more processors via a connection or bus such as PCI, PCIe, or similar, and in some cases, this connection bus can be considered part of the Switching Fabric 270. The Switching Fabric can also utilize physical PCIe connections to implement more advanced protocols such as Compute Express Link (CXL).
[0018] The connectivity technology can be provided through a single connection or a hybrid connection. For example, PCIe might handle on-chip communication, 1Gb or 10Gb copper Ethernet provides relatively short connections to a 220A-220B toroidal switch, and optical cabling enables relatively long connections to the 260 core switch. Connectivity technologies include, for example, Ultra Path Interconnect (UPI), Fibre Channel, Ethernet, Fibre Channel over Ethernet (FCoE), InfiniBand, PCIe, NVLink, and fiber optic cables, to name just a few. Some are better suited for certain applications or functions than others, and selecting the appropriate fabric for a given application is a task for a specialist.
[0019] In one embodiment, the switching elements of the Switching Fabric 270 are configured to implement switching techniques to improve network performance in high-utilization scenarios. Examples of advanced switching techniques include adaptive routing, adaptive fault resolution, and adaptive and / or telemetry-based congestion control.
[0020] Adaptive Routing allows a ToR switch 220A-220B and / or a Core switch 260 to select the outbound port to which traffic is directed based on the load of the selected port, provided unrestricted port selection is enabled. An adaptive routing table can configure the forwarding tables of the switches in the Switching Fabric 270 to select between multiple ports between switches when multiple connections exist between a given set of switches in an adaptive routing group. Adaptive fault correction (e.g., self-healing) allows the automatic selection of an alternative port if the port selected by the forwarding table is in a failed or inactive state, enabling rapid recovery in the event of a switch-to-switch port failure.A notification can be sent to neighboring switches when adaptive routing or adaptive troubleshooting becomes active on a particular switch. With adaptive congestion control, a switch is configured to send a notification to neighboring switches when the congestion of a port on that switch exceeds a configured threshold. This can then prompt those neighboring switches to adaptively switch to uncongested ports on that switch or to switches connected via an alternative route to the destination.
[0021] Telemetry-based congestion control utilizes real-time monitoring of telemetry from network devices, such as switches within the Switching Fabric 270, to detect when congestion begins to impact the Switching Fabric 270's performance and to proactively adjust the switching tables within the network devices to prevent or mitigate the impending congestion. A ToR Switch 220A-220B and / or Core Switch 260 can implement an integrated telemetry-based congestion control algorithm or provide an API for implementing a programmable telemetry-based congestion control algorithm. A continuous feedback loop can be implemented in which the telemetry-based congestion control system continuously monitors the network and adjusts traffic flow in real time based on the ongoing telemetry data.The telemetry-based congestion control system can learn and adapt by adjusting to changing network conditions and improving its congestion control strategies based on historical data and trends.
[0022] Note, however, that while high-end fabrics are provided herein for illustrative purposes, Switching Fabric 270 can generally incorporate any suitable link or bus for the specific application, including legacy links used to implement local area networks (LANs), synchronous optical networks (SONETs), asynchronous transmission mode (ATM) networks, wireless networks such as Wi-Fi and Bluetooth, 5G wireless, DSL connections, MoCA, or similar wired or wireless networks. It is also expressly assumed that new networking technologies will emerge in the future that complement or replace some of those listed herein, and that such future network topologies and technologies may be, or become, part of Switching Fabric 270.
[0023] Fig. is a block diagram of a portion of a Data Center 300 according to one or more examples in this description. In the implementations described herein, the example Data Center 300 can consist of Fig. Components for implementing prompt adaptation for AI models in programmable network interface devices as discussed below with reference to Fig. Included. The depicted part of Data Center 300 is not intended to contain all components of a data center. The depicted part can be duplicated multiple times within Data Center 300, and / or Data Center 300 can contain parts beyond those depicted, depending on the capacity and functionality to be provided by Data Center 300. Data Center 300 can incorporate components of System 200's Data Center in various configurations. Fig. It could be included or another data center.
[0024] The Data Center 300 features a set of logic elements that form a multitude of nodes, with each node potentially being provided by a physical server, a group of servers, or other hardware. The server can also host one or more virtual machines, depending on the application. A Fabric 370 is provided to connect various aspects of the Data Center 300. The Fabric 370 can be provided by any suitable interconnect technology, including but not limited to InfiniBand, Ethernet, PCIe, or CXL. The Data Center 300's Fabric 370 can be a version of the System 200's Switching Fabric 270. Fig. be and / or contain elements thereof. The Fabric 370 of the Data Center 300 can connect elements of the data center, including server nodes (e.g., memory server node 304, heterogeneous compute server node 306, CPU server node 308, storage server node 310), accelerators 330, gateways 340A-340B to other fabrics, fabric architectures or interconnect technologies, and an Orchestrator 360.
[0025] The server nodes of data center 300 can include, but are not limited to, a memory server node 304, a heterogeneous compute server node 306, a CPU server node 308, and a storage server node 310. The heterogeneous compute server node 306 and a CPU server node 308 can perform independent operations for different tenants or cooperatively perform operations for a single tenant. The heterogeneous compute server node 306 and a CPU server node 308 can also host virtual machines that provide virtual server functionality to the data center's tenants.
[0026] The server nodes can be connected to the Fabric 370 via a Fabric 372 interface. The specific type of Fabric 372 interface used depends, at least in part, on the technology or protocol used to implement the Fabric 370. For example, if the Fabric 370 is an Ethernet fabric, the Fabric 372 interface can be an Ethernet network interface controller. If the Fabric 370 is a PCIe-based fabric, the Fabric interfaces can be PCIe-based connections. If Fabric 370 is an InfiniBand fabric, the Fabric interface 372 of the heterogeneous compute server node 306 and a CPU server node 308 can be a Host Channel Adapter (HCA), while the Fabric interface 372 of the memory server node 304 and the storage server node 310 can be a Target Channel Adapter (TCA).The TCA functionality can be an implementation-specific subset of the HCA functionality. The various fabric interfaces can be implemented as IP (Intellectual Property) blocks that can be inserted as a modular unit into an integrated circuit, just like other circuits within the Data Center 300.
[0027] The heterogeneous compute server node 306 has multiple CPU sockets, each capable of accommodating a CPU 319. The CPU 319 can be, but is not limited to, an Intel® Xeon™ multi-core processor. For example, the CPU 319 could also be a data center-class ARM® multi-core CPU, such as an NVIDIA® Grace™ CPU. The heterogeneous compute server node 306 has memory devices 318 for storing runtime data and storage devices 316 for persistent data storage in non-volatile memory. The heterogeneous compute server node 306 can perform heterogeneous processing through the presence of GPUs (e.g., GPU 317), which can... B. can be used for performing high-performance computing (HPC) operations, media servers, cloud gaming servers and / or machine learning operations.In one configuration, the GPUs and CPUs of the heterogeneous compute server node 306 can be interconnected via connection technologies such as PCIe, CXL or NVLink.
[0028] The CPU server node 308 has a variety of CPUs (e.g., CPU 319), memory (e.g., memory devices 318), and storage (memory devices 316) to run applications and other program code that provide server functions, such as web servers or other types of functions that clients of the CPU server node 308 can access remotely. The CPU server node 308 can also run program code that provides services or microservices enabling complex enterprise functions.Fabric 370 is provided with sufficient throughput to allow concurrent access to CPU server node 308 by a large number of clients, while simultaneously maintaining sufficient throughput for use by heterogeneous compute server node 306 and enabling the use of memory server node 304 and storage server node 310 by both heterogeneous compute server node 306 and CPU server node 308. Furthermore, in a configuration where CPU server node 308 primarily relies on distributed services provided by memory server node 304 and storage server node 310, the memory and storage capacity of CPU server node 308 may not be sufficient for all operations it may require.Instead, a large pool of high-speed or specialized memory can be dynamically provisioned between a number of nodes, so that each node has access to a large resource pool, but these resources do not remain unused when the respective node is not using them. A distributed architecture of this type is possible due to the high speeds and low latency offered by the Fabric 370 of modern data centers and can be advantageous because resources do not need to be over-provisioned for the server nodes.
[0029] Memory server node 304 can include memory node 305 with memory technologies suitable for storing data used during program code execution by heterogeneous compute server node 306 and CPU server node 308. Memory node 305 can include volatile memory modules, such as DRAM modules, and / or non-volatile memory technologies capable of operating at speeds similar to DRAM (such as 3D XPoint memory), providing sufficient throughput and latency performance metrics to be used as a layer of system memory at runtime. Memory server node 304 can be connected to heterogeneous compute server node 306 and / or CPU server node 308 via technologies such as CXL.mem, which enable memory access from a host to a device.In such a configuration, CPU 319 of heterogeneous compute server node 306, CPU server node 308, can connect to memory server node 304 and access memory nodes 305 of memory server node 304 in a similar way to how CPU 319 of heterogeneous compute server node 306 can access the device memory of a GPU within heterogeneous compute server node 306. For example, memory server node 304 can grant remote memory access (RDMA) to memory node 305, allowing CPU server node 308, for instance, to access memory resources of memory server node 304 via structure 370 using DMA operations, much like the CPU would access its own onboard memory.
[0030] Memory server node 304 can be used by heterogeneous compute server node 306 and CPU server node 308 to extend the runtime memory available for memory-intensive activities such as training machine learning models. A multi-tiered memory system can be enabled, allowing model data to be swapped in and out of memory devices 318 of heterogeneous compute server node 306 and into memory on memory server node 304, which offers higher performance and / or lower latency than local memory (e.g., memory devices 316). During workload setup, the entire workload set can be loaded into one or more of memory nodes 305 on memory server node 304 and into memory devices 318 on heterogeneous compute server node 306 during the execution of the heterogeneous workload.
[0031] Storage Server Node 310 provides storage functionality for Heterogeneous Compute Server Node 306, CPU Server Node 308, and potentially Memory Server Node 304. Storage Server Node 310 can provide networked disk storage (NBOD), program flash memory (PFM), a redundant array of independent disks (RAID), a redundant array of independent nodes (RAIN), network attached storage (NAS), or other non-volatile memory solutions. In one configuration, Storage Server Node 310 can be coupled with Heterogeneous Compute Server Node 306, CPU Server Node 308, and / or Memory Server Node 304, such as via NVMe-oF, enabling the implementation of the NVMe protocol over Fabric 370. In such configurations, the Fabric interfaces 372 of these servers can be intelligent interfaces containing hardware to accelerate NVMe-oF operations.
[0032] The Accelerators 330 within the Data Center 300 can provide various accelerated functions, including hardware or coprocessor acceleration for functions such as packet processing, encryption, decryption, compression, decompression, network security, or other accelerated functions within the data center. In some examples, the Accelerators 330 may include deep learning accelerators, such as neural processing units (NPUs), which can offload matrix multiplication operations or other neural network operations from the heterogeneous compute server node 306 or the CPU server node 308. In some configurations, the Accelerators 330 may reside in a dedicated accelerator server or be distributed across the various server nodes of the Data Center 300.For example, an NPU can be directly connected to one or more CPU cores within the heterogeneous Compute Server Node 306 or the CPU Server Node 308. In some configurations, the Accelerators 330 can have or be integrated with intelligent network controllers, infrastructure processing units (IPUs), data processing units (DPUs), or edge processing units (EPUs), which combine network controller functions with accelerator, processor, or coprocessor functions.
[0033] In one configuration, the data center 300 can have gateways 340A-340B connecting Fabric 370 to other fabrics, fabric architectures, or interconnect technologies. For example, if Fabric 370 is an InfiniBand fabric, gateways 340A-340B can be gateways to an Ethernet fabric. If Fabric 370 is an Ethernet fabric, gateways 340A-340B can contain routers to forward data to other parts of the data center 300 or to a larger network, such as the internet. For example, a first gateway 340A can connect to another network or subnet within the data center 300, while a second gateway 340B can be a router to the internet.
[0034] Orchestrator 360 manages the provisioning, configuration, and operation of network resources within Data Center 300. Orchestrator 360 can consist of hardware or software running on a dedicated orchestration server. It can also be included in software running, for example, on CPU server node 308, which configures the software-defined networking (SDN) functionality of components within Data Center 300. In various configurations, Orchestrator 360 can enable the automated provisioning and configuration of Data Center 300 components by managing network resource allocation and template-based provisioning.Template-based provisioning is a method for deploying and managing IT resources using predefined templates. These templates can be based on standard templates used by government agencies, service providers, financial institutions, standards organizations, or customers. Service Level Agreements (SLAs) or Service Level Obligations (SLOs) can also be defined within the template. Orchestrator 360 can perform, but is not limited to, functions such as load balancing and traffic engineering, network segmentation, security automation, real-time telemetry monitoring, and adaptive switching management, including telemetry-based adaptive switching.In some configurations, Orchestrator 360 can also provide multi-tenancy and virtualization support by enabling the management of virtual networks, including the creation and deletion of virtual LANs (VLANs) and virtual private networks (VPNs), and tenant isolation for multi-tenant data centers.
[0035] Fig. This section demonstrates programmable redirect elements and adaptive routing. The implementations described herein showcase exemplary programmable redirect elements and adaptive routing from Fig. as part of the implementation of prompt adaptation for AI models in programmable network interface devices according to the discussion below with reference to Fig. be used. Fig. shows a forwarding element that has a control plane and a programmable data plane. Fig. shows a network with switching devices configured for adaptive routing and telemetry-based congestion control. Fig. shows an InfiniBand switch with IB interfaces with multiple ports.
[0036] Fig. Figure 400 shows a forwarding element that can be configured to forward data messages within a network based on a user-provided program. In some embodiments, the program includes instructions for forwarding data messages as well as for performing other processes such as firewalling, denial-of-service protection, and load balancing. The forwarding element 400 can be any type of forwarding device, including but not limited to a switch, a switch chip, a router, or a bridge. The forwarding element 400 can forward data messages associated with various technologies, such as but not limited to Ethernet, Ultra Ethernet, InfiniBand, or NVLink.
[0037] In various network configurations, the forwarding element is used as a non-edge forwarding element within the network to forward data messages from a source device to a destination device. In network configurations, the forwarding element 400 is used as an edge forwarding element at the network edge to connect to computing devices (e.g., standalone or host computers) that serve as sources and destinations of the data messages. As a non-edge forwarding element, the forwarding element 400 forwards data messages between forwarding elements within the network, for example, via an intermediate network fabric. As an edge forwarding element, the forwarding element 400 forwards data messages to and from edge computing devices, to other edge forwarding elements, and / or to non-edge forwarding elements.
[0038] The forwarding element 400 contains circuitry for implementing a data plane 402, which performs the forwarding operations of the forwarding element 400 to relay data messages received by the forwarding element to other devices. The forwarding element 400 also includes circuitry for implementing a control plane 404, which configures the data plane circuitry. Furthermore, the forwarding element 400 includes physical ports 406, which receive data messages from and send data messages to devices outside the forwarding element 400. The data plane 402 includes ports 408, which receive data messages from the physical ports 406 for processing. The data messages are processed and relayed to another port on the data plane 402, which is connected to another physical port of the forwarding element 400.In addition to the physical ports of the forwarding element 400, some of the ports 408 on data plane 402 may also be connected to other modules of data plane 402.
[0039] The data plane features programmable packet processor circuitry that provides several programmable message processing stages. These stages can be configured to perform the forwarding operations of the forwarding element 400 in order to process data messages and forward them to their destinations. These message processing stages perform these forwarding operations by processing data tuples (e.g., message headers) associated with the data messages received from the data plane 402 to determine how the messages should be forwarded. The message processing stages include matching action units (MAUs) that attempt to match data tuples (e.g., header vectors) from messages with tables that specify which actions should be performed on the data tuples.In some embodiments, the table sets are populated by the control plane 404 and are not known when the data plane is configured to execute a program provided by a network user. The programmable message-handling circuits are grouped into multiple message-handling pipelines. The message-handling pipelines can be input or output pipelines before or after the traffic management stage of the routing element, which directs the messages from the input pipelines to the output pipelines.
[0040] The hardware specifications of data plane 402 depend on the communication protocol implemented via the forwarding element 400. Ethernet switches use application-specific integrated circuits (ASICs) designed to process Ethernet frames and the TCP / IP protocol stack. These ASICs are optimized for a wide range of traffic types, including unicast, multicast, and broadcast. Ethernet switch ASICs are typically designed for a balance of cost, power consumption, and performance, although high-end Ethernet switches may also support advanced features such as deep packet inspection and enhanced quality of service (QoS). InfiniBand switches use specialized ASICs designed for extremely low latency and high throughput.These ASICs enable features such as optimized processing of the InfiniBand protocol and offer support for RDMA and other functions that utilize precise timing and high-speed data processing. However, high-end Ethernet switches can support RoCE (RDMA over Converged Ethernet), which offers similar advantages to InfiniBand but has higher latency compared to native InfiniBand RDMA.
[0041] The 400 forwarding element can also be configured as an NVLink switch (e.g., NVSwitch), which connects multiple GPUs using the NVLink connection protocol. When configured as an NVLink switch, the 400 forwarding element can provide GPU servers with higher GPU-to-GPU bandwidth compared to GPU servers connected via InfiniBand. An NVLink switch can reduce network traffic hotspots that can occur when interconnected GPU-equipped servers perform tasks such as distributed neural network training.
[0042] When data plane 402, in coordination with a program running on data plane 402 (e.g., a program written in the P4 language), performs message or packet forwarding operations on incoming data, control plane 404 generally determines how messages or packets are forwarded. The behavior of a program running on data plane 402 is partly determined by control plane 404, which populates match-action tables with specific forwarding rules. The forwarding rules used by the program running on data plane 402 are independent of the data plane program itself. In a configuration, the control plane can be coupled to an administration port 410, which allows the administrator to configure the forwarding element 400.The data connection established via management port 410 is separate from the data connections for input and output data ports. In one configuration, management ports 410 can be connected to a management level 405, which facilitates administrative access to the device, enables analysis of the device's health and integrity, and allows for device reconfiguration. Management level 405 can be part of control level 404 or communicate directly with it. In one implementation, the administrator does not have direct access to the components of control level 404. Instead, information is gathered by management level 405, and changes to control level 404 are made by management level 405.
[0043] Fig. Figure 420 shows a network 420 with switches 432A-432E supporting adaptive routing and telemetry-based congestion control. The network 420 can be implemented using a variety of the communication protocols described herein. In one embodiment, the network 420 is implemented using the InfiniBand protocol. In another embodiment, the network 420 is an Ethernet, converged Ethernet, or Ultra-Ethernet network. The network 420 can incorporate aspects of Fabric 370 from Figure 420. Fig. Included. Switches 432A-432E may contain an implementation of the forwarding element 400 from Fig. Network 420 provides packet-based communication for multiple nodes (e.g., node 424, node 446), including a source node 422 and a destination node 442, for a data transmission to be carried out over Network 420. The packets of a flow are routed through Network 420 along a route that traverses the switches (Switch 432A-432E) and links (Link 426A-426B, 427A-427B, 428, 429A-429B, 430A-430B) of Network 420. In an InfiniBand application, the switches and links belong to a specific InfiniBand subnet, which is managed by a subnet manager (SM) that may be located in one of the switches (e.g., Switch 432D). Source node 422 and destination node 442 are the source and destination nodes for a sample data flow. Depending on the configuration of network 420, packets can flow from any node to any other node via one or more paths.
[0044] The switches 432A-432E comprise a data plane 402, a control plane 404, a management plane 405 and physical ports 406, as in the forwarding element 400 from Fig. A 404 control plane processor can be used to implement adaptive routing techniques to adjust the route between source node 422 and destination node 442 based on the current network state. During network operation, the route from source node 422 to destination node 442 may become unsuitable or impaired in its ability to transmit packets due to various events such as congestion, link failures, or head-of-line blocking. Should such a scenario occur, switches 432A-432E can be configured to dynamically adjust the route of packets flowing along a compromised path.
[0045] An adaptive routing (AR) event can be detected by any switch along a compromised route, for example, when the switch attempts to output packets to a specific output port. For instance, sample data from source node 422 to destination node 442 might traverse connections across the network's switches. An AR event might be detected by switch 432D for connection 429B, for example, in response to congestion or a connection failure associated with connection 429B. Upon detection of the AR event, switch 432D, as the detecting switch, generates an adaptive routing notification (ARN) with an identifier that distinguishes an ARN packet from other packet types. In various implementations, the ARN includes parameters such as an identifier for the detecting switch, the type of AR event, the source and destination addresses of the flow that triggered the AR event, and / or other appropriate parameters.The detecting switch sends the ARN backward along the route to the preceding switches. The ARN can contain a request to the notified switches to change their route to avoid passing through the detected switch. The notified switch can then check if its routes can be modified to bypass the detected switch. Otherwise, the switch forwards the ARN to the preceding switch along the route. In this scenario, switch 432B is unable to bypass switch 432D and forwards the ARN to switch 432A. Switch 432A can determine that the route to destination node 442 should be adjusted via connection 427A to switch 432C. Switch 432C can reach switch 432E via connection 429A, allowing packets from source node 422 to reach destination node 442 while bypassing the AR event with respect to connection 429B.
[0046] In various configurations, the 420 network can also adapt to congestion scenarios via programmable data planes within switches 432A-432E. These data planes are capable of executing programs to implement network-internal congestion control algorithms (CCAs) for TCP over Ethernet-based fabrics. Using in-band network telemetry (INT), the programmable data planes in switches 432A-432E can detect when a port or link along a route is congested and proactively attempt to reroute packets along alternative paths. For example, switch 432A can balance traffic to destination node 442 between links 427A and 427B based on the degree of congestion on the routes downstream of these links.
[0047] Fig. shows an InfiniBand switch 450, which is an implementation of the forwarding element 400 from Fig. The InfiniBand Switch 450 features a programmable data plane and is configurable to perform adaptive routing and telemetry-based congestion control, as described herein. The InfiniBand Switch 450 includes the multiport IB interfaces 460A-460D and the core switch logic 480. The multiport IB interfaces 460A-460D have multiple ports. In one embodiment, a single physical interface (IB PHY 453) with input and output buffers is provided, connected to the respective ports. In another embodiment, the port has a separate physical interface. The port can be coupled, for example, to an HCA 452, a TCA 461, or another InfiniBand Switch 432. The multiport IB interfaces 460A-460D feature a coupling field 454 configured to selectively couple input and output port buffers to the local working memory 456.The 454 coupling field is a non-blocking coupling field that enables direct switching with low latency and fixed or variable packet size.
[0048] Local memory 456 contains several queues, including an outer receive queue 462, an outer send queue 463, an inner receive queue 464, and an inner send queue 465. The outer queues are used for data received at a specific Multiport IB interface and intended to be sent back over the same Multiport IB interface. The inner queues are used for data that is to be forwarded over a different Multiport IB interface than the one through which the data was received. Other types of queue configurations can be implemented in local memory 456. For example, different queues can exist to support multiple traffic classes, either based on individual ports, shared ports, or a combination thereof.The multi-port IB interfaces 460A-460D feature a power management circuit 455 that can adjust the power state of the circuits within the respective multi-port IB interface. Additionally, power management logic performing similar operations can be implemented as part of the core switch logic.
[0049] The multi-port IB interfaces 460A-460D contain the packet processing and switching logic 458, which is generally used to perform aspects of packet processing and / or switching operations that are performed locally across multiple ports, rather than across the IB switch as a whole. Depending on the implementation, the packet processing and switching logic 458 may be configured to perform a subset of the operations of the packet processing and switching logic 478 within the core switch logic 480, or it may be configured with the full functionality of the packet processing and switching logic 478 within the core switch logic 480. The processing functionality of the packet processing and switching logic 458 can vary depending on the complexity of the operations and / or the speed at which the operations need to be performed.The packet processing and switching logic of the 458 can, for example, include processors ranging from microcontrollers to multi-core processors. Different types or architectures of multi-core processors can also be used. Additionally, some of the packet processing operations can be implemented using embedded hardware logic.
[0050] The core switch logic 480 comprises a crosspoint switch 482, a memory 470, a subnet management agent (SMA 476), and packet processing and switching logic 478. The crosspoint switch 482 is a non-blocking, low-latency crosspoint switch that connects the multiport IB interfaces 460A-460D and provides a connection to the memory 470. The memory 470 contains receive queues 472 and transmit queues 474. In one embodiment, packets to be routed between the multiport IB interfaces 460A-460D can be received by the crosspoint switch 482, stored in one of the receive queues 472, processed by the packet processing and switching logic 478, and stored in a transmit queue 474 for transmission to the outgoing multiport IB interface.In implementations that do not use the multiport IB interfaces 460A-460D, the core switch logic 480 and the crosspoint switch 482 switch packets directly between the I / O buffers connected to the respective ports with the receive queues 472 and send queues 474 within the main memory 470.
[0051] The Packet Processing and Switching Logic 478 contains programmable functions and can execute data-plane programs across a variety of multi-core processor types and architectures. The Packet Processing and Switching Logic 478 represents the applicable circuitry and logic for implementing switching and packet processing operations beyond those performed at the ports themselves. The processing elements of the Packet Processing and Switching Logic 478 execute software and / or firmware commands configured to implement packet processing and switching operations. This software and / or firmware can be stored in non-volatile memory on the switch itself. Alternatively, the software can be downloaded or updated over a network in conjunction with the initialization of the InfiniBand Switch 450.
[0052] The SMA 476 is configurable for managing, monitoring, and controlling the functions of the InfiniBand 450 switch. The SMA 476 also acts as an agent of the subnet manager (SM) for the subnet connected to the InfiniBand 450 switch and communicates with it. The SM is the entity that discovers the devices within the subnet and performs periodic checks of the subnet to detect changes in the subnet topology. An SMA within a subnet can be designated as the primary SMA for that subnet and act as the SM. Other SMAs within the subnet then communicate with this SMA. Alternatively, the SMA 476 can cooperate with other SMAs in the subnet and function as a distributed SM. In some embodiments, the SMA 476 incorporates its own circuitry and logic, such as... B. a microcontroller, a single multi-core processor or a multi-core processor, or is executed on these.In other embodiments, the SMA 476 is implemented via software and / or firmware instructions executed on a multi-core processor or other processing element that is part of a processor or other processing element used to implement the packet processing and switching logic 478.
[0053] The embodiments are not specifically limited to implementations with IB multiple interfaces 460A-460D. In one embodiment, the port is connected to its own receive and transmit buffers, with the crosspoint switch 482 configured to connect these buffers to the receive queues 472 and the transmit queues 474 in the main memory 470. Packet processing and switching are then primarily performed by the packet processing and switching logic 478 of the core switch logic 480.
[0054] Fig. represent an example of network interface devices. In the implementations described herein, the exemplary network interface device can be derived from Fig. a prompt adaptation for AI models in programmable network interface devices according to the discussion below with reference to Fig. implement. Fig. shows a network interface device 500 that can be configured as a Smart Ethernet device. Fig. shows a network interface device 550 that can be configured as an InfiniBand channel adapter.
[0055] As in Fig. As shown, the network interface device 500 can, in one configuration, include a transceiver 502, a transmit queue 507, a receive queue 508, a memory 510, a bus interface 512, and a DMA engine 526. The network interface device 500 can also include a SoC / SIP 545, which has processors 505 for implementing the intelligent network interface device functionality and accelerators 506 for various accelerated functions such as NVMe-oF or RDMA. The specific characteristics of the network interface device 500 depend on the protocol implemented through the network interface device 500.
[0056] In various configurations, the Network Interface Device 500 can be configured to connect to networks, including but not limited to Ethernet, including Ultra Ethernet. However, the Network Interface Device 500 can also be configured as an InfiniBand or NVLink interface by modifying various components. For example, the Transceiver 502 can receive and send packets in accordance with the InfiniBand, Ethernet, or NVLink protocols. Other protocols can also be used. The Transceiver 502 can receive and send packets to and from a network over a network medium. The Transceiver 502 can include a PHY circuit 514 and a Media Access Control (MAC) circuit 516.The PHY circuit 514 can include encoding and decoding circuitry to encode and decode data packets according to applicable physical layer specifications or standards. The MAC circuit 516 can be configured to assemble the data to be transmitted into packets containing destination and source addresses along with network control information and error detection hash values.
[0057] The SoC / SIP 545 can contain processors that can be any combination of a CPU, a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other hardware device configurable for instruction execution. For example, an intelligent network interface can provide packet processing capabilities at the network interface using 505 processors. The configuration of the 505 processors' operation, including the programmable data plane processors, can be programmed using protocol-independent packet processors (P4), C, Python, Broadcom Network Programming Language (NPL), x86- or ARM-compatible executable binaries, or other executable binaries.
[0058] The packet distributor 524 can distribute received packets for processing by multiple CPUs or cores using time slot allocation. An interrupt coalescing circuit 522 can perform interrupt moderation, where it waits for multiple packets to arrive or a timeout before generating an interrupt to the host system to process the received packets. Receive segment coalescing (RSC) can be performed by the network interface 500, where portions of incoming packets are grouped into segments of a single packet. The network interface device 500 can then provide this grouped packet to an application.A DMA engine 526 can copy a packet header, packet payload, and / or descriptor directly from host memory to the network interface, or vice versa, instead of copying the packet to an intermediate buffer on the host and then performing another copy operation from the intermediate buffer to the destination buffer. Memory 510 can be volatile or non-volatile memory of any type and can store any queue or instructions used to program the network interface device 500. The send queue 507 can contain data or references to data to be transmitted over the network interface. The receive queue 508 can contain data or references to data received from a network over the network interface.The 520 descriptor queues can contain descriptors that point to data or packets in the 507 send queue or the 508 receive queue. The 512 bus interface can provide an interface to the host device. For example, the 512 bus interface can be PCI Express compatible, although other connection standards can also be used.
[0059] As in Fig. As shown, a network interface device 550 can be configured as an implementation of the network interface device 500 to implement an InfiniBand HCA. The network interface device 550 has network ports 552A-552B, memory 554A-554B, a PCIe interface 558, and an integrated circuit 556 containing hardware, firmware, and / or software for implementing, managing, and / or controlling the HCA functionality. In one implementation, the integrated circuit includes a hardware transport engine 560, an RDMA engine 562, congestion control logic 563, virtual endpoint logic 564, swap engines 566, QoS logic 568, GSA / SMA logic 569, and a management interface. Different implementations of the network interface device 550 may include additional components or omit some components.A network interface device 550 configured as a TCA contains an implementation-specific subset of the functionality of an HCA. The integrated circuit 556 contains programmable and fixed functional hardware for implementing the described functionality.
[0060] While the illustrated implementation of the Network Interface Device 550 uses a PCIe 558 interface, other implementations may use different interfaces. For example, the Network Interface Device 550 may use an Open Compute Project (OCP) mezzanine connector. Furthermore, the PCIe 558 interface can also be configured with a multi-host solution, allowing multiple compute or storage hosts to connect to the Network Interface Device 550. The PCIe 558 interface can also support a technology that enables direct PCIe access to multiple CPU sockets, eliminating the need for network traffic to traverse the interprocessor bus of a multi-socket server motherboard when the server includes the Network Interface Device 550.
[0061] The 550 network interface device implements endpoint elements of the InfiniBand architecture, which is based on queue pairs and RDMA. InfiniBand offloads traffic control from software by using execution queues (e.g., work queues) that are initiated by a software client and managed in hardware. The communication endpoints contain a queue pair (QP) with a send queue and a receive queue. A QP is a memory-based abstraction where communication occurs between memory-to-memory transfers between applications or between applications and devices. Communication with the QPs takes place over virtual lanes of network ports 552A-552B, which allow multiple independent data flows to use the same connection, with separate buffering and flow control for each flow.
[0062] Communication occurs via channel I / O, where a virtual channel directly connects two applications residing in separate address spaces. The Hardware Transport Engine 560 contains hardware logic for performing transport-layer operations across the QP for an endpoint. The RDMA Engine 562 leverages the Hardware Transport Engine 560 to perform RDMA operations between endpoints. Implementing RDMA operations in hardware, the RDMA Engine 562 allows an application to read and write to the memory of a remote system without kernel intervention or unnecessary data copying by permitting an endpoint on a communication channel to directly place information into the memory of another endpoint. The Virtual Endpoint Logic 564 manages the operation of a virtual endpoint for channel I / O, i.e., a virtual instance of a QP used by an application.The virtual endpoint logic 564 assigns the QPs to the virtual address space of an application connected to a virtual endpoint.
[0063] Congestion control logic 563 performs operations to reduce the occurrence of congestion on a channel. In various implementations, congestion control logic 563 can perform flow control over a channel to limit congestion at the destination of a data transmission. Congestion control logic 563 can perform flow control at the link level to manage source congestion on virtual links of network ports 552A-552B. In some implementations, congestion control logic can take actions to limit congestion at intermediate points (e.g., IB switches) along a channel.
[0064] The 566 Offloading Engines enable the offloading of network tasks that would otherwise be performed in software on the 550 Network Interface Device. The 566 Offloading Engines can support the offloading of operations, including, but not limited to, offloading receive-side scaling from a device driver or stateless network operations, such as for TCP implementations over InfiniBand, like stateless TCP / UDP / IP offloading or VXLAN offloading. The 566 Offloading Engines can also offload operations of an interrupt coalescing circuit 522 of the 500 Network Interface Device. Fig. Implement 5A. The 566 swap engines can also be configured to support offloading NVMe-oF or other storage acceleration operations from a CPU.
[0065] The QoS logic 568 can perform QoS operations, including the QoS functionality built into InfiniBand's basic service provisioning mechanism. QoS logic 568 can also implement enhanced InfiniBand QoS, such as fine-grained end-to-end QoS. It can implement queueing services and management to prioritize traffic flows and ensure service levels or bandwidth according to flow priority. For example, QoS logic 568 can configure virtual lane arbitration for network ports 552A-552B according to flow priority. QoS logic 568 can also work in conjunction with congestion control logic 563.
[0066] The GSA / SMA logic 569 implements General Services Agent (GSA) operations for managing the network interface device 550 and the InfiniBand structure, as well as for performing subnet management agent operations. GSA operations include device-specific management tasks such as querying device attributes, configuring device settings, and controlling device behavior. The GSA / SMA logic 569 can also implement SMA operations, including a subset of those available from the SMA 476 of the InfiniBand switch 450. Fig. can be executed. For example, the GSA / SMA 569 logic can handle administrative requests from the subnet administrator, including requests to reset the device, update the firmware, or change configuration parameters.
[0067] The 570 management interface provides support for a hardware interface to perform out-of-band management of the 550 network interface device, such as a connection to a Board Management Controller (BMC) or a hardware debug interface.
[0068] Fig. Figure 6 is a block diagram showing a programmable network interface 600 and a data processing unit. In the implementations described herein, the exemplary programmable network interface 600 and the data processing unit can be derived from... Fig. a prompt adaptation for AI models in programmable network interface devices according to the discussion below with reference to Fig. The Programmable Network Interface 600 is a programmable network engine that can be used to accelerate network-based computing tasks in a distributed environment. The Programmable Network Interface 600 can be coupled to a host system via the Host Interface 670. The Programmable Network Interface 600 can be used to accelerate network or memory operations for the host system's CPUs or GPUs. The host system could, for example, be a node in a distributed learning system used to perform distributed training, as described in [reference to relevant documentation]. Fig. The host system can also be a data center node within a data center.
[0069] In one embodiment, access to remote storage containing model data can be accelerated by the programmable network interface 600. For example, the programmable network interface 600 can be configured to represent remote storage devices as local storage devices to the host system. The programmable network interface 600 can also accelerate RDMA operations between GPUs of the host system and GPUs of remote systems. In one embodiment, the programmable network interface 600 can enable, but is not limited to, storage features such as NVMe-oF. The programmable network interface 600 can also accelerate encryption, data integrity, compression, and other operations for remote storage on behalf of the host system, so that remote storage approaches the latencies of storage devices directly attached to the host system.
[0070] The Programmable Network Interface 600 can also perform resource allocation and management on behalf of the host system. Memory security operations can be offloaded to the Programmable Network Interface 600 and performed along with the allocation and management of remote memory resources. Network-based operations for managing access to remote memory, which would otherwise be performed by a processor in the host system, can instead be executed by the Programmable Network Interface 600.
[0071] In one embodiment, network and / or data security operations can be offloaded from the host system to the programmable network interface 600. Data center security policies for a data center node can be managed by the programmable network interface 600 rather than by the host system's processors. For example, the programmable network interface 600 can detect and mitigate an attempted network-based attack (such as a DDoS attack) on the host system, thereby preventing the attack from impacting the host system's availability.
[0072] The programmable network interface 600 can contain a system-on-a-chip (SoC / SIP 620) that runs an operating system across multiple processor cores 622. The processor cores 622 can contain general-purpose processor cores (e.g., CPU cores). In one embodiment, the processor cores 622 can also include one or more GPU cores. The SoC / SIP 620 can execute instructions stored in a memory device 640. A memory device 650 can store local operating system data. The memory device 650 and the memory device 640 can also be used to cache remote data for the host system. The network ports 660A-660B provide connectivity to a network or fabric and facilitate network access for the SoC / SIP 620 and, via the host interface 670, for the host system.In one configuration, a first network port 660A can be connected to a first forwarding element, while a second network port 660B can be connected to a second forwarding element. Alternatively, both network ports 660A-660B can be connected to a single forwarding element via a link aggregation protocol (LAG). The programmable network interface 600 can also include an I / O interface 675, such as a USB interface. The I / O interface 675 can be used to connect external devices to the programmable network interface 600 or as a debug interface. The programmable network interface 600 also includes a management interface 630, which allows software on the host device to manage and configure the programmable network interface 600 and / or the SoC / SIP 620.In one embodiment, the programmable network interface 600 can also include one or more accelerators or GPUs 645 to accept the offloading of parallel computing tasks from the SoC / SIP 620, host system, or remote systems connected via network ports 660A-660B. For example, the programmable network interface 600 can be configured with a graphics processor and participate in general-purpose or graphics computing operations in a data center environment.
[0073] One or more aspects can be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, the machine-readable medium may contain instructions representing different logics within the processor. When read by a machine, these instructions can cause the machine to create the logic to perform the techniques described herein. Such representations, known as "IP cores," are reusable logic units for an integrated circuit that may be stored on a specific machine-readable medium as a hardware model describing the structure of the integrated circuit.The hardware model can be delivered to various customers or manufacturing facilities, which load the hardware model onto manufacturing machines that produce the integrated circuit. The integrated circuit can be manufactured such that the circuit performs operations described in connection with any of the embodiments described herein.
[0074] Fig. Figure 7 is a block diagram showing an IP Core 700 development system. The IP Core 700 development system can be used to fabricate an integrated circuit to perform the fabric and data center component operations described herein. In the implementations described herein, the exemplary IP Core 700 development system can be constructed from... Fig. manufacture an integrated circuit that provides prompt adaptation for AI models in programmable network interface devices as discussed below with reference to Fig. The IP Core Development System 700 can be used to create modular, reusable designs that can be integrated into a larger design or used to build an entire integrated circuit (e.g., an integrated SoC circuit). A Design Facility 730 can generate a software simulation 710 of an IP core design in a higher-level programming language (e.g., C / C++). The software simulation 710 can be used to design, test, and verify the behavior of the IP core using a simulation model 712. The simulation model 712 can include functional, behavioral, and / or temporal simulations. A register-transfer-level design (RTL design 715) can then be generated or synthesized from the simulation model 712.The RTL-Design 715 is an abstraction of the integrated circuit's behavior, modeling the flow of digital signals between hardware registers, including the associated logic performed using the modeled digital signals. In addition to an RTL-Design 715, subordinate designs at the logic or transistor level can also be created, designed, or synthesized. Therefore, the specific details of the initial design and simulation can vary.
[0075] The RTL design 715 or an equivalent can further be synthesized by the design facility into a hardware model 720, which may be in a hardware description language (HDL) or another representation of physical design data. The HDL can further be simulated or tested to verify the IP core design. The IP core design can be stored for delivery to a manufacturing facility 765 using non-volatile memory 740 (e.g., hard disk, flash memory, or any non-volatile storage medium). The manufacturing facility 765 can be a third-party manufacturing facility. Alternatively, the IP core design can be transmitted via a wired connection 750 or a wireless connection 760 (e.g., over the Internet). The manufacturing facility 765 can then manufacture an integrated circuit based at least partially on the IP core design.The manufactured integrated circuit can be designed to perform operations in accordance with at least one embodiment described herein. Prompt adaptation for artificial intelligence models in programmable network interface devices
[0076] In highly virtualized environments, significant amounts of server resources are dedicated to processing tasks beyond user applications. These processing tasks can include hypervisors, container engines, networking and storage functions, security, and large volumes of network traffic. To handle these diverse processing tasks, programmable network interface devices (PNDs) with accelerators and network connectivity have been introduced. These programmable NPDs are referred to as infrastructure processing units (IPUs), data processing units (DPUs), edge processing units (EPUs), extended network interface devices, programmable packet processing devices, and so on.
[0077] Programmable network interface devices (PNDs) can accelerate and manage infrastructure functions using dedicated and programmable cores deployed within the devices. By acting as the host's control point for running infrastructure applications, PNDs can offload infrastructure tasks and provide an additional layer of security. Using PNDs allows the overhead associated with running infrastructure tasks to be offloaded from the server device.
[0078] In the implementations described herein, the programmable network interface devices may be referred to, for example, generally as a programmable network interface device (PNID), a network interface device, an extended network interface device, an IPU or DPU, an EPU, or a programmable packet processing device. For the purposes of this discussion, the programmable network interface device will be referred to in abbreviated form as PNID.
[0079] Data centers utilizing PNIDs can include servers hosting artificial intelligence (AI) models, such as large language models (LLMs). Custom LLMs have become a focus area for new AI systems. With increasing data volumes, the ability to leverage human feedback for reinforcement learning (e.g., reinforcement learning with human feedback (RLHF)), and the capacity for fine-tuning, there is potential to adapt and optimize LLMs to achieve improved results.
[0080] Traditional approaches to improving the performance of AI models, such as LLM models, include model fine-tuning, prompt design, and prompt tuning, to name a few. Model fine-tuning involves retuning the LLM model itself. This approach can be expensive due to the training operations across the entire LLM. Prompt design (also known as prompt engineering) involves teaching end users what questions (e.g., prompts) they should ask the LLM and how to formulate those questions to obtain more accurate results. A drawback of the prompt design approach is its difficulty in scaling to a large number of users.
[0081] The prompt tuning approach involves using a lightweight, prompt-tuned AI model (prompt-tuned model, prompt tuning model, P-tuned model, P-tuning model) that provides intelligence to augment the prompts submitted to the LLM. The prompt-tuned model can be trained to complement the prompts, thereby generating improved results from the LLM. Traditional prompt tuning approaches are typically implemented in software. Prompt tuning software can run on dedicated GPU cores and offers the flexibility to insert continuous or virtual tokens, allowing LLM models to be adapted to a variety of tasks without the risks and computational costs associated with restructuring the entire LLM model.
[0082] However, there are limitations to providing prompt voting in software. One limitation concerns large-scale dynamics, also known as large-scale automation. Prompt voting requirements can vary significantly depending on current events or geographic locations. Traditional software prompt voting approaches are unable to dynamically and / or spontaneously calibrate prompt voting models based on prompt voting occurring concurrently with other models in the same region.
[0083] Another limitation of deploying prompt voting in software is real-time augmentation. For example, current events (such as natural disasters) can feed into information that should be used to improve the prompt voting model. This type of information is typically transmitted across the network with low latency, making it difficult to incorporate such information into the prompt voting model in real time using traditional software approaches.
[0084] Another limitation of deploying prompt voting in software is cross-data center user feedback (e.g., RLHF). With traditional approaches, real-time user feedback on the effectiveness of an augmentation prompt can be lost if there is no way to synthesize and transmit it across the network, which is lacking in traditional approaches.
[0085] The implementations described herein address the aforementioned technical challenges by enabling prompt tuning for AI models in programmable network interface devices (PNIDs). In these implementations, a lightweight AI prompt tuning and augmentation model is hosted on a PNID connected to the head nodes of an AI deployment (e.g., deploying LLMs on host devices in a data center). The PNID is capable of handling real-time network propagation, which is used in conjunction with the lightweight AI prompt tuning and augmentation model hosted on the PNID. Consequently, a prompt sent to the LLM can be intercepted (or received) by the PNID for prompt tuning and augmentation as described herein.The PNID can provide a prompt augmentation microservice that works in conjunction with the lightweight AI prompt augmentation and tuning model. This prompt augmentation microservice can enhance the prompt at the PNID to include contextual information, such as real-time trend data and real-time user feedback, which can improve the results generated by the LLM.
[0086] The implementations described herein offer technical advantages. For example, they can improve the likelihood that the LLM output will provide enhanced results useful for the submitted prompt. Further details on implementing prompt adaptation for AI models in a PNID are described below. Fig. described.
[0087] Fig. Figure 800 is a block diagram illustrating an example of a Computing Environment 800 for providing prompt adaptation for AI models in programmable network interface devices according to the implementations described herein. In an implementation, the Computing Environment 800 can contain various clusters (e.g., 840A-C) of Processing Units 845A-845C (e.g., GPUs, Tensor Flow processors, other types of accelerators, etc.). A Cluster 840 can also contain one or more PNIDs 850 to enable communication between the Processing Units 845 and the Network 830. The Network 830 can further be coupled to various Storage Devices 820A-C and the Orchestrator 810.
[0088] The elements from Fig. with the same or similar names as the elements of any other figure herein describe the same elements as in the other figures, can operate or function in a similar manner, can have the same components, and can be associated with other entities such as those described elsewhere herein, but are not limited to them. Therefore, the discussion herein of any features in combination with a graphics processor also reveals, but is not limited to, a corresponding combination with the exemplary computing environment 800.
[0089] In various embodiments, components of the computing environment 800 (including requesting, destination, and / or consuming devices) can be interconnected by one or more networks (e.g., networks) that include any number of intermediate network nodes, such as routers, switches, or other computing devices. The network, the requesting device, and / or the destination device can be part of any suitable network topology, such as a data center network, a wide area network, a local area network, an edge network, or an enterprise network.
[0090] The save command can be communicated from the requesting device to the destination device, and / or data read in response to a save command can be transferred from the destination device to the consuming device via any suitable communication protocol (or protocols), such as Peripheral Component Connection (PCI), PCI Express (PCIe), CXL, Universal Serial Bus (USB), Serial Attached SCSI (SAS), Serial ATA (SATA), InfiniBand, Fibre Channel (FC), IEEE 802.3, IEEE 802.11, Ultra Ethernet, or any other current or future signaling protocol. The save command may include, among other things, commands to write data, read data, and / or erase data.In certain embodiments, the memory instructions correspond to an interface specification for a logical device (hereinafter also referred to as a network communication protocol), such as Non-Volatile Memory Express (NVMe) or Advanced Host Controller Interface (AHCI).
[0091] A computing platform, such as the Computing Environment 800, may include one or more requesting devices, consuming devices, and / or destination devices. Such devices may include one or more processing units (e.g., Processing Units 845) to generate a memory instruction, decode and process a memory instruction, and / or consume (e.g., process) data requested by a memory instruction. As used herein, the term "processing unit," "processing unit," "processor," or "processing element" may refer to any device or any part of a device that processes electronic data from registers and / or working memory to transform such electronic data into other electronic data that can be stored in registers and / or working memory.
[0092] A processing unit can consist of one or more digital signal processors (DSPs), application-specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), general-purpose GPUs (GPGPUs), accelerated processing units (APUs), field-programmable gate arrays (FPGAs), neural network processing units (NPUs), edge processing units (EPUs), vector processing units, software-defined processing units, video processing units, data processing units (DPUs), main memory processing units, storage processing units, accelerators (e.g., graphics accelerators, compression accelerators, artificial intelligence accelerators, network accelerators), controller cryptoprocessors (specialized processors that execute cryptographic algorithms within the hardware), server processors, I / O controllers, NICs (e.g.,SmartNICs), infrastructure processing units (IPUs), microcode engines, memory controllers (e.g., cache controllers, host memory controllers, DRAM controllers, SSD controllers, hard disk drive (HDD) controllers, non-volatile memory controllers, etc.), or other suitable types of processor units. Therefore, a processor unit can be referred to as an XPU.
[0093] Components of the 800 computer environment can have any suitable properties of similar components that are related to the Fig. are described. For example, the computing environment 800 can be an Ethernet, Ultra Ethernet, CXL, a network using a proprietary network protocol, or another suitable network that uses the network interface device 500, 550 from the Fig. here or the programmable network interface 600 from Fig. used.
[0094] In some embodiments, the Computing Environment 800 can be a data center or other similar environment, in which any combination of components can be placed together in a rack or shared in a data center pod. In various embodiments, the Computing Environment 800 can be a telecommunications environment, in which any combination of components can be enclosed together in curbside furniture or an enterprise wiring cabinet.
[0095] In some embodiments, the Orchestrator 810 can act as a requesting device and send memory commands, as described herein, to memory devices 820A-C, which act as destination devices. Some of these commands can read data, which is then delivered to processing units 845, which act as consuming devices. In some embodiments, a processing unit 845 or a PNID 850 can act as the requesting device. Thus, a processing unit 845 could be both a requesting device and the consuming device. In one implementation, the PNID 850 can be identical to the network interface device 500, 550 from Fig. herein or the programmable network interface 600 and the data processing unit from Fig. herein and can, for example, be referred to as an IPU or a DPU. As previously discussed, the PNID 850 in implementations herein is designed to provide prompt adaptation for AI models in programmable network interface devices as discussed below with reference to Fig. to provide.
[0096] Fig. Figure 1 is a block diagram showing an example of a PNID 900 for providing prompt customization for AI models in programmable network interface devices according to the implementations described herein. In one implementation, the PNID 900 may be identical to the one referenced in Figure 2. Fig. described PNID 850. In some implementations, PNID 900 may be identical to network interface device 500, 550. Fig. herein and / or the programmable network interface 600 and the data processing unit from Fig. in this context and can be referred to in some examples as an IPU or a DPU.
[0097] In one configuration, the PNID 900 can include a network interface 910, a memory 912, a storage 914, an accelerator / GPU 916, a host interface 920, and a SIP / SoC 930. The management interface 918 can provide a dedicated management complex for the PNID 900, including one or more processors, such as programmable circuits, and subsystems to provide secure booting, maintenance, and upgrades. In some implementations, the SIP / SoC 930 utilizes processors 935 to implement intelligent network interface device functionality. For example, the processors 935 can include CPUs, GPUs, and / or accelerators for various functionalities, such as NVMe-of or RDMA. The specific characteristics of the PNID 900 depend on the protocol implemented through it.
[0098] In various configurations, the PNID 900 can be configured to connect to networks, including but not limited to InfiniBand, Ethernet, or NVLink. The Accelerator / GPU 916 and / or Processor(s) 935 can contain processors that may be any combination of a CPU, a graphics processing unit (GPU), an FPGA (Field Programmable Gate Array), an application-specific integrated circuit (ASIC), or other programmable hardware device, enabling programming of the PNID 900. For example, an intelligent network interface can provide packet processing functions within the network interface using processors.The configuration of the operation of the PNID 900, including the programmable data plane processors, can be programmed using protocol-independent packet processors (P4), C, Python, Broadcom Network Programming Language (NPL), x86- or ARM-compatible executable binaries, or other executable binaries.
[0099] In some implementations, the PNID 900 can be communicatively coupled (e.g., via a network) to a host device, such as the host device 970, via the host interface 920. In one implementation, the host device 970 may be identical to one of the devices listed with reference to Fig. described processing units 845. The host device 970 can contain one or more CPU(s) 972, one or more GPU(s) 974 and / or one or more accelerators 976 to perform various processing tasks.
[0100] As previously discussed, data centers utilizing PNIDs, such as PNID 900, can have servers, such as Host Device 970, that host AI models, including large language models (LLMs), such as LLM 980, hosted by Host Device 970. Although the discussion herein refers to LLMS throughout, implementations described herein can also be used with other types of AI models and are not limited to application to LLMs.
[0101] Custom LLMs have become a key focus area for new AI systems. With increasing data volumes, the ability to leverage human feedback for reinforcement learning (e.g., RLHF), and the possibility of fine-tuning, there is potential to adapt and optimize LLMs, such as the LLM 980, to achieve improved results. Therefore, the implementations described herein enable prompt tuning of AI models (such as the LLM 980) by providing prompt tuning for AI models in PNIDs, such as the PNID 900. In the implementations described herein, the PNID 900 hosts a Prompt Augmentation Circuit 940 in a SIP / SoC 930. The Prompt Augmentation Circuit 940 contains a Prompt Augmentation Model 942, which is a lightweight AI prompt augmentation and tuning model.
[0102] In some implementations, the Prompt Augmentation Model 942 is connected to the head nodes of an AI deployment (e.g., deployment of LLMS on host devices in a data center). Fig. For example, Figure 1 is a block diagram representing a data center environment 1000 that includes an AI deployment supporting prompt customization for AI models according to the implementations described herein. The data center environment 1000 contains head node 1 1010A and head node 2 1010B (hereinafter collectively referred to as head node 1010). The head nodes 1010 are communicatively coupled to a respective PNID 1030A and 1030B (hereinafter collectively referred to as PNIDs 1030). The PNIDs 1030 may be identical to those referenced in the following document: Fig. The described PNID 900. The head nodes 1010 can host multiple GPUs 1015A, 1015B (collectively referred to herein as GPUs 1015) and / or CPUs 1020A, 1020B (collectively referred to herein as CPUs 1020). In some implementations, the head nodes 1010 can also host other types of hardware accelerator devices (not shown).
[0103] In the implementations described herein, the data center environment 1000 can support an AI deployment in which one or more LLMs 1040 are deployed on GPUs 1015. The head nodes 1010 can be associated with corresponding PNIDs 1030. The PNIDs 1030 can contain an augmentation circuit 1035A, 1035B (collectively referred to herein as augmentation circuit 1035). The augmentation circuit 1035 can be identical to the prompt augmentation circuit 940 from Fig. The augmentation circuit 1035 can implement a prompt augmentation model (such as the prompt augmentation model 942 from Fig. ) that is connected to the head node 1010 of the AI deployment of the data center environment 1000.
[0104] Augmentation Circuit 1035 supports prompt matching of an LLM prompt 1005 for the LLM 1040 received at PNID 1030, as discussed in more detail below. An LLM prompt augmentation 1050 generated by Augmentation Circuit 1035 at PNID 1030 can be sent from PNID 1030 to the LLM 1040 hosted by the head node 1010. Cross-network real-time propagation feedback 1060 can be transmitted between PNIDs 1030 of the data center environment 1000 to support the prompt augmentation techniques provided by Augmentation Circuit 1035, as discussed herein.
[0105] Referring back to Fig. In the implementations described herein, the PNID 900 is capable of processing real-time network propagation, which is used in conjunction with the Prompt Augmentation Model 942 hosted on the PNID 900. As a result, a prompt sent to the LLM 980 can be received (or / or intercepted) by the PNID 900 for prompt synchronization and augmentation according to the implementations described herein. In the implementations described herein, the PNID 900 can provide a Prompt Augmentation Microservice 944 that works in conjunction with the Prompt Augmentation Model 942. The Prompt Augmentation microservice 944, together with the Prompt Augmentation model 942, can cause a received LLM prompt at the PNID 900 to be enhanced to include contextual information, including real-time trend data and real-time user feedback (may herein be referred to as network data or propagated network data).This LLM prompt, which features Augmentation 950, can improve results produced by the LLM 980.
[0106] In the implementations described herein, prompt augmentation using a prompt augmentation model (prompt voting model), such as Prompt Augmentation Model 942, can involve receiving the prompt containing a task name and input text intended for the LLM 980. Task-specific virtual tokens are then retrieved based on the task name. A token can refer to words, character sets, or combinations of words and punctuation marks used by the LLM 980 to parse text. In the implementations described herein, the prompt adjustments can be generated by using the prompt voting model to retrieve the task-specific virtual tokens based on the task name. The input text is tokenized, and token embeddings are retrieved.Virtual token embeddings are then inserted between discrete token embeddings and forwarded together to the LLM 980.
[0107] The prompt augmentation provided by the Prompt Augmentation Circuit 940 in implementations herein has capabilities for the PNID 900 to accommodate cross-network, cross-data center augmentation-based feedback, in addition to incorporating a hardware-based circuit to decide how a set of LLM prompts should be augmented.
[0108] In the implementations described herein, when an LLM prompt is sent to the LLM 980, the prompt can be intercepted by the PNID 900. The PNID 900 can then use a Prompt Augmentation microservice 944 of the Prompt Augmentation Circuit 940 to compare the prompt against any real-time network trends maintained by the Prompt Augmentation Circuit 940. For example, a particular keyword might have a different context based on a recent event, and this could affect the accuracy of the LLM 980's results. The PNID 900 is able to intercept such information in real time and augment the user query (prompt) based on real-time trends as well as on a history-based learning module maintained by the Prompt Augmentation Circuit 940.Further details of the prompt augmentation of the implementations described herein are given below with reference to . Fig. described.
[0109] The implementations described herein can also be used to prevent prompt injection attacks (e.g., by enforcing them or simply by providing hints / recommendations). These attacks involve manipulated input intended to cause the LLM 980 to perform an unauthorized operation. These attacks may attempt to compromise the system through manipulated input, potentially leading to data exfiltration, social engineering, and other problems. In one implementation, the Prompt Augmentation microservice 944 could provide prompt augmentation by removing one or more layers of prompt indirection (e.g., removing tokens within the prompt) that often enable such attacks. Because the Prompt Augmentation microservice 944 runs on the PNID 900, this provides an additional layer of security.In some implementations, a fleet of PNIDs 900 can work together to share previously detected problems or attacks, so that all their "prompt injection" attacks can be managed in the same way in decentralized deployment models.
[0110] As discussed previously, the PNID 900 can provide prompt adaptation for AI models in PNIDs, as described herein. In some implementations, the PNID 900 can provide an API through which the prompt adaptation capability for AI models described herein can be implemented. In some implementations, the API can query whether a computing system supports the prompt adaptation capabilities for AI models. For example, the API can query whether the prompt adaptation capabilities for AI models described herein are provided by the PNID 900. In some implementations, the API can enable or disable such capability. For example, the API can be designed to enable and / or disable the prompt adaptation capabilities for AI models in the PNID 900.
[0111] Fig. Figure 1 is a block diagram showing a Prompt Augmentation microservice 1100 according to the implementations described herein. In one implementation, Prompt Augmentation microservice 1100 may be identical to Prompt Augmentation microservice 944 from [reference to relevant document]. Fig. The Prompt Augmentation microservice 1100 can include a Prompt Voting History 1110, Vector Database APIs 1120, a Prompt Estimator 1130, Programmable Software APIs 1140, Additional Model APIs 1150, Reinforcement Learning 1160, and Proximal Policy Optimization 1170. The implementations described herein may contain more or fewer components than those shown in the Prompt Augmentation microservice 1100.
[0112] The Prompt Voting History 1110 can be a data structure, for example, a table maintained to record the status of prompt voting by LLM(s). Table 1 below is an example of the status that can be recorded in the Prompt Voting History 1110 by the Prompt Augmentation microservice 1100 in the PNID. TABLE 1 Modell-ID GPU-Domäne LLM-Prompt Prompt-Abstimmung-Bewerter Konfidenzfaktor Geschwister-RLHF DC-Trendfaktor
[0113] For example, the Prompt Voting History 1110 can track, for each LLM model, a model ID, a GPU domain of the model, an LLM prompt, the prompt augmentation voting history for that model, a confidence metric for the prompt, whether a model is affected by "sibling models" across the data center network, whether a model is affected by current data center trends (DC trends), and so on. This information can be logged and updated using the Prompt Voting History 1110 and is used to define a prompt augmentation policy that is applied by the Prompt Augmentation microservice 1100.
[0114] Vector database APIs 1120 can include hooks for APIs to query a vector database and retrieve RAG (Retrieval Augmented Generation) updates. The prompt estimator 1130 can determine whether an augmentation should be applied to a prompt at all. In one implementation, this can be determined by estimating an LLM prompt score and using that LLM prompt score to decide whether to apply an augmentation. In some implementations, this can be determined before a prompt augmentation model is applied to an intercepted prompt for the LLM.
[0115] The programmable software APIs 1140 and additional model APIs 1150 allow the user to program rules or hints regarding prompt augmentation. For example, certain keywords could be used as filters to trigger a set of rules or to link them to specific LLMS or models. Reinforcement learning 1160 and proximal policy optimization 1170 help to further refine and improve the prompt augmentation capabilities of the prompt augmentation microservice 1100. For example, refining and improving prompt augmentation can include the use of closed-loop feedback mechanisms, such as confidence estimate updates, etc.
[0116] Fig. This is a flowchart illustrating an embodiment of Method 1200 for providing prompt adaptation for AI models in a PNID. Method 1200 can be performed by processing logic that may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (such as instructions executed on a processing device), or a combination thereof. For the sake of brevity and clarity, the process of Method 1200 is illustrated in linear sequences; however, it is understood that any number of these may be performed in parallel, asynchronously, or in different orders. Furthermore, for the sake of brevity, clarity, and ease of understanding, many of the components and processes referred to in the Fig. The points already described will not be repeated or discussed here. In an implementation, a PNID, such as PNID 850, can be used. Fig. and / or PNID 900 from Fig. , carry out procedure 1200.
[0117] Procedure 1200 begins at processing block 1210, where a PNID can receive a prompt directed to an AI model hosted by a host device communicatively coupled to the PNID. Then, at block 1220, the PNID can apply a prompt matching model to the prompt to generate an initial augmented prompt. At block 1230, the PNID can compare the initial augmented prompt for consistency with stored data of the real-time network context and cross-network, history-based augmentation from other PNIDs in a data center hosting the PNID.
[0118] Subsequently, at block 1240, in response to identifying a match with the stored data, the PNID can generate a final augmented prompt based on the matching data from the stored data for the real-time network context and cross-network, history-based augmentation. Finally, at block 1250, the PNID can transmit the final augmented prompt to the AI model hosted by the host device.
[0119] Fig. This is a flowchart illustrating an embodiment of Method 1300 for providing a prompt augmentation tracking table to support prompt adaptation for AI models in a PNID. Method 1300 can be performed by processing logic that may include hardware (e.g., circuitry, dedicated logic, programmable logic, etc.), software (such as instructions executed on a processing device), or a combination thereof. For the sake of brevity and clarity, the process of Method 1300 is illustrated in linear sequences; however, it is understood that any number of these may be performed in parallel, asynchronously, or in different orders. Furthermore, for the sake of brevity, clarity, and ease of understanding, many of the components and processes referred to in the Fig. The points already described will not be repeated or discussed here. In an implementation, a PNID, such as PNID 850, can be used. Fig. and / or PNID 900 from Fig. , carry out procedure 1300.
[0120] Procedure 1300 begins at processing block 1310, where a PNID can maintain a prompt augmentation tracking table hosted by a host device communicatively coupled to the programmable network interface device. Then, at block 1320, the PNID can record data in the prompt augmentation tracking table corresponding to the respective AI models. In one implementation, the data in the prompt augmentation tracking table might include, for example, a prompt augmentation scoring history for each AI model, a confidence metric for each AI model, a sibling factor for reinforcing learning from human feedback (RLHF) corresponding to other AI models hosted in the PNID's data center, and data center trend data for the other AI models hosted in the data center.
[0121] Subsequently, at block 1330, the PNID can access the prompt augmentation tracking table via a prompt augmentation microservice running on the PNID. This table responds to a prompt for one of the AI models received by the PNID. Finally, at block 1340, the PNID can augment the prompt using the prompt augmentation microservice, based on the data stored in the prompt augmentation tracking table for that specific AI model.
[0122] The following examples relate to further embodiments. Example 1 is a device for facilitating prompt matching for AI models in programmable network interface devices. The device of Example 1 includes a host interface; a network interface; and a programmable circuit communicatively coupled to the host interface and the network interface, wherein the programmable circuit comprises one or more processors for implementing network interface functionality and for: receiving a prompt addressed to an artificial intelligence (AI) model hosted by a host device communicatively coupled to the device; and applying a prompt-matching model to the prompt to generate an initial augmented prompt.Comparing the initial augmented prompt for a match with stored data in a prompt augmentation tracking table containing network data from programmable network interface devices in a data center where the facility is hosted; generating, in response to the identification of a match with the stored data, a final augmented prompt based on the match; and passing the final augmented prompt to the AI model.
[0123] In Example 2, the subject of Example 1 can optionally include that the AI model is a large language model (LLM). In Example 3, the subject of any of Examples 1-2 can optionally include that the prompt augmentation tracking table has fields for an AI model identifier (ID) and / or a graphics processing unit (GPU) domain and / or an AI model prompt and / or a prompt augmentation scoring history for each AI model and / or a confidence metric for each AI model and / or a sibling factor for reinforcement learning from human feedback (RLHF) that matches other AI models hosted in the data center, and real-time data center trend data for the other AI models hosted in the data center.
[0124] In Example 4, the subject of one of Examples 1-3 may optionally include one or more processors running a prompt augmentation service to compare the initial augmented prompt and generate the final augmented prompt by performing at least one of the following procedures: adding or removing tokens from the initial augmented prompt. In Example 5, the subject of one of Examples 1-4 may optionally include the prompt augmentation service having one or more application programming interface (API) hooks to query a vector database to obtain an update to the Retrieval Augmented Generation (RAG).
[0125] In Example 6, the subject of any of Examples 1-5 may optionally include that the prompt augmentation service has an AI model prompt score estimator to estimate an AI model prompt score that determines whether to apply prompt augmentation to the initial augmented prompt. In Example 7, the subject of any of Examples 1-6 may optionally include that the prompt augmentation service has one or more rules that are programmable by an end user to specify at least one of the keywords used as filters to trigger a set of rules or links to specific AI models.
[0126] In Example 8, the subject of any of Examples 1-7 may optionally include that the prompt augmentation service has an augmentation learning model and / or a proximal policy optimization model to refine the prompt augmentation service and / or the prompt tuning model. In Example 9, the subject of any of Examples 1-8 may optionally include that the one or more processors used to generate the final augmented prompt further include the one or more processors used to remove tokens from the initial augmented prompt to remove prompt indirections to prevent a prompt injection attack. In Example 9, the subject of any of Examples 1-8 may optionally include that the network data includes real-time data center trend data and cross-network historical augmentation data.
[0127] Example 11 is a method for facilitating prompt matching for AI models in programmable network interface devices. The method from Example 10 may include: receiving, through a programmable circuit communicatively coupled to a host interface and a network interface, a prompt addressed to an artificial intelligence (AI) model hosted by a host device communicatively coupled to the programmable circuit, the programmable circuit having one or more processors for implementing network interface functionality; applying, through the programmable circuit, a prompt matching model to the prompt to generate an initial augmented prompt;Compare, by the programmable circuit, the initial augmented prompt for a match with stored data in a prompt augmentation tracking table containing network data from programmable network interface devices in a data center where the facility is hosted; generate, by the programmable circuit, in response to the identification of the match with the stored data, a final augmented prompt based on the match; and transmit, by the programmable circuit, the final augmented prompt to the AI model.
[0128] In Example 12, the subject of Example 11 may optionally include that the prompt augmentation tracking table contains fields for an AI model identifier (ID) and / or a graphics processing unit (GPU) domain and / or an AI model prompt and / or a prompt augmentation scoring history for each AI model and / or a confidence metric for each AI model and / or a sibling enhancement learning factor from human feedback (RLHF) corresponding to other AI models, and real-time data center trend data for the other AI models hosted in the data center.In Example 13, the subject of Examples 11-12 may optionally include one or more processors performing a prompt augmentation service to compare the initial augmented prompt and generate the final augmented prompt by applying at least one of the following procedures: adding or removing tokens from the initial augmented prompt.
[0129] In Example 14, the subject of Examples 11-13 may optionally include the prompt augmentation service having one or more application programming interface (API) hooks to query a vector database to obtain an update to the Retrieval Augmented Generation (RAG). In Example 15, the subject of Examples 11-14 may optionally include the prompt augmentation service having an AI model prompt score estimator to estimate an AI model prompt score that determines whether to apply prompt augmentation to the initial augmented prompt.In Example 16, the subject of Examples 11-15 may optionally include that the one or more processors for generating the final augmented prompt may further include the one or more processors for removing tokens from the initial augmented prompt to remove prompt indirections in order to prevent a prompt injection attack.
[0130] Example 17 is a non-volatile, computer-readable storage medium for facilitating prompt adaptation for AI models in programmable network interface devices. The non-volatile, computer-readable storage medium of Example 17 contains instructions which, when executed by one or more processors, cause the one or more processors to perform operations comprising: receiving, by a programmable circuit communicatively coupled to a host interface and a network interface, a prompt directed to an artificial intelligence (AI) model hosted by a host device communicatively coupled to the programmable circuit, wherein the programmable circuit includes the one or more processors for implementing network interface functionality;Apply, through the programmable circuit, a prompt-matching model to the prompt to generate an initial augmented prompt; compare, through the programmable circuit, the initial augmented prompt to a match with stored data in a prompt augmentation tracking table containing network data from programmable network interface devices in a data center where the facility is hosted; generate, through the programmable circuit, in response to the identification of the match with the stored data, a final augmented prompt based on the match; and transmit, through the programmable circuit, the final augmented prompt to the AI model.
[0131] In Example 18, the subject of Example 17 may optionally include that the prompt augmentation tracking table contains fields for an AI model identifier (ID) and / or a graphics processing unit (GPU) domain and / or an AI model prompt and / or a prompt augmentation scoring history for each AI model and / or a confidence metric for each AI model and / or a sibling enhancement learning factor from human feedback (RLHF) corresponding to other AI models, and real-time data center trend data for the other AI models hosted in the data center.
[0132] In Example 19, the subject of Examples 17-18 may optionally include one or more processors executing a prompt augmentation service to compare the initial augmented prompt and generate the final augmented prompt by performing at least one of the following procedures: adding or removing tokens from the initial augmented prompt. In Example 20, the subject of Examples 17-19 may optionally include the prompt augmentation service having an AI model prompt score estimator to estimate an AI model prompt score that determines whether to apply prompt augmentation to the initial augmented prompt.
[0133] Example 21 is a system for facilitating prompt matching for AI models in programmable network interface devices. The system from Example 21 may optionally include a cluster of processing units and a programmable network interface device coupled to the cluster of processing units, comprising: a host interface; a network interface; and a programmable circuit communicatively coupled to the host interface and the network interface, the programmable circuit comprising one or more processors for implementing network interface functionality and for: receiving a prompt addressed to an artificial intelligence (AI) model hosted by a host device communicatively coupled to the device; applying a prompt matching model to the prompt to generate an initial augmented prompt;Comparing the initial augmented prompt for a match with stored data in a prompt augmentation tracking table containing real-time data center trend data and cross-network historical augmentation data from programmable network interface devices in a data center where the facility is hosted; generating, in response to the identification of a match with the stored data, a final augmented prompt based on the match; and passing the final augmented prompt to the AI model.
[0134] In Example 22, the subject of Example 21 can optionally include that the AI model is a large language model (LLM). In Example 23, the subject of any of Examples 21-22 can optionally include that the prompt augmentation tracking table has fields for an AI model identifier (ID) and / or a graphics processing unit (GPU) domain and / or an AI model prompt and / or a prompt augmentation scoring history for each AI model and / or a confidence metric for each AI model and / or a sibling factor for reinforcement learning from human feedback (RLHF) that matches other AI models hosted in the data center, and real-time data center trend data for the other AI models hosted in the data center.
[0135] In Example 24, the subject of one of Examples 21-23 may optionally include one or more processors running a prompt augmentation service to compare the initial augmented prompt and generate the final augmented prompt by performing at least one of the following procedures: adding or removing tokens from the initial augmented prompt. In Example 25, the subject of one of Examples 21-24 may optionally include the prompt augmentation service having one or more application programming interface (API) hooks to query a vector database to obtain an update to the Retrieval Augmented Generation (RAG).
[0136] In Example 26, the subject of one of Examples 21-25 may optionally include that the prompt augmentation service has an AI model prompt score estimator to estimate an AI model prompt score that determines whether to apply prompt augmentation to the initial augmented prompt. In Example 27, the subject of one of Examples 21-26 may optionally include that the prompt augmentation service has one or more rules that are programmable by an end user to specify at least one of the keywords used as filters to trigger a set of rules or links to specific AI models.
[0137] In Example 28, the subject of any of Examples 21-27 may optionally include that the prompt augmentation service has an augmentation learning model and / or a proximal policy optimization model to refine the prompt augmentation service and / or the prompt tuning model. In Example 29, the subject of any of Examples 21-28 may optionally include that the one or more processors used to generate the final augmented prompt further include the one or more processors used to remove tokens from the initial augmented prompt to remove prompt indirections to prevent a prompt injection attack. In Example 30, the subject of any of Examples 21-29 may optionally include that the network data includes real-time data center trend data and cross-network historical augmentation data.
[0138] Example 31 is a device for facilitating prompt matching for AI models in programmable network interface devices, comprising means for: receiving, using a programmable circuit communicatively coupled to a host interface and a network interface, a prompt addressed to an artificial intelligence (AI) model hosted by a host device communicatively coupled to the programmable circuit, the programmable circuit comprising one or more processors for implementing network interface functionality; applying, using the programmable circuit, a prompt matching model to the prompt to generate an initial augmented prompt;to compare, using the programmable circuit, the initial augmented prompt for a match with stored data in a prompt augmentation tracking table containing network data from programmable network interface devices in a data center where the facility is hosted; to generate, using the programmable circuit, in response to the identification of the match with the stored data, a final augmented prompt based on the match; and to transmit, using the programmable circuit, the final augmented prompt to the AI model. In Example 32, the subject matter of Example 30 may optionally include the facility being further designed to perform the procedure of any one of Examples 12 through 16.
[0139] Example 33 is at least one machine-readable medium comprising a plurality of instructions which, in response to their execution on a computing device, cause the computing device to perform a method according to any one of Examples 11 to 16. Example 34 is a device for facilitating prompt adaptation for AI models in programmable network interface devices, designed to perform the method according to any one of Examples 11 to 16. Example 35 is a device for prompt adaptation for AI models in programmable network interface devices, comprising means for performing the method according to any one of Examples 11 to 16. The specifics in the examples may be used in one or more embodiments.
[0140] The preceding description and drawings are intended to be illustrative rather than limiting. Those skilled in the art will understand that various modifications and changes can be made to the embodiments described herein without departing from the broader spirit and scope of the features set forth in the appended claims.
Claims
[1] Establishment which has the following features: a host interface; a network interface; and a programmable circuit that is communicatively coupled to the host interface and the network interface, wherein the programmable circuit has one or more processors for implementing network interface functionality, and to: to receive a prompt addressed to an artificial intelligence (AI) model hosted by a host device coupled to the host interface; to apply a prompt voting model to the prompt in order to generate an initial augmented prompt; to compare the initial augmented prompt for a match with stored data in a prompt augmentation tracking table containing network data from programmable network interface devices in a data center where the facility is hosted; In response to the identification of a match with the stored data, to generate a final augmented prompt based on the match; and to transfer the final augmented prompt to the AI model. [2] Device according to claim 1, wherein the AI model is a large language model (LLM). [3] Device according to one of claims 1-2, wherein the prompt augmentation tracking table has fields for an AI model identifier (ID) and / or a graphics processing unit (GPU) domain and / or an AI model prompt and / or a prompt augmentation scoring history for each AI model and / or a confidence metric for each AI model and / or a sibling factor for reinforcement learning from human feedback (RLHF) corresponding to other AI models hosted in the data center, and real-time data center trend data for the other AI models hosted in the data center. [4] Device according to any one of claims 1-3, wherein the one or more processors are to perform a prompt augmentation service to compare the initial augmented prompt and to generate the final augmented prompt by applying at least one of the following methods: adding or removing tokens from the initial augmented prompt. [5] Device according to one of claims 1-4, wherein the prompt augmentation service has one or more application programming interface (API) hooks to query a vector database to obtain an update of the Retrieval Augmented Generation (RAG). [6] Device according to any one of claims 1-5, wherein the prompt augmentation service includes an AI model prompt evaluation estimator to estimate an AI model prompt evaluation that determines whether to apply prompt augmentation to the initial augmented prompt. [7] Device according to any one of claims 1-6, wherein the prompt augmentation service comprises one or more rules that are programmable by an end user to specify at least one of the keywords that are used as filters to trigger a series of rules or links to specific AI models. [8] Device according to any one of claims 1-7, wherein the prompt augmentation service comprises an augmentation learning model and / or a proximal policy optimization model to refine the prompt augmentation service and / or the prompt tuning model. [9] Device according to any one of claims 1-8, wherein the one or more processors for generating the final augmented prompt further comprise the one or more processors for removing tokens from the initial augmented prompt to remove prompt indirections in order to prevent a prompt injection attack. [10] Device according to any one of claims 1-9, wherein the network data includes real-time data center trend data and cross-network historical augmentation data. [11] Method which features the following: Receiving, through a programmable circuit communicatively coupled to a host interface and a network interface, a prompt directed to an artificial intelligence (AI) model hosted by a host device communicatively coupled to the host interface, wherein the programmable circuit includes one or more processors for implementing network interface functionality; Applying, through the programmable circuit, a prompt tuning model to the prompt to generate an initial augmented prompt; Compare, through the programmable circuit, the initial augmented prompt for a match with stored data in a prompt augmentation tracking table containing real-time data center trend data and cross-network historical augmentation data from programmable network interface devices in a data center where the programmable circuit is hosted; Generate, through the programmable circuit, in response to the identification of a match with the stored data, a final augmented prompt based on the match; and The final augmented prompt is transmitted to the AI model via the programmable circuit. [12] Method according to claim 11, wherein the prompt augmentation tracking table includes fields for an AI model identifier (ID) and / or a graphics processing unit (GPU) domain and / or an AI model prompt and / or a prompt augmentation scoring history for each AI model and / or a confidence metric for each AI model and / or a sibling enhancement learning factor from human feedback (RLHF) corresponding to other AI models, and real-time data center trend data for the other AI models hosted in the data center. [13] Method according to one of claims 11-12, wherein the one or more processors are said to perform a prompt augmentation service to compare the initial augmented prompt and to generate the final augmented prompt by applying at least one of the following methods: adding or removing tokens from the initial augmented prompt. [14] Method according to one of claims 11-13, wherein the Prompt Augmentation Service has one or more application programming interface (API) hooks to query a vector database to obtain an update of the Retrieval Augmented Generation (RAG). [15] Method according to one of claims 11-14, wherein the prompt augmentation service includes an AI model prompt evaluation estimator to estimate an AI model prompt evaluation that determines whether to apply prompt augmentation to the initial augmented prompt. [16] Method according to one of claims 11-15, wherein the one or more processors for generating the final augmented prompt further comprise the one or more processors for removing tokens from the initial augmented prompt to remove prompt indirections in order to prevent a prompt injection attack. [17] System for facilitating prompt adaptation for AI models in programmable network interface devices, the system comprising: A cluster of processing units; and a programmable network interface device that is communicatively coupled to the cluster of processing units and has the following features: a host interface; a network interface; and a programmable circuit that is communicatively coupled to the host interface and the network interface, wherein the programmable circuit has one or more processors for implementing network interface functionality, and to: to receive a prompt addressed to an artificial intelligence (AI) model hosted by a host device coupled to the facility; to apply a prompt voting model to the prompt in order to generate an initial augmented prompt; to compare the initial augmented prompt for a match with stored data in a prompt augmentation tracking table containing real-time data center trend data and cross-network historical augmentation data from programmable network interface devices in a data center where the programmable circuit is hosted; In response to the identification of a match with the stored data, to generate a final augmented prompt based on the match; and to transfer the final augmented prompt to the AI model. [18] System according to claim 17, wherein the prompt augmentation tracking table includes fields for an AI model identifier (ID) and / or a graphics processing unit (GPU) domain and / or an AI model prompt and / or a prompt augmentation scoring history for each AI model and / or a confidence metric for each AI model and / or a sibling factor for reinforcement learning from human feedback (RLHF) corresponding to other AI models hosted in the data center, and real-time data center trend data for the other AI models hosted in the data center. [19] System according to one of claims 17-18, wherein the one or more processors are to perform a prompt augmentation service to compare the initial augmented prompt and to generate the final augmented prompt by applying at least one of the following methods: adding or removing tokens from the initial augmented prompt. [20] System according to one of claims 17-19, wherein the prompt augmentation service has one or more application programming interface (API) hooks to query a vector database to obtain an update of the Retrieval Augmented Generation (RAG). [21] System according to one of claims 17-20, wherein the prompt augmentation service includes an AI model prompt evaluation estimator to estimate an AI model prompt evaluation that determines whether to apply prompt augmentation to the initial augmented prompt. [22] System according to one of claims 17-21, wherein the prompt augmentation service comprises one or more rules that are programmable by an end user to specify at least one of the keywords that are used as filters to trigger a series of rules or links to specific AI models. [23] System according to one of claims 17-22, wherein the prompt augmentation service comprises an augmentation learning model and / or a proximal policy optimization model to refine the prompt augmentation service and / or the prompt tuning model. [24] Device for prompt adaptation for AI models in programmable network interface devices, comprising means for carrying out the method according to any one of claims 11 to 16. [25] At least one machine-readable medium containing a plurality of instructions which, when executed on a computing device, cause the computing device to carry out a procedure according to one of Examples 11 to 16.