Resource deployment in expert hybrid processing

By improving the resource deployment of hybrid expert (MoE) operations and utilizing accelerators and orchestrators to manage network resources, the complexity of expert selection in MoE technology is solved, achieving efficient allocation of computing resources and simplification of network operations, thereby improving the training and prediction efficiency of models.

CN122073540APending Publication Date: 2026-05-22INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTEL CORP
Filing Date
2025-10-16
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

The MoE technology involves the complexity of selecting experts for processing, leading to inefficient allocation of computing resources and complexity of network operations.

Method used

By improving the resource deployment of expert hybrid (MoE) operations, leveraging accelerator devices and orchestrators 360 to manage network resources, adaptive routing and telemetry-based congestion control are achieved, optimizing the routing capabilities of the expert hybrid model.

Benefits of technology

It improves the efficient allocation of computing resources and simplifies network operations, reduces computational complexity, and improves the efficiency of model training and prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122073540A_ABST
    Figure CN122073540A_ABST
Patent Text Reader

Abstract

The name of the invention is resource deployment in expert hybrid processing. Resource deployment with improved expert hybrid processing is described. An example of a device includes: one or more network ports; one or more direct memory access (DMA) engines; the circuit module is used for expert mixing (MoE) processing in the network; wherein the circuit module at least comprises a circuit module used for tracking a route of a token in MoE processing; a prediction circuit module to generate a prediction regarding the MoE processing, including predicting a future token load of the MoE processing; and a routing management circuitry module to manage the routing of the token in the MoE process based at least in part on the prediction regarding the MoE process.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Hybrid Expert (MoE) is an emerging paradigm for reducing the computational requirements of large language models (LLMs) while providing comparable accuracy. Applying MoE processing allows models to be pre-trained with reduced computation, thus enabling dynamic scaling of the model or dataset size with the same computational budget as dense models. Specifically, MoE models can achieve the same or similar quality as their dense counterparts, which are achieved much faster during pre-training.

[0002] In MoE operations, instead of using dense feedforward network layers in the network, sparse mixtures of expert layers are used. The layers can contain a certain number of experts (e.g., 16, 32, or 128), where the experts themselves are either feedforward networks or mixtures of experts, thus forming a hierarchical mixture of experts.

[0003] However, MoE technology introduces certain complexities in processing, including the complexity of correctly selecting experts to utilize in an instance during operation. Attached Figure Description

[0004] The embodiments described herein are illustrated in the accompanying drawings by way of example rather than limitation, wherein similar reference numerals indicate similar elements, and wherein: Figure 1 This is a block diagram illustrating a computer system configured to implement one or more aspects of the embodiments described herein; Figure 2 It is a block diagram of a system including selected components of a data center; Figure 3 This is a block diagram of a data center as described in one or more examples of this specification. Figures 4A-4C This illustrates programmable forwarding elements and adaptive routing; Figures 5A-5B An example network interface device is shown; Figure 6 This is a block diagram showing the programmable network interface and data processing unit; Figure 7 This is a block diagram showing the IP core development system; Figure 8 This is an illustration of a computing system or device, according to some embodiments, including routing capabilities for processing mixed deployments of experts in resources; Figure 9 This demonstrates the application of expert hybridization in AI processing; Figure 10 This is a diagram illustrating the routing in expert hybrid operations; Figures 11A-11FThis is a diagram illustrating the routing algorithm in expert hybrid operations; Figure 12 This is an illustration of a system for resource deployment utilizing improved expert hybrid processing, according to some embodiments; Figure 13A This is an illustration of resource deployment utilizing improved expert hybrid processing according to some embodiments; Figure 13B This illustrates expert data maintained for expert hybrid processing according to some embodiments; and Figure 14 This is a flowchart illustrating a process of expert hybrid operation in network processing, according to some embodiments. Detailed Implementation

[0005] In some embodiments, the device, system, or process provides resource deployment utilizing improved expert hybrid (MoE) operations.

[0006] In the following description, numerous specific details are set forth to provide a more thorough understanding. However, those skilled in the art will understand that the embodiments described herein can be practiced without one or more of these specific details. In other instances, well-known features have not been described to avoid obscuring the details of this embodiment.

[0007] Figure 1 This is a block diagram of a computing system 100 configured to implement one or more aspects of the embodiments described herein. The computing system 100 includes a processing subsystem 101 having one or more processors 102 (such as a central processing unit (CPU) or other host processor) and system memory 104 communicating via interconnect paths including a memory hub 105. The memory hub 105 may be a separate component within a chipset assembly or may be integrated within one or more processors 102. The memory hub 105 is coupled to an I / O subsystem 111 via a communication link 106. The I / O subsystem 111 includes an I / O hub 107 that enables the computing system 100 to receive input from one or more input devices 108. Additionally, the I / O hub 107 enables a display controller, which may be included in one or more processors 102, to provide output to one or more display devices 110A. In one embodiment, the one or more display devices 110A coupled to the I / O hub 107 may include local, internal, or embedded display devices.

[0008] Processing subsystem 101 includes, for example, one or more parallel processors 112 coupled to memory hub 105 via a communication link 113 (such as a bus or structure). Communication link 113 can be any number of standards-based communication link technologies or protocols (such as, but not limited to, PCI Fast), or it can be a vendor-specific communication interface or communication structure. The one or more parallel processors 112 can form a parallel or vector processing system that may include a large number of processing cores and / or processing clusters, such as a many integrated core (MIC) processor. For example, the one or more parallel processors 112 form a graphics processing subsystem that can output pixels to one of one or more display devices 110A coupled via I / O hub 107. The one or more parallel processors 112 may also include a display controller and a display interface (not shown) to enable direct connections to one or more display devices 110B.

[0009] Within the I / O subsystem 111, system storage unit 114 can be connected to I / O hub 107, thereby providing a storage mechanism for computing system 100. I / O switch 116 can be used to provide an interface mechanism to enable connections between I / O hub 107 and other components, such as network adapter 118 and / or wireless network adapter 119 which can be integrated into the platform, and various other devices that can be added via one or more plug-in devices 120. The plug-in devices 120 may also include, for example, one or more external graphics processing units, graphics cards, and / or computing accelerators. Network adapter 118 may be an Ethernet adapter or another wired network adapter. Wireless network adapter 119 may include one or more of the following: Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more wireless radio devices.

[0010] The computing system 100 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., and may also be connected to the I / O hub 107. Any suitable protocol (such as a PCI (Peripheral Component Interconnect) based protocol (e.g., PCI Fast)) or any other bus or point-to-point communication interface and / or (one or more) protocols (such as NVLink High-Speed ​​Interconnect, Compute ExpressLink, etc.) may be used. TM (CXL) TMInterconnection can be implemented using wired or wireless interconnect protocols known in the art, such as CXL.mem, Infinite Architecture (IF), Ethernet (IEEE 802.3), Remote Direct Memory Access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), Internet Protocol (IP), User Datagram Protocol (UDP), Fast UDP Internet Connection (QUIC), RDMA over Converged Ethernet (RoCE), Ultra Ethernet Transport (UET), Intel Fast Path Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel System-on-Chip Architecture (IOSF), Omnipath, HyperTransport, Advanced Microcontroller Bus Architecture (AMBA) Interconnect, Open Coherent Accelerator Processor Interface (CAPI), Gen-Z, Cache Coherent Interconnect for Accelerators (CCIX), 3GPP Long Term Evolution (LTE) (e.g., 4G, 5G and variants thereof), or other wired or wireless interconnect protocols known in the art. Figure 1 The communication paths for various components within the system. In some examples, data can be copied or stored to virtualized storage nodes using protocols such as structure-based non-volatile fast memory (NVMe) (NVMe-oF) or NVMe. In one embodiment, time-aware communication protocols are supported, including time-aware RDMA, time-aware NVMe, and time-aware NVMe-oF, where precise data consumption time and rate are used to control data transmission.

[0011] One or more parallel processors 112 may be combined with circuit modules optimized for graphics and video processing (including, for example, video output circuit modules) and constitute a graphics processing unit (GPU). Alternatively or additionally, as described in more detail herein, one or more parallel processors 112 may be combined with circuit modules optimized for general-purpose processing while retaining the underlying computing architecture. Components of the computing system 100 may be integrated with one or more other system elements on a single integrated circuit. For example, one or more parallel processors 112, a memory hub 105, one or more processors 102, and an I / O hub 107 may be integrated into a system-on-a-chip (SoC) integrated circuit. Alternatively, components of the computing system 100 may be integrated into a single package to form a system-in-package (SIP) configuration. In one embodiment, at least a portion of the components of the computing system 100 may be integrated into a multi-chip module (MCM), which may be interconnected with other multi-chip modules to form a modular computing system.

[0012] In some configurations, in addition to processor(s) 102 and one or more parallel processors 112, computing system 100 includes one or more accelerator devices 130 coupled to memory hub 105. Accelerator devices 130 are configured to perform domain-specific acceleration of workloads to handle computationally intensive or high-throughput tasks. Accelerator devices 130 can alleviate the load on processor(s) 102 and / or one or more parallel processors 112 of computing system 100. Accelerator devices 130 may include, but are not limited to, intelligent network interface cards, data processing units, cryptographic accelerators, storage accelerators, artificial intelligence (AI) accelerators, neural processing units (NPUs), and / or video transcoding accelerators.

[0013] It will be appreciated that the computing system 100 shown herein is illustrative and variations and modifications are possible. The connection topology, including the number and arrangement of bridges, the number of processors(one or more) 102, and the number of parallel processors(one or more) 112, can be modified as desired. For example, system memory 104 can be connected directly to processors(one or more) 102 instead of via bridges, while other devices communicate with system memory 104 via memory hub 105 and processors(one or more) 102. In other alternative topologies, parallel processors(one or more) 112 are connected to I / O hub 107 or directly to one of the processors(one or more) 102, instead of being connected to memory hub 105. In other embodiments, I / O hub 107 and memory hub 105 can be integrated into a single chip. It is also possible to attach two or more sets of processors(one or more) 102 via multiple sockets, which can be coupled to two or more instances of parallel processors(one or more) 112.

[0014] Some of the specific components shown in this document are optional and may not be included in all implementations of the computing system 100. For example, any number of plug-in cards or peripheral devices may be supported, or some components may be omitted. Furthermore, some architectures may use different terminology for... Figure 1 The similar components described in the text.

[0015] Figure 2This is a block diagram of system 200, which includes selected components of a data center. The components of the data center shown may reside, for example, in a cloud service provider (CSP) or another data center. As a non-limiting example, the data center may be a traditional enterprise data center, an enterprise "private cloud," or a "public cloud," providing services such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), or Software as a Service (SaaS). System 200 includes a number of workload clusters, including, but not limited to, workload clusters 218A and 218B. Workload clusters may be clusters of single servers, blade servers, rack servers, or any other suitable server topology.

[0016] System 200 may include workload clusters 218A-218B. Workload clusters 218A-218B may include racks 248 that house multiple servers (e.g., server 246). The servers in racks 248 and workload clusters 218A-218B may conform to rack unit (“U”) standards, where one rack unit conforms to a 19-inch wide rack frame, and a full-size industry-standard rack can accommodate 42 units (42U) of equipment. A unit (1U) of equipment (e.g., a 1U server) may be 1.75 inches high and approximately 36 inches deep. In various configurations, computing resources such as processors, memory, storage devices, accelerators, and switches may be installed in multiple rack units within rack 248.

[0017] Server 246 can host a standalone operating system configured to provide server functionality, or the server can be virtualized. Virtualized servers can be under the control of a Virtual Machine Manager (VMM), hypervisor, and / or orchestrator, and can host one or more virtual machines, virtual servers, or virtual appliances. Workload clusters 218A-218B can be co-located in a single data center or can be situated in different geographic data centers. Depending on the contractual agreement, some servers may be dedicated to specific enterprise clients or tenants, while other servers may be shared.

[0018] Various devices within a data center can be interconnected via switching infrastructure 270, which may include one or more high-speed routing and / or switching devices. Switching infrastructure 270 can provide north-south traffic 202 (e.g., traffic to and from a wide area network (WAN) such as the Internet) and east-west traffic 204 (e.g., traffic across data centers). Historically, north-south traffic 202 has accounted for the majority of network traffic, but as web services have become more complex and distributed, the amount of east-west traffic 204 is also increasing. In many data centers, east-west traffic 204 now accounts for the majority of traffic. Furthermore, traffic volume may increase further as the capabilities of server 246 increase. For example, server 246 may provide multiple processor sockets, one of which may accommodate a processor with four to eight cores, along with sufficient memory for the cores. Therefore, the server can host multiple virtual machines (VMs), which may be the source of traffic generation.

[0019] To accommodate the high volume of traffic in a data center, a high-performance implementation of switching fabric 270 can be provided. The illustrated implementation of switching fabric 270 is an example of a flat network where server 246 can be directly connected to the top-of-rack switches (ToR switches 220A-220B) (e.g., a "star" configuration). ToR switches 220A can be connected to workload cluster 218A, while ToR switches 220B can be connected to workload cluster 218B. ToR switches 220A-220B can be coupled to core switch 260. This two-layer flat network architecture is shown as an illustrative example only, and other architectures can also be used as non-limiting examples, such as a three-layer star or leaf-ridge topology (also known as a "fat tree" topology) based on a "Clos" architecture, hub and spoke topology, mesh topology, ring topology, or three-dimensional mesh topology.

[0020] Switching fabric 270 can be provided by any suitable interconnect using any suitable interconnect protocol. For example, server 246 may include some type of fabric interface (FI), network interface card (NIC), or other host interface. The host interface itself may be coupled to one or more processors via an interconnect or bus (such as PCI, PCIe, or similar), and in some cases, the interconnect bus may be considered part of switching fabric 270. Switching fabric can also use PCIe physical interconnect to implement more advanced protocols, such as Compute Fast Link (CXL).

[0021] Interconnect technologies can be provided by a single interconnect or a hybrid interconnect; for example, PCIe provides on-chip communication, 1Gb or 10Gb copper Ethernet provides a relatively short connection to the ToR switches 220A-220B, and fiber optic cable provides a relatively long connection to the core switch 260. As non-limiting examples, interconnect technologies include Hyperpath Interconnect (UPI), Fibre Channel, Ethernet, Fibre Channel Ethernet (FCoE), InfiniBand, PCIe, NVLink, or fiber optics, to name just a few. Some of these will be better suited to certain deployments or functions than others, and choosing the appropriate architecture for the immediate application is a common skill exercise.

[0022] In one embodiment, the switching elements of switching architecture 270 are configured to implement switching technologies to improve network performance in high-usage scenarios. Exemplary advanced switching technologies include, but are not limited to, adaptive routing, adaptive fault recovery, and congestion control based on adaptation and / or telemetry.

[0023] Adaptive routing enables ToR 220A-220B switches and / or core switch 260 to select the output port for service switching based on the load on the selected port, assuming unconstrained port selection is enabled. Adaptive routing tables can configure the forwarding tables of the switches in switching fabric 270 to select among multiple ports between switches when multiple connections exist between a given set of switches in an adaptive routing group. If the port selected by the forwarding table port is faulty or inactive, adaptive fault recovery (e.g., self-healing) enables automatic selection of a backup port, thus enabling rapid recovery in the event of a switch-to-switch port failure. Notifications can be sent to neighboring switches when adaptive routing or adaptive fault recovery is active in a given switch. Adaptive congestion control configures a switch to send notifications to neighboring switches when port congestion on that switch exceeds a configured threshold, which may cause those neighboring switches to adaptively switch to an uncongested port on that switch or a switch associated with an alternative route to the destination.

[0024] Telemetry-based congestion control uses real-time monitoring of telemetry from network devices, such as switches within switching fabric 270, to detect when congestion will begin to affect the performance of switching fabric 270 and proactively adjusts switching tables within the network devices to prevent or mitigate impending congestion. ToR 220A-220B switches and / or core switches 260 can implement built-in telemetry-based congestion control algorithms or provide an application programming interface (API) through which programmable telemetry-based congestion control algorithms can be implemented. A continuous feedback loop can be implemented, where the telemetry-based congestion control system continuously monitors the network and adjusts services in real time based on ongoing telemetry data. Learning and adaptation are possible through the telemetry-based congestion control system, where the system can adapt to changing network conditions and improve its congestion control strategies based on historical data and trends.

[0025] However, please note that while this document provides a high-end architecture by way of illustration, more generally, switching architecture 270 may include any suitable interconnect or bus for a particular application, including conventional interconnects for implementing local area networks (LANs), synchronous optical networks (SONETs), asynchronous transmission mode (ATM) networks, wireless networks such as Wi-Fi and Bluetooth, 4G wireless, 5G wireless, digital subscriber line (DSL) interconnects, coaxial cable multimedia alliance (MoCA) interconnects, or similar wired or wireless network interconnects. It is also explicitly anticipated that new network technologies will emerge in the future to complement or replace some of the technologies listed herein, and any such future network topologies and technologies may be, or form part of, switching architecture 270.

[0026] Figure 3 This is a block diagram of a portion of a data center 300 according to one or more examples of this specification. The portion of data center 300 shown is not intended to include all components of the data center. Depending on the capacity and functionality expected to be provided by data center 300, the shown portion may be replicated multiple times within data center 300 and / or data center 300 may include portions extending beyond the shown portion. In various embodiments, data center 300 may include... Figure 2 The system consists of components of a 200-data center, or may be different data centers.

[0027] Data center 300 includes multiple logical elements forming multiple nodes, where nodes can be provided by physical servers, server clusters, or other hardware. Servers may also host one or more virtual machines depending on their applications. A structure 370 is provided to interconnect the various aspects of data center 300. Structure 370 can be provided by any suitable interconnect technology, including but not limited to InfiniBand, Ethernet, PCIe, or CXL. Structure 370 of data center 300 can be... Figure 2The system 200 includes a version of the switching structure 270 and / or its components. The structure 370 of the data center 300 can interconnect data center components including server nodes (e.g., storage server node 304, heterogeneous computing server node 306, CPU server node 308, storage server node 310), accelerators 330, gateways 340A-340B to other structures, architectures or interconnection technologies, and orchestrators 360.

[0028] The server nodes of data center 300 may include, but are not limited to, storage server node 304, heterogeneous computing server node 306, CPU server node 308, and storage server node 310. Heterogeneous computing server node 306 and CPU server node 308 can perform independent operations for different tenants or collaboratively perform operations for a single tenant. Heterogeneous computing server node 306 and CPU server node 308 can also host virtual machines that provide virtual server functionality to tenants in the data center.

[0029] Server nodes can connect to architecture 370 via architecture interface 372. The specific type of architecture interface 372 used depends at least in part on the technology or protocol used to implement architecture 370. For example, if architecture 370 is an Ethernet architecture, architecture interface 372 can be an Ethernet network interface controller. If architecture 370 is a PCIe-based architecture, the architecture interface can be a PCIe-based interconnect. If architecture 370 is an InfiniBand architecture, the architecture interface 372 for heterogeneous computing server node 306 and CPU server node 308 can be a host channel adapter (HCA), while the architecture interface 362 for memory server node 304 and storage server node 310 can be a target channel adapter (TCA). The TCA functionality may be an implementation-specific subset of the HCA functionality. Various architecture interfaces can be implemented as intellectual property (IP) blocks, which can be inserted as modular units into integrated circuits, as can other circuit modules within data center 300.

[0030] The heterogeneous computing server node 306 includes multiple CPU sockets that can accommodate CPUs 319, which may be, but are not limited to, multi-core Intel® Xeon™ processors. CPUs 319 may also be, for example, multi-core data center-class ARM® CPUs, such as NVIDIA® Grace™ CPUs. The heterogeneous computing server node 306 includes a memory device 318 for storing runtime execution data and a storage device 316 for enabling persistent storage of data within a non-volatile memory device. The heterogeneous computing server node 306 is enabled to perform heterogeneous processing via the presence of a GPU (e.g., GPU 317), which may be used for, for example, performing high-performance computing (HPC), media server, cloud gaming server, and / or machine learning computing operations. In one configuration, the GPU may interconnect with the CPU of the heterogeneous computing server node 306 via interconnect technologies such as PCIe, CXL, or NVLink.

[0031] CPU server node 308 includes multiple CPUs (e.g., CPU 319), memory (e.g., memory device 318), and storage devices (storage device 316) to execute application and other program code that provide server functionality, such as a web server or other types of functionality that can be remotely accessed by clients of CPU server node 308. CPU server node 308 can also execute program code that provides services or microservices that enable complex enterprise functions. Structure 370 will be supplied with sufficient throughput to allow CPU server node 308 to be accessed by a large number of clients simultaneously, while also reserving sufficient throughput for heterogeneous computing server node 306, and enabling both heterogeneous computing server node 306 and CPU server node 308 to utilize memory server node 304 and storage server node 310. Furthermore, in one configuration, CPU server node 308 may primarily rely on distributed services provided by memory server node 304 and storage server node 310, because the memory and storage devices of CPU server node 308 may be insufficient for all operations that CPU server node 308 intends to perform. Conversely, large amounts of high-speed or dedicated memory can be dynamically provisioned across multiple nodes, allowing nodes to access a wide range of resources without them becoming idle when a particular node does not require them. This type of distributed architecture is possible and potentially advantageous due to the high speed and low latency offered by the architecture of modern data centers, as it eliminates the need for excessive provisioning of resources to server nodes.

[0032] Memory server node 304 may include memory node 305, which has memory technology suitable for storing data used during the execution of program code by heterogeneous computing server node 306 and CPU server node 308. Memory node 305 may include volatile memory modules, such as DRAM modules, and / or non-volatile memory technologies that can operate at speeds similar to DRAM, such that these modules have sufficient throughput and latency performance metrics to serve as a layer of system memory during runtime. Memory server node 304 may be linked to heterogeneous computing server node 306 and / or CPU server node 308 via a technology such as CXL.mem, enabling host-to-device memory access. In this configuration, the CPU 319 of heterogeneous computing server node 306 and CPU server node 308 may be linked to memory server node 304 and access memory node 305 of memory server node 304 in a manner similar to, for example, the CPU 319 of heterogeneous computing server node 306 can access the device memory of the GPU within heterogeneous computing server node 306. For example, memory server node 304 can provide remote direct memory access (RDMA) to memory node 305, where, for example, CPU server node 308 can use direct memory access (DMA) operations to access memory resources on memory server node 304 via structure 370 in a manner similar to how a CPU accesses its own onboard memory.

[0033] Memory server node 304 can be used by heterogeneous compute server node 306 and CPU server node 308 to extend runtime memory available during memory-intensive activities such as training machine learning models. A tiered storage system can be enabled, where model data can be swapped to and from memory device 318 of heterogeneous compute server node 306 to memory device 318 of memory server node 304 with higher performance and / or lower latency than local storage (e.g., storage device 316). During workload execution setup, the entire working dataset can be loaded as needed into one or more memory nodes 305 of memory server node 304 and into memory device 318 of heterogeneous compute server node 306 during the execution of heterogeneous workloads.

[0034] Storage server node 310 provides storage functionality to heterogeneous computing server node 306, CPU server node 308, and potentially storage server node 304. Storage server node 310 can provide networked disk stacks or disk-only stacks (JBOD), programmable flash (PFM), redundant array of independent disks (RAID), redundant array of independent nodes (RAIN), network-attached storage (NAS), or other non-volatile memory solutions. In one configuration, storage server node 310 can be coupled to heterogeneous computing server node 306, CPU server node 308, and / or storage server node 304 (such as NVMe-oF), enabling the NVMe protocol to be implemented on fabric 370. In this configuration, the fabric interface 372 of these servers can be an intelligent interface that includes hardware to accelerate NVMe-oF operation.

[0035] Accelerator 330 within data center 300 can provide various acceleration functions, including hardware or coprocessor acceleration for functions such as packet processing, encryption, decryption, compression, decompression, network security, or other acceleration functions within the data center. In some examples, accelerator 330 may include a deep learning accelerator, such as a neural processing unit (NPU), which can offload matrix multiplication operations from other neural network operations from heterogeneous computing server node 306 or CPU server node 308. In some configurations, accelerator 330 may reside in a dedicated accelerator server or be distributed across the various server nodes of data center 300. For example, an NPU may be directly attached to one or more CPU cores within heterogeneous computing server node 306 or CPU server node 308. In some configurations, accelerator 330 may include or be included in an intelligent network controller, infrastructure processing unit (IPU), or data processing unit that combines network controller functionality with accelerator, processor, or coprocessor functionality. Accelerator 330 may also include an edge processing unit (EPU) to perform real-time inference operations at the network edge. In some embodiments, accelerator 330 includes hardware circuitry modules for providing expert hybrid operations, including providing expert-specific tracking, prediction, and routing, as well as providing migration and replication of certain experts, for example, such as Figure 8 and Figures 12-13B As further shown in the image.

[0036] In one configuration, data center 300 may include gateways 340A-340B connecting structure 370 to other structures, architectures, or interconnect technologies. For example, if structure 370 is an InfiniBand structure, gateways 340A-340B may be gateways to an Ethernet structure. If structure 370 is an Ethernet structure, gateways 340A-340B may include routers to route data to other parts of data center 300 or a larger network, such as the Internet. For example, a first gateway 340A may connect to different networks or subnets within data center 300, while a second gateway 340B may be a router to the Internet.

[0037] Orchestrator 360 manages the provisioning, configuration, and operation of network resources within data center 300. Orchestrator 360 may include hardware or software running on a dedicated orchestration server. Orchestrator 360 may also be embodied in software, for example, running on CPU server node 308, which configures software-defined networking (SDN) capabilities for components within data center 300. In various configurations, orchestrator 360 enables the automated provisioning and configuration of components within data center 300 by performing network resource allocation and template-based deployment. Template-based deployment is a method of provisioning and managing IT resources using predefined templates, where the templates may be based on standard templates required by governments, service providers, financial institutions, standards organizations, or customers. Templates may also specify Service Level Agreements (SLAs) or Service Level Obligations (SLOs). Orchestrator 360 may also perform functions, including but not limited to load balancing and business engineering, network segmentation, security automation, real-time telemetry monitoring, and adaptive failover management, including telemetry-based adaptive failover. In some configurations, Orchestrator 360 can also provide multi-tenancy and virtualization support by enabling virtual network management, including creating and deleting virtual LANs (VLANs) and virtual private networks (VPNs), as well as tenant isolation in multi-tenant data centers.

[0038] Figures 4A-4C This illustrates programmable forwarding elements and adaptive routing. Figure 4A The diagram shows a forwarding element comprising a control plane and a programmable data plane. Figure 4B A network with switching devices configured to perform adaptive routing and telemetry-based congestion control is shown. Figure 4C This shows an InfiniBand switch that includes a multi-port IB interface.

[0039] Figure 4AA forwarding element 400 is shown, which can be configured to forward data messages within a network based on a user-provided program. In some embodiments, the program includes instructions for forwarding data messages and performing other processes, such as firewall, denial-of-service attack protection, and load balancing operations. Forwarding element 400 can be any type of forwarding element, including but not limited to switches, routers, or bridges. Forwarding element 400 can forward data messages associated with various technologies, such as, but not limited to, Ethernet, Ultra Ethernet, InfiniBand, or NVLink.

[0040] In various network configurations, forwarding elements are deployed as non-edge forwarding elements within the network to forward data messages from source devices to destination devices. In other network configurations, forwarding element 400 is deployed as an edge forwarding element at the network edge to connect to computing devices (e.g., standalone or host computers) that serve as the source and destination of data messages. As a non-edge forwarding element, forwarding element 400 forwards data messages between forwarding elements in the network, for example, through an intermediate network structure. As an edge forwarding element, forwarding element 400 forwards data messages to edge computing devices, from edge computing devices to other edge forwarding elements, and / or non-edge forwarding elements.

[0041] The forwarding element 400 includes circuit modules for implementing a data plane 402, which performs forwarding operations to forward data messages received by the forwarding element 400 to other devices. The forwarding element 400 also includes circuit modules for implementing a control plane 404 for configuring the data plane circuitry. Furthermore, the forwarding element 400 includes a physical port 406 that receives data messages from and transmits them to devices outside the forwarding element 400. The data plane 402 includes ports 408 that receive data messages from physical port 406 for processing. The processed data messages are forwarded to another port on the data plane 402, which is connected to another physical port of the forwarding element 400. In addition to being associated with the physical ports of the forwarding element 400, some ports 408 on the data plane 402 may also be associated with other modules of the data plane 402.

[0042] The data plane includes programmable packet processor circuitry that provides several programmable message processing stages configured to perform data plane forwarding operations of forwarding element 400 to process data messages and forward them to their destinations. These message processing stages perform these forwarding operations by processing data tuples (e.g., message headers) associated with data messages received by data plane 402 to determine how to forward the messages. Each message processing stage includes a Matching Action Unit (MAU) that attempts to match the message's data tuples (e.g., header vectors) with table records specifying actions to be performed on the data tuples. In some embodiments, these table records are populated by control plane 404 and are unknown when configuring the data plane to execute programs provided by network users. The programmable message processing circuitry is grouped into multiple message processing pipelines. These pipelines can be ingress or egress pipelines preceding or following a service management phase of the forwarding element, which guides messages from the ingress pipeline to the egress pipeline.

[0043] The hardware details of data plane 402 depend on the communication protocol implemented via forwarding element 400. Ethernet switches use application-specific integrated circuits (ASICs) designed to process Ethernet frames and the TCP / IP protocol stack. These ASICs are optimized for various service types, including unicast, multicast, and broadcast. Ethernet switch ASICs are typically designed to balance cost, power consumption, and performance, although high-end Ethernet switches may support more advanced features such as deep packet inspection and advanced QoS (Quality of Service). InfiniBand switches use dedicated ASICs designed for ultra-low latency and high throughput. These ASICs enable features such as optimized processing of the InfiniBand protocol and support for RDMA and other features requiring precise timing and high-speed data processing, although high-end Ethernet switches may support RoCE (RDMA over Aggregated Ethernet), which offers similar benefits to InfiniBand but with higher latency compared to native InfiniBand RDMA.

[0044] The forwarding element 400 can also be configured as an NVLink switch (such as an NVSwitch) for interconnecting multiple graphics processors via the NVLink connectivity protocol. When configured as an NVLink switch, the forwarding element 400 can provide GPU servers with increased GPU-to-GPU bandwidth compared to GPU servers interconnected via InfiniBand. NVLink switches can reduce network hotspots that may occur when interconnected GPU-equipped servers perform operations such as distributed neural network training.

[0045] Typically, when data plane 402 collaborates with a program executing on data plane 402 (e.g., a program written in P4 language) to perform message or packet forwarding operations on incoming data, control plane 404 determines how the message or packet should be forwarded. The behavior of the program executing on data plane 402 is partly determined by control plane 404, which populates a matching action table with specific forwarding rules. The forwarding rules used by the program executing on data plane 402 are independent of the data plane program itself. In one configuration, the control plane may be coupled to management port 410, which enables administrator configuration of forwarding element 400. The data connection established via management port 410 is separate from the data connections used for ingress and egress data ports. In one configuration, management port 410 may be connected to management plane 405, which facilitates administrative access to the device, enables analysis of device status and health, and enables device reconfiguration. Management plane 405 may be part of control plane 404 or communicate directly with control plane 404. In one implementation, the administrator does not have direct access to the components of control plane 404. Instead, information is collected by management plane 405, and changes to control plane 404 are performed by management plane 405.

[0046] Figure 4B A network 420 is shown, featuring switches 432A-432E that support adaptive routing and telemetry-based congestion control. Network 420 can be implemented using various communication protocols described herein. In one embodiment, network 420 is implemented using the InfiniBand protocol. In another embodiment, network 420 is an Ethernet, converged Ethernet, or ultra-Ethernet network. Network 420 may include... Figure 3 The structure of the 370 is shown in various aspects. The 432A-432E switches can be... Figure 4A The implementation of forwarding element 400. Network 420 provides packet-based communication for multiple nodes (e.g., node 424, node 446), including source node 422 and destination node 442 for data transmission to be performed on network 420. Flowing packets are forwarded via routes traversing network 420, which pass through the network 420's switches (switches 432A-432E) and links (links 426A-426B, 427A-427B, 428, 429A-429B, 430A-430B). In InfiniBand applications, the switches and links belong to an InfiniBand subnet managed by a subnet manager (SM), which may be contained within one of the switches (e.g., switch 432D). Source node 422 and destination node 442 are exemplary source and destination nodes for a data flow. Depending on the configuration of network 420, packets can flow from any node to any other node via one or more paths.

[0047] Switches 432A-432E include a data plane 402, a control plane 404, a management plane 405, and physical ports 406, such as... Figure 4A The forwarding element 400 is shown in the diagram. The processor of the control plane 404 can be used to implement adaptive routing techniques to adjust the route between source node 422 and destination node 442 based on the current state of the network. During network operation, due to various events such as congestion, link failure, or line header blockage, the route from source node 422 to destination node 442 may become unsuitable or impaired in its ability to transmit packets at some points. If this occurs, switches 432A-432E can be configured to dynamically adapt the routing of packets flowing along the impaired path.

[0048] Adaptive Routing (AR) events may be detected by one of the switches along a route that has become impaired, for example, when the switch attempts to output a packet on a designated output port. For example, exemplary data from source node 422 to destination node 442 can traverse a link through the network's switches. For example, switch 432D may detect an AR event for link 429B in response to congestion or a link failure associated with link 429B. Upon detecting an AR event, switch 432D, acting as the detection switch, generates an Adaptive Routing Notification (ARN) with an identifier that distinguishes the ARN packet from other packet types. In various embodiments, the ARN includes parameters such as the identifier of the detection switch, the type of AR event, and the source and destination addresses of the flow that triggered the AR event, and / or any other suitable parameters. The detection switch forwards the ARN backward along the route to the preceding switch. The ARN may include a request to the notifying switch to modify the route to avoid traversing the detection switch. The notified switch can then assess whether its route can be modified to bypass the detection switch. Otherwise, the switch forwards the ARN to the preceding switch along the route. In this scenario, switch 432B cannot bypass switch 432D and relay the ARN to switch 432A. Switch 432A can determine to route to destination node 442 by using link 427A to switch 432C. Switch 432C can reach switch 432E via link 429A, allowing packets from source node 422 to reach destination node 442 while bypassing the AR events associated with link 429B.

[0049] In various configurations, network 420 can also adapt to congestion conditions via programmable data planes within switches 432A-432E. These data planes can execute data plane programs to implement in-network congestion control (CCA) algorithms based on a TCP-over-Ethernet architecture. Using in-band network telemetry (INT), the programmable data planes within switches 432A-432E can become aware of when ports or links along a route become congested and preemptively seek to route packets via alternative paths. For example, switch 432A can load balance traffic to destination node 442 between links 427A and 427B based on the congestion levels observed on routes downstream of links 427A and 427B.

[0050] Figure 4C The InfiniBand switch 450 is shown, which can be Figure 4A The implementation of the forwarding element 400. The InfiniBand switch 450 includes a programmable data plane and is configurable to perform adaptive routing and telemetry-based congestion control as described herein. The InfiniBand switch 450 includes multi-port IB interfaces 460A-460D and core switch logic 480. The multi-port IB interfaces 460A-460D include multiple ports. In one embodiment, there is a single instance of a physical interface (IB PHY 453) with input and output buffers associated with the port. In one embodiment, the port has a separate physical interface. The port may be coupled to, for example, an HCA 452, a TCA 461, or another InfiniBand switch 432. The multi-port IB interfaces 460A-460D may include a crossbar switch 454 configured to selectively couple input and output port buffers to local memory 456. The crossbar switch 454 is a non-blocking crossbar switch that provides direct and low-latency switching with fixed or variable packet sizes.

[0051] Local memory 456 includes multiple queues, including an external receive queue 462, an external transmit queue 463, an internal receive queue 464, and an internal transmit queue 465. The external queue is used for data received at a given multi-port IB interface, which will be forwarded back from the same multi-port IB interface. The internal queue is used for data forwarded from a different multi-port IB interface than the one used for receiving data. Other types of queue configurations can be implemented in local memory 456. For example, different queues can exist to support multiple service classes, whether based on a single port, a shared port, or a combination thereof. Multi-port IB interfaces 460A-460D include a power management circuit module 455, which can adjust the power state of the circuit modules within the corresponding multi-port IB interface. Furthermore, power management logic performing similar operations can be implemented as part of the core switch logic.

[0052] The multi-port IB interfaces 460A-460D include packet processing and switching logic 458, which is typically used to perform various aspects of packet processing and / or switching operations at the local multi-port level rather than across the entire IB switch. Depending on the implementation, packet processing and switching logic 458 can be configured to perform a subset of the operations of packet processing and switching logic 478 within core switch logic 480, or it can be configured to have all the functionality of packet processing and switching logic 478 within core switch logic 480. The processing capabilities of packet processing and switching logic 458 can vary depending on the complexity of the operations and / or the speed of the operations to be performed. For example, packet processing and switching logic 458 can include processors ranging from microcontrollers to multi-core processors. Various types or architectures of multi-core processors can also be used. Furthermore, a portion of the packet processing operations can be implemented using embedded hardware logic.

[0053] The core switch logic 480 includes a crossbar switch 482, a memory 470, a subnet management agent (SMA 476), and packet processing and switching logic 478. The crossbar switch 482 is a non-blocking, low-latency crossbar switch that interconnects multi-port IB interfaces 460A-460D and connects to the memory 470. The memory 470 includes a receive queue 472 and a transmit queue 474. In one embodiment, packets to be exchanged between multi-port IB interfaces 460A-460D can be received by the crossbar switch 482, stored in one of the receive queues 472, processed by the packet processing and switching logic 478, and stored in the transmit queue 474 for transmission to the outbound multi-port IB interface. In implementations that do not use multi-port IB interfaces 460A-460D, the core switch logic 480 and the crossbar switch 482 exchange packets directly between I / O buffers having receive queues 472 and transmit queues 474 within the memory 470.

[0054] The packet processing and switching logic 478 includes programmable functions and can execute data plane programs via multi-core processors of various types or architectures. The packet processing and switching logic 478 represents suitable circuit modules and logic for implementing switching operations, as well as packet processing operations that can be performed outside the port itself. The processing elements of the packet processing and switching logic 478 execute software and / or firmware instructions configured to implement packet processing and switching operations. Such software and / or firmware can be stored in the switch's own non-volatile storage. This software can also be downloaded or updated over the network in conjunction with the initialization operation of the InfiniBand switch 450.

[0055] SMA 476 can be configured to manage, monitor, and control the functions of InfiniBand switch 450. SMA 476 also acts as an agent for and communicates with the subnet manager (SM) of the subnet associated with InfiniBand switch 450. The SM is an entity that discovers devices within a subnet and periodically scans the subnet to detect changes in subnet topology. One SMA within a subnet can be elected as the primary SMA of the subnet and act as the SM. Other SMAs within the subnet will then communicate with that SMA. Alternatively, SMA 476 can operate alongside other SMAs in the subnet to act as a distributed SM. In some embodiments, SMA 476 includes or executes independent circuit modules and logic, such as a microcontroller, a single-core processor, or a multi-core processor. In other embodiments, SMA 476 is implemented via software and / or firmware instructions that execute on a processor core or other processing element that is part of a processor or other processing element used to implement packet processing and switching logic 478.

[0056] The embodiments are not specifically limited to implementations including multi-port IB interfaces 460A-460D. In one embodiment, a port is associated with its own receive and transmit buffers, wherein a crossbar switch 482 is configured to interconnect these buffers with receive queues 472 and transmit queues 474 in memory 470. Packet processing and switching are then performed primarily by packet processing and switching logic 478 of the core switch logic 480.

[0057] Figures 5A-5B An example network interface device is depicted. Figure 5A A network interface device 500 that can be configured as an intelligent Ethernet device is shown. Figure 5B A network interface device 550 that can be configured as an InfiniBand channel adapter is shown.

[0058] like Figure 5A As shown, in one configuration, the network interface device 500 may include a transceiver 502, a transmit queue 507, a receive queue 508, a memory 510, a bus interface 512, and a DMA engine 526. The network interface device 500 may also include a SoC / SiP 545, which includes a processor 505 for implementing intelligent network interface device functions, and accelerators 506 for various acceleration functions, such as NVMe-oF or RDMA. The specific configuration of the network interface device 500 depends on the protocol implemented via the network interface device 500.

[0059] In various configurations, the network interface device 500 can be configured to interface with a network, including but not limited to Ethernet, including Ethernet Ultra. However, the network interface device 500 can also be configured as an InfiniBand or NVLink interface by modifying various components. For example, transceiver 502 can receive and transmit packets conforming to InfiniBand, Ethernet, or NVLink protocols. Other protocols can also be used. Transceiver 502 can receive packets from and transmit packets to the network via the network medium. Transceiver 502 may include a PHY circuit module 514 and a media access control circuit module (MAC circuit module 516). The PHY circuit module 514 may include encoding and decoding circuit modules to encode and decode data packets according to applicable physical layer specifications or standards. The MAC circuit module 516 can be configured to assemble data packets to be transmitted into packets including destination and source addresses, network control information, and error detection hash values.

[0060] The SoC / SiP 545 may include processors, which can be any combination of a CPU processor, a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other programmable hardware devices that allow programming of the network interface device 500. For example, an intelligent network interface may use processor 505 to provide packet processing capabilities within the network interface. The configuration of operation of processor 505 (including a programmable data plane processor) can be programmed using a protocol-independent packet processor (P4), C, Python, the Broadcom Network Programming Language (NPL), x86 or ARM-compatible executable binaries, or other executable binaries.

[0061] Packet distributor 524 can use time slot allocation to distribute received packets for processing by multiple CPUs or cores. Interrupt merging circuitry 522 can perform interrupt regulation, whereby interrupt merging circuitry 522 waits for multiple packets to arrive or waits for timeouts to expire before generating an interrupt to the host system to process the received packet(s). Receive segment merging (RSC) can be performed by network interface device 500, where portions of incoming packets are combined into packets segments. Network interface device 500 can then provide the merged packets to the application. DMA engine 526 can copy packet headers, packet payloads, and / or descriptors directly from host memory to the network interface and vice versa, instead of copying packets to an intermediate buffer at the host and then using another copy operation from the intermediate buffer to the destination buffer. Memory 510 can be any type of volatile or non-volatile memory device and can store any queues or instructions used to program network interface device 500. Transmit queue 507 can include data or references to data for transmission via the network interface. Receive queue 508 can include data or references to data received from the network by the network interface. Descriptor queue 520 may include descriptors referencing data or packets in transmit queue 507 or receive queue 508. Bus interface 512 may provide an interface to a host device. For example, bus interface 512 may be PCI Express compatible, but other interconnect standards may also be used.

[0062] like Figure 5B As shown, network interface device 550 can be configured as an implementation of network interface device 500 to implement InfiniBand HCA. Network interface device 550 includes network ports 552A-552B, memories 554A-554B, a PCIe interface 558, and an integrated circuit 556. Integrated circuit 556 includes hardware, firmware, and / or software for implementing, managing, and / or controlling HCA functions. In one implementation, the integrated circuit includes a hardware transfer engine 560, an RDMA engine 562, congestion control logic 563, virtual endpoint logic 564, offload engine 566, QoS logic 568, GSA / SMA logic 569, and a management interface. Different implementations of network interface device 550 may include additional components or may exclude some components. A network interface device 550 configured as TCA will include a specific subset of implementations of HCA functions. Integrated circuit 556 includes programmable and fixed-function hardware to implement the described functions.

[0063] Although the illustrated implementation of network interface device 550 is shown with PCIe interface 558, other implementations may use other interfaces. For example, network interface device 550 may use an Open Compute Project (OCP) mezzanine connector. Furthermore, PCIe interface 558 may be configured with a multi-host solution that enables multiple compute or storage hosts to couple with network interface device 550. PCIe interface 558 may also support technologies that enable direct PCIe access to multiple CPU sockets, eliminating the need for network traffic to traverse the inter-processor bus of a multi-socket server motherboard that includes network interface device 550.

[0064] Network interface device 550 implements endpoint elements of the InfiniBand architecture based on surrounding queue pairs and RDMA. InfiniBand offloads service control from software by using execution queues (e.g., work queues) initiated by a software client and managed in hardware. Communication endpoints include queue pairs (QPs) with send and receive queues. A QP is a memory-based abstraction in which communication is achieved between applications or between applications and devices for memory-to-memory transfers. Communication with QPs is conducted via virtual channels on network ports 552A-552B, which allows multiple independent data streams to share the same link and provides separate buffering and flow control for each stream.

[0065] Communication occurs via channel I / O, where a virtual channel directly connects two applications existing in separate address spaces. Hardware transfer engine 560 includes hardware logic for performing transport-level operations on endpoints via QPs. RDMA engine 562 utilizes hardware transfer engine 560 to perform RDMA operations between endpoints. RDMA engine 562 implements RDMA operations in hardware and enables applications to read and write to the memory of a remote system without OS kernel intervention or unnecessary data copying by allowing one endpoint of the communication channel to directly place information into the memory of the other endpoint. Virtual endpoint logic 564 manages the operation of virtual endpoints of channel I / O, where channel I / O is a virtual instance of a QP to be used by the application. Virtual endpoint logic 564 maps QPs to the virtual address space of the application associated with the virtual endpoint.

[0066] Congestion control logic 563 performs operations to mitigate congestion on the channel. In various implementations, congestion control logic 563 can perform flow control on the channel to limit congestion at the data transmission destination. Congestion control logic 563 can perform link-level flow control to manage source congestion at the virtual link of network ports 552A-552B. In some implementations, congestion control logic can perform operations to limit congestion at intermediate points along the channel (e.g., IB switches).

[0067] Unloading engine 566 enables the offloading of network tasks that would otherwise be performed in software to network interface device 550. Unloading engine 566 can support the offloading of operations, including but not limited to offloading receive-side scaling or stateless network operations from device drivers, such as TCP / UDP / IP stateless offloading or Virtual Extensible LAN (VXLAN) offloading for TCP implementations on InfiniBand. Unloading engine 566 can also implement... Figure 5A The network interface device 500 operates the interrupt merging circuit 522. The offloading engine 566 can also be configured to support offloading NVME-oF or other storage acceleration operations from the CPU.

[0068] QoS logic 568 can perform QoS operations, including the QoS functions inherent in the basic service delivery mechanism of InfiniBand. QoS logic 568 can also implement enhanced InfiniBand QoS, such as fine-grained end-to-end QoS. QoS logic 568 can implement queuing service and management, prioritizing flows based on flow priority and guaranteeing service levels or bandwidth. For example, QoS logic 568 can configure virtual channel arbitration for the virtual channels of network ports 552A-552B based on flow priority. QoS logic 568 can also operate in conjunction with congestion control logic 563.

[0069] The GSA / SMA logic 569 implements General Service Agent (GSA) operations to manage network interface device 550 and the InfiniBand architecture, as well as to perform subnet management agent operations. GSA operations include device-specific management tasks such as querying device attributes, configuring device settings, and controlling device behavior. The GSA / SMA logic 569 can also implement SMA operations, including those performed by… Figure 4C The SMA 476 of the InfiniBand switch 450 performs a subset of the operations. For example, the GSA / SMA logic 569 can handle management requests from the subnet manager, including device reset requests, firmware update requests, or requests to modify configuration parameters.

[0070] Management interface 570 provides support for hardware interfaces to perform out-of-band management of network interface device 550, such as interconnection with board management controller (BMC) or hardware debug interface.

[0071] Figure 6This is a block diagram illustrating a programmable network interface 600 and a data processing unit. The programmable network interface 600 is a programmable network engine that can be used to accelerate network-based computing tasks in a distributed environment. The programmable network interface 600 can be coupled to a host system via a host interface 670. The programmable network interface 600 can be used to accelerate network or storage operations of the host system's CPU or GPU. For example, the host system can be a node in a distributed learning system used to perform distributed training, such as... Figure 6 As shown. The host system can also be a data center node within a data center.

[0072] In one embodiment, the programmable network interface 600 can accelerate access to a remote storage device containing model data. For example, the programmable network interface 600 can be configured to present the remote storage device as a local storage device to the host system. The programmable network interface 600 can also accelerate RDMA operations performed between the GPU of the host system and the GPU of the remote system. In one embodiment, the programmable network interface 600 can implement storage functions, such as, but not limited to, NVMe-oF. The programmable network interface 600 can also accelerate encryption, data integrity, compression, and other operations on the remote storage device on behalf of the host system, allowing the remote storage device to approach the latency of storage devices directly attached to the host system.

[0073] The programmable network interface 600 can also perform resource allocation and management on behalf of the host system. Storage security operations can be offloaded to the programmable network interface 600 and performed in conjunction with the allocation and management of remote storage device resources. Network-based operations for managing access to remote storage devices can be performed by the programmable network interface 600 instead of by the host system's processor.

[0074] In one embodiment, network and / or data security operations can be offloaded from the host system to the programmable network interface 600. Data center security policies for data center nodes can be handled by the programmable network interface 600 instead of the host system's processor. For example, the programmable network interface 600 can detect and mitigate network-based attacks (such as DDoS) attempted against the host system, preventing attacks from compromising the host system's availability.

[0075] The programmable network interface 600 may include a system-on-a-chip (SoC / SiP 620) that executes an operating system via multiple processor cores 622. The processor cores 622 may include general-purpose processor (e.g., CPU) cores. In one embodiment, the processor cores 622 may also include one or more GPU cores. The SoC / SiP 620 may execute instructions stored in a memory device 640. The memory device 650 may store local operating system data. The memory device 650 and memory device 640 may also be used to cache remote data of the host system. Network ports 660A-660B enable connection to a network or infrastructure and facilitate network access for the SoC / SiP 620, as well as network access to the host system via the host interface 670. In one configuration, a first network port 660A may be connected to a first forwarding element, and a second network port 660B may be connected to a second forwarding element. Alternatively, both network ports 660A-660B may be connected to a single forwarding element using a link aggregation protocol (LAG). The programmable network interface 600 may also include an I / O interface 675, such as a universal serial bus (USB) interface. I / O interface 675 can be used to couple external devices to programmable network interface 600 or as a debug interface. Programmable network interface 600 also includes management interface 630, which enables software on the host device to manage and configure programmable network interface 600 and / or SoC / SiP 620. In one embodiment, programmable network interface 600 may also include one or more accelerators or GPUs 645 to accept parallel computing tasks offloaded from the SoC / SiP 620, host system, or remote system coupled via network ports 660A-660B. For example, programmable network interface 600 may be configured with a graphics processor and participate in general-purpose or graphics computing operations in a data center environment.

[0076] One or more aspects can be implemented by representative code stored on a machine-readable medium, which represents and / or defines the logic within an integrated circuit (such as a processor). For example, the machine-readable medium may include instructions representing various logics within a processor. When read by a machine, the instructions can cause the machine to manufacture the logic to perform the techniques described herein. This representation, called an "IP core," is a reusable logic unit of the integrated circuit and can be stored on a tangible machine-readable medium as a hardware model describing the structure of the integrated circuit. The hardware model can be provided to various customers or manufacturing facilities that load the hardware model onto manufacturing machines that manufacture the integrated circuit. Integrated circuits can be manufactured such that the circuit performs the operations described in association with any of the embodiments described herein.

[0077] Figure 7This is a block diagram illustrating an IP core development system 700. The IP core development system 700 can be used to manufacture integrated circuits to perform the operations of the architectures and data center components described herein. The IP core development system 700 can be used to generate modular, reusable designs that can be incorporated into larger designs or used to build entire integrated circuits (e.g., SOC integrated circuits). Design facility 730 can generate software simulations 710 of the IP core design using a high-level programming language (e.g., C / C++). Software simulations 710 can be used to design, test, and verify the behavior of the IP core using simulation models 712. Simulation models 712 can include functional, behavioral, and / or timing simulations. Register transfer level designs (RTL designs 715) can then be created or synthesized from simulation models 712. RTL designs 715 are abstractions of the behavior of an integrated circuit that models the flow of digital signals between hardware registers, including associated logic performed using the modeled digital signals. In addition to RTL designs 715, low-level designs at the logic level or transistor level can also be created, designed, or synthesized. Therefore, the specific details of the initial design and simulation may differ.

[0078] The RTL design 715 or its equivalent can be further synthesized into a hardware model 720 by a design facility. The hardware model 720 can be some other representation of hardware description language (HDL) or physical design data. The HDL can be further simulated or tested to validate the IP core design. The IP core design can be stored using non-volatile memory 740 (e.g., hard disk, flash memory, or any non-volatile storage medium) for delivery to a manufacturing facility 765. The manufacturing facility 765 can be a third-party manufacturing facility. Alternatively, the IP core design can be transmitted via a wired connection 750 or a wireless connection 760 (e.g., via the Internet). The manufacturing facility 765 can then fabricate an integrated circuit at least partially based on the IP core design. The fabricated integrated circuit can be configured to perform operations according to at least one embodiment described herein.

[0079] Resource Deployment in Expert Hybrid Processing Tokens are typically used in processing involving multiple processors, such as systems containing a large number of accelerators (like GPUs) across multiple servers. Hybrid Expert (MoE) techniques, which can involve multiple experts hosted by GPUs or other accelerators or units, enable models to be pre-trained with significantly less computation compared to other model training methods. Experts are neural networks (e.g., feedforward networks (FFNs)), where experts are designated by routing tokens to them. Typically, gating networks or routers determine which tokens to send to which experts based on some algorithm already applied in the implementation.

[0080] Routers consist of learned parameters and are typically pre-trained simultaneously with the rest of the network. In a typical deployment, one or more experts may occupy a single GPU, and multiple GPUs across multiple servers can form clusters that include large mixtures of experts. To maximize the efficiency and accuracy of the operations processed by MoE, it is important to focus on the routing, load balancing, and inference time provisioning of experts across devices.

[0081] In some embodiments, the routing capabilities of the expert hybrid deployment are implemented in separate processing resources, which may include hardware accelerators such as IPUs (Infrastructure Processing Units), Data Processing Units (DPUs), SmartNICs (Smart Network Interface Cards), and other units.

[0082] Figure 8 This is an illustration of a computing system or device, according to some embodiments, that includes routing capabilities for handling mixed deployments of experts in processing resources. In some embodiments, the computing system or device 800 includes hardware components for supporting high-bandwidth data transmission in AI processing. Figure 8 As shown, a computing system or device 800 (which may include Figure 1 The processing system 100 shown includes processing resources 805, which may include one or more central processing units (CPUs) or other general-purpose processors 807, multiple graphics processing units (GPUs) 810 (which may be configured as a set or cluster for processing), and one or more hardware accelerators 830 (which may include infrastructure processing units (IPUs), data processing units (DPUs), or other resources). The one or more processing resources may include, but are not limited to, [examples of such resources]. Figures 1 to 8 The components shown are shown. The computing system or device 800 may also include computer memory 840, such as high-bandwidth memory (HBM), and computer storage device 845, such as solid-state drive (SSD), hard disk drive (HDD), or other storage technologies, to store data for processing.

[0083] In some embodiments, the hardware accelerator 830 includes, but is not limited to, an expert hybrid (MoE) circuit module 835 for providing routing capabilities for expert hybrid deployments. In some embodiments, the circuit module includes... Figure 13A The diagram shows one or more elements for processing resources. The hardware accelerator may also include one or more network ports for connectivity over a network, and one or more direct memory access (DMA) hardware engines for transferring data over the network.

[0084] Figure 9 This illustrates the application of expert integration in AI processing. One component of the MoE topology is a gating network or router that determines which tokens are sent to which expert in the system. Figure 9 A specific example is shown where the router determines which experts should receive tokens 1 through 8 for use with a specified word in the LLM.

[0085] How to route tokens to expert representatives is one of the key decisions to be determined when operating the system using MoE. The router itself (such as...) Figure 9 The parameters shown (as illustrated) consist of those learned through pre-training over a network. In a typical deployment, one or more experts may occupy a single GPU (or other accelerator), and multiple GPUs across multiple servers form a cluster that includes a large mix of experts.

[0086] In such infrastructures, the routing, load balancing, and inference time provisioning of experts across such devices are extremely important for providing efficient operation. In some embodiments, these aspects of the MoE are managed using processing resources that include expert tracking and management capabilities.

[0087] Figure 10 This is a diagram illustrating the routing in expert hybrid operations. Figure 10 In the simplified example 1, the route from the token from device n to the expert presented by expert 1 at device 0, expert 2 at device 1 and expert 3 at device 2 is shown (with a capacity factor of 1.0).

[0088] However, routing too many tokens to a single expert will cause some tokens to be discarded. As shown in Example 1, 3 tokens are routed to expert 1, 2 tokens are routed to expert 2, and 1 token is routed to expert 3. In this example, the routing of 3 tokens to expert 1 may result in one of the tokens being discarded because the device's capacity has been exceeded.

[0089] In contrast, Example 2 similarly provides routing from a token from device n to experts presented by expert 1 at device 0, expert 2 at device 1, and expert 3 at device 2. In this instance, the capacity factor is 1.5, thus allowing the token to be routed without the experts discarding it.

[0090] Figures 11A-11F This is a diagram illustrating routing algorithms in a hybrid expert operation. The diagram shows some algorithms applying different routing strategies that can be employed and optimized within the MoE operation. However, the embodiments are not limited to these examples.

[0091] In the Figures 11A-11F The provided illustration shows: Figure 11A This shows an example of a Top-1 route, where the maximum value is routed to the expert.

[0092] Figure 11BThis shows an example of Top-2 routing, where the two maximum values ​​are routed to the expert.

[0093] Figure 11C An example of hash routing is shown, where a hash function is used to identify the route of a token.

[0094] Figure 11D An example of implementing expert selection tokens is shown, assuming multiple tokens are provided to multiple routers, where the expert selection is to be routed to a particular expert's token.

[0095] Figure 11E An example of BASE (Balanced Assignment by Experts) routing for tokens is shown, where linear assignment is resolved.

[0096] Figure 11F An example of reinforcement learning routing for tokens is shown.

[0097] Given that token routing plays a major role in expert hybrid deployments, the operational availability and accessibility of experts within the expert ensemble can be considered. For example, during inference time, an expert might become unavailable if it is oversubscribed or if there is severe network congestion. In such cases, data could be routed to different experts, as using slower experts would slow down the rest of the batch and reduce overall GPU utilization, leading to increased costs and reduced efficiency.

[0098] In some embodiments, the device, system, or process includes one or more of the following capabilities to provide scalable, universal expert hybrid deployment for processing within the network: (1) Perform expert capacity tracking over the network with minimal overhead; (2) Perform load balancing among experts with minimal overhead; (3) Track network services over the network with minimal overhead; (4) Dynamically check the distribution of experts, which can be used, for example, to dynamically migrate or replicate a subset of experts across clusters.

[0099] Existing solutions typically lack these capabilities because they are implemented in software, which incurs high operational overhead. However, instead, implementing complex distributed routers across one or more GPUs in the network reduces the need for and utilization of the expensive available GPU resources required for data processing. In some embodiments, routing support is instead implemented in separate processing resources, such as IPUs, which contain hardware circuitry modules designed to provide routing support for MoE operations.

[0100] Figure 12 This is an illustration of a system for resource deployment utilizing improved expert hybrid processing, according to some embodiments. Figure 12 In this system 1200, routing capabilities for expert hybrid deployment are implemented in one or more processing resources, which may include IPUs (Infrastructure Processing Units) and other hardware. Compared to software implementations, hardware implementations (such as IPUs) allow for significantly greater efficiency because, among other potential advantages, hardware offers better access to telemetry, network statistics, and other data, and can be used for tracking and advanced processing to achieve improved operation for expert hybrid deployments.

[0101] like Figure 12 As shown, system 1200 includes multiple GPUs or other accelerators 1210, at least a subset of which can operate as experts in an expert hybrid implementation. In some embodiments, system 1200 also includes one or more additional processing resources, which may include the illustrated IPU 1220. IPU 1220 includes hardware circuitry modules for providing expert hybrid operation, including providing tracking, prediction, and routing for experts, as well as providing migration and replication of certain experts.

[0102] In some embodiments, the processing resource maintains a state table that can be used to track token leads on a per-expert basis, along with network congestion statistics for different experts. Furthermore, the processing resource tracks the different GPUs / nodes deployed for experts, as well as the load across expert instances. The hardware may also include circuitry modules for predicting expert load based on historical tracking, further tracking previous history, and predicting the relevance / weight of experts in scoring for a given token. If an expert is found to have consistently low relevance, it may be necessary to retrain the expert, and the processing resource can facilitate signaling for retraining. The processing resource may also include the ability to migrate or replicate experts as needed.

[0103] Figure 13A This is an illustration of resource deployment utilizing improved expert hybrid processing according to some embodiments. In some embodiments, hardware accelerator 1300 (which may include IPU, DPU, SmartNIC, or other resources) includes circuit module 1305 for expert hybrid (MoE) processing. Circuit module 1305 includes one or more of the following mechanisms: Expert token capacity and network load tracker circuit module 1310—Expert Token Capacity and Network Load Tracker Circuit Module 1310 includes the ability to process resources to track different deployments of experts in an associated system, along with GPU IDs and network load. As a result, the hardware accelerator 1300 has visibility into different instances of experts, as well as the utilization and load of those instances. This, combined with network load, provides a dynamic view of expert election and assists processing resources in making high-level routing decisions for expert selection on a dynamic basis. The data maintained by the tracker circuit module 1310 is shown as a state table 1350, which... Figure 13B This is further shown in the middle.

[0104] Expert historical load tracker circuit module 1315 — In the processing implementation, past behavior in the process typically provides a good predictor of future operations, and there is a high correlation between the sequence of tokens and a small subset of experts being utilized. For this reason, the expert history load tracker circuit module 1315 can historically track this correlation by storing hot token sequences (i.e., repeating token sequences) and their corresponding experts, wherein this tracking is performed in hardware within the hardware accelerator 1300.

[0105] Expert Token Load Prediction Circuit Module 1320 – Expert Token Load Prediction Circuit Module. The 1320 can be used to provide predictions of an expert's future token load, at least in part, based on system knowledge. Because a predictive expert capacity is established at inference time, the prediction circuit module can implement different predictors that are adapted based on learning.

[0106] Expert Relevance Weight Tracker Circuit Module 1325—Expert Relevance Weight Tracker Circuit Module 1325 tracks the relevance weights of experts processing tokens in the system. Given that some experts may have better results for one or more tokens, the expert relevance weights are tracked and updated in the hardware accelerator 1300 and can be used for token routing. Furthermore, if an expert is found to have consistently low relevance, it may be necessary to retrain that expert, and the hardware accelerator 1300 can facilitate signaling for retraining based at least in part on the relevance weights tracked by the expert relevance weight tracker circuit module 1325.

[0107] Expert routing circuit module 1330—Expert Routing Circuit Module 1330 provides management of token-to-expert routing based on collected expert data, said data including data from one or more of the following: Expert Token Capacity and Network Load Tracker Circuit Module 1310, Expert Historical Load Tracker Circuit Module 1315, Expert Token Load Prediction Circuit Module 1320, and Expert Relevance Weight Tracker Circuit Module 1325. Token routing management may include adjusting token routing to avoid problems detected in expert hybrid operations, said problems may include overload of one or more experts holding tokens.

[0108] Expert migration circuit module1330 – Expert Migration Circuit Module 1330 can be used to facilitate the migration or replication of expert operations to other hardware resources. When required to resolve load or other considerations, a specific expert can be migrated (i.e., relocated) from a first GPU to a second GPU, or a specific expert can be replicated to a second GPU. In the example, when a specific expert is oversubscribed (which may occur where there is a large inconsistency in how the expert is utilized for a given sequence of tokens), the migration circuit module facilitates the replication or migration of that specific expert to another resource to offload at least a portion of the token processing load for that expert.

[0109] Accelerator 1300 includes additional components, including but not limited to one or more network ports 1360 for connection over a network, and one or more direct memory access (DMA) engines 1365 for data transfer. Accelerator 1300 may also include memory, cache memory, and data storage devices (which are located in...). Figure 13A One or more of (not shown) are used to store data for processing, including data for MoE processing.

[0110] Figure 13B This illustrates expert data maintained for expert hybrid processing according to some embodiments. In some embodiments, the status table 1350 is generated by one or more processing resources in the system (such as...). Figure 13A The hardware accelerator 1300 shown is used to maintain this.

[0111] like Figure 13B As shown, the state table 1350 maintained by one or more processing resources may include, but is not limited to: (a) The expert’s identifier, which may include the instance identifier.

[0112] (b) The token capacity of the corresponding expert.

[0113] (c) The current token load of the corresponding expert.

[0114] (d) Identifier of the GPU hosting the expert, where the GPU can host one or more experts, depending on the specific implementation.

[0115] (e) Identification of related sibling experts and host identification.

[0116] (f) The current network load of the system.

[0117] Figure 14This is a flowchart illustrating a process of expert hybrid operation in network processing according to some embodiments. In some embodiments, process 1400 includes receiving data for processing by multiple processing resources in the network (1405), said processing resources may include multiple GPUs. The processing may include processing using a model, wherein the model may be a large language model or other AI model.

[0118] In operation, the network includes an expert hybrid operation (1410) in the processing of the model, wherein the experts are hosted by GPUs or other processing resources. The processing may include routing tokens associated with the processing of the model to the experts according to an algorithm (1415).

[0119] Process 1400 may further include: monitoring expert blending operations during processing (1420) and maintaining collected data regarding experts and expert blending operations (1425). The tracked data may include data regarding expert token capacity and network token load. The operation may further include: tracking the interrelationships between tokens and the experts used for the tokens (1430) and generating predictions of future token loads for the experts (1435). Furthermore, relevance weights for the experts processing the tokens may be determined (1440). In some cases, when it is determined that one or more experts provide low relevance during operation, signals for retraining the one or more experts may be generated (1445).

[0120] In some embodiments, the routing of tokens to experts is managed (1450) based at least in part on the collected expert data. The management of token routing may include adjusting token routes to avoid problems detected in expert fusion operations, where the problems may include overload of one or more experts holding tokens. Furthermore, process 1400 may also include migrating one or more experts from a first location to a second location, or duplicating one or more experts (1455), as needed, wherein the migration or duplication may be implemented to offload at least a portion of the token processing load for the one or more experts, or to additionally address problems in expert fusion processing.

[0121] The following examples relate to certain implementations: In Example 1, a device includes: one or more network ports; one or more direct memory access (DMA) engines; and a circuit module for hybrid expert (MoE) processing in a network, wherein the circuit module includes at least: a circuit module for tracking the routing of tokens in the MoE processing; a prediction circuit module for generating predictions about the MoE processing, the predictions including predictions of future token loads in the MoE processing; and a routing management circuit module for managing the routing of the tokens in the MoE processing based at least in part on the predictions about the MoE processing.

[0122] In Example 2, for the device of Example 1, the data maintained by the circuit module regarding the MoE processing includes at least the token capacity and token load data for the MoE processing.

[0123] In Example 3, for the device of Example 1 or 2, the circuit module further includes a circuit module for tracking the historical token load of experts in the MoE process.

[0124] In Example 4, for any of the devices in Examples 1 to 3, the circuit module further includes a circuit module for tracking the expert relevance weights of experts in the MoE process.

[0125] In Example 5, for any of the devices in Examples 1 to 4, the circuit module determines whether the expert in the MoE process needs to be retrained based on the expert's relevance weights.

[0126] In Example 6, for any of the devices in Examples 1 to 5, the circuit module further includes a circuit module for performing one or more of the following: migrating one or more experts from a first accelerator to a second accelerator, or replicating one or more experts.

[0127] In Example 7, for any of the devices in Examples 1 to 6, the device includes an Infrastructure Processing Unit (IPU).

[0128] In Example 8, a device includes: a memory; a plurality of processors, including a plurality of graphics processing units (GPUs); and one or more hardware accelerators, including a circuit module for managing expert hybrid (MoE) processing, wherein the circuit module includes at least: a circuit module for tracking the routing of tokens in the MoE processing; a prediction circuit module for generating a prediction about the MoE processing, the prediction including a prediction of future token loads in the MoE processing; and a routing management circuit module for managing the routing of the tokens in the MoE processing based at least in part on the prediction about the MoE processing.

[0129] In Example 9, for the device of Example 8, the data maintained by the circuit module regarding the MoE processing includes at least token capacity and token payload data for the MoE processing.

[0130] In Example 10, for the device of Example 8 or 9, the circuit module further includes a circuit module for tracking the historical token load of experts in the MoE process.

[0131] In Example 11, for any of the devices in Examples 8 to 10, the circuit module further includes a circuit module for tracking the expert relevance weights of experts in the MoE process.

[0132] In Example 12, for any of the devices in Examples 8 to 11, the circuit module determines whether the expert in the MoE process needs to be retrained based on the expert's relevance weights.

[0133] In Example 13, for any of the devices in Examples 8 to 12, the circuit module further includes a circuit module for performing one or more of the following: migrating one or more experts from a first GPU to a second GPU, or copying one or more experts.

[0134] In Example 14, for any of the devices in Examples 8 to 13, the one or more hardware accelerators include one or more infrastructure processing units (IPUs).

[0135] In Example 15, a method includes: receiving data for processing a model by a plurality of graphics processing units (GPUs) in a network; routing tokens associated with processing of the model; monitoring operations of MoE processing in the network; generating a prediction about the MoE processing, the prediction including a prediction of future token loads in the MoE processing; and managing the routing of the tokens in the MoE processing based at least on the generated prediction.

[0136] In Example 16, for the method of Example 15, the maintained data regarding the MoE processing includes at least one expert's token capacity and token payload data.

[0137] In Example 17, for the method of Example 15 or 16, the method further includes tracking the historical token payload processed by MoE.

[0138] In Example 18, for any of the methods in Examples 15 to 17, the method further includes tracking the expert relevance weights of experts in the MoE process.

[0139] In Example 19, for any of the methods in Examples 15 to 18, the method further includes determining whether the expert needs to be retrained in the MoE process based on the expert's relevance weights.

[0140] In Example 20, for any of the methods in Examples 15 to 19, the method further includes one or more of the following: migrating one or more experts from a first GPU to a second GPU; or replicating one or more experts.

[0141] In Example 21, one or more non-transitory computer-readable storage media have stored thereon executable computer program instructions that, when executed by one or more processors, cause the one or more processors to perform operations including: receiving data for processing a model by a plurality of graphics processing units (GPUs) in a network; routing tokens associated with processing the model; monitoring operations of MoE processing in the network; generating a prediction about the MoE processing, the prediction including a prediction of future token loads in the MoE processing; and managing the routing of the tokens in the MoE processing based at least on the generated prediction.

[0142] In Example 22, for one or more storage media of Example 21, the maintained data regarding the MoE processing includes at least one expert's token capacity and token payload data.

[0143] In Example 23, for one or more storage media of Example 21 or 22, the instructions also include instructions for tracking historical token loads of MoE processing.

[0144] In Example 24, for one or more storage media of any of Examples 21 to 23, the instructions further include instructions for tracking expert relevance weights of experts in MoE processing.

[0145] In Example 25, for one or more storage media of any of Examples 21 to 24, the instructions further include instructions for determining whether the expert needs to be retrained in the MoE process based on the expert's relevance weights.

[0146] In Example 26, for one or more storage media of any of Examples 21 to 25, the instructions further include instructions for one or more of the following: migrating one or more experts from a first GPU to a second GPU; or copying one or more experts.

[0147] In Example 27, an apparatus includes: components for receiving data for processing a model by a plurality of graphics processing units (GPUs) in a network; components for routing tokens associated with processing of the model; components for monitoring the operation of MoE processing in the network; components for generating a prediction about the MoE processing, the prediction including a prediction of future token loads in the MoE processing; and components for managing the routing of the tokens in the MoE processing based at least on the generated prediction.

[0148] In Example 28, for the device of Example 27, the maintained data regarding the MoE processing includes at least one expert's token capacity and token payload data.

[0149] In Example 29, for the device of Example 27 or 28, the device also includes a component for tracking historical token payloads of MoE processing.

[0150] In Example 30, for any of the devices in Examples 27 to 29, the device further includes a component for tracking expert relevance weights of experts in the MoE process.

[0151] In Example 31, for any of the devices in Examples 27 to 30, the device further includes a component for determining whether the expert needs to be retrained in the MoE process based on the expert's relevance weights.

[0152] In Example 32, for any of the devices in Examples 27 to 31, the device further includes components for one or more of the following: migrating one or more experts from a first GPU to a second GPU; or replicating one or more experts.

[0153] While the descriptions and illustrations of the embodiments provided herein depict specific components, those skilled in the art will recognize that such components may be combined into fewer elements as required or convenient for a particular implementation, or may be broken down into a larger number of components. For example, the components provided for sorting and buffering input may be represented as a first sorting component and a second buffering component.

[0154] In the foregoing description, numerous specific details have been set forth for purposes of explanation in order to provide a thorough understanding of the embodiments. However, those skilled in the art will understand that the embodiments may be practiced without some of these specific details. In other instances, well-known structures and apparatuses are shown in block diagram form. Intermediate structures may exist between the illustrated components. Components described or illustrated herein may have additional inputs or outputs not shown or described.

[0155] Various embodiments may include various processes. These processes may be executed by hardware components or may be implemented in computer programs or machine-executable instructions that can be used to cause a general-purpose or special-purpose processor or logic circuit programmed by said instructions to perform the processes. Alternatively, the processes may also be executed by a combination of hardware and software.

[0156] Parts of the various embodiments may be provided as computer program products, which may include computer-readable media having stored thereon computer program instructions that can be used to program a computer (or other electronic device) for execution by one or more processors to perform processes according to certain embodiments. Computer-readable media may include, but are not limited to, magnetic disks, optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or other types of computer-readable media suitable for storing electronic instructions. Furthermore, embodiments may also be downloaded as computer program products, wherein the program may be transferred from a remote computer to the requesting computer.

[0157] Many methods are described in their most basic form, but processes may be added to or removed from any method, and information may be added to or subtracted from any of the stated messages without departing from the basic scope of the presented embodiments. Those skilled in the art will understand that many further modifications and adaptations are possible. Specific embodiments are provided not to limit the concept, but to illustrate it. The scope of the embodiments is not determined by the specific examples provided above, but only by the following claims.

[0158] If element "A" is coupled to or coupled to element "B", then element A may be directly coupled to element B, or indirectly coupled through, for example, element C. When the specification or claims state that component, feature, structure, process, or characteristic A "causes" component, feature, structure, process, or characteristic B, this means that "A" is at least partly responsible for "B", but at least one other component, feature, structure, process, or characteristic may also contribute to causing "B". If the specification indicates that a component, feature, structure, process, or characteristic "may", "may", or "can" be included, it does not require that particular component, feature, structure, process, or characteristic be included. If the specification or claims refer to "a" ("a" or "an") element, this does not mean that only one of the elements is present.

[0159] An embodiment is an implementation or example. References to “an embodiment,” “one embodiment,” “some embodiments,” or “other embodiments” in this specification mean that a particular feature, structure, or characteristic associated with that embodiment is included in at least some, but not necessarily all, embodiments. Various appearances of “an embodiment,” “one embodiment,” or “some embodiments” do not necessarily refer to the same embodiment. It should be understood that in the foregoing description of exemplary embodiments, various features are sometimes combined in a single embodiment, drawing, or description thereof for the purpose of simplifying disclosure and aiding understanding of one or more novel aspects. However, this method of disclosure should not be construed as reflecting an intention to require more features than expressly recited in the claims. Rather, as reflected in the following claims, novel aspects exist in fewer than all features of a single foregoingly disclosed embodiment. Therefore, the claims are hereby expressly incorporated into this specification, while the claims themselves stand alone as separate embodiments.

[0160] The foregoing description and accompanying drawings are to be regarded in an illustrative rather than restrictive sense. Those skilled in the art will understand that various modifications and alterations can be made to the embodiments described herein without departing from the broader spirit and scope of the features described in the appended claims.

Claims

1. An apparatus comprising: One or more network ports; One or more Direct Memory Access (DMA) engines; as well as Circuit modules for hybrid expert (MoE) processing in networks; The circuit module includes at least: A circuit module used to track the routing of tokens in MoE processing. A prediction circuit module is used to generate predictions about the MoE process, including predictions of future token loads for the MoE process. A routing management circuit module for managing the routing of the tokens in the MoE process based at least in part on the predictions regarding the MoE process.

2. The device as claimed in claim 1, wherein, The data maintained by the circuit module regarding the MoE processing includes at least the token capacity and token payload data used for the MoE processing.

3. The device as claimed in claim 1, wherein, The circuit module also includes a circuit module for tracking the historical token load of experts in the MoE process.

4. The device as claimed in claim 1, wherein, The circuit module also includes a circuit module for tracking the expert relevance weights of experts in the MoE process.

5. The device as claimed in claim 4, wherein, The circuit module determines whether the expert in the MoE process needs to be retrained based on the expert's relevance weights.

6. The device as claimed in claim 1, wherein, The circuit module also includes a circuit module for performing one or more of the following: migrating one or more experts from the first accelerator to the second accelerator, or replicating one or more experts.

7. The device as claimed in claim 1, wherein, The device includes an infrastructure processing unit (IPU).

8. An apparatus comprising: Memory; Multiple processors, including multiple graphics processing units (GPUs); as well as One or more hardware accelerators, including circuit modules for managing expert hybrid (MoE) processing; The circuit module includes at least: A circuit module used to track the routing of tokens in MoE processing. A prediction circuit module is used to generate predictions about the MoE process, including predictions of future token loads for the MoE process. A routing management circuit module for managing the routing of the tokens in the MoE process based at least in part on the predictions regarding the MoE process.

9. The device as claimed in claim 8, wherein, The data maintained by the circuit module regarding the MoE processing includes at least token capacity and token load data for the MoE processing.

10. The device as claimed in claim 8, wherein, The circuit module also includes a circuit module for tracking the historical token load of experts in the MoE process.

11. The device as claimed in claim 8, wherein, The circuit module also includes a circuit module for tracking the expert relevance weights of experts in the MoE process.

12. The device as claimed in claim 11, wherein, The circuit module is required to determine, at least in part, whether the expert in the MoE processing needs to be retrained based on the expert's relevance weights.

13. The device as claimed in claim 8, wherein, The circuit module also includes a circuit module for performing one or more of the following: migrating one or more experts from a first GPU to a second GPU, or copying one or more experts.

14. The device as claimed in claim 8, wherein, The one or more hardware accelerators include one or more infrastructure processing units (IPUs).

15. A method comprising: Receive data for processing the model by multiple graphics processing units (GPUs) in the network; The token associated with the routing and processing of the model; Monitor the operations of MoE processing in the network; Generate predictions about the MoE process, including predictions of future token loads in the MoE process; and The routing of the tokens in the MoE process is managed at least based on the generated predictions.

16. The method of claim 15, wherein, The data maintained regarding the MoE processing includes at least one expert's token capacity and token load data.

17. The method of claim 15, further comprising: Track historical token payloads processed by MoE.

18. The method of claim 15, further comprising: Track the expert relevance weights of experts in MoE processing.

19. The method of claim 18, further comprising: The relevance weights of the experts are used to determine whether the experts in the MoE process need to be retrained.

20. The method of claim 15, further comprising one or more of the following: Migrate one or more experts from the first GPU to the second GPU; or Copy one or more experts.

21. An apparatus comprising: A component for receiving data for processing a model by multiple graphics processing units (GPUs) in a network; A component used for routing tokens associated with the processing of the model; Components used to monitor the operation of MoE processing in the network; Components for generating predictions about the MoE process, the predictions including predicting future token loads in the MoE process; and A component for managing the routing of the tokens in the MoE process, at least based on the generated predictions.

22. The device as claimed in claim 21, wherein, The data maintained regarding the MoE processing includes at least one expert's token capacity and token load data.

23. The apparatus of claim 21, further comprising: A component used to track historical token payloads processed by MoE.

24. The apparatus of claim 21, further comprising: A component used to track the expert relevance weights of experts in MoE processing; as well as A component used to determine whether an expert needs to be retrained in the MoE process based on the expert's relevance weights.

25. The device of claim 21, further comprising one or more of the following: Components used to migrate one or more experts from a first GPU to a second GPU; or Used to replicate components of one or more experts.