Network pipeline abstraction layer (NPAL) split interface

By implementing flexible boot rules and configurable dynamic interface mapping for hardware acceleration on the SFC architecture of the DPU, combined with NPAL-optimized network acceleration, the problems of insufficient flexibility and efficiency in the existing SFC architecture are solved, and efficient network traffic processing and management are achieved.

CN121940152APending Publication Date: 2026-04-28MELLANOX TECHNOLOGIES LTD(IL)
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MELLANOX TECHNOLOGIES LTD(IL)
Filing Date
2025-10-27
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The existing SFC architecture does not support the creation of flexible bootstrapping rules and configurable dynamic interface mappings in a single accelerated data plane on the DPU, and does not always support acceleration for all operations of the SFC architecture.

Method used

It provides flexible boot rules for hardware acceleration on the SFC architecture of the DPU, configurable dynamic SFC interfaces, and network pipeline abstraction layer (NPAL) to optimize network acceleration. By enabling virtual bridges and virtual ports on the DPU, it supports accelerated processing of multiple network protocols and functions.

Benefits of technology

It enables flexible and efficient processing of network traffic data on the DPU, improving the scalability and management efficiency of network functions, and reducing operational complexity and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940152A_ABST
    Figure CN121940152A_ABST
Patent Text Reader

Abstract

The invention relates to a network pipeline abstraction layer (NPAL) split interface. Techniques for creating optimized and accelerated network pipelines using a network pipeline abstraction layer (NPAL) for splitting interfaces are described. The DPU includes: a physical port configured to couple to a branch cable, the branch cable physically coupled to a set of a plurality of devices; a DPU hardware; and a memory operably coupled to the DPU hardware. The NPAL supports a plurality of logical split ports, each logical split port corresponding to one of the plurality of devices, where the network pipeline includes a set of tables organized in a particular order and logic to be accelerated by the acceleration hardware engine. The acceleration hardware engine is used to process network traffic data using a network pipeline.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application is a continuation-in-part of U.S. Patent Application No. 18 / 649,319, filed April 29, 2024, the entire contents of which are incorporated herein by reference. This application relates to U.S. Patent Application No. 18 / 649,295, filed April 29, 2024, which is jointly assigned; U.S. Patent Application No. 18 / 649,334, filed April 29, 2024; U.S. Patent Application No. 18 / 928,781, filed concurrently with this application, entitled “Network Pipeline Abstraction Layer (NPAL) Fast Link Recovery”; U.S. Patent Application No. 18 / 928,794, filed concurrently with this application, entitled “Hardware-accelerated Policy-Based Routing (PBR) over Service Function Chaining (SFC)”; and U.S. Patent Application No. 18 / 928,772, filed concurrently with this application, entitled “Network Pipeline Abstraction Layer (NPAL) Emulation”. Technical Field

[0003] At least one embodiment relates to processing resources used to perform and facilitate operations for providing a split interface for the Network Pipeline Abstraction Layer (NPAL). For example, at least one embodiment relates to a processor or computing system used by an acceleration hardware engine to provide and enable the split interface for processing network traffic data in a single accelerated data plane, according to various novel techniques described herein. Background Technology

[0004] In traditional network architectures, various security and performance functions are managed by dedicated hardware devices known as middleboxes, each playing a distinct role. Firewalls, operating as independent physical facilities, act as the primary defense mechanism at the network edge, carefully scrutinizing incoming and outgoing traffic based on predefined rules to block or allow data transmission, thus protecting the internal network from external threats. Load balancers, operating as separate hardware units, intelligently distribute incoming network and application traffic across multiple servers to prevent overload and ensure efficient resource utilization, thereby enhancing application availability and performance. Strategically located intrusion detection systems (IDS) within the network monitor and analyze network traffic to detect anomalies, attacks, or signs of security policy violations, serving as a security component in identifying potential security vulnerabilities.

[0005] Additionally, the network utilizes other middlebox functions such as Data Loss Prevention (DLP) systems to monitor and prevent unauthorized data exfiltration, Virtual Private Network (VPN) gateways to establish secure, encrypted connections on the network, and WAN optimization facilities designed to improve data transmission efficiency over wide area networks. While indispensable, these middleboxes present several challenges: they require significant capital investment, occupy valuable data center space, and demand professional operation and maintenance. Expanding these network capabilities typically means acquiring and integrating more physical devices, increasing the complexity and cost of the network infrastructure. Attached Figure Description

[0006] Various embodiments according to this disclosure will be described with reference to the accompanying drawings, in which:

[0007] Figure 1 It is a block diagram of an integrated circuit having SFC logic for generating virtual bridges and interface mappings in a Service Function Chain (SFC) architecture, according to at least one embodiment.

[0008] Figure 2 This is a block diagram of an example DPU-based Service Function Chain (SFC) infrastructure for providing an SFC architecture according to at least one embodiment.

[0009] Figure 3 It is a block diagram of an SFC architecture having a first virtual bridge, a second virtual bridge, a virtual port, and network services according to at least one embodiment.

[0010] Figure 4 It is a block diagram of an SFC architecture having a first virtual bridge, a second virtual bridge, a virtual port, and network services according to at least one embodiment.

[0011] Figure 5It is a block diagram of a non-SFC architecture having a first virtual bridge and network services according to at least one embodiment.

[0012] Figure 6 This is a flowchart of an example method for configuring an SFC architecture with multiple virtual bridges and interface mappings according to at least one embodiment.

[0013] Figure 7 This is a block diagram of an example DPU-based SFC infrastructure for providing hardware acceleration rules for an SFC architecture 220 according to at least one embodiment.

[0014] Figure 8 It is a block diagram of an SFC architecture having flexible hardware acceleration rules for a single accelerated data plane according to at least one embodiment.

[0015] Figure 9 It is a block diagram of an SFC architecture having flexible hardware acceleration rules for a single accelerated data plane according to at least one embodiment.

[0016] Figure 10 This is a flowchart of an example method for configuring an SFC architecture with flexible hardware acceleration rules for acceleration on a single accelerated data plane of the DPU, according to at least one embodiment.

[0017] Figure 11 It is a block diagram of an example computing system with a DPU according to at least one embodiment, the DPU having a Network Pipeline Abstraction Layer (NPAL) for providing an optimized and accelerated network pipeline to be accelerated by an acceleration hardware engine.

[0018] Figure 12 This is a network diagram of an example network pipeline optimized and accelerated on an acceleration hardware engine with a DPU having NPAL, according to at least one embodiment.

[0019] Figure 13 This is a network diagram of an example network pipeline optimized and accelerated on an acceleration hardware engine with a DPU having NPAL, according to at least one embodiment.

[0020] Figure 14 This is a network diagram of an example network pipeline optimized and accelerated on an acceleration hardware engine with a DPU having NPAL, according to at least one embodiment.

[0021] Figure 15 This is a flowchart of an example method for creating an optimized and accelerated network pipeline using the Network Pipeline Abstraction Layer (NPAL) according to at least one embodiment.

[0022] Figure 16It is a block diagram of the software stack of a DPU having NPAL with support for a split interface, according to at least one embodiment.

[0023] Figure 17 This is a network diagram of an example network pipeline optimized and accelerated on an acceleration hardware engine of a DPU with NPAL supporting split interfaces, according to at least one embodiment.

[0024] Figure 18 This is a flowchart of a method for operating a DPU having a split interface according to at least one embodiment.

[0025] Figure 19 It is a block diagram of a software stack of a DPU having NPAL supporting fast link recovery according to at least one embodiment.

[0026] Figure 20 This is a network diagram of an example network pipeline optimized and accelerated on an acceleration hardware engine of a DPU with NPAL supporting fast link recovery, according to at least one embodiment.

[0027] Figure 21 It is a flowchart of equal-cost multipath (ECMP) operations before, after and after link failure, according to at least one embodiment.

[0028] Figure 22 This is a flowchart of a method for operating a DPU with fast link recovery according to at least one embodiment.

[0029] Figure 23 This is a block diagram of an SFC architecture with a PBR strategy according to at least one embodiment.

[0030] Figure 24 This is a flowchart of a method for supporting a PBR-based DPU on an SFC architecture according to at least one embodiment.

[0031] Figure 25 It is a block diagram of a simulated network pipeline on a simulated DPU with simulated NPAL according to at least one embodiment.

[0032] Figure 26 It is a block diagram of an emulated SFC architecture 2600 having an emulated host device and an emulated DPU according to at least one embodiment.

[0033] Figure 27 This is a flowchart of a method for operating a simulated DPU according to at least one embodiment.

[0034] Figure 28 It is a block diagram of a computing system having two processing devices coupled to each other and multiple networks according to at least one embodiment.

[0035] Figure 29 It is a block diagram of a computing system having a central processing unit (CPU) and a graphics processing unit (GPU) in a single integrated circuit according to at least one embodiment.

[0036] Figure 30 It is a block diagram of a computing system having a tensor kernel graphics processing unit (GPU) according to at least one embodiment. Detailed Implementation

[0037] This document describes techniques for providing flexible bootstrapping rules to deliver hardware acceleration on a Service Function Chain (SFC) architecture. It also describes techniques for optimizing network acceleration using a Network Pipeline Abstraction Layer (NPAL). Furthermore, it describes techniques for providing configurable, dynamic SFC interfaces on Data Processing Units (DPUs). The DPU is described in more detail below. Additionally, it describes techniques for providing NPAL splitting interfaces. Furthermore, it describes techniques for providing fast NPAL link recovery. Finally, it describes policy-based routing (PBR) techniques for hardware acceleration on SFC architectures. Finally, it describes techniques for NPAL emulation.

[0038] As described above, in traditional network architectures, various security and performance functions are managed by dedicated hardware devices known as middleboxes (e.g., firewalls, load balancers, IDS, etc.). Traditional networks are designed under the assumption that all resources are housed within locally deployed data centers and are typically characterized by a centralized model.

[0039] Modern networks are increasingly cloud-centric, designed to support cloud services and applications. This includes the use of public, private, and hybrid cloud infrastructures, demanding greater network flexibility and scalability. Unlike traditional network architectures that heavily rely on physical hardware (i.e., each network function requires its own dedicated equipment), current network architectures leverage virtualization technologies such as Software-Defined Networking (SDN) and Network Functions Virtualization (NFV). These allow network resources to be abstracted from hardware, providing greater flexibility, easier management, and lower costs. Modern networks increasingly use automation and orchestration tools to efficiently manage network resources, reduce operational overhead, and enable faster deployment of network services. Modern networks are designed for scalability and high performance, utilizing technologies such as edge computing to process data closer to the source and reduce latency. Current network architectures are more flexible, scalable, and efficient than traditional network architectures, designed to support the dynamic, distributed nature of modern computing resources and work practices. They integrate advanced technologies such as cloud services, virtualization, and automation to meet the needs of today's digital environment.

[0040] Service Function Chain (SFC) Architecture

[0041] A network concept and architecture used in SDN and NFV environments is Service Function Chaining (SFC). SFC can be used to define and orchestrate the order of network services through a series of interconnected network nodes. SFC aims to virtualize network services (e.g., firewalls, load balancers, IDS, and other middleware functions) and define the sequence through which network traffic data is passed to achieve specific processing or control. Each network service is represented as a Service Function (SF). These SFs can be implemented as instances of virtualized software running on physical or virtual infrastructure. A service chain defines the sequence of SFs through which network traffic data is passed. For example, a service function chain might specify that network traffic data first passes through a firewall, then a load balancer, and finally an IDS using a Service Function Path (SFP) and a Service Function Forwarder (SFF). An SFP is a defined sequence of extensible functions (SFs) that steers network traffic data through in a specific order. An SFP is a logical representation of the path that network traffic data will follow as it travels through the network, passing through various service functions such as firewalls, load balancers, IDS, etc. An SFP specifies the traffic flow and ensures that the traffic flow is delivered in the correct order through each specified service function. SFPs can be used to implement policy-based routing and network services in a flexible and dynamic manner. SFFs are components within an SFC (Service Function Group) architecture responsible for actually forwarding network traffic data to the specified service function as defined by the SFP. The SFF acts as a router or switch that directs traffic between different service functions and ensures that network traffic data follows the prescribed path defined by the SFP. The SFF makes decisions about where to send network traffic data next based on SFC encapsulation information and the SFP. It handles routing and forwarding between service functions and performs any traffic encapsulation and decapsulation used for SFC operations. For example, when a packet enters the network, it is classified based on its attributes (such as source / destination Internet Protocol (IP) address, protocol, port, etc.) and the appropriate SFP is selected to determine the path through the appropriate SF. The SFF then guides the packet along the SFP path.

[0042] Service function chains offer several benefits, including increased flexibility, scalability, and agility in deploying and managing network services. They enable the dynamic creation of service chains based on application requirements, traffic conditions, or policy changes, resulting in more efficient and customizable network service delivery.

[0043] The current solution in the SFC architecture does not support the creation and use of flexible bootstrapping rules within a single accelerated data plane on the DPU. The current solution in the SFC architecture does not support configurable dynamic interface mappings on the DPU. Furthermore, the current solution does not always support acceleration for all operations within the SFC architecture.

[0044] The aspects and embodiments of this disclosure address these and other problems by providing flexible boot rules for hardware acceleration on the SFC architecture of the DPU, providing a configurable dynamic SFC interface on the DPU, and / or using a network pipeline abstraction layer to optimize network acceleration, as described in more detail below. The aspects and embodiments of this disclosure can provide and enable virtual bridges with different boot rules to the acceleration hardware engine to process network traffic data in a single accelerated data plane using a combined set of network rules from different boot rules of the different virtual bridges. The aspects and embodiments of this disclosure can provide and enable a network pipeline abstraction layer (NPAL) that supports multiple network protocols and network functions in the network pipeline, wherein the pipeline includes a set of tables and logic organized in a specific order to be accelerated by the acceleration hardware engine. The aspects and embodiments of this disclosure can provide and enable a first virtual bridge, a second virtual bridge, and a virtual port between the first and second virtual bridges, wherein the first virtual bridge is controlled by a first network service hosted on the DPU, and the second virtual bridge is controlled by user-defined logic. The aspects and embodiments of this disclosure can provide and enable an NPAL that supports split interfaces. Various aspects and embodiments of this disclosure may provide and enable NPAL supporting fast link recovery. Various aspects and embodiments of this disclosure may provide and enable a first virtual bridge, a second virtual bridge, and a virtual port between the first and second virtual bridges, wherein the first virtual bridge is controlled by a first network service hosted on the DPU, and the second virtual bridge is controlled by a policy-based routing policy (PBR policy).

[0045] Data Processing Unit (DPU)

[0046] In modern network architectures, Data Processing Units (DPUs) can be used to provide a suite of software-defined networking, storage, security, and management services at the data center scale, with the ability to offload, accelerate, and isolate data center infrastructure. DPUs can offload processing tasks normally handled by a server's central processing unit (CPU), such as encryption / decryption, firewalling, Transmission Control Protocol / Internet Protocol (TCP / IP) and Hypertext Transfer Protocol (HTTP) processing, and any combination of network operations. A DPU can be an integrated circuit or a System-on-a-Chip (SoC) considered as on-chip data center infrastructure. The CPU can include DPU hardware and DPU software (e.g., a software framework with acceleration libraries). DPU hardware may include a CPU (e.g., a single-core or multi-core CPU), one or more hardware accelerators, memory, one or more physical host interfaces operatively coupled to one or more host devices (e.g., the CPU of a host device), and one or more physical network interfaces operatively coupled to a network (e.g., a public network (e.g., the Internet), a private network (e.g., a local area network (LAN) or a wide area network (WAN)), a wired network (e.g., an Ethernet network), a wireless network (e.g., an 802.11 network or a Wi-Fi network), a cellular network (e.g., a Long Term Evolution (LTE) network), a router, a hub, a switch, a server computer, a network adapter, an NVLink switch, and / or combinations thereof). The DPU can handle network data path processing for network traffic data, while the host device can control path initialization and exception handling. An acceleration hardware engine (e.g., the DPU hardware) can be used to offload and filter network traffic based on predefined filters, utilizing the hardware capabilities of the acceleration hardware engine. The software framework and acceleration library may include one or more hardware acceleration services, including hardware acceleration services (e.g., NVIDIA DOCA), hardware acceleration virtualization services, hardware acceleration networking services, hardware acceleration storage services, hardware acceleration artificial intelligence / machine learning (AI / ML) services, hardware acceleration services, and hardware acceleration management services.

[0047] A DPU can provide accelerated network services (also known as Host-Based Networking (HBN) services) to one or more host devices. DPU network services can be used to accelerate Layer 2 (L2) protocols, Layer 3 (L3) protocols, tunneling protocols, etc., on the DPU hardware. The HBN infrastructure is based on an SFC topology, where a single virtual bridge (e.g., an OpenvSwitch (OVS) bridge) is controlled by the HBN service, providing all accelerated network capabilities. The HBN service can support various protocols and network capabilities, such as Access Control Lists (ACLs), Equal Cost Multipath (ECMP), tunneling, Connection Tracking (CT), Quality of Service (QoS) rules, Spanning Tree Protocol (STP), Virtual Local Area Network (VLAN) mapping, Network Address Translation (NAT), Software-Defined Networking (SDN), Multiprotocol Label Switching (MPLS), etc.

[0048] Configurable dynamic SFC interface mapping on DPU

[0049] In addition to the first virtual bridge controlled by the HBN service, various aspects and embodiments of this disclosure may also provide a second virtual bridge that can be controlled by user-defined logic. The second virtual bridge can be programmed by users, clients, or controllers (such as Open Virtual Network (OVN) controllers). OVN is an open-source project designed to provide network virtualization to virtual machines (VMs) and container instances. OVN acts as an extension of OVS, a virtual switch primarily used to automate networks in large-scale network environments. OVN complements OVS by adding native support for virtual network abstractions such as virtual L2 and L3 overlays and security groups. Various aspects and embodiments of this disclosure may support configurable dynamic interface mapping on DPUs based on SFC infrastructure. This configuration can be supported as part of the DPU's operating system (OS) installation, and also dynamically supported for production DPUs. Configuration can be completed within the deployed DPU without reinstalling the DPU OS. Interface configurations in the configuration file can support different use cases for network acceleration on the DPU.

[0050] In at least one embodiment, the DPU includes a memory for storing configuration files specifying a plurality of virtual bridges (such as the first and second virtual bridges described above). The configuration files also specify interface mappings for the plurality of virtual bridges. The DPU includes a processing device operatively coupled to the memory. The processing device generates the first and second virtual bridges according to the configuration files. The first virtual bridge is controlled by a first network service hosted on the DPU, while the second virtual bridge is controlled by user-defined logic. The processing device adds one or more host interfaces to the second virtual bridge, adds a first service interface to the first virtual bridge to operatively couple to the first network service, and adds one or more virtual ports between the first and second virtual bridges, all according to the configuration files. The second virtual bridge provides flexibility to users, clients, or controllers to define additional or different network functions compared to the network functions performed by the first network service. In one implementation, the second network service includes user-defined logic. The processing device adds a second service interface to the second virtual bridge to operatively couple to the second network service. Alternatively, the user-defined logic may be implemented in the second virtual bridge itself or in logic operatively coupled to the second virtual bridge.

[0051] Flexible boot rules for hardware acceleration of the DPU's SFC architecture

[0052] Various aspects and embodiments of this disclosure can provide a second virtual bridge to allow users, clients, or controllers to specify flexible boot rules on the SFC architecture of the DPU. Using SFC on the DPU, users (or controllers) can create flexible and dynamic network boot rules that are hardware-accelerated by the DPU as a single data plane on the DPU. Specifically, as described in more detail herein, user-defined rules can be accelerated together with existing networking rules in the HBN service within a single accelerated data plane. Users (or controllers) can program different boot rules on the SFC in a flexible manner in parallel with the HBN service, thereby generating a single accelerated data plane through the DPU hardware and DPU software. The hardware acceleration service for the DPU may include an OVS infrastructure based on open-source OVS with additional features (new acceleration capabilities). For example, the hardware acceleration service may include OVS-DOCA technology developed by Nvidia, Inc. of Santa Clara, California. OVS-DOCA, as the OVS infrastructure for the DPU, is based on open-source OVS with additional features (new acceleration capabilities), and the OVS backend is purely based on DOCA. Hardware acceleration services also support the OVS kernel and OVS-DPDK, two common modes. All three operating modes utilize flow offloading for hardware acceleration, but the OVS-DOCA mode offers the most efficient performance and feature set due to its architecture and the use of the DOCA library. The OVS-DOCA mode leverages the DOCA flow library to configure and use hardware offloading mechanisms and application technologies to generate a combined set of network rules used by the acceleration hardware engine to process network traffic data in a single accelerated data plane. Using the SFC infrastructure defined in the configuration file, users and customers can use the DPU as a network accelerator on edge devices, eliminating the need for complex smart switches in different network topologies within data center (DC) and service provider (SP) networks.

[0053] In at least one embodiment, the DPU includes an acceleration hardware engine for providing a single accelerated data plane. The DPU includes a memory storing a configuration file that specifies at least a first virtual bridge, a second virtual bridge, and a virtual port between the first and second virtual bridges. A processing device of the DPU is operatively coupled to the memory and the acceleration hardware engine. The processing device generates the first and second virtual bridges according to the configuration file. The first virtual bridge is controlled by a first network service hosted on the DPU and has a first set of one or more network rules. The second virtual bridge has a second set of one or more user-defined network rules. The processing device adds virtual ports between the first and second virtual bridges according to the configuration file. The processing device generates a combined set of network rules based on the first set of one or more network rules and the second set of one or more user-defined network rules. The acceleration hardware engine can use the combined set of network rules to process network traffic data in a single accelerated data plane.

[0054] Network Pipeline Abstraction Layer (NPAL) Optimization Pipeline for Network Acceleration

[0055] Various aspects and embodiments of this disclosure can provide an NPAL as a software programmable layer for an optimized network pipeline that supports different accelerated network capabilities, such as L2 bridging, L3 routing, tunnel encapsulation, tunnel decapsulation, hash computation, ECMP operations, static and dynamic ACLs, CT, etc. In other words, an NPAL is an accelerated programmable network pipeline that provides an abstraction of the underlying functionality of a network pipeline optimized for hardware acceleration on DPU hardware. An NPAL can be a Database Abstraction Layer (DAL) or something similar. A DAL is a programming concept used in software engineering to provide an abstraction of the underlying database system, allowing applications to interact directly with different databases, lower-level software layers, or hardware without changing the application code. A DAL typically includes a set of classes or application programming interfaces (APIs) that provide a unified interface for performing common database operations, such as querying, inserting, updating, and deleting data. By using a DAL, developers can write database-independent code, thereby reducing the coupling between the application and specific database implementations. Similarly, NPAL can include a collection of classes or APIs that provide a unified interface for performing common network operations in a network pipeline optimized for hardware acceleration on DPU hardware. Specifically, NPAL can provide a unified interface to one or more applications, network services, etc., executed by the DPU or host device. NPAL can include an optimized network pipeline supporting multiple network protocols and functions. The network pipeline can include a set of tables and logic ordered in a specific order, optimized for acceleration by DPU hardware, thereby providing customers and users with a rich set of capabilities and high performance.

[0056] Using NPAL in a DPU offers various benefits, including operational independence, logic encapsulation, performance, code reusability, and platform independence. For example, developers can write agnostic code, allowing applications (e.g., network services) to work with different underlying access logic and network functions. NPAL can encapsulate access or network function-related logic, making it easier to manage and maintain the codebase system. Changes to the schema or underlying technology can be isolated within the NPAL implementation. NPAL can provide optimized, high-performance pipelines to address diverse network requirements and functions. By separating access logic from application logic, developers can reuse NPAL components across multiple parts of the application (network service), promoting code reuse and maintainability. NPAL can abstract platform-specific differences, data types, and other access or network function-related characteristics, enabling applications (network services) to run seamlessly on different platforms and environments. Overall, NPAL provides a powerful tool for building flexible, scalable, and maintainable network function-driven applications, offering a level of abstraction that simplifies interactions between network functions and drives code efficiency and portability.

[0057] In at least one embodiment, the DPU includes DPU hardware, which includes a processing device and an acceleration hardware engine. The DPU includes memory operatively coupled to the DPU hardware. The memory may store DPU software, which includes NPALs supporting multiple network protocols and network functions in the network pipeline. The network pipeline includes a set of tables and logic organized in a specific order to be accelerated by the acceleration hardware engine. The acceleration hardware engine can use the network pipeline to process network traffic data. The network pipeline can be optimized for network services running on the DPU.

[0058] OVS and OVS bridge

[0059] Open Virtual Switch (OVS) is an open-source, multi-tiered virtual switch used to manage network traffic in virtualized environments (specifically, data centers and cloud computing platforms). OVS provides network connectivity between virtual machines (VMs), containers, and physical devices. OVS is widely used in virtualization and cloud technologies and is a typical component of many software-defined networking (SDN) and network virtualization solutions.

[0060] Virtual switches, typically found in virtualized computing environments, are software applications that allow virtual machines (VMs) on a single physical host to communicate with each other and with external networks. Virtual switches provide network connectivity between VMs, containers, and physical devices. Virtual switches can emulate the functionality of a physical network switch but operate at a software level within a hypervisor or host operating system. Virtual switches manage network traffic by routing packets between VMs on the same host or between VMs and the physical network using ports. These ports can be configured for various policies, such as security settings, Quality of Service (QoS) rules, etc. Virtual switches can segment network traffic to provide isolation between different virtual networks. Virtual switches provide an interface between the virtualized environment and the physical network, allowing VMs to communicate outside their hosts. Virtual switches can support standard network protocols and features, such as Virtual LAN (VLAN) tagging, Layer 2 forwarding, Layer 3 capabilities, etc. OVS can support the OpenFlow protocol, allowing network controllers to control virtual switches to make decisions about how traffic should be routed across the network. Network controllers, such as Software-Defined Networking (SDN) controllers, are centralized entities that manage flow control over network devices. It is the "brain" of the network, maintaining a comprehensive understanding of the network and making decisions about where to send packets. The OpenFlow (OF) protocol enables the controller to interact directly with both physical network devices and virtual network devices (such as switches and routers) on the forwarding plane. OF configuration refers to setting up and managing network behavior using the OpenFlow protocol in an SDN environment. It involves defining flow rules and actions to control how network devices, typically centrally managed by an SDN controller, handle traffic. OF configuration can include flow tables containing rules on how packets should be handled. Each flow table contains a set of flow entries. Flow entries define how to handle packets that match certain criteria. An entry can have three parts: a match field, an action, and a counter. The match field defines the packet attributes to match, such as source / destination Internet Protocol (IP) addresses, Media Access Control (MAC) addresses, port numbers, VLAN tags, etc. Actions can define how to handle matching packets, such as forwarding the matching packet to a specific port, modifying fields in the packet, or dropping it. Counters can be used to track the number of packets and bytes in each flow. The network controller can use control messages to manage flow entries in switches. It can add, update, or delete flow entries. Optional configurations can include group tables for more advanced forwarding actions such as multicast, load balancing, etc. It should be noted that OVS is one type of virtual switching technology, but other virtual switching technologies exist, such as SDN-based switches.

[0061] An OVS bridge functions like a virtual network switch at the software level, allowing the connection and management of multiple network interfaces as if they were ports on a physical switch. OVS bridges enable the creation and management of virtual networks within or across multiple servers in a data center or cloud environment. An OVS bridge connects virtual and physical network interfaces, facilitating communication between them. This can include interfaces from VMs, containers, physical network interfaces, or even other virtual bridges. Similar to a physical Ethernet switch, an OVS bridge operates at Layer 2 (L2) of the Open Systems Interconnection model (known as the OSI model), forwarding, filtering, and managing traffic based on Media Access Control (MAC) addresses. OVS bridges can support advanced features such as Virtual LAN (VLAN) tagging, Quality of Service (QoS), traffic mirroring, and Access Control Lists (ACLs). OVS bridges can be controlled by a controller using protocols such as OpenFlow (OF), enabling dynamic, programmable network configuration.

[0062] This document describes aspects and embodiments of the present disclosure with respect to OVS, and these aspects and embodiments include OVS- and OpenFlow-specific terminology. However, some aspects and embodiments of the present disclosure can be used with other virtual switching and bridging technologies. Similarly, while the various embodiments are described in the context of a DPU, they can also be used in other virtual switching environments, including virtual bridges, switches, network interface cards (NICs) (also referred to as network interface controllers), smart NICs, network interface devices, network switches, network adapters, intelligent processing units (IPUs), or other dedicated computing devices designed to offload specific tasks from the CPU of a computer or server.

[0063] DPUs are specialized semiconductor devices designed to offload and accelerate networking, security, and storage tasks traditionally run on server CPUs. By taking over these functions, DPUs aim to significantly improve the overall efficiency and performance of data centers. They are equipped with their own processors and memory, enabling them to handle complex data processing tasks independently of the host CPU. DPUs are embedded in data center infrastructure, where they manage data movement and processing across the network, freeing up CPU resources to focus more on application and workload processing. This architectural shift allows for increased workload density, improved data throughput, and enhanced security measures at the hardware level. DPUs play a crucial role in software-defined networking (SDN), providing hardware acceleration for advanced features such as encryption, traffic management, and virtualization. By optimizing these critical operations, DPUs help create more agile, secure, and efficient data centers.

[0064] An Integrated Processing Unit (IPU) is a dedicated hardware accelerator designed to optimize the performance of machine learning algorithms and artificial intelligence (AI) workloads. Unlike general-purpose CPUs or graphics processing units (GPUs), which are general-purpose but not optimized for AI tasks, IPUs are specifically modified to handle the high computational demands and data throughput requirements of deep learning models and neural network processing. They achieve this by implementing highly parallel computing architectures and memory systems that can efficiently handle the large amounts of data associated with AI applications. IPUs aim to reduce latency and increase the speed of AI computation, thereby enabling the training of more complex models faster and more efficiently.

[0065] Intelligent NICs are advanced network interface cards (NICs) equipped with built-in processing capabilities to offload network tasks from the CPU, thereby enhancing the efficiency and performance of data processing within servers. Unlike traditional NICs, which primarily act as data pipes between servers and the network, intelligent NICs can perform a wide range of network functions directly on the card, such as traffic management, encryption / decryption, and network virtualization tasks. These capabilities allow intelligent NICs to significantly reduce CPU load, freeing up resources to improve the server's overall processing capacity for application workloads. By handling complex network functions, intelligent NICs can achieve lower latency and higher throughput in data center environments, making them particularly valuable in scenarios requiring real-time processing and high-speed networking, such as cloud computing, high-performance computing (HPC), and enterprise data centers. The intelligence and programmability of intelligent NICs provide a flexible solution to meet the evolving needs of modern network infrastructure, facilitating more efficient and customizable network operations.

[0066] NPAL split interface

[0067] As described above, NPAL is a powerful acceleration network pipeline for building flexible, scalable, and maintainable database-driven applications, providing a level of abstraction that simplifies database interactions and drives code efficiency and portability. Various aspects and embodiments of this disclosure can provide NPAL support for both hardware and software port splitting. Physical ports can be physically split into multiple ports, which requires advanced software support. Splitting a host's NIC physical port into multiple ports (also known as port splitting or physical NIC port branching) involves taking a single high-bandwidth physical NIC port and splitting it into multiple low-bandwidth logical or physical ports. This can be done at the physical level or via software configuration. Splitting NIC physical ports has multiple use cases and advantages, especially in environments where network bandwidth and redundancy requirements vary.

[0068] One use case for splitting NIC physical ports includes bandwidth optimization and efficient resource utilization, network redundancy, and high availability. When a system has a high-speed NIC (e.g., 100Gbps), many workloads or applications may not require the full bandwidth of the NIC. In such cases, splitting a 100Gbps port into four 25Gbps ports allows hosts to use the available bandwidth more efficiently. This is particularly useful when multiple applications or services have different bandwidth requirements but do not require the full capacity of a single port.

[0069] Another use case and benefit of splitting NIC physical ports includes network redundancy and high availability. Splitting a NIC's physical ports into multiple ports enables better redundancy and failover. With multiple physical links, operators can configure active-active or active-backup redundancy strategies. This setup ensures that if one link fails, traffic can be rerouted via another link, thus enhancing reliability and uptime.

[0070] Another use case and benefit of splitting NIC physical ports includes better network segmentation and isolation. By physically splitting NIC ports, different physical interfaces can be dedicated to different types of traffic (e.g., management, production, backup). This physical isolation improves security by preventing traffic from one network type (e.g., management traffic) from mixing with another type (e.g., production traffic).

[0071] Another use case and benefit of splitting physical NIC ports is increasing port density. Splitting physical NIC ports can increase the overall port density available to the host without requiring additional NIC hardware. If the number of available PCIe slots in a server is limited, splitting a high-speed port into multiple low-speed ports helps maximize the number of network connections available to the host.

[0072] Another use case and benefit of splitting NIC physical ports includes cost savings. Splitting a physical NIC port into multiple logical ports reduces the need to purchase additional physical NICs. Hosts can split existing ports to achieve the same result at a lower cost, rather than purchasing additional NIC cards to provide more network ports.

[0073] Another use case and benefit of splitting NIC physical ports includes enhanced control and traffic shaping. For split physical ports, different priority levels or traffic shaping policies can be applied to each port separately. This is useful in environments where different services require different Quality of Service (QoS) levels or bandwidth control.

[0074] Another use case and benefit of splitting NIC physical ports includes multi-homing and diversified paths. Splitting NIC ports allows servers to be multi-homed to separate networks using different uplinks. This improves fault tolerance and provides diversified paths for outbound and inbound traffic, thereby ensuring better load balancing and failover mechanisms.

[0075] When a NIC's physical ports are split into multiple ports, the result is the creation of multiple Physical Functions (PFs) for the host. Each PF can be managed independently by the host operating system. This improves resource allocation. Each PF functions like a separate NIC interface, meaning you can allocate resources (e.g., bandwidth, CPU, memory) more efficiently across the system. Multiple PFs can ensure that specific applications or virtual machines (VMs) have dedicated resources and do not contend for the same network interface, thus improving performance isolation. Furthermore, network administrators can assign different policies and configurations to each PF. This can include VLAN tags, firewall rules, or QoS settings for different workloads, traffic types, or security zones. Additionally, when a physical NIC is split into multiple PFs, Single Root I / O Virtualization (SR-IOV) can be used to assign each PF to different VMs or containers. SR-IOV achieves near-native I / O performance by allowing VMs to bypass the hypervisor for network communication, thereby reducing latency and increasing throughput.

[0076] To support NIC / DPU port splitting, the NIC / DPU requires both hardware (HW) and software (SW) support. From a hardware perspective, special cables are needed to physically split physical ports. These cables are called branch cables. In high-speed network environments, especially in data centers, it is common to use branch cables that allow a single high-speed port (such as 40Gbps or 100Gbps) to be "split" into multiple low-speed ports (such as 4×10Gbps or 4×25Gbps). These cables physically separate data streams and allow multiple devices or ports to connect to a single high-speed interface. From a software perspective, the NIC / DPU needs to support port splitting across the entire software stack, including the NIC / DPU firmware (FW), drivers, virtual switches, and the NPAL on top of the virtual switches. Generally, the firmware and drivers can present multiple physical functions (PFs) to the virtual switches and NPAL. The NPAL can be used to configure different policies for different PFs, isolate networks between PFs, improve resource utilization, and ultimately support multiple PFs instead of one or two PFs. As described in this document, NPAL comprises a set of tables and logic ordered in a specific order, optimized for acceleration by NIC / DPU hardware, thereby providing customers and users with a rich set of capabilities and high performance. Once physical ports are split into multiple logical ports, the network pipeline can send network traffic data to any PF (as part of the output port logic). Furthermore, each PF can be configured with different policies, such as different QoS or different traffic management.

[0077] NPAL Fast Link Recovery

[0078] Various aspects and embodiments of this disclosure can provide an NPAL that offers an optimized network pipeline that supports rapid link recovery when link failures occur and equal-cost multipath (ECMP) groups need to be updated to reflect the new network topology. A link failure in a network topology refers to a situation where a communication link between two network devices (such as routers, NICs, switches, or hosts) becomes unavailable due to various reasons such as hardware failure, cable breakage, or network congestion. Link failures can significantly impact the performance, availability, and reliability of network services, especially in large-scale or critical environments such as data centers or enterprise networks. Several factors can lead to link failures, such as physical link failures (e.g., cable damage or breakage due to fiber optic cable severance, Ethernet cable breakage, etc.), device failures (e.g., hardware failures that cause interface downtime), congestion or overload (i.e., network links becoming overloaded due to excessive traffic, resulting in timeouts or packet drops), software vulnerabilities or misconfigurations that cause network devices to drop connections, power failures, maintenance activities (e.g., planned interruptions to maintenance or upgrades can also lead to temporary link failures), etc.

[0079] Link failures can have different implications. When a link fails, packets being transmitted over that link are lost. This can cause transmission delays and impact real-time applications (e.g., VoIP, streaming). Connections between devices relying on the failed link become unavailable until rerouting occurs. In the event of a link failure, traffic may need to be routed through longer alternative paths, which introduce higher latency and negatively impact application performance. If the failed link is part of a redundant configuration (such as in ECMP, Spanning Tree Protocol (STP), or dual-homed devices), traffic can be rerouted to the backup link. However, this may reduce the level of redundancy. In networks running dynamic routing protocols (e.g., Open Shortest Path First (OSPF), Border Gateway Protocol (BGP), Intermediate System to Intermediate System (IS-IS)), when a link fails, the network must reconverge by recalculating new routes. This convergence can take time, during which affected parts of the network may be unreachable. If the rerouted traffic exceeds the capacity of the alternative paths, rerouting traffic from the failed link can cause congestion on those paths. Mission-critical applications, such as financial services or online games, may experience downtime or performance degradation due to link failures. TCP connections may undergo retransmissions, increasing latency, while UDP connections may suffer data loss without retries. Point-to-point communication may cease entirely until network convergence or the failed link recovers, potentially leading to power outages. Services relying on high availability (HA) solutions may be affected if there are no failover paths or mechanisms that take a considerable amount of time to activate.

[0080] One routing strategy is called ECMP. ECMP is a routing strategy that allows multiple forwarding paths for packets toward their destination when multiple paths have the same routing cost (or metric). By utilizing ECMP, networks can balance traffic across multiple paths, improve bandwidth utilization, enhance redundancy, and provide fault tolerance. Due to the lessons learned from link failures mentioned above, rapid recovery is crucial when a link goes down. Specifically, when a link goes down, the ECMP group must be updated immediately to reflect that the specific link is no longer available. Furthermore, when a link comes online, the ECMP group must be updated immediately with the recovered link to improve network utilization.

[0081] As described in more detail below, NPAL can operate as an accelerated network pipeline and virtual switch hardware offloading mechanism for periodic link monitoring. In at least one embodiment, the user can configure NPAL to support fast link recovery, thereby enabling link monitoring by the virtual switch. For example, inter-process communication (IPC) messages, such as Linux Netlink messages from the Linux kernel, can be used to monitor all ports in a specific ECMP group.

[0082] If a link goes down (i.e., experiences a link failure), the virtual switch identifies the link as down and immediately updates the ECMP group by removing it from the ECMP group. In other words, the virtual switch updates the ECMP group tables in the bridges and / or routers of the network pipeline to remove the failed link. Once the ECMP group is updated, traffic can be distributed to other links in the ECMP group.

[0083] If a link comes online (i.e., no longer experiences link failure), the virtual switch identifies the link as online and immediately updates the ECMP group by adding the link to it. In other words, the virtual switch updates the ECMP group tables in the bridges and / or routers of the network pipeline to add the recovered link (also known as the new link). Once the ECMP group is updated, traffic can be distributed to the new link along with other links in the ECMP group.

[0084] Hardware-accelerated policy-based routing (PBR) on SFC

[0085] Various aspects and embodiments of this disclosure can provide hardware-accelerated policy-based routing (PBR) on the SFC architecture of the DPU. Using SFC on the DPU, users (or controllers) can add different PBR policies, which are accelerated by the DPU hardware as a single data plane on the DPU.

[0086] As described above, a DPU can provide accelerated network services (also known as HBN services or network services) to one or more host devices. DPU network services can be used to accelerate L2, L3, tunneling, and other protocols on the DPU hardware. The network service infrastructure is based on an SFC topology, where the first virtual bridge (e.g., an Open Virtual Switch (OVS) bridge) is controlled by the network service to provide all accelerated network functions, and the second virtual bridge (e.g., an OVS bridge) can be programmed by the user or any other controller. Network services can support different protocols and network capabilities, such as ACLs, ECMP, tunneling, CT, QoS, STP, VLAN mapping, NAT, SDN, MPLS, etc.

[0087] Policy-based routing (PBR) is a technology that allows network administrators to make routing decisions based on policies set by the network administrator, rather than relying on a default routing table that uses the destination IP address to determine the next hop. Traditional routing uses the destination IP address and routing table to determine how to forward packets. PBR allows routers to make routing decisions based on other criteria, such as: source IP address or subnet; IP protocol type; port number; incoming interface; packet size; QoS parameters, etc. PBR is very useful for controlling the path traffic takes through the network. It provides network administrators with greater flexibility to implement routing rules that are not solely dependent on the destination IP address.

[0088] Common use cases for PBR include traffic engineering, load balancing, network preferences, security and compliance, and QoS. For example, network administrators might want to route specific types of traffic via preferred or better paths to achieve better performance. PBR can be used to distribute traffic across multiple network links to balance the load on network resources. PBR can be used to route traffic to certain websites (e.g., YouTube or Netflix) via cheaper internet connections, while routing mission-critical applications via more reliable or faster connections (e.g., MPLS). PBR can be used to enforce security policies by ensuring that traffic from specific users or networks is routed through security devices such as firewalls or IDS. PBR can be used to help enforce QoS policies by routing traffic based on certain QoS tags.

[0089] PBR typically works by providing policy definitions. A policy definition is a set of rules or conditions used to classify traffic. These rules typically match on criteria such as source address, destination address, protocol type, or port number. The defined policy is applied to incoming traffic on a specific interface for policy enforcement. When a packet arrives at that interface, the router checks if the packet matches the conditions defined in the policy. For policy-based data path traffic routing, if the packet matches the policy, it is routed according to the policy's routing table or next-hop information, rather than the default routing table. Here are some PBR examples:

[0090] Example 1: Match source IP and set next hop

[0091] Match: SRC_IP = 1.1.1.1

[0092] Action: Next_Hop = 2.2.2.2

[0093] Example 2: Matching the protocol (HTTP) and routing through a specific interface

[0094] Match: PROTOCOL = TCP

[0095] Action: OUTPUT_PORT = PORT1

[0096] Example 3: Matching the ingress interface and marking QoS

[0097] o Matching: INGRESS_INTERFACE = PORT1

[0098] Action: SET_QOS = VALUE1

[0099] Specifically, users (or controllers) can program PBR policies on the SFC in parallel with network services, which generates a single accelerated data plane through virtual switches and DPU hardware. That is, PBR policies can use existing network rules from the network services for acceleration within a single accelerated data plane. As described above, the hardware acceleration service of the DPU can include OVS infrastructure (e.g., OVS-DOCA technology) for configuring and using hardware offloading mechanisms and application technologies to generate a combined set of network rules that are used by the acceleration hardware engine to process network traffic data in a single accelerated data plane. Using the SFC infrastructure defined in the configuration file, users and customers can use the DPU as a network accelerator on edge devices, eliminating the need for complex smart switches in different network topologies within data center (DC) networks and service provider (SP) networks.

[0100] NPAL simulation

[0101] Various aspects and embodiments of this disclosure can provide NPAL emulation. Generally, emulation involves replicating the behavior of one system on another. The goal of emulation is to make the second system behave like the original system, typically to replace or recreate the environment of the original system. NPAL emulation is used to provide a simulated network pipeline, rather than the actual hardware device running NPAL. Specifically, NPAL emulation is used to simulate a DPU running network services with NPAL, as well as a complete DPU environment including virtual bridges, SFCs, etc. To support the emulated NPAL with network services and a complete environment including virtual switches and SFCs, software-based modules can be used to support all the different components, such as network services, NPAL (used by the network services), virtual bridges, SFCs, etc. Each of these components is implemented in software to simulate behavior exactly the same as a real hardware device with a hardware-accelerated network pipeline as described herein. The simulated network pipeline can be similar to the hardware network pipeline described herein.

[0102] Figure 1This is a block diagram of an integrated circuit 100 having SFC logic 102 for generating a virtual bridge 104 and interface mapping 106 in an SFC architecture, according to at least one embodiment. The integrated circuit 100 may be a DPU, NIC, smart NIC, network interface device, or network switch. The integrated circuit 100 includes a memory 108, a processing device 110, an acceleration hardware engine 112, a network interconnect 114, and a host interconnect 116. The processing device 110 is coupled to the memory 108, the acceleration hardware engine 112, the network interconnect 114, and the host interconnect 116. The processing device 110 hosts the virtual bridge 104 generated by the SFC logic 102. The virtual bridge 104 (also called a virtual switch) is software that operates within a computer network to connect different segments or devices, much like a physical bridge but in a virtualized environment. It is a core component in network virtualization, enabling virtual machines (VMs), containers, and other virtual network interfaces to connect to each other and to the physical network, thereby simulating traditional Ethernet functionality purely in software. Virtual Bridge 104 allows the creation and management of isolated network segments within a single physical infrastructure, facilitating communication, enforcing security policies, and providing bandwidth management while offering the flexibility and scalability required in dynamic virtualization and cloud environments. Virtual Bridge 104 can be an Open Virtual Switch (OVS) bridge. An OVS bridge acts as the virtual switch at the heart of an Open Virtual Switch architecture, enabling advanced network management and connectivity in virtualized environments. It manages traffic flows between VMs on the same physical host and between external networks by aggregating multiple network interfaces into a single logical interface. Unlike traditional virtual bridges, OVS bridges support a wide range of network features such as VLAN tagging, traffic monitoring with sFlow and NetFlow, Quality of Service (QoS), and Access Control Lists (ACLs), providing network administrators with enhanced flexibility and control. OVS bridges efficiently direct network traffic based on predefined policies and rules, providing an essential tool for building complex multi-tenant cloud and data center networks.

[0103] Specifically about Figure 1Virtual bridge 104 can provide network connectivity between VMs running on the same integrated circuit 100 or on a separate host device, container, and / or physical device. In short, virtual bridge 104 allows VMs on a single physical host to communicate with each other and with an external network 118. Virtual bridge 104 can emulate the functionality of a physical network switch but operates at a software level within integrated circuit 100. Virtual bridge 104 can manage network traffic data 120, thereby directing packets between VMs on the same host or between VMs and the physical network using ports. These ports can be configured for various policies, such as security settings, QoS rules, etc. Virtual bridge 104 can segment network traffic to provide isolation between different virtual networks. Virtual bridge 104 can provide an interface between the virtualization environment and the physical network, allowing VMs to communicate outside the host. Virtual Bridge 104 supports standard network protocols and features such as VLAN tagging, Layer 2 (L2) forwarding, Layer 3 (L3) capabilities, tunneling protocols (e.g., Virtual Extended LAN (VXLAN), Generic Routing Encapsulation (GRE), and Geneve), flow-based forwarding, OpenFlow support, integration with virtualization platforms (e.g., integration with VMware, KVM, Xen, etc., enabling network connectivity between virtual machines and containers), scalability, traffic monitoring and mirroring, security, and multi-platform support (e.g., Linux, FreeBSD, Windows, etc.). For Layer 2 switching, one or more virtual bridges in Virtual Bridge 104 act as Layer 2 Ethernet switches, enabling the forwarding of Ethernet frames between different network interfaces (including virtual and physical ports). For Layer 3 routing, one or more virtual bridges in Virtual Bridge 104 support Layer 3 IP routing, allowing them to route traffic between different IP subnets and perform IP-based forwarding. Virtual Bridge 104 can support VLAN tagging and allows the segmentation of network traffic into different VLANs using VLAN tags. Virtual Bridge 104 can use flow-based forwarding, where network flows are classified based on their characteristics and packet forwarding decisions are made based on flow rules; it can also enforce security policies and access controls. OVS is commonly used in data center and cloud environments to provide network agility, flexibility, and automation. It plays a crucial role in creating and managing virtual networks, enabling network administrators to adapt to the ever-changing needs of modern dynamic data centers.

[0104] In at least one embodiment, the virtual bridge 104 may use OVS and OF technologies. The virtual bridge 104 may be controlled by a network controller (also called a network service) to make decisions about how traffic should be routed through the network. As described herein, a network controller (e.g., an SDN controller) is a centralized entity that manages flow control over network devices. The OF protocol may be used to interact directly with the forwarding plane of network devices, such as virtual or physical switches and routers. In at least one embodiment, the virtual bridge 104 may use flow tables containing rules for how packets should be handled. Each flow table contains a set of flow entries. Flow entries define how to handle packets that match certain criteria. An entry may have three parts: a matching field, an action, and a counter. The matching field defines the packet attributes to be matched, such as source / destination Internet Protocol (IP) address, Media Access Control (MAC) address, port number, VLAN tag, etc. The action may define how to handle the matching packet, such as forwarding the matching packet to a specific port, modifying fields in the packet, or discarding the matching packet. The counter may be used to track the number of packets and bytes in each flow. Because Virtual Bridge 104 is virtualized, it can create rules at the software, data path (DP), and hardware levels. Rules created at the software level are called software (SW) rules or OF rules. Rules created at the DP level are called DP rules. Rules created at the hardware level are called hardware (HW) rules. When a software rule is created, corresponding DP and HW rules are also created. The network controller can add, update, or delete flow entries, thereby changing configuration settings.

[0105] In another embodiment, virtual bridge 104 is a standard virtual switch or a distributed virtual switch. In another embodiment, virtual bridge 104 is an SDN-based switch integrated with an SDN controller. Integrated circuit 100 can be used in data centers, cloud computing environments, development and testing environments, network function virtualization (NFV) environments, etc. Virtual bridge 104 can be used in data centers where server virtualization is common for efficiently facilitating communication within and between servers. Virtual bridge 104 in a cloud computing environment can enable multi-tenant networks, allowing different clients to have isolated network segments. Virtual bridge 104 can allow the connection and management of network function virtualization (e.g., NFV) within a virtual infrastructure. Some advantages of virtual bridge 104 include: virtual bridge 104 can be easily configured or reconfigured without physical intervention, reducing the need for physical network hardware and related maintenance, and providing the ability to create isolated networks for different applications or tenants. In summary, the Virtual Bridge 104 is a software-based device that performs the networking functions of a physical switch in a virtualized environment (such as a data center and cloud computing environment) and provides flexibility, isolation, and efficient network management in that virtualized environment.

[0106] In at least one embodiment, the integrated circuit 100 may also host one or more hypervisors and one or more virtual machines (VMs). Network traffic data 120 may be directed to the corresponding VMs by the virtual bridge 104.

[0107] During operation, SFC logic 102 can use configuration file 124 to generate virtual bridge 104 and interface mapping 106 between virtual bridge 104, network interconnect 114, and host interconnect 116. Configuration file 124 can specify the configuration of virtual bridge 104, interface mapping 106, and each of them. SFC logic 102 can generate a first virtual bridge and a second virtual bridge according to configuration file 124. The first virtual bridge is used to be controlled by a first network service 130 hosted on integrated circuit 100, and the second virtual bridge is used to be controlled by user-defined logic 126. SFC logic 102 can add one or more host interfaces to the second virtual bridge and add a first service interface to the first virtual bridge to be operatively coupled to the first network service 130. SFC logic 102 can add one or more virtual ports between the first virtual bridge and the second virtual bridge.

[0108] In at least one embodiment, user-defined logic 126 is part of user-defined service 132 hosted on integrated circuit 100, such as a user-defined network service. SFC logic 102 can add a second service interface to the second virtual bridge according to configuration file 124 to be operatively coupled to user-defined service 132. User-defined service 132 may be a user-defined security service, a user-defined telemetry service, a user-defined storage service, etc.

[0109] In at least one embodiment, integrated circuit 100 stores operating system 122 (OS122) in memory 108. Integrated circuit 100 can execute on processing device 110. In at least one embodiment, as part of installing OS122 on integrated circuit 100, SFC logic 102 generates virtual bridge 104 and interface mapping 106. In another embodiment, as part of the runtime of integrated circuit 100 and without reinstalling OS122 on integrated circuit 100, SFC logic 102 can generate virtual bridge 104 and interface mapping 106. In at least one embodiment, SFC logic 102 can configure OS attributes (e.g., page size) associated with OS122 in one of the virtual bridges 104 according to configuration file 124.

[0110] In at least one embodiment, SFC logic 102 can perform and facilitate operations for recognizing changes to the configuration settings of virtual bridge 104 in configuration file 124 (or a new configuration file). Thus, SFC logic 102 can configure virtual bridge 104 and interface mapping 106 during installation or runtime and without requiring a reinstallation of operating system 122.

[0111] like Figure 1 As illustrated, SFC logic 102 is implemented in an integrated circuit 100 having a memory 108, a processing device 110, an acceleration hardware engine 112, a network interconnect 114, and a host interconnect 116. In other embodiments, SFC logic 102 may be implemented in a processor, computing system, CPU, DPU, smart NIC, IPU, etc. The underlying hardware may host the virtual bridge 104 and interface mapping 106.

[0112] In at least one embodiment, the integrated circuit 100 can be deployed in a data center (DC) network or a service provider (SP) network. A data center (DC) network is the fundamental infrastructure that facilitates communication, data exchange, and connectivity between different computing resources, storage systems, and network devices within a data center. It is designed to support high-speed data transmission, reliable access to distributed resources, and efficient management of data flows across various physical and virtual platforms. At its core, a DC network integrates a large number of switches, routers, firewalls, and load balancers orchestrated by advanced network protocols and software-defined networking (SDN) technologies to ensure optimal performance, scalability, and security. The architecture of a DC network typically includes both a physical backbone and a virtual overlay. The physical backbone has high-capacity cabling and switches to ensure bandwidth and redundancy; and the virtual overlay enables flexibility, rapid configuration, and resource optimization through virtual networks. A well-designed DC network supports a range of applications, from enterprise services to cloud computing and big data analytics, by providing the infrastructure to handle large volumes of data, complex computing, and application workloads typical of modern data centers. It plays a crucial role in disaster recovery, data replication, and high availability strategies, ensuring that data center services remain resilient and efficient under varying loads. Service Provider (SP) networks are extensive, high-capacity communications infrastructures operated by organizations that provide a wide range of telecommunications, internet, cloud computing, and digital services to businesses, residential customers, and other entities. These networks are designed to provide broad coverage, connecting numerous geographic locations, including city centers, remote areas, and international destinations, to facilitate global communications and data exchange. SP networks have a multi-layered architecture, incorporating a hybrid of technologies such as fiber optics, wireless transmission, satellite links, and broadband access to enable broad connectivity. At the heart of these networks is a high-performance backbone responsible for transmitting large volumes of data at high speeds over long distances. On top of the physical infrastructure, SP networks deploy advanced networking technologies, including MPLS, Software-Defined Networking (SDN), and Network Functions Virtualization (NFV), to enhance the efficiency, flexibility, and scalability of service delivery. Service Provider networks are designed to support a wide range of services, from routine voice and data services to modern cloud-based applications and streaming services, similarly addressing the evolving needs of consumers and businesses. They are crucial for enabling the internet, mobile communications, enterprise networking solutions, and the emerging Internet of Things (IoT) ecosystem, ensuring the connectivity and accessibility of digital resources and services globally.

[0113] In at least one embodiment, the virtual bridge 104 and interface mapping 106 are part of a Service Function Chain (SFC) architecture implemented in at least one of the following: a DPU, a NIC, a smart NIC, a network interface device, or a network switch. In at least one embodiment, the SFC logic 102 may be implemented as part of a hardware acceleration service on an agentless hardware product (such as a DPU), as described below. Figure 2 As illustrated and described, integrated circuit 100 can be a DPU. The DPU can be a programmable data center infrastructure on a chip. The hardware acceleration service can be part of NVIDIA OVS-DOCA, developed by NVIDIA Corporation in Santa Clara, California. OVS-DOCA, as a new OVS infrastructure for the DPU, is based on open-source OVS with additional features (new acceleration capabilities), and the OVS backend is purely based on DOCA. Alternatively, SFC logic 102 can be part of other services.

[0114] Service Function Chain (SFC) Infrastructure

[0115] SFC infrastructure refers to a network architecture and framework that enables the creation, deployment, and management of service function chains within a network. A service function chain is a technology used to define an ordered list of network services (such as firewalls, load balancers, and intrusion detection systems) through which traffic is systematically routed. This ordered list is called a "chain," and each service in the chain is called a "service function." SFC infrastructure is designed to ensure that network traffic flows through these service functions in a specified sequence, thereby improving the efficiency, security, and flexibility of network service delivery. SFC infrastructure can include service function forwarders, service functions (SFs), service function paths (SFPs), etc. An SFF is a network device responsible for forwarding traffic to the desired service function according to the defined service chain. An SFF ensures that packets are directed through the correct sequence of service functions. An SF is the actual network service that processes packets. These can be physical or virtual network functions, such as firewalls, WAN optimizers, load balancers, intrusion detection / prevention systems, etc. An SFP is a defined path that traffic takes through the network, which includes a specific sequence of service functions it passes through. SFPs are established based on policy rules and can be dynamically adjusted to respond to changing network conditions or needs. An SFC (Service Container Registry) infrastructure can use one or more SFC descriptors, which are policies or templates describing service chains, including service function sequences, performance requirements, and other relevant metadata. One or more SFC descriptors can act as blueprints for instantiating and managing service chains within the network. An SFC infrastructure may include classification functions responsible for initial inspection and classification of incoming packets to determine the appropriate service chain to which traffic should be routed. Classification can be based on various packet attributes such as source and destination IP addresses, port numbers, and application identifiers. Typically, as part of a larger Software-Defined Networking (SDN) or NFV framework, one or more network controllers can manage the SFC infrastructure. They can be responsible for orchestrating and deploying service chains, configuring network elements, and ensuring real-time adjustment and optimization of traffic flows. SFC infrastructure offers numerous benefits, including enhanced network flexibility, optimized resource utilization, and improved overall security. By decoupling the network's control plane from its data plane and leveraging virtualization, SFC infrastructure can dynamically adapt to changing network needs, enabling a more efficient and scalable service delivery model. (See also: [link to relevant information]) Figure 2 As illustrated and described, the SFC infrastructure can be deployed as a DPU-based SFC infrastructure 200.

[0116] Figure 2This is a block diagram of an example DPU-based SFC infrastructure for providing an SFC architecture 220 according to at least one embodiment. The DPU-based SFC infrastructure 200 includes a DPU 204 coupled between a host device 202 and a network 210. In at least one embodiment, the DPU 204 is a system-on-a-chip (SoC), considered as on-chip data center infrastructure. The DPU 204 is a dedicated processor designed to offload and accelerate networking, storage, and security tasks from the central processing unit (CPU) of the host device 202, thereby enhancing overall system efficiency and performance. The DPU 204 can be used in data center and cloud computing environments to manage data traffic more efficiently and securely.

[0117] DPU 204 may include network interconnects (e.g., one or more Ethernet ports) operatively coupled to network 210. These network interconnects may be high-speed network interfaces that enable direct connection to the data center network infrastructure. These interfaces may support various speeds (e.g., 10Gbps, 25Gbps, 40Gbps, or higher) depending on the model and deployment requirements. Network 118 may include public networks (e.g., the Internet), private networks (e.g., local area networks (LANs) or wide area networks (WANs)), wired networks (e.g., Ethernet networks), wireless networks (e.g., 802.11 networks or Wi-Fi networks), cellular networks (e.g., LTE networks), routers, hubs, switches, server computers, and / or combinations thereof.

[0118] DPU 204 can be coupled to the CPU (or multiple host devices or servers) of host device 202 via one or more host interconnects (e.g., Peripheral Component Interconnect High Speed ​​(PCIe)). PCIe provides a high-speed connection between DPU 204 and the CPU of the host device, allowing for rapid transfer of data and instructions. This connection is used to offload tasks from the CPU and to ensure that DPU 204 can efficiently access system memory and storage resources. To enable communication between host device 202 and DPU 204, dedicated software drivers and firmware are installed on host device 202. These software components allow the host's operating system and applications to interact with DPU 204, thereby offloading specific tasks to DPU 204 and retrieving processed data. In a virtualization environment, DPU 204 can also interface with a hypervisor or container management system. This allows DPU to support multiple virtual machines (VMs) or containers by providing virtualization networking capabilities, network isolation, and security features without burdening the CPU of the host device. DPU 204 can directly read from and write to the host device's memory using Direct Memory Access (DMA), bypassing the CPU to reduce latency and freeing up CPU resources for other tasks. This enables efficient data movement between host memory, DPU 204, and network 210. In at least one embodiment, DPU 204 includes a Direct Memory Access (DMA) controller coupled to the host interface. Figure 2 (Not illustrated). The DMA controller can read data from the host's physical memory via a host interface. In at least one embodiment, the DMA controller uses PCIe technology to read data from the host's physical memory. Alternatively, other technologies can be used to read data from the host's physical memory. In other embodiments, the DPU 204 can be any computing system or computing device capable of performing the techniques described herein.

[0119] Once physically connected, DPU 204 is configured to communicate with network 210. This involves setting the IP address, VLAN tag (if a virtual network is used), and routing information to ensure that DPU 204 can send and receive data packets to and from other devices on network 210. As described herein, DPU 204 performs network-related tasks such as packet forwarding, encryption / decryption, load balancing, and Quality of Service (QoS) implementation. By doing so, it effectively becomes an enhanced intelligent network interface controller capable of performing complex data processing and traffic management.

[0120] In at least one embodiment, DPU 204 includes DPU hardware 208 and DPU software 206 (e.g., a software framework with acceleration libraries). DPU hardware 208 may include one or more CPUs (e.g., single-core or multi-core CPUs), acceleration hardware engines 214 (or multiple hardware accelerators), memory 218, and network interconnects and host interconnects. In at least one embodiment, DPU 204 includes DPU software 206, which includes a software framework and acceleration libraries. The software framework and acceleration libraries may include one or more hardware acceleration services, including hardware acceleration services (e.g., NVIDIA DOCA), hardware acceleration virtualization services, hardware acceleration network services, hardware acceleration storage services, hardware acceleration artificial intelligence / machine learning (AI / ML) services, hardware acceleration services, and hardware acceleration management services.

[0121] In at least one embodiment, memory 218 stores configuration file 124. Configuration file 124 specifies virtual bridge 104, interface mapping 106 (host interfaces and network ports) between virtual bridges 104, and network functions 222 in SFC architecture 220. For example, one of one or more CPUs 216 can generate a first virtual bridge and a second virtual bridge according to configuration file 124, the first virtual bridge being controlled by a first network service hosted on DPU 204, and the second virtual bridge being controlled by user-defined logic. The CPU can add one or more host interfaces to the second virtual bridge, add a first service interface to the first virtual bridge to be operatively coupled to the first network service, and add one or more virtual ports between the first and second virtual bridges according to the configuration file. Figure 1 As described, SFC logic 102 can be implemented using DPU software 206 to generate and manage SFC architecture 220. SFC logic 102 can utilize acceleration hardware engine 214 (e.g., DPU hardware 208) to offload and filter network traffic data 212 based on predefined filters using the hardware capabilities of acceleration hardware engine 214. DPU hardware 208 can receive network traffic data 212 from a second device (or multiple devices) on network 210 via a network port.

[0122] In at least one embodiment, the DPU software 206 may perform several actions when creating the virtual bridge 104 and the corresponding interface mapping 106 to ensure proper configuration and integration within the virtualization environment. The DPU software 206 can initialize the creation of the virtual bridge by allocating resources and setting initial configuration parameters. These configurations can be stored in a configuration file 124. Configuration parameters can define the bridge name, supported network protocols, and any specific settings related to performance or security. A virtual network interface is created to act as the virtual bridge. This interface serves as the anchor point for all virtual and physical interfaces that will be connected to the virtual bridge. The DPU software 206 can identify specified physical interfaces (e.g., Ethernet ports) and virtual interfaces (e.g., virtual machine network adapters) and link those physical and virtual interfaces to the newly created virtual bridge. This action involves configuring the settings for each interface to ensure compatibility and optimal communication within the virtual bridge. The DPU software 206 can configure network protocols. Network protocols and services, such as Spanning Tree Protocol (STP) for preventing loops, are configured on the virtual bridge. The DPU software 206 can also configure VLAN tags for traffic segmentation, QoS policies for traffic prioritization, and security features such as Access Control Lists (ACLs). The DPU software 206 can assign IP addresses to bridge interfaces. If the virtual bridge acts as a Layer 3 (L3) switch, the DPU software 206 assigns IP addresses to the bridge interfaces, enabling them to participate in IP routing between different connected networks or devices. The DPU software 206 provides a unified interface that allows for centralized control and monitoring of the network. Network administrators can manage the virtual bridge and other virtual network components through this unified interface. The DPU software 206 enables the monitoring and management of virtual bridge characteristics, allowing network administrators to observe traffic flows, identify potential problems, and make adjustments as needed to optimize network performance and security. In these steps, the software ensures that the virtual bridge integrates seamlessly into the existing network architecture, providing a flexible and efficient way to connect various network segments within a virtualized environment.

[0123] In addition to generating virtual bridge 104, the DPU software 206 can also generate one or more virtual ports between virtual bridges 104. Virtual ports, often referred to as patch ports in the context of virtual networks, are software-defined networking components that facilitate connectivity and communication between different virtual devices or between virtual devices and physical devices within a network. Unlike physical ports on network switches or routers, virtual ports are not bound to specific hardware interfaces; instead, they are created and managed through software, providing a flexible and efficient way to route traffic in a virtualized environment. Virtual ports play a crucial role in creating complex network topologies in VMs, containers, and virtual networks. They can be used to configure virtual switches (vSwitch) or bridges, allowing virtual machines on the same or different hosts to communicate as if they were connected to the same physical network switch. Additionally, patch ports can connect virtual networks to physical networks, enabling VMs to access external network resources. Virtual ports can be dynamically created, configured, and deleted based on network requirements, making them more adaptable to changes in network topology or workload demands. By optimizing the use of the underlying physical network infrastructure, virtual ports can help improve overall network efficiency, thereby reducing the need for additional physical hardware. Virtual ports support advanced network features such as VLAN tagging, QoS settings, and ACL configuration, enabling precise management of network traffic. Virtual ports also provide visualization of virtual network traffic, allowing for detailed monitoring, logging, and troubleshooting activities.

[0124] In addition to generating virtual bridge 104, DPU software 206 can also configure link state propagation for virtual bridge 104. In the context of virtual bridges or virtual switches (such as Open Virtual Switches (OVS)), link propagation refers to the process by which changes in the state of physical or virtual network interfaces are transmitted across the network. This ensures that the entire network topology is aware of the connection status and can adjust routing and switching behavior accordingly. Link propagation is used to maintain the accuracy of the network's operational state, enabling efficient data flow and ensuring high availability and reliability of network services. In OVS, OVS monitors the state of the physical ports and virtual interfaces connected to it. This includes tracking when a port comes online (becomes active) or goes offline (becomes inactive) due to changes in physical link state or virtual interface configuration. Upon detecting a change in port state, OVS propagates this information throughout the network. This is done by sending notifications to relevant components within the network infrastructure, such as other switch instances, network controllers, or virtual machines connected to the virtual switch. Based on the propagated link state information, network devices and protocols can adjust their operations. This can involve recalculating routes, redistributing network traffic, or initiating failover processes to alternative paths or interfaces to maintain network connectivity and performance. Link propagation helps maintain a consistent network view across the topology. By ensuring that all elements of the network have up-to-date information about link states, it enables coherent and coordinated network behavior, especially in dynamic environments with frequent changes. OVS can integrate link propagation with standard network protocols and mechanisms such as Spanning Tree Protocol (STP) for preventing loops and Link Layer Discovery Protocol (LLDP) for network discovery. This integration enhances the ability of switches to participate in and benefit from the broader network ecosystem, conforming to established network management practices. Link propagation plays a fundamental role in the adaptive and resilient behavior of networks utilizing virtual bridges or virtual switches (such as OVS), ensuring that changes to the network infrastructure are reflected quickly and accurately across the entire network. This capability is particularly important in virtualized and cloud environments, where topologies can be highly dynamic and the efficiency and reliability of network connectivity are also critical.

[0125] In at least one embodiment, the DPU software 206 can configure link state propagation in a virtual bridge by setting mechanisms to monitor and transmit the operational status (such as UP or DOWN) of links across the network. This allows the virtual bridge and its connected entities to be dynamically adjusted to adapt to changes in network topology, such as when interfaces are added, removed, or when interfaces experience failures. The DPU software 206 can activate monitoring capabilities for all connected physical and virtual interfaces on the virtual bridge. This typically involves enabling the detection of link state changes so that the bridge can identify when a port becomes active or inactive. Once monitoring is enabled, the system needs to be configured to notify relevant components within the network of any changes. This may involve setting up event listeners or subscribers that can respond to notifications about link state changes. For virtual bridges managed by a controller (in an SDN environment), this may also mean configuring communication between the bridge and the controller to ensure that it receives timely updates about the network state. Configuring link propagation also involves specifying the actions that a change in link state should trigger. For example, this could include automatically recalculating routing tables, redistributing traffic to available paths, or even triggering alerts and logging events for network administrators. The forwarding database (FDB) or MAC table of a virtual bridge may need to be dynamically updated based on link state changes to ensure that traffic is routed efficiently within the network. This ensures that packets are not sent to a downed interface.

[0126] In at least one embodiment, the DPU software 206 can configure the virtual bridge to filter network traffic data 212. For example, configuration file 124 can specify what data the virtual bridge should extract from network traffic data 212. Configuration file 124 can specify one or more filters that extract specified types of data from network traffic data 212 for inclusion or removal. Network traffic that meets the filtering criteria can be structured and streamed to a network function in network function 222 for processing. For example, configuration file 124 can specify to extract all Hypertext Transfer Protocol (HTTP) traffic from network traffic data 212 and route it to a network function in network function 222. Configuration file 124 can specify that all traffic on a specific port should be extracted from network traffic data 212 for processing by network function 222, which will be described in more detail below.

[0127] As described in this document, the SFC architecture 220 can support various network protocols and capabilities within network function 222. These different network protocols and capabilities serve as the backbone of modern networks, enabling a wide range of functions from basic connectivity to advanced security and traffic optimization. Network function 222 can include network functions that include sets of tables and logic for performing corresponding network functions such as ACLs, ECMP routing, tunneling, connection tracking (CT), NAT, QoS, etc. An ACL is a fundamental network security feature that allows or denies traffic based on a set of rules. These lists are applied to network interfaces to control packet flow at ingress or egress points. These rules can specify various parameters, such as source and destination IP addresses, port numbers, and protocol types, to fine-tune the traffic filtering process, thereby enhancing security and compliance. ECMP is a routing strategy used to distribute outgoing network traffic across multiple paths with equal costs. By evenly balancing the load across these paths, ECMP can significantly improve network bandwidth and reliability. This protocol is particularly useful in data center and cloud environments where high availability and scalability are critical. Tunneling encapsulates one protocol or session within another, allowing data to traverse networks with incompatible address spaces or architectures. It is widely used to implement Virtual Private Networks (VPNs), where secure tunnels over the Internet enable private communication. Protocols such as IPsec and GRE are common examples of tunneling facilitated for security and protocol encapsulation purposes. CT refers to the ability of a network device (such as a firewall or router) to maintain state information about network connections passing through it. This capability allows devices to make more informed decisions about which packets to allow or block based on the context of the sessions they belong to. CT is crucial for implementing stateful firewalls and NAT (Network Address Translation) functionality. QoS capabilities refer to mechanisms that prioritize certain types of traffic to ensure application performance, especially in congested network scenarios. Network functions 222 can include other types of network functions such as Segmentation Routing (SR), Multiprotocol Label Switching (MPLS), network virtualization, Software-Defined Networking (SDN), etc. SR allows the source of a packet to define the path the packet takes through the network using a segment list, improving routing efficiency and flexibility. MPLS is a method for accelerating and shaping network traffic flows, where data packets are labeled and routed quickly along predetermined paths within the network. Network virtualization involves abstracting physical network devices and resources into virtual networks, enabling more flexible and efficient resource management. Software-defined networking (SDN) decouples network control and forwarding functions, enabling the management of programmable networks and the efficient orchestration of network services.These protocols and capabilities represent only a small fraction of the many technologies that form the foundation of modern networks, each playing a specific role in ensuring the efficient, secure, and reliable transmission of data across digital infrastructure.

[0128] Therefore, integrating DPU 204 into network 210 and host device 202 represents a powerful approach to optimizing data processing tasks and significantly enhancing the performance and security of data center and cloud computing environments. By handling the majority of network, storage, and security workloads, DPU 204 allows the CPU to focus more on application processing, thereby improving overall system efficiency and throughput. For example, DPU 204 can handle network data path processing for network traffic data 212. The CPU can control path initialization and exception handling. DPU 204 can be part of a data center and includes one or more data stores, one or more server machines, and other components of the data center infrastructure. It should be noted that, unlike CPUs or GPUs, the DPU 204 is a new type of programmable processor that combines three key elements, including, for example: 1) an industry-standard, high-performance, software-programmable CPU (single-core or multi-core CPU) tightly coupled with other SoC components; 2) a high-performance network interface capable of parsing, processing, and efficiently transferring data to GPUs and CPUs at line rates or the speed of the rest of the network; and 3) a rich collection of flexible programmable acceleration engines that can offload and enhance the performance of applications used for AI and machine learning, security, telecommunications, and storage. These capabilities enable isolated, bare-metal, cloud-native computing platforms for cloud-scale computing. In at least one embodiment, the DPU 204 can be used as a standalone embedded processor. In at least one embodiment, the DPU 204 can be incorporated into a network interface controller (also known as a SmartNIC) used as a component of a server system. A DPU-based network interface card (network adapter) can offload processing tasks typically handled by the CPU of a server system. Using its processor, a DPU-based SmartNIC can perform any combination of encryption / decryption, firewall, Transmission Control Protocol / Internet Protocol (TCP / IP), and Hypertext Transfer Protocol (HTTP) processing. For example, a SmartNIC can be used in a high-traffic web server.

[0129] In at least one embodiment, DPU 204 can be configured for modern cloud workloads and high-performance computing for traditional enterprises. In at least one embodiment, DPU 204 can deliver a collection of software-defined networking, storage, security, and management services at data center scale with the ability to offload, accelerate, and isolate data center infrastructure. In at least one embodiment, DPU 204 can deliver these software services to multi-tenant cloud-native environments. In at least one embodiment, DPU 204 can deliver data center services up to hundreds of CPU cores, thereby freeing up valuable CPU cycles to run business-critical applications. In at least one embodiment, DPU 204 can be considered a new class of processors designed to handle data center infrastructure software to offload and accelerate computational workloads related to virtualization, networking, storage, security, cloud-native AI / ML services, and other management services.

[0130] In at least one embodiment, DPU 204 may include connectivity to packet-based interconnects (e.g., Ethernet), switched-structure interconnects (e.g., InfiniBand, Fibre Channel, Omni-Path), etc. In at least one embodiment, DPU 204 can provide an accelerated, fully programmable, and securely configured (e.g., zero-trust security) data center to protect against data breaches and cyberattacks. In at least one embodiment, DPU 204 may include a network adapter, a processor core array, and an infrastructure offloading engine with full software programmability. In at least one embodiment, DPU 204 may be located at the edge of a server to provide flexible, secure, high-performance cloud and AI workloads. In at least one embodiment, DPU 204 can reduce total cost of ownership and improve data center efficiency. In at least one embodiment, DPU 204 can provide a software framework and acceleration libraries (e.g., NVIDIA DOCA). TM These acceleration libraries enable developers to quickly create applications and services for the DPU 204, such as security services, virtualization services, network services, storage services, AI / ML services, and management services. In at least one embodiment, the software framework and acceleration libraries make it easy to leverage the hardware accelerator of the DPU 204 to deliver data center performance, efficiency, and security. In at least one embodiment, the DPU 204 may be coupled to a GPU. The GPU may include one or more accelerated AI / ML pipelines.

[0131] In at least one embodiment, DPU 204 can provide network services with virtual switch (vSwitch), virtual router (vRouter), network address translation (NAT), load balancing, and network virtualization (NFV). In at least one embodiment, DPU 204 can provide storage services, including structurally NVME. TM (NVMe-oF TM Technologies such as NVM Express include elastic storage virtualization, hyperconverged infrastructure (HCI) encryption, data integrity, compression, and data deduplication. TM It is an open logic device interface specification for accessing interconnected peripheral components quickly. Non-volatile storage media attached to a PCIe interface. NVMe-oF TM It provides an efficient mapping of NVMe commands to several network transport protocols, enabling one computer (“initiator”) to access block-level storage devices attached to another computer (“target”) with very high efficiency and minimal latency. The term “Fabric” is a generalization of the more specific concepts of network and input / output (I / O) channels. It essentially refers to the N:M interconnect of elements, typically in a peripheral context. NVMe-oF TMThe technology enables the transmission of NVMe command sets over various interconnect infrastructures, including networks (e.g., Internet Protocol (IP) / Ethernet) and I / O channels (e.g., Fibre Channel). In at least one embodiment, the DPU 204 can provide hardware-accelerated services using next-generation firewalls (NGFW), intrusion detection systems (IDS), intrusion prevention systems (IPS), roots of trust, micro-segmentation, distributed denial-of-service (DDoS) defense technologies, and ML detection. NGFW is a network security appliance that provides capabilities beyond stateful firewalls, such as application awareness and control, integrated intrusion prevention, and cloud-delivered threat intelligence. In at least one embodiment, one or more network interfaces may include Ethernet interfaces (single-port or dual-port) and InfiniBand interfaces (single-port or dual-port). In at least one embodiment, one or more host interfaces may include PCIe interfaces and PCIe switches. In at least one embodiment, one or more host interfaces may include other memory interfaces. In at least one embodiment, the CPU may include multiple cores (e.g., up to eight 64-bit core pipelines) and a DDR4 dynamic random access memory (DRAM) controller having an L2 cache for every two cores or two cores, and an L3 cache with an eviction policy supporting Double Data Rate (DDR) dual in-line memory modules (DIMMs) (e.g., Double Data Rate 4 (DDR4) DIMM support). The memory may be onboard DDR4 memory with error correction code (ECC) error protection support. In at least one embodiment, the CPU may include a single core with L2 and L3 caches and a DRAM controller. In at least one embodiment, one or more hardware accelerators may include security accelerators, storage accelerators, and network accelerators.In at least one embodiment, the security accelerator can provide secure boot with a hardware root of trust, secure firmware updates, Cerberus compliance, regular expression (RegEx) acceleration, IP security (IPsec) / Transport Layer Security (TLS) data-in-motion encryption, Advanced Encryption Standard Galois / Counter Mode (AES-GCM) 512 / 256-bit keys for data-at-rest encryption (e.g., Advanced Encryption Standard (AES) with Ciphertext Theft (XTS) (e.g., AES-XTS256 / 512), 256-bit hardware acceleration of Secure Hash Algorithm (SHA), hardware public key accelerators (e.g., Rivest-Shamir-Adleman (RSA), Diffie-Hellman, Digital Signal Algorithm (DSA), ECC, Elliptic Curve Cryptography Digital Signal Algorithm (ECC-DSA), Elliptic Curve Diffie-Hellman (EC-DH)), and a True Random Number Generator (TRNG). In at least one embodiment, the storage accelerator can provide BlueField SNAP-NVMe. TM And VirtIO-blk, NVMe-oF TM Acceleration, compression, and decompression, as well as data hashing and deduplication. In at least one embodiment, the network accelerator can provide Converged Ethernet (RoCE) Remote Direct Memory Access (RDMA) over RoCE, Zero Touch RoCE, stateless offloading of TCP, IP, and User Datagram Protocol (UDP), Large Receive Offload (LRO), Large Segment Offload (LSO), checksums, Total Sum of Squares (TSS), Residual Sum of Squares (RSS), HTTP Dynamic Streaming (HDS), and Virtual LAN (VLAN) insertion / stripping, Single Root I / O Virtualization (SR-IOV), virtual Ethernet cards (e.g., VirtIO-net), multi-function per port, VMware NetQueue support, virtualization hierarchy, and ingress and egress Quality of Service (QoS) levels (e.g., 1K ingress and egress QoS levels). In at least one embodiment, the DPU 204 may also provide boot options including (RSA certified) secure boot, remote boot over Ethernet, remote boot over Internet Small Computer System Interface (iSCSI), Preboot Execution Environment (PXE), and Unified Extensible Firmware Interface (UEFI).

[0132] In at least one embodiment, DPU 204 can provide management services including a 1GbE out-of-band management port, a network controller sideband interface (NC-SI), a management component transport protocol (MCTP) on the system management bus (SMBus) and a monitoring and control table (MCT) on PCIe, a platform-level data model (PLDM) for monitoring and control, a PLDM for firmware updates, an inter-integrated circuit (I2C) interface for device control and configuration, a serial peripheral interface (SPI) to flash memory, an embedded multimedia card (eMMC) memory controller, a universal asynchronous receiver / transmitter (UART), and a universal serial bus (USB).

[0133] Host device 202 may be a desktop computer, laptop computer, smartphone, tablet computer, server, or any suitable computing device capable of performing the techniques described herein. In some embodiments, host device 202 may be a computing device of a cloud computing platform. For example, host device 202 may be a server machine of a cloud computing platform or a component of a server machine. In such embodiments, host device 202 may be coupled to one or more edge devices (not shown) via network 210. An edge device is a computing device that enables communication between computing devices at the boundary of two networks. For example, an edge device may be connected to host device 202, one or more data stores, one or more server machines via network 210, and may be connected to one or more endpoint devices (not shown) via another network. In such an example, the edge device may enable communication between host device 202, one or more data stores, one or more server machines, and one or more client devices. In other or similar embodiments, host device 202 may be an edge device or a component of an edge device. For example, host device 202 can facilitate communication between one or more data storage devices, one or more server machines connected to host device 202 via network 210, and one or more client devices connected to host device 202 via another network.

[0134] In other or similar embodiments, host device 202 may be an endpoint device or a component of an endpoint device. For example, host device 202 may be a device such as a television, smartphone, cellular phone, data center server, data DPU, personal digital assistant (PDA), portable media player, netbook, laptop, e-book reader, tablet, desktop computer, set-top box, game console, computing device for autonomous vehicles, monitoring device, etc., or may be a component of such devices. In such embodiments, host device 202 may be connected to DPU 204 via network 210 through one or more network interfaces. In other or similar embodiments, host device 202 may be connected to an edge device (not shown) via another network, and the edge device may be connected to DPU 204 via network 210.

[0135] In at least one embodiment, host device 202 executes one or more computer programs. The one or more computer programs can be any process, routine, or code executed by host device 202, such as a host OS, an application, a guest OS of a virtual machine, or a guest application, such as a guest application executed in a container. Host device 202 may include one or more CPUs with one or more cores, one or more multi-core CPUs, one or more GPUs, one or more hardware accelerators, etc.

[0136] As described above, DPU 204 can use configurable dynamic SFC interface mappings of multiple virtual bridges 104 to generate and configure the SFC architecture 220 of network function 222. The following section discusses... Figure 3 , Figure 5 and Figure 4 An example of the SFC architecture is explained and described.

[0137] Figure 3This is a block diagram of an SFC architecture 300 according to at least one embodiment, having a first virtual bridge 302 (labeled "BR-EXT" for an external bridge), a second virtual bridge 304 (labeled "BR-INT" for an internal bridge), a virtual port 306, and a network service 308. As described herein, SFC logic 102 can generate the first virtual bridge 302, the second virtual bridge 304, and the virtual port 306 within the SFC architecture 300. SFC logic 102 can configure the first virtual bridge 302 to be controlled by the network service 308 hosted on DPU 204, and configure the second virtual bridge 304 to be controlled by user-defined logic 126. SFC logic 102 adds a service interface to the first virtual bridge 302 to operatively couple the network service 308 to the first virtual bridge 302. SFC logic 102 adds the virtual port 306 between the first virtual bridge 302 and the second virtual bridge 304. Network service 308 can provide one or more network service rules 318 to the first virtual bridge 302. SFC logic 102 can add one or more host interfaces 310 to the second virtual bridge 304.

[0138] As illustrated, three separate host interfaces can be added to connect the second virtual bridge 304 to hosts, such as three separate VMs hosted on host device 202. For example, one VM could host a firewall application, another a load balancer application, and yet another an IDS application. SFC logic 102 can add one or more network interfaces 312 to the first virtual bridge 302. Specifically, SFC logic 102 can add a first network interface to the first virtual bridge 302 to operatively couple to a first network port 314 (labeled PORT1) of DPU 204; and add a second network interface to operatively couple to a second network port 316 (labeled PORT2) of DPU 204. The first virtual bridge 302 can receive network traffic data from the first network port 314 and the second network port 316. The first virtual bridge 302 can redirect network traffic data to the second virtual bridge 304 via virtual port 806. The second virtual bridge 304 can direct network traffic data to the corresponding host via the host interface 310.

[0139] In at least one embodiment, user-defined logic 126 is part of the second virtual bridge 304. In at least one embodiment, user-defined logic 126 is part of a user-defined service hosted on DPU 204, such as a user-defined network service, a user-defined security service, a user-defined telemetry service, a user-defined storage service, etc. SFC logic 102 may add another service interface to the second virtual bridge 304 to operatively couple the user-defined service to the second virtual bridge 304.

[0140] In at least one embodiment, SFC logic 102 can configure a first link state propagation between the first host interface and virtual port 306 and a second link state propagation between the second host interface and virtual port 306 in the second virtual bridge 304. Similarly, SFC logic 102 can configure a third link state propagation between the third host interface and virtual port 306 in the second virtual bridge 304. Similar link state propagation can be configured in the first virtual bridge 302 for the link between virtual port 306 and network interface 312.

[0141] In at least one embodiment, SFC logic 102 can configure operating system (OS) properties in the second virtual bridge 304. In at least one embodiment, SFC logic 102 can configure OS properties for each host interface in host interface 310.

[0142] As described in this document, the SFC architecture 300 can be created either as part of installing the OS on the DPU 204 or as part of the DPU 204's runtime. This can be accomplished using a second configuration file or modifications to the original configuration file. Reconfiguring the DPU as part of the runtime can be done without reinstalling the OS on the DPU 204.

[0143] It should be noted that SFC logic 102 can generate different combinations of virtual bridges and interface mappings in different SFC architectures, such as... Figure 4 As shown in the diagram.

[0144] Figure 4 This is a block diagram of an SFC architecture 400 having a first virtual bridge 302, a second virtual bridge 304, a virtual port 306, and a network service 308 according to at least one embodiment. Except that the SFC architecture 400 includes an additional host interface 402, which is described in more detail below, the SFC architecture 400 is similar to the SFC architecture 300 as indicated by similar reference numerals.

[0145] As described above, SFC logic 102 can generate a first virtual bridge 302, a second virtual bridge 304, and a virtual port 306 within the SFC architecture 400. SFC logic 102 can configure the first virtual bridge 302 to be controlled by a network service 308 hosted on DPU 204, and configure the second virtual bridge 304 to be controlled by user-defined logic 126 (as described above, on the second virtual bridge 304 or on a user-defined service). SFC logic 102 adds a service interface to the first virtual bridge 302 to operatively couple the network service 308 to the first virtual bridge 302. SFC logic 102 adds a virtual port 306 between the first virtual bridge 302 and the second virtual bridge 304. Network service 308 can provide one or more network service rules 318 to the first virtual bridge 302. SFC logic 102 can add one or more host interfaces 310 to the second virtual bridge 304 and one or more host interfaces 402 to the first virtual bridge 302.

[0146] As illustrated, two separate host interfaces can be added to connect the second virtual bridge 304 to one or more hosts (e.g., VMs or containers hosted on host device 202), and a host interface can be added to connect the first virtual bridge 302 to a host (e.g., a VM or container hosted on host device 202). In other embodiments, different numbers of host interfaces can be added to multiple virtual bridges according to configuration file 124. SFC logic 102 can add one or more network interfaces 312 to the first virtual bridge 302. Specifically, SFC logic 102 can add a first network interface to the first virtual bridge 302 to be operatively coupled to a first network port 314 (labeled PORT1) of DPU 204, and add a second network interface to be operatively coupled to a second network port 316 (labeled PORT2) of DPU 204. The first virtual bridge 302 can receive network traffic data from the first network port 314 and the second network port 316.

[0147] In at least one embodiment, SFC logic 102 can configure a first link state propagation between the first host interface and virtual port 306 and a second link state propagation between the second host interface and virtual port 306 in the second virtual bridge 304. Similarly, SFC logic 102 can configure a third link state propagation between the third host interface and network interface 312 in the first virtual bridge 302. Similar link state propagation can be configured for the link between virtual port 306 and network interface 312 in the first virtual bridge 302.

[0148] In at least one embodiment, SFC logic 102 can configure operating system (OS) features in the second virtual bridge 304. In at least one embodiment, SFC logic 102 can configure OS features for each host interface in host interface 310 and for each host interface in host interface 402.

[0149] In at least one embodiment, DPU 204 can support configurable dynamic interface mapping on DPU 204 based on SFC infrastructure. This configuration can be supported as part of the DPU's OS installation and is dynamically supported for production DPUs. Interface configuration can support different use cases for network acceleration on DPU 204. As described above, the SFC architecture can consist of two main bridges: an external virtual bridge (BR-EXT) controlled by the network service running on DPU 204; and a user controller (e.g., as described below regarding...). Figure 9 The OVN controller described herein controls the internal virtual bridge (BR-INT). Interface configuration can be tailored to support different requirements based on customer use cases. For example, interface configuration can support uplink interfaces to the external virtual bridge (BR-EXT) or the internal virtual bridge (BR-INT), host interfaces to the external virtual bridge or the internal virtual bridge, and additional services connected to the internal virtual bridge (BR-INT), such as security services, telemetry services, etc. Interface configuration can support Scalable Functionality (SF) configuration. Interface configuration can support link propagation and different OS attributes (e.g., HUGEPAGE_SIZE, HUGEPAGE_COUNT, etc.).

[0150] The following is a sample configuration file.

[0151] ENABLE_EX_VB = yes

[0152] #Enable external virtual bridge - Default is No

[0153] ENABLE_INT_VB = yes

[0154] #Enable internal virtual bridge - default is no

[0155] EX_VB_UPLINKS = "port1,port2"

[0156] #Optional, define the uplink - defaults to "port1,port2"

[0157] INT_VB_UPLINKS = ""

[0158] #The uplink port can be attached to only one VB.

[0159] INT_VB_REPS="hostIF0,hostIF1,hostIF2"

[0160] EXT_VB_REPS="hostIF3,hostIF4"

[0161] #Depending on the configuration, the host interface is attached to the first virtual bridge / second virtual bridge.

[0162] #Each host interface is either on the first virtual bridge or on the second virtual bridge.

[0163] INT_VB_SFS = "service1"

[0164] EXT_VB_SFS = "service2"

[0165] #Connect to the service interfaces on the first / second virtual bridge

[0166] EXT_INT_VB_VPORTS="vport0","vport1"

[0167] #Create a patch port and set it as a peer on the first / second virtual bridge.

[0168] LINK_PROPAGATION_1="hostIF0:vport0","hostIF1:vport1"

[0169] LINK_PROPAGATION_2="hostIF2:vport1","hostIF3:vport2"

[0170] #Link propagation between different interfaces

[0171] HUGEPAGE_SIZE = 2048

[0172] #Optional, unit is kilobytes (KB)

[0173] HUGEPAGE_COUNT = 4096

[0174] #Optional, number

[0175] As described in this document, the SFC architecture 400 can be created either as part of installing the OS on the DPU 204 or as part of the DPU 204's runtime. This can be accomplished using a second configuration file or modifications to the original configuration file. Reconfiguring the DPU as part of the runtime can be done without reinstalling the OS on the DPU 204.

[0176] It should be noted that SFC logic 102 can generate different combinations of virtual bridges and interface mappings in different SFC architectures, such as... Figure 3 and Figure 4 As shown in the illustration. For comparison, Figure 5 The diagram illustrates a non-SFC architecture 500 that allows only one network service to control a single virtual bridge between the host and the network.

[0177] Figure 5 This is a block diagram of a non-SFC architecture 500 having a single virtual bridge 503 and a single network service 504 according to at least one embodiment. The single virtual bridge 502 may include multiple host interfaces 506, a service interface operatively coupled to the single network service 504, and two network interfaces 508. The single virtual bridge 502 is controlled by the single network service 504 using one or more network service rules 510. The network service 308 may provide one or more network service rules 318 through the service interface for configuring the first virtual bridge 302. The single virtual bridge 502 may route network traffic data between a host device 202 (e.g., multiple VMs) and a first network port 314 and a second network port 316 of the DPU 204. The single virtual bridge 502 and the single network service 504 are limited in providing more network functionality than the single network service 504 provides.

[0178] Figure 6 This is a flowchart of an example method 600 configured with an SFC architecture having multiple virtual bridges and interface mappings according to at least one embodiment. The processing logic can be a combination of hardware, firmware, software, or any combination thereof. In at least one embodiment, the processing logic is implemented in a DPU, switch, network device, GPU, NIC, CPU, etc. In at least one embodiment, the processing logic is implemented in an acceleration hardware engine coupled to the switch. In at least one embodiment, method 600 can be executed by multiple processing threads, each thread executing one or more individual functions, routines, subroutines, or operations of the method. In at least one embodiment, the processing threads implementing method 600 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization logic). Alternatively, the processing threads implementing method 600 can be executed asynchronously with respect to each other. This can be done in a manner similar to... Figure 6 Various operations of method 600 are executed in different orders as shown. Some operations of these methods can be executed concurrently with other operations. In at least one embodiment, Figure 6 One or more of the operations shown are not always performed.

[0179] refer to Figure 6The processing logic begins with a configuration file (box 602) specifying multiple virtual bridges and their interface mappings. At box 604, the processing logic generates a first virtual bridge and a second virtual bridge from the multiple virtual bridges based on the configuration file. The first virtual bridge is controlled by a first network service hosted on the DPU, and the second virtual bridge is controlled by user-defined logic. At box 606, the processing logic adds one or more host interfaces to the second virtual bridge based on the configuration file. At box 608, the processing logic adds a first service interface to the first virtual bridge based on the configuration file to operatively couple it to the first network service. At box 610, processing logic 102 adds one or more virtual ports between the first and second virtual bridges based on the configuration file. At box 612, the processing logic adds one or more network ports (uplink ports) to the first and / or second virtual bridges based on the configuration file.

[0180] In another embodiment, the method 600 further includes: processing logic adding a first network interface to a first virtual bridge to be operatively coupled to a first network port of the DPU, adding a second network interface to the first virtual bridge to be operatively coupled to a second network port of the DPU, and adding a first host interface from one or more host interfaces to a second virtual bridge to be operatively coupled to a host device; and adding a second host interface from one or more host interfaces to the second virtual bridge to be operatively coupled to a host device, according to a configuration file. The method 600 further includes: processing logic configuring a first link state propagation in the second virtual bridge between the first host interface and one or more virtual ports, and a second link state propagation in the second virtual bridge between the second host interface and one or more virtual ports, according to a configuration file.

[0181] In another embodiment, the method 600 further includes: processing logic adding a first network interface to a first virtual bridge to be operatively coupled to a first network port of the DPU, adding a second network interface to the first virtual bridge to be operatively coupled to a second network port of the DPU, adding a first host interface from one or more host interfaces to a second virtual bridge to be operatively coupled to a first host device, and adding a second host interface from one or more host interfaces to the second virtual bridge to be operatively coupled to a second host device. The first host device is at least one of a virtual machine or a container, and the second host device is at least one of a virtual machine or a container. In another embodiment, the method 600 further includes: processing logic configuring a first link state propagation in the second virtual bridge between the first host interface and one or more virtual ports, and a second link state propagation in the second virtual bridge between the second host interface and one or more virtual ports, according to a configuration file.

[0182] In another embodiment, the method 600 further includes: processing logic installing an OS for execution on the processing device of the DPU. The processing logic may generate multiple virtual bridges and interface mappings for the multiple virtual bridges as part of installing the OS on the DPU.

[0183] In another embodiment, the method 600 further includes: installing an OS on the processing logic to execute on the processing device of the DPU. The processing logic can be part of the DPU's runtime and can generate multiple virtual bridges and interface mappings for the multiple virtual bridges without installing an OS on the DPU.

[0184] Flexible bootstrapping rules in SFC architecture

[0185] Bootstrapping rules are a crucial component of network and traffic management, specifying how data packets are directed across the network based on particular criteria. These rules can be applied in various contexts, including load balancing, security, compliance, and optimization of network resources. The following describes some common types of bootstrapping rules.

[0186] Source IP-based routing focuses on routing traffic based on its originating IP address. This is crucial for managing traffic from specific regions or known addresses and is very useful for location, imposing geographic constraints, or enhancing security. On the other hand, destination IP-based routing targets traffic to its intended endpoint, allowing networks to route traffic to specific data centers or cloud regions based on the destination IP address.

[0187] Port-based bootstrapping uses TCP or UDP port numbers to direct specific types of traffic (such as HTTP or File Transfer Protocol (FTP)) to the resources best suited to handle them, thus optimizing both performance and security. Application-aware bootstrapping goes further, examining packets to identify the applications generating the traffic and routing different types of application traffic through paths optimized for their specific needs (such as low latency or high bandwidth).

[0188] Load-based routing is typically combined with load balancers to direct traffic based on the current load or capacity of a network path, distributing the load evenly and preventing any resource from becoming a bottleneck. Time-based routing is effective for managing network load at different times of the day or week, routing traffic to backup resources during peak periods to maintain performance.

[0189] Protocol-based routing makes routing decisions based on the specific protocol used (such as HTTP or HTTPS), ensuring that traffic is processed according to its specific requirements. Content- or data-based routing examines the content in packets and directs types such as video or text to processing services optimized for those data types, thereby enhancing content delivery.

[0190] User-identity-based routing directs traffic based on a user's identity or role, allowing the network to provide differentiated services or implement security policies tailored to specific user profiles.

[0191] The combination of these guiding rules can form a comprehensive approach to managing traffic in complex environments such as data centers, cloud networks, and large enterprises, thereby ensuring efficient resource utilization and maintaining robust performance and security standards across the network.

[0192] Existing solutions cannot provide flexible boot rules within a single accelerated data plane on the DPU. Various aspects and embodiments of this disclosure can provide flexible boot rules within a single accelerated data plane on the DPU. Various aspects and embodiments of this disclosure can provide hardware-accelerated flexible boot rules on the SFC architecture, as described below regarding... Figures 7 to 10 As described.

[0193] Figure 7 This is a block diagram of an example DPU-based SFC infrastructure 700 for providing hardware acceleration rules for SFC architecture 220 according to at least one embodiment. Except that the DPU-based SFC infrastructure 700 includes hardware acceleration rules 708 derived from network rules from different sources in SFC architecture 220 and accelerated on a single accelerated data plane 702 of the accelerated hardware engine 214, as described in more detail below, the DPU-based SFC infrastructure 700 is similar to (as indicated by similar reference numerals) the DPU-based SFC infrastructure 200.

[0194] Acceleration hardware engine 214 can provide a single accelerated data plane 702 for SFC architecture 220. Memory 218 can store configuration file 124, which specifies at least a first virtual bridge, a second virtual bridge, and a virtual port between the first and second virtual bridges. SFC logic 102 can generate a first virtual bridge based on configuration file 124, which is controlled by a first network service hosted on DPU 204 and has a first set of one or more network rules 704. The first set of one or more network rules 704 may include Layer 2 (L2) protocol rules, Layer 3 (L3) protocol rules, tunneling protocol rules, access control list (ACL) rules, equal cost multipath (ECMP) rules, tunneling encapsulation rules, tunneling decapsulation rules, connection tracking (CT) rules, virtual local area network (VLAN) rules, network address translation (NAT) rules, etc.

[0195] SFC logic 102 can generate a second virtual bridge according to configuration file 124, which has a second set of one or more user-defined network rules 706. In at least one embodiment, the user-defined network rules 706 can be programmed by a user or controller. User-defined network rules 706 may include L2 protocol rules, L3 protocol rules, tunneling protocol rules, ACL rules, ECMP rules, tunneling encapsulation rules, tunneling decapsulation rules, CT rules, VLAN rules, NAT rules, etc. In at least one embodiment, network rule 704 may include a first set of bootstrapping rules for the first virtual bridge, and user-defined network rule 706 may include a second set of bootstrapping rules for the second virtual bridge. Bootstrapping rules may be application-based bootstrapping rules, policy-based bootstrapping rules, geolocation-based bootstrapping rules, load balancing rules, QoS rules, failover rules, redundancy rules, security-based bootstrapping rules, cost-based routing rules, software-defined wide area network (SD-WAN) path bootstrapping rules, software-defined networking (SDN) rules, etc.

[0196] SFC logic 102 can add virtual ports between the first virtual bridge and the second virtual bridge according to configuration file 124. SFC logic 102 can also generate a combined set of network rules based on a first set of one or more network rules 704 and a second set of one or more user-defined network rules 706. This combined set of rules can be added to a single accelerated data plane 702 as hardware acceleration rules 708. Accelerated hardware engine 214 can use the hardware acceleration rules 708 (i.e., the combined set of network rules) to process network traffic data 212 in a single accelerated data plane 702.

[0197] In at least one embodiment, the virtual bridge 104, which includes the first and second virtual bridges described above, is an Open Virtual Switch (OVS) bridge. The DPU 204 can execute an OVS application with a hardware offload mechanism to provide a single accelerated data plane 702 in the acceleration hardware engine 214, thereby processing network traffic data 212 using hardware acceleration rules 708 (i.e., a combined set of network rules).

[0198] In at least one embodiment, SFC logic 102 may add one or more host interfaces to the second virtual bridge to operatively couple to one or more host devices operatively coupled to the DPU; and add one or more network interfaces to the first virtual bridge to operatively couple to one or more network ports of the DPU. SFC logic 102 may add a first service interface to the first virtual bridge to operatively couple to a first network service, and add a second service interface to the second virtual bridge to operatively couple to a second network service. The first and second network services may be part of the SFC architecture 220 of the DPU-based SFC infrastructure 700. The first and second network services may provide accelerated network capabilities within a single accelerated data plane 702 using hardware acceleration rules 708 (i.e., a combined set of network rules).

[0199] As described above, DPU 204 can generate and configure SFC architecture 220 with hardware acceleration rules for network function 222 within a single accelerated data plane of the SFC infrastructure. The following section discusses... Figure 8 and Figure 9 An example of the SFC architecture is explained and described.

[0200] Figure 8 This is a block diagram of an SFC architecture 800 with flexible hardware acceleration rules for a single accelerated data plane, according to at least one embodiment. The SFC architecture 800 is similar to the SFC architecture 300, but uses different reference numerals. The SFC architecture 800 includes a first virtual bridge 802 (labeled "OVS BR-1"), a second virtual bridge 804 (labeled "OVS BR-2"), a virtual port 806, and a network service 808. As described herein, SFC logic 102 can generate the first virtual bridge 802, the second virtual bridge 804, and the virtual port 806 in the SFC architecture 800. The SFC logic 102 can configure the first virtual bridge 802 to be controlled by network service rules 814 provided by the network service 808 hosted on the DPU 204, and configure the second virtual bridge 804 to be controlled by user-defined network rules 816. The SFC logic 102 adds a service interface to the first virtual bridge 802 to operatively couple the network service 808 to the first virtual bridge 802. SFC logic 102 adds a virtual port 806 between the first virtual bridge 802 and the second virtual bridge 804. Network service 808 can provide one or more network service rules 814 to the first virtual bridge 802. User-defined logic 126 can provide one or more user-defined network rules 816 to the second virtual bridge 804. SFC logic 102 can add one or more host interfaces 310 to the second virtual bridge 804.

[0201] As illustrated, three separate host interfaces can be added to connect the second virtual bridge 804 to hosts, such as three separate VMs hosted on host device 202. For example, one VM can host a firewall application, another a load balancer application, and yet another an IDS application. SFC logic 102 can add one or more network interfaces 812 to the first virtual bridge 802. Specifically, SFC logic 102 can add a first network interface to the first virtual bridge 802 to operatively couple a first network port 314 (labeled PORT1) to the DPU 204, and add a second network interface to operatively couple a second network port 316 (labeled PORT2) to the DPU 204. The first virtual bridge 802 can receive network traffic data from the first network port 314 and the second network port 316. The first virtual bridge 802 can redirect network traffic data to the second virtual bridge 804 via virtual port 806. The second virtual bridge 804 can redirect network traffic data to the corresponding host via host interface 810.

[0202] In at least one embodiment, user-defined network rules 816 may be provided by user 818. User 818 may provide user-defined network rules 816 using user-defined logic 126. User 818 may program the second virtual bridge 804 using user-defined network rules 816. Alternatively, user-defined logic 126 may receive user input from user 818, and user-defined logic 126 may generate user-defined network rules 816 and provide them to the second virtual bridge 804. In another embodiment, user-defined network rules 816 may be provided by a user-defined service or another network service separate from network service 808. The other network service (or user-defined service) may be a user-defined network service, user-defined security service, user-defined telemetry service, user-defined storage service, etc., of an application hosted on DPU 204 or as host device 202. When user-defined network rules 816 are provided by a second network service, SFC logic 102 may add another service interface to the second virtual bridge 804 to operatively couple the second network service to the second virtual bridge 804.

[0203] DPU 204 can combine network rules corresponding to different network services to obtain a set of combined network rules that can be accelerated in a single accelerated data plane 702 of DPU 204. The set of combined network rules becomes the hardware acceleration rules accelerated by DPU 204 of SFC architecture 800.

[0204] In at least one embodiment, DPU 204 provides DPU services supporting Host-Based Networking (HBN) as Network Service 808 for accelerating L2 / L3 / tunneling protocols on DPU 204. The HBN infrastructure is based on an SFC topology, where one OVS bridge is controlled by the HBN service, thus providing all accelerated network capabilities. As described above, the second OVS bridge (second virtual bridge 804) can be controlled by user 818 or any other controller (such as...). Figure 9 (As illustrated in the diagram) Programming. The HBN service can support different protocols and network capabilities, such as ACL, ECMP, tunneling, CT, etc. User 818 can program different boot rules on the SFC architecture 800 in a flexible manner in parallel with the HBN service, which produces hardware acceleration rules 708 for a single accelerated data plane 702 provided by OVS-DOCA and DPU hardware. Using the SFC infrastructure, users and customers can use DPU 204 as a network accelerator on edge devices without the need for complex smart switches in different network topologies in DC or SP networks.

[0205] It should be noted that SFC logic 102 can generate different combinations of hardware acceleration rules 708 in different SFC architectures, such as... Figure 9 As shown in the diagram.

[0206] Figure 9This is a block diagram of an SFC architecture 900 with flexible hardware acceleration rules for a single accelerated data plane, according to at least one embodiment. The SFC architecture 900 is similar to the SFC architecture 800 as indicated by similar reference numerals, except that it receives user-defined network rules 816 from a controller 902 (such as an Open Virtual Network (OVN) controller). OVN is an open-source project designed to provide network virtualization solutions that enable the creation and management of virtual network infrastructure in cloud and data center environments. It is an extension of the OVS project, leveraging its underlying technology to provide advanced network automation and scalability capabilities for virtualized networks. OVN aims to simplify the process of setting up and managing virtual network components such as virtual switches, routers, firewalls, and load balancers. It allows these components to be created dynamically through software without the need for manual configuration of physical network hardware. This facilitates the deployment of highly flexible and scalable networks that can easily adapt to the changing needs of applications and services running in a virtualized environment. OVN abstracts the physical network, allowing users to define logical networks that are mapped to the underlying physical infrastructure. This abstraction layer simplifies network design and management by enabling the use of high-level architectures such as logical switches and routers. Through integration with orchestration systems (e.g., OpenStack), OVN supports the automatic provisioning and configuration of network resources based on the automated network management requirements of deployed applications and services. OVN provides a range of network services, including L2 / L3 virtual networks, access control policies, NAT, and more, thus providing the functionality needed to support complex network topologies with advanced network services. With OVN, isolated virtual networks can be created, allowing security policies and rules to be applied at the logical level to ensure that only authorized traffic can flow between different parts of the network. OVN is designed to scale efficiently with the size of the network and the number of virtualized workloads, aiming to minimize the performance impact as the network grows. OVN is a tool for organizations looking to leverage the benefits of network virtualization, providing an efficient and flexible approach to managing virtual network infrastructure in modern cloud and data center environments. In another embodiment, controller 902 can be other types of controllers, such as controllers for specific network services provided in the SFC architecture 900.

[0207] Figure 10This is a flowchart of an example method 1000 of an SFC architecture configured according to at least one embodiment, having flexible hardware acceleration rules for acceleration on a single accelerated data plane of a DPU. The processing logic can be a combination of hardware, firmware, software, or any combination thereof. In at least one embodiment, the processing logic is implemented in a DPU, switch, network device, GPU, NIC, CPU, etc. In at least one embodiment, the processing logic is implemented in an acceleration hardware engine coupled to the switch. In at least one embodiment, method 1000 can be executed by multiple processing threads, each thread executing one or more individual functions, routines, subroutines, or operations of the method. In at least one embodiment, the processing threads implementing method 1000 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization logic). Alternatively, the processing threads implementing method 1000 can be executed asynchronously with respect to each other. This can be done in parallel with... Figure 10 The operations of method 1000 are executed in different orders as shown. Some operations of these methods can be executed concurrently with other operations. In at least one embodiment, Figure 10 One or more of the operations shown are not always performed.

[0208] refer to Figure 10 The processing logic begins by storing a configuration file (box 1002) that specifies at least a first virtual bridge, a second virtual bridge, and virtual ports between the first and second virtual bridges. At box 1004, the processing logic generates the first and second virtual bridges based on the configuration files. The first virtual bridge is controlled by a first network service hosted on the DPU and has a first set of one or more network rules; the second virtual bridge has a second set of one or more user-defined network rules. At box 1006, the processing logic adds virtual ports between the first and second virtual bridges based on the configuration files. At box 1008, the processing logic adds one or more network ports (uplink ports) to the first and / or second virtual bridges based on the configuration files. At box 1010, the processing logic generates a combined set of network rules based on the first set of one or more network rules and the second set of one or more user-defined network rules, according to the configuration files. At box 1012, the processing logic utilizes the accelerated hardware engine to process network traffic data in a single accelerated data plane using the combined set of network rules.

[0209] In another embodiment, the method 1000 may further include: processing logic adding one or more network interfaces to the first virtual bridge according to a configuration file to be operatively coupled to one or more network ports of the DPU, and adding a first service interface to the first virtual bridge to be operatively coupled to a first network service. The first network service may use a first set of one or more network rules to provide accelerated network capabilities, wherein the first set of one or more network rules includes any one or more of the following: L2 protocol rules, L3 protocol rules, tunneling protocol rules, ACL rules, ECMP rules, tunneling encapsulation rules, tunneling decapsulation rules, CT rules, VLAN rules, NAT rules, etc.

[0210] In another embodiment, the method 1000 may further include: processing logic adding one or more host interfaces to the second virtual bridge according to a configuration file to operatively couple one or more host devices, the one or more host devices being operatively coupled to the DPU; and receiving user input from a user or controller, the user input specifying a second set of one or more user-defined network rules, the second set of one or more user-defined network rules including one or more bootstrapping rules. The processing logic may add one or more second sets of user-defined network rules to the second virtual bridge according to a configuration file. In at least one embodiment, the one or more bootstrapping rules may include any one or more of the following: application-based bootstrapping rules, policy-based bootstrapping rules, geolocation-based bootstrapping rules, load balancing rules, QoS rules, failover rules, redundancy rules, security-based bootstrapping rules, cost-based routing rules, SD-WAN path bootstrapping rules, software-defined networking (SDN) rules, etc.

[0211] In another embodiment, the method 1000 may further include: processing logic adding one or more network interfaces to a first virtual bridge according to a configuration file to be operatively coupled to one or more network ports of the DPU, and adding a first service interface to the first virtual bridge to be operatively coupled to a first network service, the first network service being used to provide accelerated network capabilities. The first set of one or more network rules includes at least one of the following: ACL rules, ECMP rules, tunneling rules, CT rules, QoS rules, STP, VLAN rules, NAT rules, SDN rules, MPLS rules, etc.

[0212] In another embodiment, the method 1000 may further include: processing logic 102 adding one or more host interfaces to a second virtual bridge according to a configuration file to operatively couple one or more host devices, the one or more host devices being operatively coupled to the DPU; and adding one or more network interfaces to a first virtual bridge to operatively couple one or more network ports of the DPU. In another embodiment, the method 1000 may further include: processing logic adding a first service interface to the first virtual bridge according to a configuration file to operatively couple a first network service. In another embodiment, the method 1000 may further include: processing logic adding a second service interface to the second virtual bridge according to a configuration file to operatively couple a second network service, wherein the first network service and the second network service are part of an SFC infrastructure for providing accelerated network capabilities in a single accelerated data plane using a combined set of network rules.

[0213] Create optimized and accelerated network pipelines using the Network Pipeline Abstraction Layer (NPAL).

[0214] A Database Abstraction Layer (DAL) is a software component that provides a unified interface for interacting with different types of database systems, enabling applications to perform database operations without using database-specific syntax. The DAL acts as an intermediary between the application and the database, translating the application's data access requests into appropriate queries to the underlying database. By abstracting the details of the database system, the DAL allows developers to write database-independent code, enhancing the portability and scalability of applications. This layer can support a wide range of database operations, including creating, reading, updating, and deleting records, and can be implemented in various forms, such as Object Relational Mapping (ORM) libraries, which further simplify data manipulation by allowing data to be disposed of according to objects.

[0215] As described herein, integrated circuits (e.g., DPUs) can provide a Network Pipeline Abstraction Layer (NPAL) similar to the DAL. NPAL is a software-programmable layer that provides an optimized and accelerated network pipeline supporting various accelerated network capabilities, such as L2 bridging, L3 routing, tunnel encapsulation, tunnel decapsulation, hash computation, ECMP operations, static and dynamic ACLs, CT, etc. NPAL may include a set of classes or APIs providing a unified interface for performing common network operations within the network pipeline, optimized for hardware acceleration on an acceleration hardware engine. The network pipeline may include a set of tables and logic ordered in a specific order, optimized for acceleration by the DPU hardware's acceleration hardware engine, thereby providing customers and users with a rich set of capabilities and high performance.

[0216] Using NPAL in a DPU offers various benefits, including operational independence, encapsulation of logic, performance, code reusability, and platform independence. For example, developers can write agnostic code, allowing applications (e.g., network services) to work with different underlying access logic and network functions. NPAL can encapsulate access or network function-related logic, making it easier to manage and maintain the codebase system. Changes to the schema or underlying technology can be isolated in NPAL implementations. NPAL can provide optimized, high-performance pipelines to address diverse network requirements and functions. By separating access logic from application logic, developers can reuse NPAL components across multiple parts of the application (network services), promoting code reuse and maintainability. NPAL can abstract platform-specific differences, data types, and other access or network function-related characteristics, enabling applications (network services) to run seamlessly across different platforms and environments. Overall, NPAL can be a powerful tool for building flexible, scalable, and maintainable network function-driven applications, providing a level of abstraction that simplifies interactions between network functions and promotes code efficiency and portability.

[0217] In at least one embodiment, the DPU includes DPU hardware, which includes a processing device and an acceleration hardware engine. The DPU includes memory operatively coupled to the DPU hardware. The memory may store DPU software, which includes NPALs supporting multiple network protocols and network functions in the network pipeline. The network pipeline includes a collection and logic of tables organized in a specific order for acceleration by the acceleration hardware engine. The acceleration hardware engine can use the network pipeline to process network traffic data. The network pipeline can be optimized for network services running on the DPU.

[0218] Figure 11 This is a block diagram of an example computing system 1100 with a DPU 1104 according to at least one embodiment, the DPU 1104 having an NPAL 1114 for providing an optimized and accelerated network pipeline to be accelerated by an acceleration hardware engine 1116. The computing system 1100 includes the DPU 1104 coupled between a host device 1102 and a network 1110. Except as explicitly indicated, the host device 1102 and DPU 1104 may be similar to the host device 202 and DPU 204 of the DPU-based SFC infrastructure 200 and DPU-based SFC infrastructure 700 described above.

[0219] In at least one embodiment, DPU 1104 includes DPU hardware 1108 and DPU software 1106 (e.g., a software framework with acceleration libraries). DPU hardware 1108 may include one or more CPUs (e.g., single-core or multi-core CPUs), an acceleration hardware engine 1116 (or multiple hardware accelerators), memory, and network and host interconnects. In at least one embodiment, DPU 1104 includes DPU software 1106, which includes a software framework and acceleration libraries. The software framework and acceleration libraries may include one or more hardware acceleration services, including hardware acceleration services (e.g., NVIDIA DOCA), hardware acceleration virtualization services, hardware acceleration network services, hardware acceleration storage services, hardware acceleration AI / ML services, hardware acceleration services, and hardware acceleration management services. DPU software 1106 also includes NPAL 1114. NPAL 1114 may include a set of classes or APIs that provide a unified interface for performing common network operations in a network pipeline optimized for hardware acceleration on the acceleration hardware engine 1116. This category or set of APIs can provide a unified interface to one or more applications, network services, or other logic executed by DPU 1104 or host device 1102. The network pipeline may include a set of tables and logic in a specific order, and the network pipeline is optimized for acceleration by the acceleration hardware engine 1116 of DPU hardware 1108.

[0220] During operation, DPU hardware 1108 can receive network traffic data 1112 from network 1110 and process the network traffic data 112 using an optimized and accelerated network pipeline programmed by NPAL 1114. As described herein, NPAL 1114 supports multiple network protocols and network functions in the network pipeline. The network pipeline includes a collection and logic of tables organized in a specific order to be accelerated by acceleration hardware engine 1116. Acceleration hardware engine 1116 can use the network pipeline to process network traffic data 1112. In at least one embodiment, the network pipeline includes input ports, ingress dynamic or static access control lists (ACLs), bridges, routers, egress dynamic or static ACLs, and output ports. The following section discusses... Figure 12 , Figure 13 and Figure 14 Examples of optimized and accelerated network pipelines are illustrated and described.

[0221] Figure 12This is a network diagram of an example network pipeline 1200 optimized and accelerated on an acceleration hardware engine with a DPU having NPAL, according to at least one embodiment. The network pipeline 1200 includes an input port 1202, a filtering network function 1204, an ingress port 1206, a first network function 1208, a bridge 1210, a Switched Virtual Interface (SVI) ACL 1212, a router 1214, a second network function 1216, an egress port 1218, and an output port 1220. The input port 1202 can receive network traffic data and provide network traffic data to the filtering network function 1204, which is operatively coupled to the input port 1202. The filtering network function 1204 can filter network traffic data. The ingress port 1206 is operatively coupled to the filtering network function 1204 and is also operatively coupled to the first network function 1208. The first network function 1208 can process network traffic data using one or more ingress ACLs. Bridge 1210 is operatively coupled to a first network function 1208. Bridge 1210 can perform Layer 2 (L2) bridging operations. One or more SVI ACLs are operatively coupled to bridge 1210 and router 1214. Router 1214 can perform Layer 3 (L3) routing operations. A second network function 1216 is operatively coupled to router 1214. Second network function 1216 can process network traffic data using one or more egress ACLs. Egress port 1218 is operatively coupled to second network function 1216. Egress port 1218 is operatively coupled to output port 1220. Output port 1220 can output network traffic data.

[0222] Figure 13This is a network diagram of an example network pipeline optimized and accelerated on an acceleration hardware engine of a DPU with NPAL, according to at least one embodiment. The network pipeline 1300 includes an input port 1302, a filtering network function 1304, an ingress port with a first network function 1306, a second network function 1308, a bridge 1310, an SVIACL 1312, a router 1314, a third network function 1316, an egress port with a fourth network function 1318, and an output port 1320. The input port 1302 can receive network traffic data and provide network traffic data to the filtering network function 1304, which is operatively coupled to the input port 1302 having the first filtering network function 1304. The filtering network function 1304 can filter network traffic data, thereby providing network traffic data to the ingress port having the first network function 1306. The first network function 1306 can perform a first Virtual Local Area Network (VLAN) mapping on network traffic data and provide the data to the second network function 1308 (or alternatively, bridge 1310). The second network function can process network traffic data using one or more ingress access control lists (ACLs). The ingress ACLs can be dynamic or static ACLs. Bridge 1310 is operatively coupled to the first and second network functions 1308. Bridge 1310 can perform Layer 2 (L2) bridging operations. An SVIACL is operatively coupled to bridge 1310 and router 1314. Router 1314 can perform Layer 3 (L3) routing operations. The third network function 1316 is operatively coupled to router 1314. The third network function 1316 can process network traffic data using one or more egress ACLs. The egress ACLs can be dynamic or static. An egress port with a fourth network function 1318 is operatively coupled to the third network function 1316. The fourth network function can perform a second VLAN mapping on network traffic data. Output port 1320 is operatively coupled to an egress port having a fourth network function 1318. Output port 1320 can output network traffic data.

[0223] Figure 14This is a network diagram of an example network pipeline 1400 optimized and accelerated on an acceleration hardware engine with a DPU having NPAL, according to at least one embodiment. Network pipeline 1400 includes an input port 1402, a first network function 1404, a second network function 1406, a third network function 1408, a fourth network function 1410, a fifth network function 1412, a sixth network function 1414, a seventh network function 1416, and an output port 1418. Network pipeline 1400 may include any combination of network functions in a specified order, each network function including logic and / or tables for implementing the corresponding network function and passing network traffic data to subsequent network functions. In at least one embodiment, the first network function 1404 can perform Layer 2 (L2) bridging, the second network function 1406 can perform Layer 3 (L3) routing, the third network function 1408 can perform tunnel encapsulation or tunnel decapsulation, the fourth network function 1410 can perform hash calculation, the fifth network function 1412 can perform ECMP operation, the sixth network function 1414 can perform CT operation, and the seventh network function 1416 can perform NAT operation. Alternatively, the network pipeline 1400 can include different numbers and types of network operations between the input port 1402 and the output port 1418.

[0224] Figure 15 This is a flowchart of an example method 1500 for creating an optimized and accelerated network pipeline using a Network Pipeline Abstraction Layer (NPAL) according to at least one embodiment. The processing logic can be hardware, firmware, software, or any combination thereof. In at least one embodiment, the processing logic is implemented in a DPU, switch, network device, GPU, NIC, CPU, etc. In at least one embodiment, the processing logic is implemented in an acceleration hardware engine coupled to the switch. In at least one embodiment, method 1500 can be executed by multiple processing threads, each thread executing one or more individual functions, routines, subroutines, or operations of the method. In at least one embodiment, the processing threads implementing method 1500 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization logic). Alternatively, the processing threads implementing method 1500 can be executed asynchronously with respect to each other. Figure 15 Various operations of method 1500 are executed in different orders as shown. Some operations of these methods can be executed concurrently with other operations. In at least one embodiment, Figure 15 One or more of the operations shown are not always performed.

[0225] refer to Figure 15The processing logic begins by executing one or more instructions from the Network Pipeline Abstraction Layer (NPAL) that supports multiple network protocols and network functions within the network pipeline (box 1502). The network pipeline comprises a collection of tables and logic organized in a specific order to be accelerated by the DPU's accelerated hardware engine. At box 1504, the processing logic receives network traffic data over the network. At box 1506, the processing logic utilizes the DPU's accelerated hardware engine to process the network traffic data using the network pipeline.

[0226] In at least one embodiment, the network pipeline includes the above-mentioned... Figure 12 , Figure 13 and Figure 14 The described port and network functions.

[0227] NPAL supports splitting interfaces

[0228] Figure 16 This is a block diagram of the software stack 1600 of a DPU with an NPAL 1602 supporting port splitting, according to at least one embodiment. As described above, to support port splitting, the DPU needs to support it in both hardware and software. From a hardware perspective, the DPU hardware includes a network interconnect including one or more physical ports coupled to the network, a host interconnect coupled to host devices, and an acceleration hardware engine. The network interconnect includes physical ports 1614 (or more physical ports) configured to couple to a branch cable 1616, which physically couples to a collection of multiple devices. The branch cable 1616 physically separates the data flow associated with individual devices. In high-speed network environments, especially in data centers, it is common to use branch cables that allow a single high-speed port (such as 40Gbps or 100Gbps) to be “split” into multiple low-speed ports (such as 4×10Gbps or 4×25Gbps). These branch cables physically separate data flows and allow multiple devices or ports to connect to a single high-speed interface. DPU hardware may include additional hardware such as processing devices (e.g., one or more CPU cores), one or more GPUs, switches, memory, etc. It should also be noted that the DPU hardware may have a second physical port, which is also split using a second branch cable physically coupled to a second set of devices.

[0229] From a software perspective, software stack 1600 can support logical split ports via a single physical port 1614. Software stack 1600 can be stored in the DPU's memory. Software stack 1600 includes firmware 1606, driver 1608, virtual switch 1612, and NPAL 1602. Firmware 1606 can interact with DPU hardware 1604 and driver 1608. Driver 1608 can interact with firmware 1606 and virtual switch 1612. NPAL 1602 can reside on top of virtual switch 1612. Software stack 1600 can consist of instructions that, when executed by the DPU hardware, can provide NPAL 1602 with multiple network protocols and network functions supporting the set of logical split ports 1610 in the network pipeline. Each logical split port 1610 corresponds to one device in a set of devices. As described above, the network pipeline includes a set of tables and logic organized in a specific order by the acceleration hardware engine.

[0230] In at least one embodiment, firmware 1606 can be configured to generate multiple logically split ports 1610 as physical functions (PFs) and map the physical channels of physical port 1614 to different logically split ports 1610. In at least one embodiment, firmware 1606 and driver 1608 can present multiple logically split ports 1610 as PFs to virtual switch 1612 and NPAL 1602. NPAL 1602 can be configured to manage each PF as if it were a separate physical port, even if these PFs are supported on a single physical port 1614.

[0231] NPAL 1602 can configure different policies for different PFs. In at least one embodiment, a first logical split port is configured using a first policy, and a second logical split port is configured using a second policy different from the first policy. As described herein, NPAL 1602 includes a set of classes or APIs that provide a unified interface to one or more applications executed by a processing device or a host device coupled to a DPU. NPAL 1602 can be used to define different policies for different logical split ports. NPAL 1602 can be used to configure a network pipeline to perform a first network function for a first logical split port (first PF) and a second network function for a second logical split port (second PF), which is different from the first network function. For example, different policies (such as different QoS requirements, different traffic management, etc.) can be used to configure the network pipeline. In at least one embodiment, different network functions may include L2 bridging, L3 routing, tunnel encapsulation, tunnel decapsulation, hash calculation, ECMP operation, static and dynamic ACLs, CT, etc. Similarly, different network functions can be performed for other logical split ports (other PFs). Once configured, the DPU hardware can use the network pipeline to process network traffic data.

[0232] like Figure 16 As illustrated, because branch cable 1616 couples to four separate devices, firmware 1606 generates a first logical split port, a second logical split port, a third logical split port, and a fourth logical split port. For example, physical port 1614, which can support 100Gbps (or 40Gbps) of network traffic data, will be "split" into multiple low-speed logical ports of 4 × 25Gbps (or 4 × 10Gbps). Once physical port 1614 is physically and logically split, the network pipeline can send network traffic data to any of the logical split ports 1610 (i.e., to any PF) that are part of the output port logic.

[0233] Generating and using multiple PFs for a single physical port 1614 can be used to isolate networks between PFs, thereby providing better resource utilization and supporting more than two PFs instead of one or more PFs. In at least one embodiment, NPAL 1602 is configured to configure a first logical split port using a first policy and a second logical split port using a second policy different from the first policy. Similarly, NPAL 1602 can configure a third logical split port using a third policy and a fourth logical split port using a fourth policy. NPAL 1602 can configure multiple logical split ports using the same policy. In at least one embodiment, NPAL 1602 is configured to isolate networks between multiple PFs. In at least one embodiment, the set of logical split ports 1610 includes two or more logical splits (i.e., more than two logical split ports). When a branch cable is inserted into the physical port of the DPU, NPAL 1602 can provide customers or users with a rich set of capabilities, high performance, and flexibility.

[0234] In at least one embodiment, the network pipeline includes input ports, ingress dynamic or static ACLs, bridges, routers and egress dynamic or static ACLs, and output ports with multiple logically split ports 1610.

[0235] In at least one embodiment, the network pipeline includes an input port for receiving network traffic data via physical port 1614. The network pipeline includes a filtering network function operatively coupled to the input port for filtering network traffic data. The network pipeline includes an ingress port operatively coupled to the filtering network function and a first network function operatively coupled to the ingress port for processing network traffic data using one or more ingress access control lists (ACLs). The network pipeline includes a bridge operatively coupled to the first network function for performing L2 bridging operations; and one or more SVIACLs (e.g., static or dynamic) operatively coupled to the bridge. The network pipeline includes a router operatively coupled to the SVIACLs (e.g., static or dynamic) for performing L3 routing operations. The network pipeline includes a second network function operatively coupled to the router for processing network traffic data using one or more egress ACLs. The network pipeline includes an egress port operatively coupled to the second network function and an output port having a logically split set of ports 1610 for outputting network traffic data. As described above, once physical port 1614 is logically and physically split, network traffic data can be sent to any of the logically split ports in logically split port 1610.

[0236] In at least one embodiment, the network pipeline includes an input port for receiving network traffic data via physical port 1614. The network pipeline includes a filtering network function operatively coupled to the input port for filtering network traffic data. The network pipeline includes an ingress port operatively coupled to the filtering network function, the ingress port having a first network function for performing a first VLAN mapping on the network traffic data. The network pipeline includes a second network function operatively coupled to the ingress port for processing network traffic data using one or more ingress ACLs. The network pipeline includes a bridge operatively coupled to the second network function for performing L2 bridging operations; and one or more SVI ACLs (e.g., static or dynamic) operatively coupled to the bridge. The network pipeline includes a router operatively coupled to the SVI ACLs (e.g., static or dynamic) for performing Layer 3 (L3) routing operations. The network pipeline includes a third network function operatively coupled to the router for processing network traffic data using one or more egress ACLs. The network pipeline includes an egress port operatively coupled to a third network function, which has a fourth network function for performing a second VLAN mapping on network traffic data. The network pipeline also includes output ports with a set of logically split ports 1610 for outputting network traffic data. As described above, once physical port 1614 is logically and physically split, network traffic data can be sent to any of the logically split ports 1610.

[0237] In at least one embodiment, the network pipeline includes two or more of the following network functions: a first network function for performing L2 bridging, a second network function for performing L3 routing, a third network function for performing tunnel encapsulation or tunnel decapsulation, a fourth network function for performing hash calculation, a fifth network function for performing ECMP operations, a sixth network function for performing CT operations, or a seventh network function for performing NAT operations. The following section discusses... Figure 17 The example network pipeline is described and illustrated.

[0238] Figure 17 This is a network diagram of an example network pipeline 1700 optimized and accelerated on an acceleration hardware engine of a DPU with NPAL supporting split interfaces, according to at least one embodiment. In addition to the network pipeline 1700 being able to send network traffic data to any of the logical split ports 1610 of the output port 1220, the network pipeline 1700 is connected to... Figure 12The network pipeline is similar to the 1200. The NPAL 1602 can be configured with different policies for different logical splits. The network pipeline can process network traffic data and send it to the appropriate logical split port 1610.

[0239] like Figure 17 As illustrated, network pipeline 1700 includes an input port 1202 that receives network traffic data and provides it to a filtering network function 1204 operatively coupled to the input port 1202. The filtering network function 1204 filters network traffic data. An ingress port 1206 is operatively coupled to the filtering network function 1204 and is also operatively coupled to a first network function 1208. The first network function 1208 can process network traffic data using one or more ingress ACLs. A bridge 1210 is operatively coupled to the first network function 1208. The bridge 1210 can perform L2 bridging operations. One or more SVI ACLs are operatively coupled to the bridge 1210 and router 1214. Router 1214 can perform L3 routing operations. A second network function 1216 is operatively coupled to router 1214. The second network function 1216 can process network traffic data using one or more egress ACLs. Egress port 1218 is operatively coupled to a second network function 1216. Egress port 1218 is operatively coupled to output port 1220. Output port 1220 can output network traffic data on any of the logical split ports 1610.

[0240] Figure 18 This is a flowchart of a method 1800 for operating a DPU with a split interface according to at least one embodiment. The processing logic can be hardware, firmware, software, or any combination thereof. In at least one embodiment, the processing logic is implemented in a DPU, switch, network device, GPU, NIC, CPU, etc. In at least one embodiment, the processing logic is implemented in an acceleration hardware engine coupled to the switch. In at least one embodiment, method 1800 can be executed by multiple processing threads, each thread executing one or more individual functions, routines, subroutines, or operations of the method. In at least one embodiment, the processing threads implementing method 1800 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization logic). Alternatively, the processing threads implementing method 1800 can be executed asynchronously with respect to each other. Figure 18 The operations of method 1800 are executed in different orders as shown. Some operations of these methods can be executed concurrently with other operations. In at least one embodiment, Figure 18 One or more of the operations shown are not always performed.

[0241] refer to Figure 18 The processing logic begins by executing one or more instructions (box 1802) that support multiple network protocols and network functions in the network pipeline for multiple logical split ports. Each logical split port corresponds to one of multiple devices. The network pipeline includes a collection of tables and logic that are accelerated by the acceleration hardware engine and organized in a specific order. At box 1804, the processing logic receives first network data from a first device via a branch cable through a physical port. At box 1806, the processing logic uses the network pipeline to process the first network traffic data using the DPU's accelerated hardware engine. At box 1808, the processing logic sends the first network data to the first logical split port of the multiple logical split ports. At box 1810, the processing logic receives second network data from a second device via a branch cable through a physical port. At box 1812, the processing logic uses the network pipeline to process the second network traffic data using the DPU's accelerated hardware engine. At box 1814, the processing logic sends the first network data to the first logical split port of the multiple logical split ports.

[0242] In another embodiment, the processing logic configures a first logical split port using a first strategy and configures a second logical split port using a second strategy different from the first strategy.

[0243] In another embodiment, processing logic (e.g., firmware) maps the physical channels of a physical port to multiple logically split ports and presents these multiple logically split ports as multiple physical functions (PFs) to the virtual switch and NPAL. The processing logic (e.g., NPAL) manages each of the multiple PFs as if that PF were a separate physical port.

[0244] In at least one embodiment, the processing logic configures a network pipeline to perform a first network function on a first PF among a plurality of PFs, and configures a network pipeline to perform a second network function on a second PF among a plurality of PFs, the second network function being different from the first network function.

[0245] NPAL supports fast link recovery

[0246] As described above, a link failure in a network topology refers to a situation where the communication link between two network devices (such as routers, NICs, switches, or hosts) becomes unavailable due to various reasons such as hardware failure, cable breakage, or network congestion. Link failures can significantly impact the performance, availability, and reliability of network services, especially in large-scale or critical environments such as data centers or enterprise networks. Several factors can lead to link failures, such as physical link failures (e.g., cable damage or breakage due to fiber optic cable severance, Ethernet cable breakage, etc.), device failures (e.g., hardware failures causing interface downtime), congestion or overload (i.e., network links overloaded by excessive traffic, resulting in timeouts or packet drops), software vulnerabilities or misconfigurations that cause network devices to drop connections, power outages, and maintenance activities (e.g., planned interruptions for maintenance or upgrades can also cause temporary link failures). As described above, link failures can have different implications.

[0247] One routing strategy is called ECMP. As described above, NPAL can support rapid link recovery when link failures occur and the Equal Cost Multipath (ECMP) group needs to be updated to reflect the new network topology. Due to the implications of link failures mentioned above, rapid recovery is crucial when links fail. Specifically, when a link fails, the ECMP group must be updated immediately to reflect that the specific link is no longer available. Furthermore, when a link comes online, the ECMP group must be updated immediately with the recovered link to improve network utilization.

[0248] As described in more detail below, NPAL can operate as an accelerated network pipeline and virtual switch hardware offloading mechanism for periodic link monitoring. In at least one embodiment, a user can configure NPAL to support rapid link recovery, enabling link monitoring by the virtual switch. For example, inter-process communication (IPC) messages, such as Linux Netlink messages from the Linux kernel, can be used to monitor all ports in a specific ECMP group. If a link fails (i.e., experiences a link failure), the virtual switch identifies the link as failed and immediately updates the ECMP group by removing the link from the ECMP group. That is, the virtual switch updates the ECMP group in the tables of bridges and / or routers in the network pipeline to remove the failed link. Once the ECMP group is updated, traffic can be distributed to other links in the ECMP group. If a link comes online (i.e., no longer experiences a link failure), the virtual switch identifies the link as online and immediately updates the ECMP group by adding the link to the ECMP group. That is, the virtual switch updates the ECMP group in the tables of bridges and / or routers in the network pipeline to add the recovered link (also called a new link). Once the ECMP group is updated, traffic can be distributed to the new link as well as other links in the ECMP group.

[0249] In at least one embodiment, the DPU includes DPU hardware, which includes a processing device and an acceleration hardware engine, and memory operatively coupled to the DPU hardware. The memory may store instructions that, when executed by the DPU hardware, provide the fast link recovery to the virtual switch and NPAL. The following section discusses… Figure 19 An example software stack for the DPU is described and illustrated.

[0250] Figure 19 This is a block diagram of a software stack 1900 for a DPU with NPAL supporting fast link recovery, according to at least one embodiment. As described above, to support fast link recovery, the DPU needs to support it in both hardware and software. From a hardware perspective, the DPU hardware includes network interconnects coupled to the network, host interconnects coupled to host devices, and acceleration hardware engines. The DPU hardware may include additional hardware such as processing devices (e.g., one or more CPU cores), one or more GPUs, switches, memory, etc.

[0251] From a software perspective, software stack 1900 supports fast link recovery. Software stack 1900 can be stored in the DPU's memory. Software stack 1900 includes firmware 1906, driver 1908, virtual switch 1910, and NPAL 1902. Firmware 1906 can interact with DPU hardware 1904 and driver 1908. Driver 1908 can interact with firmware 1906 and virtual switch 1910. NPAL 1902 can reside on top of virtual switch 1910. Software stack 1900 can consist of instructions that, when executed by the DPU hardware, can support multiple network protocols and network functions in the network pipeline, as provided by NPAL 1902. As described above, the network pipeline includes a collection and logic of tables organized in a specific order and accelerated by the acceleration hardware engine.

[0252] In at least one embodiment, the virtual switch 1910 may include link monitoring logic 1912 that monitors the link availability of each of a plurality of links to the destination. The plurality of links may be part of or specified in an initial identifier group in routing table 1918. That is, the link monitoring logic 1912 can monitor link availability by monitoring all ports associated with the initial link identifier group. The link monitoring logic 1912 may detect a link failure of the first link among the plurality of links. The link monitoring logic 1912 may notify ECMP logic 1914 in NPAL 1902. ECMP logic 1914 may remove the first link identifier associated with the first link from the initial link identifier group to obtain a modified link identifier group in ECMP group table 1916. ECMP logic 1914 may also cause an update to routing table 1918 in NPAL 1902 to remove the first link identifier. The acceleration hardware engine in DPU hardware 1904 can process network traffic data using network pipelines and distribute network traffic data only to the remaining links among the plurality of links corresponding to the modified identifier group.

[0253] In at least one embodiment, the virtual switch is controlled by a network service hosted on the DPU. The virtual switch can send a notification of a link failure on a first link to the network service. The network service can update the routing table in the NPAL in response to the notification. The routing table stores configuration information associated with the initial identifier group. In at least one embodiment, the notification is a message from the kernel (e.g., a NetLink message from the Linux kernel).

[0254] In at least one embodiment, ECMP logic 1914 may receive a first packet before a link failure occurs on the first link, and perform an IP address lookup on the first packet to identify an ECMP group identifier associated with multiple links. ECMP logic 1914 may hash the ECMP group identifier to identify any one of the multiple links. ECMP logic 1914 may receive a second packet after a link failure occurs on the first link, and perform an IP address lookup on the second packet to identify an ECMP group identifier. ECMP logic 1914 may hash the ECMP group identifier to identify any one of the remaining links.

[0255] In at least one embodiment, the virtual switch 1910 can monitor the link availability of each of the multiple links to the destination and detect the link recovery of the first link. ECMP logic 1914 can add a first link identifier to a modified link identifier group to obtain an initial link identifier group in ECMP group table 1916. ECMP logic 1914 can cause the routing table 1918 in NPAL 1902 to be updated to add the first link identifier. An acceleration hardware engine is used to process subsequent network traffic data using network pipelines and distribute the subsequent network traffic data to the multiple links corresponding to the initial identifier group. In at least one embodiment, ECMP logic 1914 can receive a first packet before the link recovery of the first link and perform an IP address lookup on the first packet to identify the ECMP group identifier associated with the multiple links. ECMP logic 1914 can hash the ECMP group identifier to identify any of the remaining links. ECMP logic 1914 can receive a second packet after the link recovery of the first link and perform an IP address lookup on the second packet to identify the ECMP group identifier. ECMP logic 1914 can hash the ECMP group identifier to identify any of the multiple links.

[0256] In at least one embodiment, virtual switch 1910 is controlled by a network service hosted on the DPU. Virtual switch 1910 can send a first notification of a link failure of a first link to the network service, which updates the routing table in NPAL, storing configuration information associated with the initial identifier group, in response to the first notification. Virtual switch 1910 can send a second notification of link recovery of the first link to the network service, which updates the routing table 1918 in NPAL 1902 in response to the second notification. In at least one embodiment, virtual switch 1910 can remove the first link identifier from the ECMP group in virtual switch 1612 and modify the routing table 1918 in NPAL in parallel.

[0257] In at least one embodiment, the network pipeline includes input ports, ingress dynamic or static ACLs, bridges, routers, egress dynamic or static ACLs, and output ports. Routing table 1918 may be stored in bridges, routers, or both. The following section discusses... Figure 20 The example network pipeline is described and illustrated.

[0258] Figure 20 This is a network diagram of an example network pipeline 2000 optimized and accelerated on an accelerated hardware engine of a DPU supporting NPAL, according to at least one embodiment. In addition to the network pipeline 2000 including a routing table 1918 stored at bridge 1210, router 1214, or both, the network pipeline 2000 and... Figure 12 The network traffic data is similar to 120. As described above, the virtual switch can monitor the link availability of each of the multiple links to the destination. Once the virtual switch detects a link failure of the first link, it can remove the first link identifier associated with the first link from the initial link identifier group to obtain the modified link identifier group in ECMP group table 1916. The virtual switch causes the routing table 1918 to be updated to remove the first link identifier. Once the routing table 1918 is updated, the acceleration hardware engine can use the network pipeline to process network traffic data and distribute the network traffic data only to the remaining links of the multiple links corresponding to the modified identifier group. Similarly, the virtual switch can continue to monitor the link availability of each link. The virtual switch can detect the link recovery of the first link. The virtual switch can add the first link identifier to the modified link identifier group to obtain the initial link identifier group and cause the routing table 1918 to be updated to add the first link identifier. Once the routing table 1918 is updated, the acceleration hardware engine can use the network pipeline to process subsequent network traffic data and distribute subsequent network traffic data to the multiple links corresponding to the initial identifier group.

[0259] like Figure 20As illustrated, the network pipeline 2000 includes an input port 1202, a filtering network function 1204, an ingress port 1206, a first network function 1208, a bridge 1210 storing a routing table 1918, an SVI ACL 1212, a router 1214 storing a routing table 1918, a second network function 1216, an egress port 1218, and an output port 1220. Input port 1202 receives network traffic data and provides it to the filtering network function 1204, which is operatively coupled to input port 1202. Filtering network function 1204 filters network traffic data. Ingress port 1206 is operatively coupled to filtering network function 1204 and is also operatively coupled to the first network function 1208. The first network function 1208 can process network traffic data using one or more ingress ACLs. Bridge 1210 is operatively coupled to the first network function 1208. Bridge 1210 can perform ECMP bridging operations using the current routing table 1918, excluding any failed links as described above. One or more SVI ACLs are operatively coupled to bridge 1210 and router 1214. Router 1214 can perform ECMP routing operations using the current routing table 1918, excluding any failed links as described above. Second network function 1216 is operatively coupled to router 1214. Second network function 1216 can process network traffic data using one or more egress ACLs. Egress port 1218 is operatively coupled to second network function 1216. Egress port 1218 is operatively coupled to output port 1220. Output port 1220 can output network traffic data.

[0260] As described above, NPAL may include ECMP logic 1914, which performs ECMP operations on incoming packets before, after, and after a link failure, as described below. Figure 21 The illustrations and descriptions are as follows.

[0261] Figure 21 This is a flowchart illustrating ECMP operations before, after, and after link recovery, according to at least one embodiment. In the first process 2102, the ECMP logic receives a first packet 2108 and performs an IP lookup operation 2110 to return to ECMP group 2112. ECMP group 2112 may be a 5-tuple hash that identifies the first link 2114 (Link0) and the second link 2116 (Link1) as part of ECMP group 2110. The first process 2102 occurs before a link failure is detected.

[0262] In the second process 2104, if a link failure is detected in the first link 2114, the ECMP logic receives the second packet 2118 and performs an IP lookup operation 2110 returning to ECMP group 2112. In these instances, because the first link 2114 has a link failure, the 5-tuple hash of ECMP group 2112 only identifies the second link 2116. As a result, in this case, all traffic will flow only to the second link 2116 (Link1). The second process 2104 occurs after the link failure is detected.

[0263] In the third process 2106, if link recovery of the first link 2114 is detected, the ECMP logic receives the third packet 2120 and performs an IP lookup operation 2110 returning to ECMP group 2112. In these instances, ECMP group 2112 can be a 5-tuple hash that identifies the first link 2114 (Link0) and the second link 2116 (Link1) as part of ECMP group 2112. As a result, in this case, all traffic will flow to both the first link 2114 (Link0) and the second link 2116 (Link1). The third process 2106 occurs after link recovery is detected.

[0264] Figure 22 This is a flowchart of a method 2200 for operating a DPU with fast link recovery according to at least one embodiment. The processing logic can be hardware, firmware, software, or any combination thereof. In at least one embodiment, the processing logic is implemented in a DPU, switch, network device, GPU, NIC, CPU, etc. In at least one embodiment, the processing logic is implemented in an acceleration hardware engine coupled to the switch. In at least one embodiment, method 2200 can be executed by multiple processing threads, each thread executing one or more individual functions, routines, subroutines, or operations of the method. In at least one embodiment, the processing threads implementing method 2200 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization logic). Alternatively, the processing threads implementing method 2200 can be executed asynchronously with respect to each other. Figure 22 Various operations of method 2200 are executed in different orders as shown. Some operations of these methods can be executed concurrently with other operations. In at least one embodiment, Figure 22 One or more of the operations shown are not always performed.

[0265] refer to Figure 22The processing logic begins by executing one or more instructions from the virtual switch and the NPAL (Network Access Protocol Array) supporting multiple network protocols and functions in the network pipeline (box 2202). The network pipeline includes a set of tables and logic organized in a specific order by the acceleration hardware engine. At box 2204, the processing logic monitors the link availability of each of the multiple links to the destination via the virtual switch, which are specified in the initial identifier group. At box 2206, the processing logic detects a link failure of the first link among the multiple links via the virtual switch. At box 2208, the processing logic removes the first link identifier associated with the first link from the initial link identifier group via the virtual switch to obtain a modified link identifier group. At box 2210, the processing logic causes the routing table in the NPAL to be updated via the virtual switch to remove the first link identifier. At box 2212, the processing logic uses the DPU's acceleration hardware engine to process network traffic data using the network pipeline. At box 2214, the processing logic uses the DPU's acceleration hardware engine to distribute network traffic data only to the remaining links among the multiple links corresponding to the modified identifier group.

[0266] In another embodiment, the initial link identifier group is an equal-cost multipath group (ECMP group). The processing logic removes the first link identifier from the initial group by removing the first link identifier from the ECMP. The processing logic causes the routing table in NPAL to be updated in parallel with the removal of the first link identifier from the ECMP.

[0267] In another embodiment, the processing logic receives a first packet before the first link fails. The processing logic performs an IP address lookup on the first packet to identify an ECMP group identifier associated with multiple links. The processing logic hashes the ECMP group identifier to identify any one of the multiple links. The processing logic receives a second packet after the first link fails. The processing logic performs an IP address lookup on the second packet to identify an ECMP group identifier. The processing logic hashes the ECMP group identifier to identify any one of the remaining links.

[0268] In another embodiment, the processing logic monitors the link availability of each of the multiple links to the destination. The processing logic detects the recovery of the first link. The processing logic adds the first link identifier to the modified link identifier group to obtain an initial link identifier group. The processing logic causes the routing table in the NPAL to be updated to add the first link identifier. The acceleration hardware engine can use the network pipeline to process subsequent network traffic data and distribute the subsequent network traffic data to the multiple links corresponding to the initial identifier group.

[0269] In another embodiment, the processing logic receives a first packet before the first link is restored. The processing logic performs an Internet Protocol (IP) address lookup on the first packet to identify an ECMP group identifier associated with multiple links. The processing logic hashes the ECMP group identifier to identify any of the remaining links. The processing logic receives a second packet after the first link is restored. The processing logic performs an IP address lookup on the second packet to identify an ECMP group identifier. The processing logic hashes the ECMP group identifier to identify any of the multiple links.

[0270] NPAL supporting PBR on SFC architecture

[0271] In at least one embodiment, NPAL can provide hardware-accelerated policy-based routing (PBR) on the SFC architecture of the DPU. Using SFC on the DPU, users (or controllers) can add different PBR policies, which are accelerated by the DPU hardware as a single data plane on the DPU. As described above, PBR is a technology that allows network administrators to make routing decisions based on policies set by the network administrator rather than relying on a default routing table that uses the destination IP address to determine the next hop. With traditional routing, routers decide how to forward packets based on the destination IP address and the routing table. PBR allows routers to make routing decisions based on other criteria, such as: source IP address or subnet; IP protocol type; port number; ingress interface; packet size, QoS parameters, etc. PBR is very useful for controlling the path taken by traffic through the network. It provides network administrators with greater flexibility to implement routing rules that are not solely dependent on the destination IP address, as discussed below. Figure 23 and Figure 24 As described.

[0272] Figure 23 This is a block diagram of an SFC architecture 2300 with a PBR policy 2302 according to at least one embodiment. The SFC architecture 2300 differs from the SFC architecture 800 in that it uses a PBR policy 2302 instead of a user-defined network rule 816. Figure 8 SFC architecture 800 and Figure 3The SFC architecture 2300 is similar to that of SFC architecture 300. SFC architecture 2300 includes a first virtual bridge 802 (labeled "OVS BR-1"), a second virtual bridge 804 (labeled "OVS BR-2"), a virtual port 806, and a network service 808. As described herein, SFC logic 102 can generate the first virtual bridge 802, the second virtual bridge 804, and the virtual port 806 within SFC architecture 800. SFC logic 102 can configure the first virtual bridge 802 to be controlled by network service rules 814 provided by the network service 808 hosted on DPU 204, and configure the second virtual bridge 804 to be controlled by PBR policy 816. SFC logic 102 adds a service interface to the first virtual bridge 802 to operatively couple the network service 808 to the first virtual bridge 802. SFC logic 102 adds a virtual port 806 between the first virtual bridge 802 and the second virtual bridge 804. Network service 808 can provide one or more network service rules 814 to the first virtual bridge 802. User-defined logic 126 can provide one or more PBR policies 2302 to the second virtual bridge 804. As described herein, user-defined logic 126 can be a user or a controller. SFC logic 102 can add one or more host interfaces 310 to the second virtual bridge 804.

[0273] In at least one embodiment, the first virtual bridge 802 and the second virtual bridge 804 are OVS bridges. The processing device can execute an OVS application with a hardware offloading mechanism to provide a single accelerated data plane 702 in the acceleration hardware engine for routing network traffic data using PBR policy 2302 and processing network traffic data using a set of one or more network rules.

[0274] As illustrated, three separate host interfaces can be added to connect the second virtual bridge 804 to hosts, such as three separate VMs hosted on host device 202. For example, one VM could host a firewall application, another a load balancer application, and yet another an IDS application. SFC logic 102 can add one or more network interfaces 812 to the first virtual bridge 802. Specifically, SFC logic 102 can add a first network interface to the first virtual bridge 802 to be operatively coupled to a first network port 314 (labeled PORT1) of DPU 204, and add a second network interface to be operatively coupled to a second network port 316 (labeled PORT2) of DPU 104. The first virtual bridge 802 can receive network traffic data from the first network port 314 and the second network port 316. The first virtual bridge 802 can direct network traffic data to the second virtual bridge 804 via virtual port 806. The second virtual bridge 804 can direct network traffic data to the corresponding host via host interface 810.

[0275] In at least one embodiment, the PBR policy 2302 may be provided by user 2304. User 2304 may provide the PBR policy 2302 using user-defined logic 126. User 2304 may program the second virtual bridge 804 using the PBR policy 2302. Alternatively, user-defined logic 126 may receive user input from user 2304, and user-defined logic 126 may generate PBR policies 2302 and provide them to the second virtual bridge 804. In another embodiment, the PBR policy 2302 may be provided by a user-defined service or another network service separate from network service 808. The other network service (or user-defined service) may be a user-defined network service, user-defined security service, user-defined telemetry service, user-defined storage service, etc., hosted on DPU 204 or as an application on host device 202. When the PBR policy 2302 is provided by a second network service, SFC logic 102 may add another service interface to the second virtual bridge 804 to operatively couple the second network service to the second virtual bridge 804. In another embodiment, the controller may provide PBR policy 2302 to the second virtual bridge 804.

[0276] DPU 204 can combine network rules corresponding to different network services to obtain a set of combined network rules that can be accelerated in a single accelerated data plane 702 of DPU 204. The set of combined network rules becomes the hardware acceleration rules accelerated by DPU 204 of SFC architecture 800.

[0277] In at least one embodiment, DPU 204 provides DPU services supporting Host-Based Networking (HBN) as Network Service 808 for accelerating L2 / L3 / tunneling protocols on DPU 204. The HBN infrastructure is based on an SFC topology, where one OVS bridge is controlled by the HBN service providing all accelerated network capabilities. As described above, the second OVS bridge (second virtual bridge 804) can be controlled by user 2304 or any other controller (such as...). Figure 9 (As illustrated in the diagram) Programming. The HBN service can support different protocols and network capabilities, such as ACL, ECMP, tunneling, CT, etc. User 2304 can flexibly program different routing rules in parallel with the HBN service on the SFC architecture 2300 according to one or more PBR policies. This produces hardware-accelerated rules 708 for a single accelerated data plane 702 provided by OVS-DOCA and DPU hardware. Using the SFC infrastructure, users and customers can use DPU 204 as a network accelerator on edge devices without the need for complex smart switches in different network topologies in DC and SP networks.

[0278] It should be noted that SFC logic 102 can generate different combinations of hardware acceleration rules 708 in different SFC architectures, such as those illustrated and described herein.

[0279] In at least one embodiment, the acceleration hardware engine of DPU 204 provides a single accelerated data plane 702. DPU 204 may include memory for storing configuration files that at least specify a first virtual bridge, a second virtual bridge, and a virtual port between the first and second virtual bridges. The processing device of the DPU may generate the first and second virtual bridges, the first virtual bridge being controlled by a first network service hosted on the DPU and having a set of one or more network rules; and the second virtual bridge having a PBR policy 2302. The processing device may add a PBR policy to the second virtual bridge. The processing device may add virtual ports between the first and second virtual bridges. In the single accelerated data plane 702, the acceleration hardware engine may use the PBR policy to route network traffic data and process network traffic data using a set of one or more network rules.

[0280] In at least one embodiment, the processing device may receive user input from user 2304 (or controller). The user input may specify PBR policy 2302. SFC architecture 2300 may include one or more routing rules, each routing rule including matching conditions (also referred to as one or more matching criteria) and corresponding actions. It should be noted that matching conditions may be conditions that do not include the destination address, such as those used in conventional routing. In at least one embodiment, the matching condition specifies at least one of the following: source IP address, source port or destination port, protocol identifier, Virtual LAN (VLAN) label, Differential Service Code Point (DSCP) or Type of Service (ToS) value, application or type of service, time of day, or similar conditions. The action may include at least one of the following: forwarding action, dropping action, rerouting action, mirroring action, load balancing action, rate limiting action, Quality of Service (QoS) marking action, traffic shaping action, encapsulation action, redirection action, or other similar actions.

[0281] In at least one embodiment, the processing device may add one or more host interfaces 810 to the second virtual bridge 804 according to a configuration file to operatively couple one or more host devices 202, which are operatively coupled to the DPU 204. The processing device may add one or more network interfaces 812 to the first virtual bridge 802 to operatively couple to one or more network ports 314 and 316 of the DPU 204. The processing device may add a first service interface to the first virtual bridge to operatively couple to a first network service for providing accelerated network capabilities using a set of one or more network service rules 814 (e.g., L2 protocol rules, L3 protocol rules, tunneling protocol rules, ACL rules, ECMP rules, tunneling encapsulation rules, tunneling decapsulation rules, CT rules, VLAN rules, or NAT rules). Network service rule 814 may include one or more bootstrapping rules (e.g., application-based bootstrapping rules, policy-based bootstrapping rules, geolocation-based bootstrapping rules, load balancing rules, QoS rules, failover rules, redundancy rules, security-based bootstrapping rules, cost-based routing rules, SD-WAN path bootstrapping rules, or SDN rules).

[0282] In at least one embodiment, the processing device may add one or more host interfaces 810 to the second virtual bridge 804 according to a configuration file to operatively couple one or more host devices 202, which are operatively coupled to the DPU 204. The processing device may add one or more network interfaces 812 to the first virtual bridge 802 to operatively couple one or more network ports of the DPU 204. The processing device may add a first service interface to the first virtual bridge 802 to operatively couple a first network service. The processing device may add a second service interface to the second virtual bridge 804 to operatively couple a second network service, wherein the first and second network services are part of an SFC infrastructure for providing accelerated network capabilities in a single accelerated data plane 702 using a combined set of network rules. The combined set of rules includes a set of one or more network rules associated with the first network service and a second set of one or more network rules associated with the second network service.

[0283] In at least one embodiment, an operating system (OS) can be installed and executed on the processing device of the DPU 204. Generating multiple virtual bridges and their interface mappings can be part of installing the OS on the DPU. In at least one embodiment, generating multiple virtual bridges and their interface mappings can be part of the runtime of the DPU 204, and it is not necessary to reinstall the OS on the DPU 204.

[0284] Figure 24 This is a flowchart of method 2400 for a DPU supporting PBR on an SFC architecture, according to at least one embodiment. The processing logic can be hardware, firmware, software, or any combination thereof. In at least one embodiment, the processing logic is implemented in a DPU, switch, network device, GPU, NIC, CPU, etc. In at least one embodiment, the processing logic is implemented in an acceleration hardware engine coupled to the switch. In at least one embodiment, method 2400 can be executed by multiple processing threads, each thread executing one or more individual functions, routines, subroutines, or operations of the method. In at least one embodiment, the processing threads implementing method 2400 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization logic). Alternatively, the processing threads implementing method 2400 can be executed asynchronously with respect to each other. Figure 24 The operations of method 2400 are executed in different orders as shown. Some operations of these methods can be executed concurrently with other operations. In at least one embodiment, Figure 24 One or more of the operations shown are not always performed.

[0285] refer to Figure 24The processing logic begins with a processing logic storage configuration file that specifies at least a first virtual bridge, a second virtual bridge, and a virtual port between the first and second virtual bridges (box 2402). At box 2404, the processing logic generates the first and second virtual bridges based on the configuration file. The first virtual bridge is controlled by a first network service hosted on the DPU and has a set of one or more network rules; the second virtual bridge has a PBR policy. At box 2406, processing logic 102 adds a virtual port between the first and second virtual bridges based on the configuration file. At box 2408, the processing logic utilizes an acceleration hardware engine in a single accelerated data plane to route network traffic data using the PBR policy. At box 2410, the processing logic utilizes an acceleration hardware engine in a single accelerated data plane to process network traffic data using the set of network rules.

[0286] In another embodiment, the processing logic receives user input from a user or controller, which specifies a PBR policy. The PBR policy includes one or more routing rules, each including a matching condition and a corresponding action. The processing logic adds the PBR policy to the second virtual bridge. The matching condition can be any of the examples described above. Similarly, the action can be any of the examples described above.

[0287] In another embodiment, the processing logic may add one or more host interfaces to the second virtual bridge according to a configuration file to operatively couple to one or more host devices, which are operatively coupled to the DPU. The processing logic may add one or more network interfaces to the first virtual bridge according to a configuration file to operatively couple to one or more network ports of the DPU. The processing logic may add a first service interface to the first virtual bridge according to a configuration file to operatively couple to a first network service, which provides accelerated network capabilities using a set of one or more network rules.

[0288] In another embodiment, the processing logic may add one or more host interfaces to the second virtual bridge according to a configuration file to operatively couple one or more host devices to the DPU. The processing logic may add one or more network interfaces to the first virtual bridge according to a configuration file to operatively couple one or more network ports to the DPU. The processing logic may add a first service interface to the first virtual bridge according to a configuration file to operatively couple a first network service for providing accelerated network capabilities using a set of one or more network rules (e.g., L2 protocol rules, L3 protocol rules, tunneling protocol rules, ACL rules, ECMP rules, tunneling encapsulation rules, tunneling decapsulation rules, CT rules, VLAN rules, or NAT rules). Network service rule 814 may include one or more bootstrapping rules (e.g., application-based bootstrapping rules, policy-based bootstrapping rules, geolocation-based bootstrapping rules, load balancing rules, QoS rules, failover rules, redundancy rules, security-based bootstrapping rules, cost-based routing rules, SD-WAN path bootstrapping rules, or SDN rules).

[0289] In another embodiment, the processing logic may add one or more host interfaces to the second virtual bridge according to a configuration file to operatively couple one or more host devices to the DPU. The processing logic may add one or more network interfaces to the first virtual bridge according to a configuration file to operatively couple one or more network ports to the DPU. The processing logic may add a first service interface to the first virtual bridge according to a configuration file to operatively couple a first network service. The processing logic may add a second service interface to the second virtual bridge according to a configuration file to operatively couple a second network service, wherein the first and second network services are part of an SFC infrastructure for providing accelerated network capabilities in a single accelerated data plane using a combined set of network rules. The combined set of rules includes a set of one or more network rules associated with the first network service and a second set of one or more network rules associated with the second network service.

[0290] NPAL simulation

[0291] As described herein, emulated NPAL can be used to emulate a hardware DPU with NPAL as an emulated network pipeline. An emulated network pipeline refers to a software-based framework that mimics or simulates the behavior of a hardware-based network pipeline, typically found in physical network devices such as switches, routers, NICs, and DPUs. The emulated pipeline is implemented in software, typically in an environment such as a virtual machine (VM) or containerized environment, to simulate the same packet processing, forwarding, and routing decisions as the hardware pipeline, without relying on the physical network hardware. The emulated network pipeline plays the role of traditional network hardware (such as packet switching, traffic management, firewall processing, etc.) but operates entirely as software. This is particularly useful in scenarios where physical network hardware is unavailable or when network functionality is required in test, development, or virtualization environments. In at least one embodiment, the computing system includes one or more processors and one or more memories storing instructions that, when executed by one or more processors, perform the operations of the Emulated Network Pipeline Abstraction Layer (NPAL) of the emulated data processing unit (DPU). The DPU includes emulated processing devices and an emulation acceleration hardware engine. The emulation NPAL supports multiple network protocols and functions within the emulated network pipeline. The emulated network pipeline comprises a collection and logic of tables organized in a specific order to be accelerated by the emulation acceleration hardware engine. The emulation acceleration hardware engine is used to process network traffic data using the emulated network pipeline. The following section discusses... Figure 25 The example simulated network pipeline is explained and described.

[0292] Figure 25This is a block diagram of an emulated network pipeline 2500 on an emulated DPU with emulated NPAL according to at least one embodiment. The emulated network pipeline 2500 includes a collection and logic of tables organized in a specific order for acceleration by the emulation acceleration hardware engine. The emulation acceleration hardware engine can use the emulated network pipeline 2500 to process network traffic data. As illustrated, the emulated network pipeline 2500 includes an input port 2502, a filtering network function 2504, an ingress port 2506, a first network function 2508, a bridge 2510, an SVI ACL 2512, a router 2514, a second network function 2516, an egress port 2518, and an output port 2520. The input port 2502 can receive network traffic data and provide network traffic data to the filtering network function 2504, which is operatively coupled to the input port 2502. The filtering network function 2504 can filter network traffic data. Ingress port 2506 is operatively coupled to filtering network function 2504, and ingress port 2506 is operatively coupled to first network function 2508. First network function 2508 can process network traffic data using one or more ingress ACLs. Bridge 2510 is operatively coupled to first network function 2508. Bridge 2510 can perform L2 bridging operations. One or more SVI ACLs are operatively coupled to bridge 2510 and router 2514. Router 2514 can perform L3 routing operations. Second network function 2516 is operatively coupled to router 2514. Second network function 2516 can process network traffic data using one or more egress ACLs. Egress port 2518 is operatively coupled to second network function 2516. Egress port 2518 is operatively coupled to output port 2520. Output port 2520 can output network traffic data.

[0293] With Figure 2 Part of DPU 204 Figure 12 Unlike the 1200 network pipeline, the 2500 network pipeline is implemented in the emulated DPU, such as... Figure 27 As illustrated, the simulated network pipeline 2500 can be used in conjunction with various embodiments described herein, including configurable dynamic SFC interfaces in the DPU, flexible boot rules for hardware acceleration of the SFC architecture in the DPU, NPAL in the DPU, OVS bridge, NPAL split interface, NPAL fast link recovery, PBR on SFC, etc.

[0294] In at least one embodiment, one or more processors may also emulate emulated physical ports of the emulated DPU, configured to couple to emulated branch cables that are physically coupled to a set of emulated devices. One or more processors may also emulate an emulated NPAL supporting multiple network protocols and network functions in a network pipeline for multiple logical split ports, each logical split port corresponding to one of the multiple devices. The emulated network pipeline 2500 can send network traffic data to any of the multiple logical split ports. In at least one embodiment, a first logical split port among the multiple logical split ports is configured using a first policy, and a second logical split port among the multiple logical split ports is configured using a second policy different from the first policy. An emulation acceleration hardware engine can use the emulated network pipeline 2500 to process network traffic data. In at least one embodiment, the emulated DPU includes firmware configured to map physical channels of physical ports to multiple logical split ports. The firmware can present the multiple logical split ports as a set of PFs to the emulated NPAL. The NPAL simulation can configure the simulated network pipeline 2500 to perform a first network function on the first PF among a plurality of PFs and a second network function on the second PF among a plurality of PFs, the second network function being different from the first network function.

[0295] In at least one embodiment, one or more processors may also emulate a virtual switch of the emulated DPU. The virtual switch can monitor the link availability of each of a plurality of links to the destination, which are specified in an initial identifier group. The virtual switch can detect a link failure of a first link among the plurality of links. The emulated NPAL can remove the first link identifier associated with the first link from the initial link identifier group to obtain a modified link identifier group. The emulated NPAL can cause the routing table in the emulated NPAL to be updated to remove the first link identifier. The emulation acceleration hardware engine can use the emulated network pipeline 2500 to process network traffic data and distribute the network traffic data only to the remaining links among the plurality of links corresponding to the modified identifier group. In at least one embodiment, the virtual switch may be controlled by a network service hosted on the emulated DPU. The virtual switch may send a notification of a link failure of the first link to the network service. The network service may update the routing table in the emulated DPU in response to the notification. The routing table may store configuration information associated with the initial identifier group.

[0296] In at least one embodiment, the simulation processing device can generate a first virtual bridge and a second virtual bridge. The first virtual bridge is controlled by a first network service hosted on the simulation DPU and has a set of one or more network rules; the second virtual bridge has a policy-based routing policy (PBR policy). The simulation processing device can add virtual ports between the first and second virtual bridges. The simulation acceleration hardware engine in a single accelerated data plane can route network traffic data using the PBR policy and process the network traffic data using a set of one or more network rules. In at least one embodiment, the simulation processing device can receive user input from a user or controller specifying the PBR policy. The PBR policy can include one or more routing rules, each including a matching condition and a corresponding action. The simulation processing device can add the PBR policy to the second virtual bridge.

[0297] The Simulation Network Pipeline 2500 can be used in a variety of scenarios. In at least one embodiment, the Simulation Network Pipeline 2500 can be used for network testing and simulation. Before deploying network policies, rules, or configurations to physical hardware, network engineers typically use the Simulation Network Pipeline to simulate how these changes will affect the network. This can be used to debug, test new protocols, or troubleshoot network behavior in a virtual environment. The Simulation Network Pipeline 2500 can eliminate the need to invest in physical hardware during the early stages of deployment or testing, thereby reducing costs for network labs, developers, or organizations that heavily rely on virtualization infrastructure.

[0298] Compared to hardware-based pipelines that require physical modifications, the Emulated Network Pipeline 2500 can be quickly adapted or reconfigured. This allows developers to experience different configurations, test various topologies, or simulate failures without touching the physical infrastructure. For example, small companies, startups, or development environments can use the Emulated Network Pipeline 2500 when deploying expensive switches and routers is not feasible. In cloud environments, scalability is critical. The Emulated Network Pipeline 2500 allows for dynamically scaling network behavior up or down as needed, something that physical hardware struggles to achieve due to its limited capacity.

[0299] As described above, the emulated network pipeline 2500 can be emulated in an emulation environment that includes other virtualization components, such as one or more network services, NPALs used by one or more network services, virtual bridges, SFCs, etc. Each emulation component is implemented in software to simulate behavior identical to that of a real hardware device. The emulated network pipeline 2500 performs identically to the hardware network pipeline. The following section discusses... Figure 26 The example simulation environment is explained and described.

[0300] Figure 26This is a block diagram of a simulated SFC architecture 2600 having a simulated host device 2602 and a simulated DPU 2604 according to at least one embodiment. Except that the hardware and software components are simulated hardware and simulated software, the simulated SFC architecture 2600 is similar to... Figure 3 The SFC architecture 300 is similar. In this embodiment, the physical ports of the emulated DPU 2604 are emulated as a first emulated network port 2618 and a second emulated network port 2620. As described above, according to at least one embodiment, the emulated DPU 2604 can generate and configure an emulated SFC architecture 2600 with emulated software components such as a first emulated virtual switch 2606, a second emulated virtual switch 2608, a virtual emulated port 2610, and an emulated network service 2612.

[0301] The emulated DPU 2604 may include SFC logic, which can generate a first emulated virtual switch 2606, a second emulated virtual switch 2608, and a virtual emulated port 2610 within the emulated SFC architecture 2600. The SFC logic can configure the first emulated virtual switch 2606 to be controlled by an emulated network service 2612 hosted on the emulated DPU 2604, and configure the second emulated virtual switch 2608. The SFC logic can add service interfaces to the first emulated virtual switch 2606 to operatively couple the emulated network service 2612 to the first emulated virtual switch 2606. The SFC logic can add a virtual emulated port 2610 between the first emulated virtual switch 2606 and the second emulated virtual switch 2608. The emulated network service 2612 can provide one or more emulated network service rules 2622 to the first emulated virtual switch 2606. The SFC logic can add one or more emulated host interfaces 2614 to the second emulated virtual switch 2608.

[0302] As illustrated, three separate host interfaces can be added to connect the second emulated virtual switch 2608 to hosts, such as three separate VMs hosted on host device 202. For example, one VM can host a firewall application, another VM can host a load balancer application, and yet another VM can host an IDS application. The SFC logic can add one or more emulated network interfaces 2616 to the first emulated virtual switch 2606. Specifically, the SFC logic 102 can add a first network interface to the first emulated virtual switch 2606 to operatively couple to the first emulated network port 2618 (labeled PORT1) of the emulated DPU 2604, and add a second network interface to operatively couple to the second emulated network port 2620 (labeled PORT2) of the emulated DPU 2604. The first emulated virtual switch 2606 can receive network traffic data from the first emulated network port 2618 and the second emulated network port 2620. The first emulated virtual switch 2606 can redirect network traffic data to the second emulated virtual switch 2608 via virtual port 806. The second emulated virtual switch 2608 can redirect network traffic data to the corresponding host via emulated host interface 2614.

[0303] In at least one embodiment, the emulated DPU 2604 may include user-defined logic as part of the second emulated virtual switch 2608. In at least one embodiment, the user-defined logic may be part of user-defined services hosted on the emulated DPU 2604, such as user-defined network services, user-defined security services, user-defined telemetry services, user-defined storage services, etc. The SFC logic may add another service interface to the second emulated virtual switch 2608 to operatively couple the user-defined service to the second emulated virtual switch 2608.

[0304] In at least one embodiment, the SFC logic can configure a first link state propagation between the first host interface and the virtual emulation port 2610, and a second link state propagation between the second host interface and the virtual emulation port 2610, in the second emulated virtual switch 2608. Similarly, the SFC logic can configure a third link state propagation between the third host interface and the virtual emulation port 2610 in the second emulated virtual switch 2608. A similar link state propagation can be configured in the first emulated virtual switch 2606 for the link between the virtual emulation port 2610 and the emulated network interface 2616.

[0305] In at least one embodiment, the SFC logic can configure OS features in the second emulated virtual switch 2608. In at least one embodiment, the SFC logic 102 can configure OS features for each emulated host interface in the emulated host interface 2614.

[0306] As described in this document, the emulated SFC architecture 2600 can be created either as part of installing an OS on the emulated DPU 2604 or as part of the runtime of the emulated DPU 2604. This can be accomplished using a second configuration file or modifications to the original configuration file. Reconfiguring the emulated DPU 2604 as part of the runtime can be done without reinstalling the OS on the emulated DPU 2604.

[0307] In at least one embodiment, the emulated SFC architecture 2600 is a set of instructions executed by a computing system having a processing device and memory operatively coupled to the processing device. When executed by the processing device, these instructions cause the processing device to perform a first operation on the emulated host device 2602 and a second operation on the NPAL of the emulated DPU 2604, which includes the emulated processing device and the emulation acceleration hardware engine. The emulated NPAL supports multiple network protocols and network functions in the emulated network pipeline. The emulated network pipeline includes a collection and logic of tables organized in a specific order to be accelerated by the emulation acceleration hardware engine, such as... Figure 25 As shown in the diagram. The simulation acceleration hardware engine can use the Simulated Network Pipeline 2500 to process network traffic data.

[0308] In at least one embodiment, the processing device can emulate the emulated physical virtual emulation port 2610 of the emulated DPU 2604. The emulated DPU 2604 is configured to couple to an emulated branch cable, which is physically coupled to a collection of multiple emulated devices. The emulated NPAL supports multiple network protocols and network functions in the emulated network pipeline 2500 for multiple logical split ports, each logical split port corresponding to one of the multiple devices. The emulated network pipeline can send network traffic data to any of the multiple logical split ports. A first logical split port among the multiple logical split ports can be configured using a first policy, and a second logical split port among the multiple logical split ports can be configured using a second policy different from the first policy. The emulation acceleration hardware engine can use the emulated network pipeline 2500 to process the network traffic data.

[0309] In at least one embodiment, the emulated DPU 2604 includes firmware configured to map physical channels of physical ports to multiple logically split ports. In at least one embodiment, the firmware can present the multiple logically split ports as multiple PFs to the emulated NPAL. The emulated NPAL can configure the emulated network pipeline to perform a first network function on a first PF of the multiple PFs and a second network function on a second PF of the multiple PFs, the second network function being different from the first network function.

[0310] In at least one embodiment, the processing device may emulate a virtual switch, such as a first emulated virtual switch 2606 or a second emulated virtual switch 2608 emulating the DPU 2604. The second emulated virtual switch 2608 may monitor the link availability of each of a plurality of links to the destination, which are specified in an initial identifier group. The second emulated virtual switch 2608 may detect a link failure of the first link among the plurality of links. The emulated NPAL may remove the first link identifier associated with the first link from the initial link identifier group to obtain a modified link identifier group. The emulated NPAL may cause the routing table in the emulated NPAL to be updated to remove the first link identifier. The emulated acceleration hardware engine may use the emulated network pipeline to process network traffic data and distribute the network traffic data only to the remaining links among the plurality of links corresponding to the modified identifier group. In at least one embodiment, the second emulated virtual switch 2608 is controlled by an emulated network service 2612 hosted on the emulated DPU 2604. The second emulated virtual switch 2608 can send a notification of a link failure of the first link to the emulated network service 2612, which updates the routing table in the emulated DPU 2604 in response to the notification. The routing table stores configuration information associated with the initial identifier group.

[0311] In at least one embodiment, the processing device can emulate a physical processing device. The emulated processing device can generate a first virtual bridge and a second virtual bridge. The first virtual bridge is controlled by an emulated network service 2612 hosted on the emulated DPU 2604 and has a set of one or more network rules; the second virtual bridge has a policy-based routing policy (PBR policy). The emulated processing device can add virtual ports between the first and second virtual bridges. The emulated acceleration hardware engine in a single accelerated data plane can route network traffic data using the PBR policy and process the network traffic data using a set of one or more network rules.

[0312] In at least one embodiment, the simulation processing device can receive user input from a user or controller, which specifies a PBR policy. The PBR policy includes one or more routing rules, each including a matching condition and a corresponding action. The simulation processing device can add PBR policies to the second virtual bridge.

[0313] It should be noted that SFC logic can generate different combinations of virtual bridge and interface mappings in different SFC architectures, such as those illustrated and described in this article.

[0314] Figure 27 This is a flowchart of a method 2700 for operating a simulated DPU according to at least one embodiment. The processing logic can be hardware, firmware, software, or any combination thereof. In at least one embodiment, the processing logic is implemented in a DPU, switch, network device, GPU, NIC, CPU, etc. In at least one embodiment, the processing logic is implemented in an acceleration hardware engine coupled to the switch. In at least one embodiment, method 2700 can be executed by multiple processing threads, each thread executing one or more individual functions, routines, subroutines, or operations of the method. In at least one embodiment, the processing threads implementing method 2700 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization logic). Alternatively, the processing threads implementing method 2700 can be executed asynchronously with respect to each other. Figure 27 The operations of method 2700 are executed in different orders as shown. Some operations of these methods can be executed concurrently with other operations. In at least one embodiment, Figure 27 One or more of the operations shown are not always performed.

[0315] refer to Figure 27 The processing logic begins by executing one or more instructions from a simulated NPAL that supports multiple network protocols and functions within the simulated network pipeline of a simulated DPU, which includes a simulated processing device and a simulated acceleration hardware engine. The simulated network pipeline includes a collection of tables and logic organized in a specific order to be accelerated by the simulated acceleration hardware engine. At block 2704, the processing logic receives simulated network traffic data over the network. At block 2706, the processing logic utilizes the simulated network pipeline to process the simulated network traffic data using the simulated network pipeline and the simulated DPU's simulated acceleration hardware engine.

[0316] In another embodiment, the processing logic emulates the emulated physical port of the emulated DPU, which is configured to couple to an emulated branch cable that is physically coupled to a collection of multiple emulated devices. The emulated NPAL supports multiple network protocols and network functions in the emulated network pipeline for the multiple logical split ports, each logical split port corresponding to one of the multiple devices. The emulated network pipeline can send network traffic data to any one of the multiple logical split ports. A first logical split port among the multiple logical split ports is configured using a first policy, and a second logical split port among the multiple logical split ports is configured using a second policy different from the first policy.

[0317] In another embodiment, the processing logic maps the physical channels of physical ports to multiple logically split ports, and presents the multiple logically split ports as multiple PFs to the simulated NPAL. The processing logic configures the simulated network pipeline to perform a first network function on a first PF among the multiple PFs and a second network function on a second PF among the multiple PFs, the second network function being different from the first network function.

[0318] In at least one embodiment, the processing logic simulates the virtual switch of the emulated DPU by monitoring the link availability of each of a plurality of links to the destination and detecting a link failure of a first link among the plurality of links, which are specified in an initial identifier group. The processing logic removes the first link identifier associated with the first link from the initial link identifier group to obtain a modified link identifier group. The processing logic causes the routing table in the emulated NPAL to be updated to remove the first link identifier. The emulation acceleration hardware engine can use the emulated network pipeline to process network traffic data and distribute the network traffic data only to the remaining links among the plurality of links corresponding to the modified identifier group.

[0319] In at least one embodiment, the virtual switch is controlled by an emulated network service hosted on the emulated DPU, and emulating the virtual switch includes: sending a notification of a link failure of a first link to the emulated network service, the emulated network service updating a routing table in the emulated DPU in response to the notification, and the routing table storing configuration information associated with an initial identifier group.

[0320] In at least one embodiment, the processing logic emulates the processing device by: generating a first virtual switch and a second virtual switch, the first virtual switch being controlled by an emulated network service hosted on the emulated DPU and having a set of one or more network rules, and the second virtual switch having a PBR policy; adding a virtual port between the first virtual switch and the second virtual switch; routing network traffic data using the PBR policy; and processing the network traffic data using a set of one or more network rules.

[0321] Figure 28 This is a block diagram of a computing system 2800 having two processing devices coupled to each other and multiple networks according to at least one embodiment. The computing system 2800 is designed to have multiple integrated circuits (referred to as processing devices), each including a CPU and two GPUs, forming a powerful and flexible architecture. These processing devices are interconnected via NVLink (or other high-speed interconnects) to enable high-speed communication between the processing devices; and are also connected via network interface cards (NICs) or data processing units (DPUs) to ensure efficient data transfer on the computing system 2800. Coupling the processing devices via NVLink enables seamless data exchange and parallel processing, thereby enhancing overall computing performance. Additionally, these processing devices are connected to multiple networks via one or more network interface cards (NICs) or DPUs, enabling the system to handle complex multi-network tasks with high bandwidth and low latency. This configuration makes the computing system 2800 ideal for demanding applications requiring high processing power, such as artificial intelligence (AI), machine learning (ML), and data-intensive computing, while ensuring robust connectivity and scalability across various network environments. The integrated circuits of the computing system 2800 may include one or more CPUs and one or more GPUs. Figure 28 The diagram illustrates an example architecture for a multi-GPU architecture.

[0322] like Figure 28 As illustrated, the computing system 2800 includes a processing device 2802 with a multi-GPU architecture. Specifically, the processing device 2802 includes a CPU 2806, a GPU 2808, and a GPU 2810. The CPU 2806 can be coupled to the GPU 2808 via die-to-die (D2D) or chip-to-chip (C2C) interconnects 2812 (such as a Ground Reference Signalling Interconnect (GRS interconnect)). The CPU 2806 can be coupled to the GPU 2810 via a D2D or C2C interconnect 2814. The CPU 2806 can also be coupled to the GPU 2808 and GPU 2810 via a PCIe interconnect. The CPU 2806 can be coupled to one or more network interface cards (NICs) or data processing units (DPUs), which are coupled to one or more networks. For example, as Figure 28As illustrated, CPU 2806 is coupled to a first NIC / DPU 2826, which is coupled to network 2830. CPU 2806 is also coupled to a second NIC / DPU 2828, which is coupled to network 2830. NIC / DPU 2826 and NIC / DPU 2828 can be coupled to network 2830 via Ethernet (ETH) or InfiniBand (IB) connections.

[0323] The computing system 2800 also includes a processing device 2804 with a multi-GPU architecture. Specifically, the processing device 2804 includes a CPU 2816, a GPU 2818, and a GPU 2820. The CPU 2816 can be coupled to the GPU 2818 via a D2D or C2C interconnect 2822. The CPU 2816 can be coupled to the GPU 2820 via a D2D or C2C interconnect 2824. The CPU 2816 can also be coupled to the GPU 2818 and GPU 2820 via a PCIe interconnect. The CPU 2816 can be coupled to one or more NICs or DPUs, which in turn are coupled to one or more networks. For example, as... Figure 28 As illustrated, CPU 2816 is coupled to a first NIC / DPU 2832, which is coupled to network 2836. CPU 2816 is also coupled to a second NIC / DPU 2834, which is coupled to network 2836. NIC / DPU 2832 and NIC / DPU 2834 can be coupled to network 2836 via Ethernet (ETH) or InfiniBand (IB) connections.

[0324] In at least one embodiment, processing device 2802 and processing device 2804 can communicate with each other via NIC / DPU 2838 (such as via PCIe interconnect). Processing device 2802 and processing device 2804 can also communicate with each other via high-bandwidth communication interconnect 2840 (such as NVLink interconnect or other high-speed interconnect).

[0325] Figure 28 The NIC / DPU can be the one mentioned in this article. Figures 1 to 26 Various embodiments of the described DPU.

[0326] In at least one embodiment, the computing system 2800 is used for high-speed network communication and includes a processing unit (e.g., CPU 2806, GPU 2808, GPU 2810, CPU 2816, GPU 2818, GPU 2820, NIC / DPU 2826, NIC / DPU 2828, NIC / DPU 2832, NIC / DPU 2834, or NIC / DPU 2838) and a network interface coupled to the processing unit. The network interface may include the operation and functionality of the DPU described herein.

[0327] In at least one embodiment, the computing system 2800 includes a host device and an auxiliary device. The auxiliary device includes a device memory and a processor communicatively coupled to the device memory. The auxiliary device performs the functions described herein. Figures 1 to 26 The described operation. The auxiliary device may include a GPU. The auxiliary device may include a DPU. The auxiliary device may include a DPU. The auxiliary device may include accelerator hardware.

[0328] Figure 29This is a block diagram of a computing system 2900 having a CPU 2902 and a GPU 2904 on a single integrated circuit, according to at least one embodiment. The computing system 2900 can be a highly integrated design in which the CPU 2902 and GPU 2904 are connected on a single integrated circuit, thereby utilizing NVLink C2C (chip-to-chip) interconnect 2906 to achieve fast, low-latency communication between the two processing units. This tight integration allows for efficient data transfer and parallel processing between the CPU 2902 and GPU 2904, thereby optimizing the performance of complex computing tasks. The GPU elements within the computing system 2900 can be interconnected using an NVLink network, allowing for scalability of up to 256 GPU elements, creating an ideal and powerful unified processing environment for large-scale AI, ML, and high-performance computing applications. The NVLink network can be a GPU architecture with high-bandwidth communication interconnect 2910. Additionally, the computing system 2900 can be designed to interface with high-speed I / O via PCIe interconnect 2908, thereby ensuring rapid data transfer to and from external devices, further enhancing the system's ability to handle data-intensive tasks and providing robust connectivity to peripheral components. It should be noted that since the CPU 2902 and GPU 2904 reside on the same integrated circuit, the C2C interconnect 2906 can be considered a D2D interconnect. The integrated circuit may include CPU memory (also called main memory) and GPU memory, which the CPU 2902 and GPU 2904 can access respectively via the high-speed interconnect. The computing system 2900 can combine the performance of the GPU 2904 with the versatility of the CPU 2902. The CPU 2902 can be connected in a single integrated circuit with the high-bandwidth and memory-coherent C2C interconnect 2906. The computing system 2900 can support a link-switching system.

[0329] The computing system 2900 can be used in this article regarding... Figures 1 to 26 The various embodiments described.

[0330] In at least one embodiment, the computing system 2900 is used for high-speed network communication and includes a processing unit and a network interface coupled to the processing unit. The network interface may include the operation and functions of the DPU described herein.

[0331] In at least one embodiment, the computing system 2900 includes a host device and an auxiliary device. The auxiliary device includes a device memory and a processor communicatively coupled to the device memory. The auxiliary device performs the functions described herein. Figures 1 to 26 The described operation. The auxiliary device may include a GPU. The auxiliary device may include a DPU. The auxiliary device may include a DPU. The auxiliary device may include accelerator hardware.

[0332] Figure 30 This is a block diagram of a computing system 3000 having a tensor core GPU 3008 according to at least one embodiment. The computing system 3000 may be a DGX H100 system, which is a high-performance computing platform designed to meet the needs of AL, ML, and deep learning (DL) workloads. The computing system 3000 may include multiple tensor core GPUs 3008 (e.g., NVIDIA H100 tensor core GPUs). Each tensor core GPU 3008 may be one described above. Figure 29 One of the integrated circuits described is the Tensor Core GPU 3008, which is optimized for AI / ML / DL applications, delivering superior performance for deep learning training, inference, and high-performance computing tasks. The Tensor Core GPUs 3008 within the Computing System 3000 interconnect using high-speed communication interfaces such as NVLinks, enabling rapid data transfer between them, crucial for handling large-scale AI models and datasets with low latency. The Computing System 3000 is designed for scalability, allowing for the integration of additional GPUs as needed, making it versatile enough for research, development, and deployment in data centers used for production AI workloads. Each GPU is equipped with a Tensor Core (a dedicated processing unit for accelerating matrix operations), a fundamental component of AI and deep learning algorithms. These Tensor Cores enable the system to perform mixed-precision computations efficiently, balancing speed and accuracy. Considering the power consumption and heat generation of multiple Tensor Core GPUs 3008, the Computing System 3000 may include advanced cooling solutions and power management features to ensure safe operation while maintaining peak performance. It is supported by a comprehensive software ecosystem that includes NVIDIA's CUDA programming model, AI frameworks such as TensorFlow and PyTorch, and other HPC and AI software tools that enable developers and researchers to fully leverage the capabilities of the TensorCore GPU 3008 for their specific applications. The Computing System 3000 is ideally suited for large-scale AI model training, real-time inference, scientific simulations, data analysis, and other computationally intensive tasks requiring significant parallel processing power.

[0333] The Tensor Core GPU 3008 can be coupled to multiple CPUs, such as CPU 3002 and CPU 3004, using a switch 3006 (e.g., a CX7 HCA / NIC with a PCIe switch). The Tensor Core GPU 3008 can be coupled to each other via a switch 3010 (e.g., an NVSwitch). Switches 3006 and 3010 can be coupled to a high-speed transceiver module 3012. The high-speed transceiver module 3012 can be an Octal Small Pluggable (OSFP) module. OSFP modules are high-speed transceiver modules designed for fast data communication, especially in environments requiring high bandwidth, such as data centers and high-performance computing systems. These modules support extremely high data rates, typically up to 400 Gbps per module, with future expansion capabilities to 800 Gbps or higher. OSFP modules interface with the system via a PCIe interface, enabling fast and efficient data transfer between integrated CPU-GPU components and external networks or other connected systems. Their hot-swappable nature allows for easy insertion or removal without shutting down the system, providing flexibility and maintainability crucial in critical-uptime environments. Additionally, OSFP modules are designed for high density, maximizing the number of high-speed connections within limited space, such as in densely packed server racks. By adhering to the latest networking standards, OSFP modules ensure the Compute System 3000 remains capable of meeting ever-growing data demands and can be upgraded to support future increases in network speed, thus contributing to overall system performance and scalability.

[0334] In at least one embodiment, the computing system 3000 can be viewed as a data network configuration with full-bandwidth in-server NVLink. In this example, all eight tensor core GPUs 3008 can simultaneously saturate eighteen NVLinks to other GPUs within the server. Bandwidth is limited by oversubscription from multiple other GPUs. In another embodiment, the data network configuration can be half-bandwidth in-server NVLink. In this example, all eight tensor core GPUs 3008 can half-subscribe to eighteen NVLinks to GPUs in other servers. Four tensor core GPUs 3008 can saturate eighteen NVLinks to GPUs in other servers. This is equivalent to full bandwidth on AllReduce with Scalable Hierarchical Aggregation and Reduce Protocol (SHARP). The reduction of all-to-all (All2All) bandwidth is a trade-off between server complexity and cost. In at least one embodiment, all eight Tensor Core GPU 3008s can independently transmit data using the Remote Direct Memory Access (RDMA) protocol via their own dedicated switch (e.g., a 400Gb / s HCA / NIC) in a multi-track InfiniBand / Ethernet configuration. In this example, the aggregated full-duplex bandwidth is 800Gbps for non-NVLink network devices.

[0335] The NIC / switch of the computing system 3000 may include the features described in this article. Figures 1 to 26 The various embodiments described.

[0336] In at least one embodiment, the computing system 3000 is used for high-speed network communication and includes a processing unit (e.g., CPU 3002, CPU 3004, switch 3006, Tensor Core GPU 3008, switch 3010, high-speed transceiver module 3012) and a network interface coupled to the processing unit. The network interface may include the operation and functions of the DPU described herein.

[0337] In at least one embodiment, the computing system 3000 includes a host device and an auxiliary device. The auxiliary device includes device memory and a processor communicatively coupled to the device memory. The auxiliary device performs the functions described herein. Figures 1 to 26 The described operation. The auxiliary device may include a GPU. The auxiliary device may include a DPU. The auxiliary device may include a DPU. The auxiliary device may include accelerator hardware.

[0338] Other variations are within the spirit of this disclosure. Therefore, while the disclosed technology is readily adaptable to various modifications and alternative configurations, certain embodiments thereof are illustrated in the accompanying drawings and have been described in detail above. However, it should be understood that the disclosure is not intended to be limited to one or more specific forms disclosed, but rather, it is intended to cover all modifications, alternative configurations, and equivalents falling within the spirit and scope of this disclosure as defined in the appended claims.

[0339] Unless otherwise stated or obviously contradicted by the context, the terms “a,” “an,” and “the,” and similar references, used in the context of describing the disclosed embodiments (particularly in the context of the appended claims), should be interpreted as encompassing both singular and plural forms, rather than as definitions of the terms. Unless otherwise stated, the terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (meaning “including, but not limited to”). The term “connection” (referring to a physical connection where not modified) should be interpreted as partially or wholly contained, attached to, or joined together, even with some intervention. Unless otherwise indicated herein, references to numerical ranges herein are intended only as a way of abbreviating each individual value falling within that range, and each individual value is incorporated into the specification as if it were separately described herein. Unless otherwise indicated or contradicted by the context, the use of the terms “set” (e.g., “item set”) or “subset” should be interpreted as a non-empty set comprising one or more members. Furthermore, unless otherwise indicated or contradicted by the context, the term “subset” of the corresponding set does not necessarily mean an appropriate subset of the corresponding set, but rather that the subset and the corresponding set can be equal.

[0340] Unless otherwise explicitly stated or clearly contradicted by the context, connective phrases such as “at least one of A, B, and C” or “at least one of A, B, and C” are understood in the context to generally refer to items, terms, etc., which can be A or B or C, or any non-empty subset of the set A, B, and C. For example, in an exemplary example of a set with three members, the connective phrases “at least one of A, B, and C” and “at least one of A, B, and C” refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Therefore, such connective language is generally not intended to imply that some embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise stated or contradicted by the context, the term “multiple” indicates a plural state (e.g., “multiple items” means multiple items). In at least one embodiment, multiple refers to at least two items, but more may be indicated if explicitly stated or by the context. Furthermore, unless otherwise stated or clearly understood from the context, the phrase “based on” means “at least partially based on” rather than “based on only”.

[0341] Unless otherwise indicated herein or clearly contradicted by the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations thereof and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more application programs) that is executed jointly on one or more processors via hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues). In at least one embodiment, code (e.g., executable code or source code) is stored on one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, which, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described herein. In at least one embodiment, the set of non-transitory computer-readable storage media comprises multiple non-transitory computer-readable storage media, and one or more of the individual non-transitory storage media lack all the code, but the multiple non-transitory computer-readable storage media collectively store all the code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors, for example, the instructions are stored on non-transitory computer-readable storage media, and the main central processing unit (“CPU”) executes some instructions while the graphics processing unit (“GPU”) executes other instructions. In at least one embodiment, different components of the computer system have separate processors, and the different processors execute different subsets of the instructions.

[0342] Therefore, in at least one embodiment, the computer system is configured to implement one or more services that perform the operations of the processes described herein, either individually or collectively, and such a computer system is configured with suitable hardware and / or software to enable the implementation of the operations. Furthermore, the computer system implementing at least one embodiment of this disclosure is a single device, and in another embodiment it is a distributed computer system comprising multiple devices operating in different ways, such that the distributed computer system performs the operations described herein, and that a single device does not perform all the operations.

[0343] The use of any and all examples or exemplary language (e.g., “such as”) provided herein is intended only to better illustrate embodiments of this disclosure and does not constitute a limitation on the scope of the disclosure unless otherwise required. No language in the specification should be construed as indicating that any unclaimed element is essential to the practice of the disclosure.

[0344] All references cited in this article, including publications, patent applications and patents, are incorporated herein by reference as if each reference were individually and specifically indicated to be incorporated herein by reference and the entire contents of which are described herein.

[0345] In the specification and claims, the terms “coupled” and “connected”, as well as their derivatives, may be used. It should be understood that these terms may not be intended to be synonyms with each other. Rather, in specific examples, “connected” or “coupled” may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. “Coupled” may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0346] Unless otherwise expressly stated, it will be understood that throughout this specification, terms such as “processing,” “computing,” “determining,” etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that process and / or convert data represented as physical quantities (e.g., electrons) in the registers and / or memory of the computing system into other data represented as physical quantities in the memory, registers, or other such information storage, transmission, or display devices of the computing system.

[0347] In a similar manner, the term "processor" can refer to any device or part of memory that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" can be a CPU or a GPU. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Similarly, each process can refer to multiple processes that execute instructions sequentially or intermittently, sequentially, or in parallel. The terms "system" and "method" are used interchangeably herein, provided that a system can embody one or more methods, and a method can be considered a system.

[0348] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. Obtaining, acquiring, receiving, or inputting analog and digital data can be accomplished in various ways, such as by receiving data as a parameter to a function call or a call to an application programming interface (API). In some implementations, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In another implementation, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. Reference may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be implemented by transmitting data as an input or output parameter to a function call, an API, or an inter-process communication mechanism.

[0349] While the discussion above illustrates example implementations of the described technologies, other architectures can be used to implement the described functionality and are intended to fall within the scope of this disclosure. Furthermore, although specific assignments of responsibilities have been defined above for discussion purposes, various functions and responsibilities can be assigned and divided in different ways depending on the circumstances.

[0350] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or actions described. Rather, specific features and actions are disclosed as exemplary forms for implementing the claims.

Claims

1. A data processing unit (DPU), comprising: A physical port configured to couple to a branch cable, which physically couples to a collection of multiple devices; DPU hardware, which includes processing devices and acceleration hardware engines; as well as A memory, operatively coupled to the DPU hardware, stores instructions that, when executed by the DPU hardware, provide a Network Pipeline Abstraction Layer (NPAL). The NPAL supports multiple network protocols and functions within a network pipeline for multiple logical split ports, each logical split port corresponding to one of the plurality of devices. The network pipeline includes a set of tables and logic organized in a specific order to be accelerated by the acceleration hardware engine. The network pipeline is used to send network traffic data to any of the multiple logical split ports. A first logical split port is configured using a first strategy, and a second logical split port is configured using a second strategy different from the first strategy. The acceleration hardware engine processes the network traffic data using the network pipeline.

2. The DPU of claim 1, wherein the memory is further configured to store firmware, wherein the firmware is configured to map the physical channels of the physical ports to the plurality of logical split ports, and wherein the firmware is configured to present the plurality of logical split ports as a plurality of physical function (PF) ports to the NPAL, and wherein the NPAL is configured to configure the network pipeline to perform a first network function on a first PF among the plurality of PFs and to perform a second network function on a second PF among the plurality of PFs, the second network function being different from the first network function.

3. The DPU of claim 1, wherein the NPAL includes a class or set of application programming interface (API) classes that provide a unified interface to one or more applications executed by the processing device or a host device coupled to the DPU.

4. The DPU of claim 1, wherein the memory is further configured to store firmware, wherein the firmware is configured to map physical channels of the physical ports to the plurality of logically split ports.

5. The DPU of claim 4, wherein the instructions also provide a driver and a virtual switch, wherein the firmware and the driver present the plurality of logically split ports as a plurality of physical function (PF) ports to the virtual switch and the NPAL, and wherein the NPAL is configured to manage each of the plurality of PF ports as if the PF were a separate physical port.

6. The DPU of claim 4, wherein the NPAL is configured to configure the first logical split port using the first strategy and the second logical split port using a second strategy different from the first strategy.

7. The DPU of claim 1, wherein the NPAL is configured to isolate the network between the plurality of logical split ports.

8. The DPU of claim 1, wherein the plurality of logical split ports includes more than two logical split ports.

9. A computing system, comprising: Host equipment; as well as An integrated circuit coupled to the host device and the network, wherein the integrated circuit includes: A network interconnect coupled to the network, the network interconnect including a physical port configured to couple to a branch cable, the branch cable physically coupling to a collection of multiple devices; Host interconnect, which is coupled to the host device; Accelerated hardware engine; and A memory for storing instructions that, when executed by the integrated circuit, provide a Network Pipeline Abstraction Layer (NPAL) supporting multiple network protocols and network functions for multiple logical split ports in the network pipeline, each logical split port corresponding to one of the plurality of devices. The network pipeline includes a set of tables and logic organized in a specific order to be accelerated by the acceleration hardware engine. The network pipeline is used to send network traffic data to any of the plurality of logical split ports. A first logical split port is configured using a first strategy, and a second logical split port is configured using a second strategy different from the first strategy. The acceleration hardware engine is used to process the network traffic data using the network pipeline.

10. The computing system of claim 9, wherein the integrated circuit is at least one of: a data processing unit (DPU), a network interface card (NIC), a network interface device, or a switch, wherein the DPU is an on-chip programmable data center infrastructure.

11. The computing system of claim 9, wherein the NPAL includes a class or set of application programming interface (API) classes that provide a unified interface to one or more applications executed by the computing device or a host device coupled to the integrated circuit.

12. The computing system of claim 9, wherein the memory is further configured to store firmware configured to map physical channels of the physical ports to the plurality of logically split ports.

13. The computing system of claim 12, wherein the instructions are further configured to provide a driver and a virtual switch, wherein the firmware and the driver present the plurality of logically split ports as a plurality of physical function (PF) ports to the virtual switch and the NPAL, and wherein the NPAL is configured to manage each of the plurality of PF ports as if the PF were a separate physical port.

14. The computing system of claim 12, wherein the NPAL is configured to configure the first logical split port using the first strategy and to configure the second logical split port using a second strategy different from the first strategy.

15. The computing system of claim 9, wherein the plurality of logical split ports includes more than two logical split ports.

16. The computing system of claim 9, wherein the memory is further configured to store firmware, the firmware being configured to map physical channels of the physical ports to the plurality of logically split ports, and wherein the firmware is used to present the plurality of logically split ports as a plurality of physical function (PF) ports to the NPAL, and wherein the NPAL is configured to configure the network pipeline to perform a first network function on a first PF of the plurality of PFs and to perform a second network function on a second PF of the plurality of PFs, the second network function being different from the first network function.

17. A method for operating a data processing unit (DPU), the method comprising: Execute one or more instructions of the Network Pipeline Abstraction Layer (NPAL), which supports multiple network protocols and network functions in the network pipeline for multiple logical split ports, each logical split port corresponding to one of multiple devices, wherein the network pipeline includes a set of tables and logic organized in a specific order by an acceleration hardware engine. Receive first network data from the first device via a branch cable through a physical port; The acceleration hardware engine of the DPU uses the network pipeline to process the first network data; Send the first network data to the first logical split port among the plurality of logical split ports; Receive second network data from the second device via the branch cable through the physical port; The acceleration hardware engine of the DPU uses the network pipeline to process the second network data; as well as The first network data is sent to the first logical split port among the plurality of logical split ports.

18. The method of claim 17, further comprising: Configure the first logical split port using the first strategy; as well as Configure the second logical split port using a second strategy that is different from the first strategy.

19. The method of claim 17, further comprising: The physical channels of the physical ports are mapped to the multiple logical split ports using firmware. The firmware presents the multiple logically split ports as multiple physical function (PF) ports to the virtual switch and the NPAL. as well as The NPAL is used to manage each of the multiple PFs as if the PF were a separate physical port.

20. The method of claim 17, further comprising: Configure the network pipeline to perform a first network function for a first PF among a plurality of physical function PFs; as well as The network pipeline is configured to perform a second network function on a second of the plurality of PFs, the second network function being different from the first network function.

Citation Information

Patent Citations

  • Configurable and dynamic service function chaining (SFC) interface mapping on a data processing unit (DPU)

    US20250337613A1

  • Network pipeline abstraction layer (NAPL) fast link recovery

    US20250337679A1

  • Hardware-accelerated flexible steering rules over service function chaining (SFC)

    US20250337684A1

  • Hardware-accelerated policy-based routing (PBR) over service function chaining (SFC)

    US20250337688A1

  • Network pipeline abstraction layer (NAPL) emulation

    US20250337698A1