Active load balancing for network traffic handling of workloads
By using a reinforcement learning training strategy model in a virtual router, dynamically allocating and rebalancing network traffic processing tasks, the problem of poor load balancing in the virtual router in the prior art is solved, and the overall performance of the system is improved.
Patent Information
- Application Number
- CN202410336401.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-03-22
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to achieve effective load balancing in network traffic processing for workloads in virtual routers, resulting in the processing cores that may be hungry or overloaded, affecting the overall system performance.
Adopting a reinforcement learning training strategy model, network traffic processing tasks are dynamically assigned to specific processing cores of computing devices, and actively rebalances the allocation based on real-time traffic load and other attributes.
Through dynamic allocation and rebalancing, the hunger or overload of the processing core is effectively avoided, and the overall utilization of the system and network processing performance are improved.
Smart Images

Figure CN120200976A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Application No. 18 / 394,526, filed on Dec. 22, 2023, the entire content of which is incorporated herein by reference. Technical Field
[0003] The present disclosure relates to virtualized computing infrastructure, and more particularly, to load - balancing network traffic processing of virtual routers. Background Art
[0004] Virtualized data centers are becoming the core foundation of modern information technology (IT) infrastructure. In particular, modern data centers have widely utilized virtualized environments to deploy and execute workloads such as virtual hosts or containers on the underlying computing platforms of physical computing devices. A virtual router creates a virtual overlay network (“virtual network”) over the physical underlying network and uses the virtual network to process and forward data traffic between workloads. Summary of the Invention
[0005] Generally, techniques are described herein for actively load - balancing between processing cores of a virtual router in a computing device for processing network traffic associated with different workloads executed on the computing device. The virtual router is responsible for processing network traffic generated by and destined for the workloads; such processing may include routing and forwarding, encapsulation and decapsulation for overlay tunnel network traffic, and other network traffic processing. As described herein, the virtual router assigns (or instructs the assignment of) network traffic processing tasks for a given workload to a particular processing core of the computing device according to a policy model trained using reinforcement learning. The reinforcement learning algorithm can update the policy model based on the predicted traffic load for the workload and other attributes that affect the overall state of the computing device.
[0006] The virtual router processes network traffic for the workload using the assigned processing core. In some examples, the virtual router can re - balance the assignment based on new (or deleted) workloads, changes in the amount of network processing required for the corresponding workloads, or other factors.
[0007] These techniques can provide one or more technical advantages and enable one or more practical applications. For example, compared with a static round-robin allocation scheme, a virtual router can allocate network traffic processing for workloads according to a policy model trained by reinforcement learning, thereby more effectively balancing the traffic processing load for workloads among the processing cores of a computing device. Active rebalancing can reduce the phenomena of processing core starvation (too low network processing load) or overload (too high network processing load), improve overall utilization, and / or reduce network processing bottlenecks and the corresponding latency that may occur due to over-allocation of processing cores. Compared with conventional supervised / unsupervised machine learning, the techniques described herein do not determine the allocation of workloads to processing cores only based on real-time metrics, but update the policy model by applying reinforcement learning, which can be based on the predicted traffic load for the workload and only requires a minimal amount of data. In this way, the techniques described herein can allocate workloads to the processing cores of a virtual router in a more efficient manner while consuming fewer computing resources (e.g., memory, processing for constant metric monitoring, etc.).
[0008] In one example, a computing system includes processing circuitry capable of accessing a storage device. The processing circuitry is configured to apply, by a reinforcement learning agent, a policy model to a predicted network traffic load associated with a workload to allocate the workload to a first processing core among a plurality of processing cores of the computing device. The processing circuitry is configured to use the first processing core to process network traffic for the workload through a virtual router and based on the allocation of the workload to the first processing core.
[0009] In one example, the method includes applying, by a reinforcement learning agent, a policy model to a predicted network traffic load associated with a workload to allocate the workload to a first processing core among a plurality of processing cores of a computing device. The method further includes using the first processing core to process network traffic for the workload through a virtual router and based on the allocation of the workload to the first processing core.
[0010] In another example, a computer-readable storage medium includes instructions that, when executed, cause processing circuitry to apply, by a reinforcement learning agent, a policy model to a predicted network traffic load associated with a workload to allocate the workload to a first processing core among a plurality of processing cores of a computing device. The instructions further cause the processing circuitry to use the first processing core to process network traffic for the workload through a virtual router and based on the allocation of the workload to the first processing core.
[0011] Details of one or more examples of the present disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A block diagram of an example computing infrastructure that can implement examples of the techniques described herein.
[0013] Figure 2 A block diagram of an example computing device that can implement examples of the techniques described herein.
[0014] Figure 3 A block diagram of example components of an example computing device that executes a virtual router for a virtual network according to the techniques described herein.
[0015] Figure 4 A conceptual diagram of example operations for updating a policy model for generating predicted allocations according to the techniques described herein.
[0016] Figure 5 A block diagram of an example computing device that executes a workload and example components of a virtual router for a virtual network according to the techniques described herein.
[0017] Figure 6 A conceptual diagram of example data for generating predicted allocations according to the techniques described herein.
[0018] Figure 7 A flowchart of example operations of a method according to the techniques described herein.
[0019] Throughout the description and the figures, like reference numerals represent like elements. Detailed Description
[0020] Figure 1 A block diagram of an example computing infrastructure 8 in which examples of the techniques described herein can be implemented. Generally, a data center 10 is used to provide an operating environment for applications and services at a customer site 11 (illustrated as "Customer 11"), the customer site having one or more customer networks coupled to the data center via a service provider network 7. For example, the data center 10 can host infrastructure devices (such as, network and storage systems, redundant power environmental controls). The service provider network 7 is coupled to a public network 15, which can represent one or more networks managed by other providers and can thus form part of a large-scale public network infrastructure (such as, the Internet). For example, the public network 15 can represent a local area network (LAN), a wide area network (WAN), the Internet, a virtual LAN (VLAN), an enterprise LAN, a layer 3 virtual private network (VPN), an Internet protocol (IP) intranet operated by the service provider operating the service provider network 7, an enterprise IP network, or some combination thereof.
[0021] Although customer site 11 and public network 15 are mainly illustrated and described as edge networks of service provider network 7, in some examples, one or more customer sites 11 and public network 15 can be tenant networks within data center 10 or another data center. For example, data center 10 can host multiple tenants (customers), each tenant being associated with one or more virtual private networks (VPNs), and each virtual private network can implement a customer site 11.
[0022] Service provider network 7 provides packet-based connectivity to the connected customer sites 11, data center 10, and public network 15. Service provider network 7 can represent a network owned and operated by a service provider to interconnect multiple networks. Service provider network 7 can implement Multiprotocol Label Switching (MPLS) forwarding and, in such cases, can be referred to as an MPLS network or MPLS backbone. In some examples, service provider network 7 represents multiple interconnected autonomous systems, such as the Internet, which provides services from one or more service providers.
[0023] In some examples, data center 10 can represent one of many geographically distributed network data centers. As Figure 1 shown in the example, data center 10 can be a facility that provides network services to customers. Customers of the service provider can be collective entities, such as enterprises and governments or individuals. For example, a network data center can host network services for several enterprises and end users. Other example services can include data storage, virtual private networks, traffic engineering, file services, data mining, scientific or supercomputing, etc. Although shown as an independent edge network of service provider network 7 in the figure, elements of data center 10 (e.g., one or more physical network functions (PNFs) or virtualized network functions (VNFs)) can be included within the core of service provider network 7.
[0024] In this example, data center 10 includes storage and / or computing servers (or "nodes") having servers 12A to 12X (referred to herein as "servers 12") depicted as being coupled to top-of-rack (TOR) switches 16A to 16N, which are interconnected via a switching fabric 14 provided by one or more layers of physical network switches and routers. Servers 12 are computing devices and can also be referred to herein as "hosts" or "host devices" because servers 12 host workloads 35 for execution by servers 12. Although only server 12A coupled to TOR switch 16A is shown in detail in Figure 1 data center 10 can include many additional servers coupled to other TOR switches 16 of data center 10.
[0025] The switching fabric 14 in the illustrated example includes interconnected top-of-rack (TOR) (or other "leaf") switches 16A through 16N (collectively referred to as "TOR switches 16"), which are coupled to the distribution layer of the rack-mounted (or "spine" or "core") switches 18A through 18M (collectively referred to as "rack switches 18"). Although not shown, data center 10 may also include, for example, one or more non-edge switches, routers, hubs, gateways, security devices (such as firewalls, intrusion detection and / or prevention devices), servers, computer terminals, laptops, printers, databases, wireless mobile devices (such as cellular phones or personal digital assistants), wireless access points, bridges, cable modems, application accelerators, or other network devices. Data center 10 may also include one or more physical network functions (PNFs) (such as physical firewalls, load balancers, routers, route reflectors, broadband network gateways (BNGs), mobile core network elements, and other PNFs).
[0026] In this example, the TOR switches 16 and the rack switches 18 provide redundant (multi-homed) connections for the servers 12 to the IP fabric 20 and the service provider network 7. The rack switches 18 aggregate traffic flows and provide connections between the TOR switches 16. The TOR switches 16 can be network devices that provide Layer 2 (MAC) and / or Layer 3 (such as IP) routing and / or switching functions. The TOR switches 16 and the rack switches 18 can each include one or more processors and memories and can execute one or more software processes. The rack switches 18 are coupled to the IP fabric 20, which can perform Layer 3 routing to route network traffic between the data center 10 and the customer sites 11 through the service provider network 7. The switching architecture of the data center 10 is merely an example. For example, other switching architectures can have more or fewer switching layers. The IP fabric 20 can include one or more gateway routers.
[0027] The term "packet flow", "traffic flow", or simply "flow" refers to a set of packets that originate from a specific source device or endpoint and are sent to a specific destination device or endpoint. For example, a single packet flow can be identified by a 5-tuple: <source network address, destination network address, source port, destination port, protocol>. This 5-tuple typically identifies the packet flow corresponding to the received packet. An n-tuple refers to any n items extracted from the 5-tuple. For example, a 2-tuple for a packet can refer to a combination of <source network address, destination network address> or <source network address, source port> for the packet.
[0028] Servers 12 may each represent a computing server or a storage server. For example, each server 12 may represent a computing device (such as an x86 processor-based server) configured to operate in accordance with the techniques described herein. The servers 12 may provide a network function virtualization infrastructure (NFVI) for an NFV architecture.
[0029] The servers 12 may include one or more processing cores. Server 12A includes processing core 30A, and server 12X includes processing core 30X (not shown), and so on. Each core of the processing cores 30 is an independent execution unit (“core”) to execute instructions that conform to an instruction set architecture, i.e., instructions stored in a storage medium. The execution units may be implemented as separate integrated circuits (ICs), or may be combined within one or more multi-core processors (or “many-core” processors), and each multi-core processor may be implemented using a single IC (i.e., a chip multi-processor).
[0030] Any of the servers 12 may be configured with a workload 35 by virtualizing the resources of the server to provide isolation between one or more processes (applications) executed on the server. “Hypervisor-based” or “hardware-level” or “platform” virtualization refers to creating virtual machines, each virtual machine including a guest operating system for executing one or more processes. Generally, the virtual machine provides a virtualized / guest operating system for executing applications in an isolated virtual environment. Since the virtual machine is virtualized from the physical hardware of the host server, the executed applications are isolated from both the host and the hardware of other virtual machines. Each virtual machine may be configured with one or more virtual network interfaces for communicating on a corresponding virtual network.
[0031] A virtual network is a logical structure implemented on top of a physical network. Virtual networks may be used to replace VLAN-based isolation and provide multi-tenancy in a virtualized data center (such as one of the data centers 10). Each tenant or application may have one or more virtual networks. Each virtual network may be isolated from all other virtual networks unless a security policy explicitly permits it.
[0032] The virtual network may use a data center 10 gateway router ( Figure 1 not shown in the figure) to connect to and extend across a physical multi-protocol label switching (MPLS) Layer 3 virtual private network (L3VPN) and an Ethernet virtual private network (EVPN) network. Virtual networks may also be used to implement network function virtualization (NFV) and service chaining.
[0033] Virtual networks can be implemented using a variety of mechanisms. For example, each virtual network can be implemented as a Virtual Local Area Network (VLAN), a Virtual Private Network (VPN), etc. Virtual networks can also be implemented using two networks - a physical underlying network consisting of an IP fabric 20 and a switching fabric 14, and a virtual overlay network. The role of the physical underlying network is to provide an "IP fabric" that provides unicast IP connectivity from any physical device (server, storage device, router, or switch) to any other physical device. The underlying network can provide a unified, low-latency, non-blocking, high-bandwidth connection from any point in the network to any other point in the network.
[0034] Virtual routers (such as virtual router 21) running in server 12 create a virtual overlay network on top of the physical underlying network using a dynamic "tunnel" mesh among themselves. For example, these overlay tunnels can be MPLS tunnels over GRE / UDP, or VXLAN tunnels, or NVGRE tunnels. The underlying physical routers and switches may not store any per-tenant state (such as any Media Access Control (MAC) address, IP address, or policy) for virtual machines or other workloads. For example, the forwarding tables of the underlying physical routers and switches may only contain the IP prefixes or MAC addresses of physical servers 12. The gateway routers or switches that connect the virtual network to the physical network are an exception and may contain tenant MAC or IP addresses.
[0035] The virtual routers 21 of server 12 typically contain per-tenant state. For example, they may contain separate forwarding tables (routing instances) for each virtual network. This forwarding table contains the IP prefixes (in the case of layer 3 overlay) or MAC addresses (in the case of layer 2 overlay) of virtual machines or other workloads (such as pods of containers). No single virtual router 21 needs to contain all the IP prefixes or all the MAC addresses of all the virtual machines in the entire data center. A given virtual router 21 only needs to contain those routing instances that exist on the local server 12 (i.e., there is at least one workload on server 12).
[0036] Server 12 hosts virtual network endpoints for one or more virtual networks that operate over a physical network, which in this document is represented by the IP fabric 20 and the switching fabric 14. Although described primarily for data center-based switching networks, other physical networks such as service provider network 7 can also serve as the basis for one or more virtual networks.
[0037] Each server 12 can host one or more workloads 35, and each workload 35 has at least one virtual network endpoint for one or more virtual networks configured in a physical network. The virtual network endpoints for a virtual network can represent one or more workloads of virtual network interfaces sharing the virtual network. For example, a virtual network endpoint can be a virtual machine, one or more sets of containers (e.g., deployed using pods), or another type of workload (such as a Layer 3 endpoint for a virtual network). The term "workload" includes virtual machines, containers, and other virtualized computing resources, as well as native processes that provide at least partially independent execution environments for applications. The term "workload" can also include pods of one or more containers. As Figure 1 shown, server 12A hosts one or more virtual network endpoints as workloads 35 in the form of pods 38A through 38N (collectively referred to as "pods 38") and virtual machines 36A through 36N (collectively referred to as "virtual machines 36"). Pods 38 can be Kubernetes pods or another container deployment structure. Server 12 can host only pods, VMs, different numbers of pods and VMs, and / or can host other types of workloads.
[0038] Due to the hardware resource limitations of a given server 12A, server 12A can execute as many workloads 35 as possible. Each virtual network endpoint can use one or more virtual network interfaces to perform packet I / O, or otherwise process packets. For example, a virtual network endpoint can use one virtual hardware component (such as an SR-IOV virtual function) enabled by NIC 13A to perform packet I / O and receive / send packets on one or more communication links with TOR switch 16A. Other examples of virtual network interfaces are described below.
[0039] Each of the servers 12 includes at least one network interface card (NIC) 13, and each NIC includes at least one interface to exchange data packets with a TOR switch 16 via a communication link. For example, server 12A includes NIC 13A. Any NIC 13 can provide one or more virtual hardware components 21 for virtualized input / output (I / O). The virtual hardware components for I / O can be virtualizations of physical NICs (“physical functions”). For example, in single root I / O virtualization (SR-IOV) described in the Peripheral Component Interconnect Special Interest Group SR-IOV specification, the PCIe physical functions of a network interface card (or “network adapter”) are virtualized to represent one or more virtual network interfaces as “virtual functions” for use by individual endpoints executing on the server 12. In this way, virtual network endpoints can share the same PCIe physical hardware resources, and this virtual function is an example of a virtual hardware component 21. As another example, one or more of the servers 12 can implement Virtio, a para-virtualization framework that can be used, for example, with the Linux operating system, to provide emulated NIC functionality as a virtual hardware component to provide a virtual network interface to a virtual network endpoint. As another example, one or more of the servers 12 can implement Open vSwitch to perform distributed virtual multi-layer switching between one or more virtual network interface cards (vNICs) for hosted virtual machines, where the vNICs can also represent a virtual hardware component that provides a virtual network interface to a virtual network endpoint here. In some examples, the virtual hardware component is a virtual I / O (e.g., NIC) component. In some examples, the virtual hardware component is an SR-IOV virtual function. In some examples, any of the servers 12 can implement a Linux bridge, which emulates a hardware bridge and forwards data packets between the virtual network interfaces of the server, or between the virtual network interface of the server and the physical network interface of the server. For a Docker implementation of containers hosted by a server, a Linux bridge or other operating system bridge that exchanges data packets between containers executing on the server can be referred to as a “Docker bridge”. The term “virtual router” as used herein can include a Contrail or Tungsten Fabric virtual router, Open vSwitch (OVS), an OVS bridge, a Linux bridge, a Docker bridge, or other devices and / or software located on a host device and performing packet exchange, bridging, or routing between virtual network endpoints of one or more virtual networks, where the virtual network endpoints are hosted by one or more of the servers 12.
[0040] Each of the servers 12 includes one or more processing cores. The workload 35 of the server 12A is executed by the processing core 30A according to the allocation, which is generally referred to as "pinning" because the underlying processing of a given workload 35 is "pinned" to a specific core of the processing core 30A to execute the software instructions for that workload. For example, the VM 36A can be pinned to the first core of the processing core 30A, which executes the VM 36A and any applications therein. The pod 38N can be pinned to the second core of the processing core 30A, which executes the pod 38N and any containerized applications therein. In the present disclosure, the processing core assigned to execute the workload can be different from the processing core assigned to process the network traffic for that workload through the virtual router.
[0041] One or more servers 12 can each include a separate virtual router 21 ("vRouter 21"), and the virtual router 21 executes one or more routing instances for the corresponding virtual network within the data center 10 to provide a virtual network interface and route data packets between virtual network endpoints. Each routing instance can be associated with a network forwarding table. Each routing instance can represent a virtual routing and forwarding instance (VRF) for an Internet Protocol virtual private network (IP-VPN). For example, the data packet received by the virtual router 21 of the server 12A from the underlying physical network structure of the data center 10 (i.e., the IP structure 20 and the switching structure 14) can include an external header to allow the physical network structure to tunnel the payload or "internal data packet" to the physical network address of the network interface card 13A of the server 12A where the virtual router is executed. This external header can include not only the physical network address of the network interface card 13A of the server, but also a virtual network identifier, such as a VxLAN tag or a Multiprotocol Label Switching (MPLS) tag identifying one of the virtual networks and the corresponding routing instance executed by the virtual router 21. The internal data packet includes an internal header with a destination network address that conforms to the virtual network addressing space for the virtual network identified by the virtual network identifier.
[0042] The virtual router 21 terminates the virtual network overlay tunnel, determines the virtual network of the received packet based on the packet-oriented tunnel encapsulation header, and forwards the packet to the appropriate destination virtual network endpoint. For example, for server 12A, for each packet outbound from a virtual network endpoint hosted by server 12A (e.g., VM 36A or pod 38A), the virtual router 21A attaches a tunnel encapsulation header that indicates the virtual network of the packet to generate an encapsulated packet or "tunnel" packet, and the virtual router 21A outputs the encapsulated packet to a physical destination computing device (e.g., another device in server 12) via the overlay tunnel of the virtual network. As used herein, the virtual router 21 can perform the operations of a tunnel endpoint to encapsulate internal packets initiated by a virtual network endpoint to generate tunnel packets, and to decapsulate tunnel packets to obtain internal packets for routing to other virtual network endpoints.
[0043] Each server 12 provides an operating environment for executing one or more application workloads 35. As used herein, the terms "application workload" or "workload" can be used interchangeably to refer to an application workload. The workload 35 can be deployed using a virtualized environment, such as virtual machines 36A through 36N (collectively "virtual machines 36"), pods 38A through 38N (collectively "pods 38"), or other types of virtualized computing instances, or in some examples deployed on a bare metal server that directly rather than indirectly executes the workload 35 in a virtualized environment. Some or all of the servers 12 can be bare metal servers (BMS). A BMS can be a physical server dedicated to a particular customer or tenant.
[0044] Each workload 35 is executed by one of the processing cores 30A. That is, the processing core 30A executes the corresponding software instructions for the workload 35, and the workload 35 uses a virtual network that is partially managed by the controller 24 and the virtual router 21 to send and receive traffic.
[0045] The processing core 30A also executes the virtual router 21A. The virtual router 21 uses the processing core 30A to service the networking requirements of each workload, such as to process (e.g., by routing, forwarding, or applying services) network traffic sent by or destined for a workload and thus associated with the workload. Different workloads 35 can be associated with different amounts of network traffic, which also vary from time to time for each workload.
[0046] The virtual router 21A can load balance between processing cores 30A that the virtual router 21A uses to process network traffic associated with different workloads 35. The virtual router 21 can use one or more processing cores 30A to service the network capabilities of a workload by processing network traffic associated with the workload. The virtual router typically assigns processing cores to process network traffic for different workloads in a round-robin fashion, and the assigned processing cores continue to process network traffic for the workload until the workload is deleted. In other words, the virtual router statically assigns cores of the processing cores 30A to process network traffic for each workload. Conventional virtual routers can assign network traffic processing for network traffic associated with a workload to a processing core that is overloaded in terms of the existing load that has already been assigned. In addition, some processing cores may be "starved" due to having minimal load and would be better equipped to process the network capabilities for a workload.
[0047] According to the techniques described herein, the controller 24 can provide intelligent dynamic balancing for one or more processing cores that process network traffic for a workload. The controller 24 can generate an assignment for the virtual router 21A to selectively assign different processing cores 30A to process network traffic for each workload 35. In some cases, the controller 24 can reactively generate an assignment based on real-time metrics indicative of network traffic processing associated with each workload 35. However, generating an assignment by the controller 24 based on real-time metrics may require continuous use of computing and network resources. Generating an assignment by the controller 24 based only on real-time metrics may result in an assignment being generated only after any of the processing cores 30A are overworked or starved. The techniques described herein can include one or more reinforcement learning agents (e.g., reinforcement learning agent 156) that apply a policy model (e.g., policy model 154) to generate an assignment or network traffic processing mapping for a workload 35 to a processing core 30A. The techniques described herein can generate a workload assignment based on a reinforcement learning agent applying the policy model to a predicted (or "anticipated") network traffic load to be processed by the processing core 30A. The controller 24 can determine the predicted network traffic load based on attributes and / or historical data associated with the workload 35 and / or the processing core 30A, and then process it by a reinforcement learning agent that applies the policy model. In this way, the controller 24 can proactively optimize network traffic processing for the workload 35.
[0048] In Figure 1In the example, the reinforcement learning agent 156 of the controller 24 (also referred to herein as the "RL agent 156") includes a policy model 154 and an RL algorithm 155. The RL agent 156 can apply the policy model 154 to predict network traffic load to predict the network traffic load assigned to one of the processing cores 130A. For example, the RL agent 156 can input the predicted network traffic load and other attributes related to the workload 35 and the processing core 30A into a machine learning model (e.g., a neural network) of the policy model 154 to determine which of the processing cores 30A is assigned to handle the network traffic for the workload. The RL agent 156 can perform operations by generating assignments associated with selecting a processing core assignment for a newly instantiated workload, reassigning a processing core to handle the network traffic for an existing workload, and / or rebalancing the assignment of processing cores to workload network traffic. The RL agent 156 can include the generated assignments in the assignment data 162.
[0049] The policy model 154 of the RL agent 156 can specify the operations or instructions that the RL agent 156 can perform to generate assignments based on the version of the machine learning model included in the policy model 154. The RL agent 156 can update the version of the machine learning model of the policy model 154 by providing a reward signal to the reinforcement learning algorithm 155. The RL agent 156 can determine the reward signal based on the state data 160 associated with network traffic processing performance (e.g., a set of values received by the RL agent 156), some of which may result from the assignments previously generated by the RL agent 156. The RL agent 156 can determine the reward signal by providing the state data 160 to a reward function. The RL agent 156 can determine the reward signal using a reward function configured to maximize a long-term reward signal (e.g., the reward signal over a period of time) (also referred to herein as the "reward value") based on criteria related to optimizing network traffic processing for the workload 35.
[0050] The policy model 154 may include one or more machine learning models that can be used to generate assignments or map the workload 35 to the processing cores 30A. The one or more machine learning models can be trained by simulation (e.g., Monte-Carlo based methods). The one or more machine learning models can include one or more classification algorithms (e.g., multi-layer perceptron, deep neural network, etc.). The one or more machine learning models can be trained using static data, such as historical assignment data, historical throughput data, usage and / or type of the processing cores 30A, requirements and / or type of the workload 35, profile of the workload 35, profile of the processing cores 30A, etc. The type of machine learning models included in the policy model 154 can be based on the type and / or quantity of available training data (e.g., historical data associated with the processing cores 30A and / or the workload 35).
[0051] The controller 24 can predict the network traffic load and provide it to the policy model 154 by the RL agent 156 to generate an assignment. The controller 24 can use machine learning algorithms to predict the network traffic load, and these algorithms can include time series prediction methods, which can include but are not limited to autoregressive integrated moving average (ARIMA), triangular seasonal box-cox transformed ARIMA error trend seasonal component (TBATS), regression methods (Bayesian ridge regression, random forest regressor, neural networks (long short-term memory (LSTM), recurrent neural network (RNN), etc.).
[0052] The RL agent 156 can apply an initial version of the policy model 154 to predict the network traffic load in order to allocate the network traffic of the workload 35 to the processing cores 30A. The RL agent 156 can provide the predicted network traffic load and other attributes (e.g., profile and status of the processing cores 30A, current assignment of the processing cores 30A, current utilization of the processing cores 30A, throughput and queue number of each core in the processing cores 30A, current throughput of the workload 35, priority of the workload 35, profile of the workload 35, etc.) to an initial version of one or more machine learning models of the policy model 154 in order to allocate the network traffic processing of the workload 35 to the processing cores 30A. The RL agent 156 can monitor the generated assignment results by collecting the state values 160. The RL agent 156 can update the machine learning models of the policy model 154 based on the state values 160.
[0053] The RL agent 156 can determine a reward signal for a machine learning model used to update the policy model 154. The RL agent 156 can determine the reward signal based on state data 160, which can include a set of values regularly obtained related to the network traffic of the processing core 30A and the workload 35. The state data 160 represents the state of the server 12A environment and is available for use by the reinforcement learning agent. The state data 160 can include attribute values related to the processing core 30A at one or more specific time points. For example, the state data 160 can include values related to the utilization rate of each processing core of the processing core 30A at a specific time, the throughput of each processing core of the processing core 30A at a specific time, and / or the number of queues of each processing core of the processing core 30A. In some cases, the state data 160 can include values related to the workload 35 at a specific time point. For example, the state data 160 can include values such as tail packet loss associated with the network traffic of the workload 35 over a period of time, jitter associated with the network traffic of the workload 35 over a period of time, and / or packet delay over a period of time. In some examples, the state data 160 can include values related to the processing core 30A and values related to the workload 35 at any one or more time periods. The virtual router 21A generates the state data 160, and the controller 24 can receive the state data 160 from the virtual router 21A through an application programming interface (e.g., Prometheus), a telemetry interface, or other interfaces.
[0054] The RL agent 156 can determine the reward signal by inputting the state data 160 into a reward function configured according to a maximization criterion of a reward value (e.g., a long-term reward signal). For example, the RL agent 156 can configure the reward function according to the following criteria: maximizing the utilization rate of the overall processing core 30A, minimizing the packet delay related to the average time for the processing core 30A to process the network traffic for the workload 35, minimizing the number of idle processing cores of the processing core 30A, minimizing the number of packet losses or tail packet losses related to the network traffic processing of the workload 35, preferentially processing the network traffic of certain workloads for the workload 35, or a certain combination of the above.
[0055] The RL agent 156 can implement a reward function to determine a reward signal for updating the policy model 154. For example, the RL agent 156 can implement a reward function to reward or punish the actions taken by the RL agent 156 in applying the policy model 154. The RL agent 156 can encourage or discourage the actions taken by the RL agent 156 based on whether the allocation generated by the RL agent 156 causes the state data 160 to meet or be within a certain threshold range of meeting a predefined criterion. When the allocation generated by the RL agent 156 causes the state data 160 to exceed the allowed value defined in the criterion, the RL agent 156 can execute the reward function to punish or discourage the actions taken by the RL agent 156 in applying the policy model 154. In some examples, software executed on the server 12A (e.g., the virtual router 21A) can calculate a reward signal based on the reward function and the state data 160. In this example, the server 12A can send the reward signal (e.g., the output of the reward function provided together with the state data 160) to the RL agent 156.
[0056] The RL algorithm 155 of the RL agent 156 can update the machine learning model of the policy model 154. The RL algorithm 155 can update the policy model 154 based on the actions taken by the RL agent 156 in applying the policy model 154, one or more reward signals determined as a result of the actions taken by the RL agent 156, or observations (e.g., state data 160) perceived by the RL agent 156 as a result of the actions taken by the RL agent 156. The RL algorithm 155 can update the policy model 154 by adjusting the parameters of the machine learning model of the policy model 154 based on the reward signal determined by the RL agent 156. For example, in response to a positive reward signal, the RL algorithm 155 can update the policy model 154 by strengthening the parameters of the machine learning model of the policy model 154 that are related to generating a network traffic allocation of the workload 35 to the processing core 30A. In response to a negative reward signal, the RL algorithm 155 can update the policy model 154 by changing the parameters of the machine learning model of the policy model 154 that are related to generating the allocation. The RL algorithm 155 can include any type of reinforcement learning algorithm, such as a policy gradient method. The RL agent 156 can continuously or periodically generate an improved network traffic allocation for the processing core 30A of the workload 35 by applying the updated policy model.
[0057] The RL agent 156 may include the generated assignments in the assignment data 162. The RL agent 156 may generate the assignment data 162 to map or queue the network traffic of the workload 35 to each processing core 30A. In some cases, the controller 24 may include another module to generate the assignment data 162 based on the assignments generated by the RL agent 156. The controller 24 may include a module that includes the assignments generated by the RL agent 156 in the assignment data 162 based on a determination of whether the assignments generated by the RL agent 156 would result in a significant deviation from predefined criteria or cause network performance issues.
[0058] To instruct the virtual router 21A to assign the generated network processing for the workload 35 to the individual processing cores 30A, the controller 24 may send the assignment data 162 to the virtual router 21A. The assignment data 162 may include mapping the workload to a specific one of the processing cores 30A to instruct the virtual router 21A to assign the network traffic processing of the workload to the processing core. A workload identifier, name, or other identifier may be used in the assignment data 162 to identify the workload. The processing core may be identified in the assignment data 162 using a processor ID, core ID, a combination thereof, or other identifier that uniquely identifies the processing core 30A. The assignment data 162 may include mappings for multiple workloads. The assignment data 162 may include predicted assignments generated by the policy model of the controller 24. In response to the policy model of the controller 24 being updated via reinforcement learning, the controller 24 may update the assignment data 162 with the predicted assignments generated by the updated policy model.
[0059] The controller 24 may receive the status data 160, implement a reward function to update the policy model, determine whether the network traffic processing for the workload should be assigned to a different processing core, and instruct the virtual router 21A to assign or reassign the network traffic processing for the workload at regular intervals (e.g., every ten minutes).
[0060] The virtual router 21A may assign or reassign the workload 35 according to the assignment data 162 from the controller 24. The virtual router 21A then services each workload 35 (i.e., processes the network traffic of each workload) using the correspondingly assigned processing core 30A.
[0061] To utilize the one processing core 30A allocated to process network traffic for a specific workload from workload 35, virtual router 21A may use parameters such as: the name and identifier for the virtual router interface; one or more software queues associated with the network traffic of the workload; one or more hardware queues allocated to one or more corresponding forwarding cores of virtual router 21A associated with the allocated processing core; and the identifier for each of the multiple processing cores. For example, based on the allocation as described above, the network traffic of a workload (e.g., pod 38A or VM 36A) may be enqueued with the software queue into the hardware queue mapped to a specific processing core. The thread of virtual router 21A executing on the processing core (e.g., virtual router 21A allocates a forwarding core to process the corresponding hardware queue) services the software queue of the workload to process the network traffic for the workload. Other threads of virtual router 21A may operate similarly for other processing cores to service the corresponding hardware queues of these processing cores.
[0062] The techniques described herein may provide one or more technical advantages and enable one or more practical applications. For example, compared with a static round-robin allocation scheme, controller 24 may more efficiently and effectively achieve network processing load balancing of workloads among processing cores 30A. The RL agent 156 of controller 24 may apply and continuously update one or more machine learning models of policy model 154, which are trained using historical data and updated through reinforcement learning, to proactively rebalance the network processing for workload 35 and / or allocate it to processing core 30A, rather than passively rebalancing the network processing for workload 35 or allocating it to processing core 30A based on real-time metrics. In this way, controller 24 may more effectively and efficiently determine the allocation of the network traffic processing for workload 35 to processing core 30A based on the state or attributes related to workload 35 and / or processing core 30A.
[0063] Figure 2 A block diagram of an example computing device is shown in which the techniques described herein may be implemented. Computing device 200 may represent Figure 1 any of servers 12 or another device (e.g., any of TOR switches 16).
[0064] In this example, computing device 200 includes a system bus 342 that couples the hardware components of the computing device 200's hardware environment. The system bus 342 couples a memory 344, a network interface card 330, a storage disk 346, and a multi-core computing environment 102 having a plurality of processing cores 108A through 108J (collectively referred to as "cores 108"). The network interface card 330 includes an interface configured to exchange data packets using a link of an underlying physical network. The multi-core computing environment 102 may include any number of processors and any number of hardware cores (e.g., from 4 to thousands). Each core 108 includes independent execution units to execute instructions that conform to the core instruction set architecture. The cores 108 may each be implemented as separate integrated circuits (ICs), or may be combined within one or more multi-core processors (or "multi-core" processors), each implemented using a single IC (i.e., a chip multi-processor), package, or die.
[0065] The storage disk 346 represents a computer-readable storage medium that includes volatile and / or non-volatile, removable and / or non-removable media implemented in any method or technology for storing information such as processor-readable instructions, data structures, process modules, or other data. Computer-readable storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), EEPROM, flash memory, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by the cores 108.
[0066] The main memory 344 includes one or more computer-readable storage media and may include random access memory (RAM) such as various forms of dynamic RAM (DRAM), e.g., DDR2 / DDR3 SDRAM, or static RAM (SRAM), flash memory, or any other form of fixed or removable storage media that can be used to execute or store desired program code and program data in the form of instructions or data structures and that can be accessed by a computer. The main memory 344 provides a physical address space composed of addressable memory locations.
[0067] In some examples, the memory 344 may exhibit a non-uniform memory access (NUMA) architecture for the multi-core computing environment 102. That is, the cores 108 may not have equal memory access times for the various storage media that make up the memory 344. In some examples, the cores 108 may be configured to use the portion of the memory 344 that provides the lowest memory latency for the cores to reduce the overall memory latency.
[0068] In some examples, a physical address space for a computer-readable storage medium (i.e., shared memory) can be shared among one or more cores 108. For example, cores 108A, 108B can be connected via a memory bus (not shown) to one or more DRAM packages, modules, and / or chips (also not shown) that present a physical address space accessible by cores 108A, 108B. Although this physical address space can provide the lowest memory access time to cores 108A, 108B for any part of memory 344, at least some of the remainder of memory 344 can be directly accessed by cores 108A, 108B. One or more cores 108 can also include an L1 / L2 / L3 cache or a combination thereof. The respective caches for cores 108 provide the lowest latency memory access to any storage medium. When rebalancing, controller 224 can apply load balancing module 252 to rebalance the network traffic handling of a particular workload to a different core that has a similar memory access time for the workload as the previous core (e.g., from core 108A to core 108B in the above example). Load balancing module 252 that incorporates this rebalancing factor (i.e., similar memory access times) can improve performance.
[0069] Memory 344, network interface card 330, storage disk 346, and multi-core computing environment 102 provide an operating environment for the software stack that executes virtual router 221 and one or more virtual machines 236A through 236N (collectively referred to as "virtual machines 236"). Virtual machines 236 can represent Figure 1 example instances of any virtual machine 36. Computing device 200 divides the virtual and / or physical address space provided by main memory 344 and, in the case of virtual memory, the virtual and / or physical address space provided by storage disk 346 into user space 345 (allocated for running user processes) and kernel space 314 (protected and generally not accessible to users). An operating system kernel (not shown) can execute in kernel space and can include, for example, Linux, Berkeley Software Distribution (BSD), another Unix variant kernel, or a Windows Server operating system kernel provided by Microsoft Corporation. In some examples, computing device 200 can execute a hypervisor to manage virtual machines 236 (also not shown). Examples of hypervisors include the Kernel-based Virtual Machine (KVM) for Linux kernels, Xen, ESXi provided by VMware, Windows Hyper-V provided by Microsoft, and other open-source and proprietary hypervisors. Additionally, virtual machines 236 can also execute application 237.
[0070] The pods 238A through 238N (collectively "pods 238") can include containers 239A through 239N (collectively "containers 239"). The containers 239 can include virtualization of an operating system to run multiple isolated systems on a single machine (virtual or physical). Examples of containers 239 include containers provided by the open-source DOCKER container application or by CoreOS Rkt ("Rocket"). Like virtual machines, each container is virtualized and can be isolated from the host and other containers. However, unlike virtual machines, each container can omit a separate operating system and provide an application suite and application-specific libraries. Generally, containers are executed by the host as isolated user-space instances and can share the operating system and common libraries with other containers executing on the host. Thus, compared to virtual machines, containers may require fewer processing capabilities, storage, and network resources. One or more sets of containers can be configured to share one or more virtual network interfaces for communication on corresponding virtual networks.
[0071] In some examples, containers are managed by their host cores to allow for the limiting and prioritization of resources (CPU, memory, block I / O, network, etc.) without the need to start any virtual machines. In some examples, namespace isolation features are used that allow for a completely isolated view of the operating environment by the application, which includes process trees, networks, user identifiers, and mounted file systems. In some examples, containers can be deployed according to Linux Containers (LXC), which is an operating system-level virtualization method for running multiple isolated Linux systems (containers) on a control host using a single Linux core.
[0072] In this example, the computing device 200 of the virtual router includes a kernel space 314 module: virtual router 221, and a user space 345 module: virtual router agent 223. The virtual router 221 performs the "forwarding plane" or packet forwarding function of the virtual router, while the virtual router agent 223 performs the "control plane" function of the virtual router. Some instances of the virtual router forwarding plane 221 can execute in user space (e.g., as a DPDK virtual router) or on a SmartNIC. More descriptions of the virtual router agent and the virtual router forwarding plane for implementing the virtual router can be found in U.S. Patent 9,571,394, issued on February 14, 2017, which is incorporated herein by reference in its entirety.
[0073] When the virtual router agent 223 of the virtual router receives an indication of a new workload deployed to the computing device 200, the virtual router agent 223 can initially allocate cores of the core 108 to handle the network traffic of the new workload (e.g., pod 238 or virtual machine 236).
[0074] According to the techniques described herein, the controller 224 can include a processing circuit 250, a load balancing module 252, a reinforcement learning (RL) agent 256, and a prediction module 258. The load balancing module 252 can be executed on the processing circuit 250 to generate allocation data 162 indicating which processing cores of the core 108 of the virtual router 221 should be used to handle the network traffic for a workload (e.g., virtual machine 236 or pod 238). The load balancing module 252 can generate the allocation data 162, which can include a mapping of the workload 35 to the core 108. The RL agent 256 is configured to be executed on the processing circuit 250 or the computing device 200 to apply a policy model 254 to generate or determine an allocation of the workload 35 to the core 108, and these workloads can be included in the allocation data 162. The prediction module 258 is configured to be executed on the processing circuit 250 to determine predicted attribute values, e.g., a predicted network traffic load associated with the workload 35 and the core 108. The prediction module 258 can predict the network traffic load at any future time using any machine learning prediction method (e.g., ARIMA, TBATS, regression methods, neural networks, etc.). The prediction module 258 can provide the predicted network traffic load to the RL agent 256.
[0075] The RL agent 256 can perform operations by allocating the core 108 to the workload 35. The RL agent 256 can perform operations by generating an allocation, i.e., selecting or picking one of the cores 108 to allocate to the network traffic of a new or existing workload. The RL agent 256 can perform operations by generating a rebalance or reallocation of the core 108 that has been allocated to handle the network traffic of a specific workload. The RL agent 256 can apply the policy model 254 to the predicted attribute values determined by the prediction module 258, thereby generating an allocation of the workload 35 to the core 108. The RL agent 256 can provide the predicted allocation to the load balancing module 252. The load balancing module 252 can determine whether to include the predicted allocation generated by one of the agents of the RL agent 256 in the allocation data 162.
[0076] In Figure 2In the example, the RL agent 256 can include a policy model 254 and a reinforcement learning (“RL”) algorithm 255. The policy model 254 can include a policy for generating a predicted assignment of network traffic of the workload 35 to the cores 108. The RL agent 256 can generate an initial set of predicted assignments by applying an initial policy of the policy model 254. The policy model 254 can include a policy composed of one or more machine learning models configured to instruct the RL agent 256 to perform operations for generating assignments. The policy model 254 can include machine learning models trained with static data, such as historical assignment data (e.g., assignment data generated by the load balancing module 252), the usage and / or type of the cores 108, the requirements and / or type of the workload 35, the profile of the workload 35, the profile of the cores 108, etc. In some cases, the policy model 254 can include machine learning models trained by simulation (e.g., Monte-Carlo-based methods). In some cases, the RL agent 256 can use one or more classification algorithms (e.g., multi-layer perceptron, deep neural network, etc.) to perform operations of selecting one of the cores 108 to assign to the workload to classify the cores 108 and / or the workload 35.
[0077] The controller 224 can generate or receive a profile of the type of the workload. The controller 224 can generate the profile based on one or more real-time metrics associated with processing network traffic for the type of the workload. In some examples, the controller 224 can apply a machine learning model trained with real-time metrics to generate a profile based on network traffic patterns associated with the type of the workload. For example, relevant patterns of network traffic that can be discerned from the real-time metrics can include the amount of traffic or the periodicity of the traffic (e.g., the highest traffic during evenings or weekends), which can be used to predict CPU utilization patterns through traffic also included in the real-time metrics. For example, the controller 224 can learn the typical CPU core usage of various instances of the virtual router 221 used within the system to process network traffic for the type of the workload based on one or more real-time metrics. The controller 224 can generate a profile of the type of the workload based on its typical CPU core usage. The RL agent 256 can obtain the profile of the workload and apply the policy model 254 to the profile to generate an assignment.
[0078] In some cases, the controller 224 may not receive metrics for classifying the workload and may not generate a profile of the workload type that can be used to generate an initial policy model. In such cases, the RL agent 256 can apply the policy model 254 to the predicted network traffic associated with the workload 35. The prediction module 258 of the controller 224 can determine the predicted network traffic based on the historical throughput associated with the workload network traffic. The prediction module 258 can collect historical throughput data (e.g., packets per second) of the unclassified workload. The prediction module 258 can use machine learning prediction methods to determine the predicted network traffic load. For example, the prediction module 258 can use time series prediction methods to predict the network traffic load, such as ARIMA, TBATS, or regression methods (e.g., Bayesian ridge regression, random forest regression, neural networks, etc.). The prediction module 258 can determine the predicted network traffic load based on the attributes and / or historical data associated with the workload 35 and / or the core 108. For example, the prediction module 258 can predict the peak or increment of the processing core utilization of a specific workload within a future period (e.g., 30 seconds, 60 seconds, etc.) based on the collected historical throughput data. The RL agent 256 can obtain the predicted network traffic load associated with the workload 35 from the prediction module 258. The RL agent 256 can apply the policy model 254 to the predicted network traffic load to allocate the workload 35 to the core 108. In some examples, the RL agent 256 can apply the policy model 254 to other attributes, such as the processing core state (e.g., active, inactive, available, reserved, etc.), packet or tail loss, workload priority, packet delay, etc., to generate an allocation. The RL agent 256 can send the generated allocation to the load balancing module 252 to generate the allocation data 162.
[0079] The RL agent 256 can perform the operation of generating an allocation according to the policy maintained by the policy model 254. In some cases, the RL agent 256 can implement supervised learning algorithms that provide direct user feedback on the performance of the RL agent 256 in generating prediction tasks. Over time, the RL agent 256 can learn the best possible operations (e.g., mapping or allocating the network traffic of the workload 35 to the core 108) to maximize the reward signal calculated by the reward function.
[0080] The RL agent 256 can apply a reward function to calculate a reward signal for an action taken by the RL agent 256. The RL agent 256 can apply a reward function that is defined as a combination of all expected values of attributes that the computing device 200 realizes from an allocation generated by the RL agent 256. In some cases, the computing device 200 can apply a reward function to determine a reward signal for an action taken by the RL agent 256. The computing device 200 can send the reward signal to the RL algorithm 255 of the RL agent 256.
[0081] In some examples, the reward function can be defined according to a predefined criterion. The reward function can be defined according to a criterion such as maximizing the overall utilization of the core 108, minimizing the latency associated with the average time for the core 108 to process network traffic for the workload 35, minimizing the number of idle cores 108, and minimizing the number of packet or tail drops. In some cases, the reward function can be defined as a weighted sum of all attributes (such as core 108 throughput, core 108 utilization, tail drops, jitter, packet latency, etc.) input into a machine learning model of the policy model 254. The reward function can be defined to minimize an attribute (such as minimizing packet latency or minimizing the number of idle cores 108) by attaching a negative weight to the attribute value and / or taking the reciprocal of the attribute value. By attaching a positive weight to an attribute according to the importance of maximizing the attribute, the reward function can be defined to maximize the attribute (such as maximizing the overall utilization of the core 108). The reward function can be defined such that each attribute is associated with a corresponding bias specific to the predefined criterion.
[0082] The RL agent 256 or the computing device 200 can calculate a reward signal for the actions taken by the RL agent 256 according to a reward function. For example, the RL agent 256 can calculate a reward signal for the actions taken by the RL agent 256 (e.g., rebalancing the core 108 to the workload 35) based on the allocated state or result generated by the RL agent 256. The RL agent 256 can obtain the generated allocated state (e.g., the state data 160) and input the attribute values related to the obtained state into the reward function. The RL agent 256 can calculate the reward according to the output of the reward function. For example, the RL agent 256 can periodically collect the state data 160 related to the current attribute values related to the workload 35 and / or the core 108. In some cases, the state data 160 can include the values related to the processing core 108 at a specific time point. For example, the state data 160 can include the values related to the utilization rate of each processing core of the processing core 108, the throughput of each processing core of the processing core 108, and / or the number of queues of each processing core of the processing core 108 at a specific time. In some cases, the state data 160 can include the values related to the workload 35 at a specific time point. For example, the state data 160 can include values such as tail loss related to the network traffic of the workload 35 over a period of time, jitter related to the network traffic of the workload 35 over a period of time, and / or packet delay related to the network traffic of the workload 35 over a period of time. In some examples, the state data 160 can include the values related to the core 108 and the values related to the workload 35. The RL agent 256 can receive the state data 160 from the virtual router agent 223 through an application programming interface (e.g., Prometheus), a telemetry interface, or other interfaces.
[0083] The RL agent 256 can update the policy model 254 through the reinforcement learning algorithm 255. The RL algorithm 255 can update the policy model 254 according to the reward signal calculated for the actions taken by the RL agent 256. For example, the RL algorithm 255 can slightly adjust the parameters of the machine learning model of the policy model 254 in response to the average reward signal value. In contrast, the RL algorithm 255 can significantly change (e.g., penalize) the parameters of the machine learning algorithm of the policy model 254 in response to a low reward signal value. For example, the reward function can be defined to prioritize the processing of network traffic for the workload 35A. The RL agent 256 can take actions by generating a predicted allocation that does not prioritize the processing of network traffic for the workload 35A. The RL algorithm 255 can penalize or otherwise punish the actions taken by the RL agent 256 by changing the parameters, biases, equations, etc. of the machine learning model on which the policy of the policy model 254 is based. The RL agent 256 can track the reward value according to the calculated reward signal over a period of time.
[0084] In some cases, the load balancing module 252 may generate allocation data 162 based on the prediction assignment provided by the RL agent 256. For example, the RL agent 256 may determine to allocate the workload to the first processing core by applying the policy model 254, while the load balancing module 252 may determine the allocation based on real-time metrics (e.g., the current load of core 108 or the network traffic processing requirements of workload 35). The load balancing module 252 may select the allocation determined by the RL agent 256 instead of the allocation determined based on real-time metrics. The load balancing module 252 may include the allocation determined by the RL agent 256 in the allocation data 162. In some cases, the load balancing module 252 may generate the allocation data 162 based on real-time metrics obtained from the virtual router agent 223 in response to the load balancing module 252 determining that the allocation generated based on real-time metrics will result in a more balanced distribution of the network traffic of workload 35 to the cores 108 compared to the predicted allocation generated by the RL agent 256. The load balancing module 252 may send the allocation data 162 to the virtual router agent 223.
[0085] The virtual router agent 223 may apply the allocation data 162 to allocate or re-allocate which of the cores 108 the virtual router 221 uses to process the network traffic for workloads 35A and 35B. In this example, according to the allocation data 162, the virtual router 221 may use core 108A when processing the network traffic for workload 35A and use core 108C when processing the network traffic for workload 35B.
[0086] Figure 3 A block diagram of example components of an example computing device 242 that executes the virtual router 221 for a virtual network in accordance with the techniques described herein is shown. The computing device 242 may be Figure 2 the computing device 200, or Figure 1 any one of the servers 12, or an example of another computing device described in the present disclosure.
[0087] The example computing device 242 includes a network interface card (NIC) 106 that is configured to direct data packets received by the NIC 106 to the processing core 108A for processing. As shown, the NIC 106 receives data packet streams 240A to 240C (collectively referred to as "data packet streams 240"), and may store the data packet streams 240 in the memory 144 for final processing by the virtual router 221. Similarly, the network traffic output by the workload to the virtual router 221 via the virtual network interface 390 is stored in the memory 144 for final processing by the virtual router 221.
[0088] The virtual router 221 includes multiple routing instances 122A to 122C (collectively referred to as "routing instances 122") for corresponding virtual machines 236A to 236K (collectively referred to as virtual machines 236). The virtual router 221 uses the routing instances 122 to process the data packet flow 240 originating from or destined for a specific workload 35 based on the workload distribution 382 for workload allocation. The workload distribution 382 maps each of the workloads 35 to the cores allocated in the cores 108. The virtual router 221 can enqueue the network traffic of the workload 35 using the queue 380. Further descriptions of the hardware queue, DPDK, and software queue for the workload interface (virtual interface) are all included in U.S. Patent No. 2022 / 0278927 published on September 1, 2022, which is incorporated herein by reference in its entirety.
[0089] The queue 380 can enqueue the network traffic based on the mapping of the workload 35 to the cores 108 provided by the distribution 382. The virtual router agent 223 can update the distribution 382 based on the allocation data 162 to update the mapping. The distribution 382 can be based on the allocation data 162 sent by the controller 224. The allocation data 162 can be a mapping of which core in the cores 108 should process the network traffic for each workload of the workload 35.
[0090] In this example, the network interface card 106 receives the data packet flow 240. The network interface card 106 can store the data packet flow 240 in the memory 144 for final processing by the virtual router 221 that executes one or more cores 108. In this example, the data packet flow 240A is destined for VM 36A, the data packet flow 240B is destined for VM 36B, and the data packet flow 240C is destined for the pod 38A. The virtual router 221 can use the mapping of the workload 35 to the cores 108 included in the distribution 382 to determine which core in the cores 108 will process the data packet flow corresponding to each workload (e.g., VM 36A, VM 36B, and the pod 38A). For example, the virtual router 221 can use the queue 380 to enqueue the data packet flow 240A to be processed with the core 108A, enqueue the data packet flow 240B to be processed with the core 108A, and enqueue the data packet flow 240C to be processed with the core 108C. A similar process can be applied to the data packet flows originating from each of the workloads.
[0091] According to the techniques described herein, the controller 224 may generate allocation data 162 for an allocation 382 according to an allocation generated by the RL agent 256 of the application policy model 254. The RL agent 256 may include a policy model 254 that includes policies for predicting a mapping of a workload 35 to a core 108 to meet one or more criteria. For example, the policy model 254 may manage policies that include one or more machine learning algorithms (e.g., neural networks) for predicting a mapping of a workload 35 to a core 108 according to established criteria to maximize overall throughput while minimizing utilization per core of the core 108, throughput per core of the core 108, the number of idle cores 108, packet delay, jitter, tail drop, etc.
[0092] The RL agent 256 may apply the policy model 254 to a predicted network traffic load determined by the prediction module 258. The RL agent 256 may obtain a predicted network traffic load related to the workload 35 from the prediction module 258. The prediction module 258 may determine the predicted network traffic load using a machine learning prediction method. For example, the prediction module 258 may predict the network traffic load using a time series prediction method, such as ARIMA, TBATS, or a regression method (e.g., Bayesian ridge regression, random forest regression, neural networks, etc.). The prediction module 258 may determine the predicted network traffic load according to attributes and / or historical data related to the workload 35 and / or the core 108 (e.g., workload profile, processing core profile, historical throughput, etc.). The RL agent 256 may provide the predicted network traffic load to the policy of the policy model 254, thereby generating an allocation of the workload 35 to the core 108. The RL agent 256 may send the allocation to the load balancing module 252 to generate the allocation data 162.
[0093] The RL agent 256 may update the policy model 254 through a reinforcement algorithm 255 to generate an optimized mapping of the network traffic of the workload 35 to the core 108. The RL agent 256 and / or the virtual router 221 may determine a reward signal for an operation taken by the RL agent 256. For example, the RL agent 256 may calculate a reward signal for an allocation generated by the RL agent 256 based on the state data 160, where the state data includes a set of values corresponding to the state or result of the environment, in response to the virtual router 221 implementing the allocation generated by the RL agent 256. The RL algorithm 255 may update the policy model 254 according to one or more calculated reward signals. The RL agent 256 may generate a network traffic allocation or mapping of the workload 35 to the core 108 by applying an updated version of the policy model 254.
[0094] The RL agent 256 can send the mapping to the load balancing module 252. In some cases, the load balancing module 252 can decide whether to include the prediction mapping when generating the allocation data 162. In some examples, the policy model 254 and / or the load balancing module 252 can monitor real-time metrics, such as the utilization of the current active core 108, in order to generate the allocation data 162 in a timely manner, which can include the mapping of the network traffic of the workload 35 to the core 108, and these mappings can respond to any core 108 being overworked or starved. The load balancing module 252 can send the allocation data 162 to the virtual router agent 223. The virtual router agent 223 can reconfigure the allocation 382 to map the network processing (associated with the data packet flow 240B) of the virtual machine 36B to the core 108B. The virtual router 221 will apply the reconfigured allocation 382 to enqueue the data packet flow 240B into the queue 380 for processing by the virtual router 221 executing on the core 108B. The virtual router 221 can obtain the network traffic of the workload based on the queue before processing the network traffic for the workload 35.
[0095] Figure 4 A conceptual diagram showing an example operation of a reinforcement learning agent according to the techniques described herein for allocating network traffic processing for a workload to processing cores. Figure 4 discussed in Figures 1 to 3 for example purposes only.
[0096] The policy model 254 can be instantiated with an initial policy for allocating the network traffic of the workload 35 to the core 108 (402). The policy model 254 can include a policy composed of one or more machine learning models (e.g., neural networks), and the model is trained using historical data related to the workload 35 and / or the core 108. In some examples, the policy model 254 can include a policy that contains a machine learning model trained using a simulation method (e.g., Monte-Carlo-based method). The RL agent 256 applies the policy model 254, which defines the allocation of the network traffic load of the workload 35 to the core 108 for processing by the virtual router 221 (404). In some examples, the RL agent 256 can apply the policy model 254 to rebalance the allocation of the network traffic for the workload 35 to the core 108. In some examples, the RL agent 256 can apply the policy model 254 to determine the allocation of the network traffic processing of a new workload to one of the cores 108. The RL agent 256 can send the generated allocation to the load balancing module 252 to generate the allocation data 162. The allocation 382 of the virtual router 221 can execute the allocation data 162. In some cases, the RL agent 256 can send the generated allocation directly to the allocation 382.
[0097] The RL agent 256 may collect the state of the computing environment 8 (406). The RL agent 256 may collect the state of the computing environment 8, such as, for example, the utilization rate of the cores 108, the throughput of each core 108 processing cores, the number of queues of each core 108 processing cores, tail packet loss, jitter, and / or packet delay. The RL agent 256 may calculate a reward signal (408) using a reward function. The RL agent 256 may calculate a reward signal based on the collected state of the computing environment 8 (e.g., state data 160). The RL agent 256 may calculate a reward signal using a reward function configured to minimize the utilization rate of each core 108, minimize the throughput of each core 108, minimize the number of idle processing cores of the cores 108, minimize packet delay, minimize jitter, minimize tail packet loss, and / or maximize overall throughput.
[0098] The RL agent 256 may apply the RL algorithm 255 to update the policy of the policy model 254 based on the calculated reward signal (410). The RL agent 256 may provide the determined reward signal to the RL algorithm 255 (e.g., a policy gradient method) to update the policy. Then, the RL agent 256 may repeat steps 404-410 using the updated policy and future-collected state data.
[0099] Figure 5 A block diagram shows an example computing device 500 for executing a workload according to the techniques described herein and example components of a virtual router for a virtual network. The virtual router 521 may represent Figure 3 an example case of the virtual router 221 in Figure 3 The virtual router 521 may include a queue 580 and an assignment 582, which may respectively represent Figure 3 example cases of the queue 380 and the assignment 382 in Figure 3 The queue 580 may include one or more forwarding cores (e.g., lcore 584A and lcore 584B), where each forwarding core is associated with a hardware core used by an NIC (e.g.,
[0100] The virtual router 521 may have forwarding cores lcore 584A and lcore 584B, which are associated with the hardware queues assigned by the NIC. The NIC is connected to the virtual router and configured with virtual router interfaces. When creating workloads (e.g., workloads 535A to E), the virtual interfaces of the workloads may have one or more software queues (e.g., software queue 537), where each software queue is assigned to the forwarding core of the virtual router 521 via the assignment 582. Typically, the assignment 582 assigns the forwarding cores of the assigned queues 580 to the software queues of the workloads (e.g., software queue 537) in a round-robin manner. For example, the software queue Q0 of workload 535A can be assigned to lcore 584A, the software queue Q0 of workload 535B can be assigned to lcore 584B, the software queue Q0 of workload 535C can be assigned to lcore 584A, the software queue Q1 of workload 535C can be assigned to lcore 584B, the software queue Q0 of workload 535D can be assigned to lcore 584A, the software queue Q0 of workload 535E can be assigned to lcore 584B, the software queue Q1 of workload 535E can be assigned to lcore 584A, the software queue Q2 of workload 535E can be assigned to lcore 584B, etc.
[0101] In another example, the virtual router may have already assigned four forwarding cores corresponding to four hardware queues to serve the network traffic of multiple workloads each associated with one or more software queues. The following example of the assignment of software queues to the forwarding cores of the virtual router is different from Figure 5 the example provided in, and this example demonstrates how software queues are typically assigned to forwarding cores in a round-robin manner. First, the virtual router obtains the physical function from the NIC and assigns the same number of hardware queues (identified by hardware queue IDs) in the NIC as the number of forwarding cores (identified by Lcore) of the virtual router (identified by the virtual router interface). In this example, the NIC has four hardware queues, which results in the virtual router assigning four forwarding cores, one core for each hardware queue.
[0102] lcore 584C:
[0103] Virtual router interface: 0000:17:01.1 Hardware queue ID: 0
[0104] lcore 584D:
[0105] Virtual router interface: 0000:17:01.1 Hardware queue ID: 1
[0106] lcore 584E:
[0107] Virtual router interface: 0000:17:01.1 Hardware queue ID: 2
[0108] lcore 584F:
[0109] Virtual router interface: 0000:17:01.1 Hardware queue ID: 3
[0110] Second, when creating or instantiating a workload, one or more software queues (identified by the SW queue ID) of each workload (identified by the workload interface) are assigned to each forwarding core of the virtual router in a round-robin manner.
[0111] lcore 584C:
[0112] Virtual router interface: 0000:17:01.1 Hardware queue ID: 0
[0113] Workload interface: vhostnet1-XXX-a1 Software queue ID: 0
[0114] Workload interface: vhostnet1-XXX-b1 Software queue ID: 0
[0115] Workload interface: vhostnet1-XXX-c1 Software queue ID: 0
[0116] Workload interface: vhostnet1-XXX-d1 Software queue ID: 0
[0117] lcore 584D:
[0118] Virtual router interface: 0000:17:01.1 Hardware queue ID: 1
[0119] Workload interface: vhostnet1-XXX-a2 Software queue ID: 0
[0120] Workload interface: vhostnet1-XXX-b1 Software queue ID: 1
[0121] Workload interface: vhostnet1-XXX-c2 Software queue ID: 0
[0122] Workload interface: vhostnet1-XXX-d2 Software queue ID: 0
[0123] lcore 584E:
[0124] Virtual router interface: 0000:17:01.1 Hardware queue ID: 2
[0125] Workload Interface: vhostnet1-XXX-a3 Software Queue ID: 0
[0126] Workload Interface: vhostnet1-XXX-b1 Software Queue ID: 2
[0127] Workload Interface: vhostnet1-XXX-c2 Software Queue ID: 1
[0128] lcore 584F:
[0129] Virtual Router Interface: 0000:17:01.1 Hardware Queue ID: 3
[0130] Workload Interface: vhostnet1-XXX-a3 Software Queue ID: 1
[0131] Workload Interface: vhostnet1-XXX-b2 Software Queue ID: 0
[0132] Workload Interface: vhostnet1-XXX-c2 Software Queue ID: 2
[0133] However, when creating and / or deleting workloads, the forwarding cores of virtual router 21A may become unbalanced because virtual router 21A does not actively reallocate the software cores of the workloads to the forwarding cores. For example, a workload with workload interfaces of vhostnet1-XXX-a3, vhostnet1-XXX-b2, and vhostnet1-XXX-c2 can be deleted, and a workload with a workload interface of vhostnet1-XXX-e1 can be created using only one software queue, which results in the following:
[0134] lcore 584C:
[0135] Virtual Router Interface: 0000:17:01.1 Hardware Queue ID: 0
[0136] Workload Interface: vhostnet1-XXX-a1 Software Queue ID: 0
[0137] Workload Interface: vhostnet1-XXX-b1 Software Queue ID: 0
[0138] Workload Interface: vhostnet1-XXX-c1 Software Queue ID: 0
[0139] Workload Interface: vhostnet1-XXX-d1 Software Queue ID: 0
[0140] lcore 584D:
[0141] Virtual router interface: 0000:17:01.1 Hardware queue ID: 1
[0142] Workload interface: vhostnet1-XXX-a2 Software queue ID: 0
[0143] Workload interface: vhostnet1-XXX-b1 Software queue ID: 1
[0144] Workload interface: vhostnet1-XXX-c2 Software queue ID: 0
[0145] Workload interface: vhostnet1-XXX-d2 Software queue ID: 0
[0146] lcore 584E:
[0147] Virtual router interface: 0000:17:01.1 Hardware queue ID: 2
[0148] Workload interface: vhostnet1-XXX-b1 Software queue ID: 2
[0149] Workload interface: vhostnet1-XXX-c2 Software queue ID: 1
[0150] Workload interface: vhostnet1-XXX-e1 Software queue ID: 0
[0151] lcore 584F:
[0152] Virtual router interface: 0000:17:01.1 Hardware queue ID: 3
[0153] This example demonstrates how the dynamic nature of deleting and creating workloads can lead to an uneven distribution of software queues to forwarding cores. Here, no software queues are assigned to lcore 584F, while four software queues are assigned to lcore 584C and lcore 584D. If software queues continue to be assigned to forwarding queues in a round-robin fashion, the inequality of software queues assigned to forwarding cores may never be resolved.
[0154] According to the technology of the present disclosure, the assignment of software queue 573 to the forwarding cores of lcores 584 can be actively redefined based on a policy model. In some cases, an external controller (e.g., Figure 3The controller 224 therein can be configured with a reinforcement learning agent (e.g., RL agent 256), which applies a policy model (e.g., policy model 254) to generate a predicted allocation of network traffic of workload 535 (e.g., software queue 537) to lcores 584, and the lcores 584 can correspond to the cores of core 108. The controller 224 can obtain the state (e.g., state data 160) related to the computing environment 8, such as the utilization rate of core 108, the throughput of each core 108, the number of hardware queues of core 108, the tail drop related to the network traffic of workload 535, the jitter related to the network traffic of workload 535, the packet delay related to the average time for core 108 to process the network traffic for workload 535, etc. The policy model 252 can include a policy that contains a machine learning model trained with static data (e.g., real-time metrics related to workload 535 and / or core 108, a profile for workload 535, the attributes of core 108, historical throughput data, etc.).
[0155] The reinforcement learning agent 256 can apply the policy model 254 to output the allocation of the network traffic of workload 535 to core 108. In Figure 5 the example, the reinforcement learning agent 256 can apply the policy model 254 to output the allocation of the software queue 537 of workload 535 to the forwarding core 584 of the virtual router 521. The reinforcement learning agent 256 can output the allocation by providing the predicted network traffic load determined by the prediction module 258. The prediction module 258 can predict the network traffic load related to workload 535 based on one or more machine learning prediction methods. In some cases, the controller 224 can obtain the state data 160, which indicates the number of hardware queues corresponding to each lcores 584 and the historical throughput data related to each lcores 584. The reinforcement learning agent 256 can input the state data 160 and the historical throughput data into one or more machine learning models of the policy model 254 to generate the allocation. In some examples, the reinforcement learning agent 256 can input the profile related to workload 535 as an input to one or more machine learning models of the policy model 254 to generate the allocation.
[0156] In Figure 5In the example of, the controller 224 can obtain the status data 160 and the configuration file related to the workload 535 to train one or more machine learning models of the policy model 254. For example, the controller 224 can obtain the status data 160 indicating that lcore 584A has four hardware queues and is expected to complete the network traffic processing related to the software queue 537 in the near future. The controller 224 can obtain the status data 160 indicating that lcore 584B has six hardware queues with lower throughput and is expected to continue to process the network traffic related to the software queue 537 for some time in the future. The controller 224 can also obtain the configuration file of the workload 535. For example, the controller 224 can obtain the configuration file for the workload 535A indicating that the workload 535A has a software queue with a currently low throughput but a high predicted throughput.
[0157] The reinforcement learning agent 256 can apply the policy model 254 to generate an allocation based on the output of the machine learning model of the policy model 254. For example, the reinforcement learning agent 256 can apply the policy model 254 to determine a new workload to create with high predicted throughput based on the workload configuration file provided to the machine learning model of the policy model 252. The reinforcement learning agent 256 can generate a new workload allocation for lcore 584A by applying the policy model 252 to the predicted network traffic load indicating that lcore 584A will complete the current network traffic processing in the near future. The reinforcement learning agent 256 will not simply allocate a new workload to lcore 584B because real-time metrics indicate that lcore 584B has lower throughput and more hardware queues. The reinforcement learning agent may generate an allocation to allocate the new workload to lcore 584A because the reinforcement learning agent 256 predicts that lcore 584A may be starved in the near future based on the output of the machine learning algorithm of the policy model 254. In this way, the reinforcement learning agent 256 can apply the policy model 252 to provide a proactive method to generate the allocation of the software queue 537 to lcores 584.
[0158] Generally, any operation of the above external controller (e.g., Figure 3 the controller 224 in ) can be completed by the virtual router (e.g., Figure 3 the virtual router 221 in ), the virtual router agent (e.g., Figure 3 the virtual router agent 223 in ) or any other server component included in the techniques described herein.
[0159] Figure 6 A conceptual diagram showing an example data 660 for generating a predicted allocation according to the techniques described herein. Figure 6 The discussion of is related toFigures 1 to 7 For example purposes, data 660 can be related to Figure 1 an example of the status data 160 in
[0160] The controller 224 can obtain the data 660 to train the reinforcement learning agent 256, so as to apply and update the policy model 252 to generate a predicted allocation of network traffic for the workload 535 to the lcores 584. In Figure 6 the example of, the data 660 can represent the number of hardware queues for each lcore 584 (x-axis) and the throughput (in packets per second (PPS)) of each hardware queue for the lcores 584 (y-axis). The lcore 584A can have four hardware queues with the average throughput value for each queue. The lcore 584B can have three hardware queues, where two queues have high throughput and one queue has low throughput. The lcore 584C can have three hardware queues for the average throughput of each queue. The lcore 584D can have six hardware queues with lower throughput for each queue.
[0161] According to the techniques described herein, the reinforcement learning agent 256 can receive an indication of the time when the lcore 584 completes the processing of the workload network traffic. For example, the reinforcement learning agent 256 can receive an indication that the lcore584A is expected to continue processing network traffic related to the indicated throughput for a period of time. The reinforcement learning agent 256 can receive an indication that the lcore 584B is expected to continue processing network traffic associated with the indicated throughput for a period of time. The reinforcement learning agent 256 can receive an indication that the lcore 584C is expected to complete the processing of network traffic related to the indicated throughput in the near future. The reinforcement learning agent 256 can receive an indication that the lcore 584D is expected to continue processing network traffic related to the indicated throughput for a period of time. In some cases, the reinforcement learning agent 256 can generate an indication of how long the forwarding core processing may take based on the predicted network traffic load determined by the prediction module 258.
[0162] The reinforcement learning agent 256 can input the data 660 along with these instructions into one or more machine learning models of the policy model 254 to output the network traffic for the workload to be assigned to the lcores 584. For example, the reinforcement learning agent 256 can receive an indication of a new workload to be assigned to the forwarding cores of the lcores 584. Instead of simply assigning the network traffic for the new workload to the lcore 584D due to low total throughput, the reinforcement learning agent 256 applies the policy model 254 to generate an assignment of the network traffic for the new workload to the lcore 584C because the predicted network traffic load indicates that the processing of the lcore 584C may be completed in the near future. In this way, the reinforcement learning agent 256 can predict that the lcore 584C may be in a "starving" state in the near future and generate an assignment to provide work to the lcore 584C before the lcore 584C is in a "starving" state.
[0163] Figure 7 A flowchart illustrating an example operation of a method according to the techniques described herein is shown. Figure 7 The discussion of Figures 1 to 6 is for example purposes.
[0164] A computing device (e.g., the controller 224) can include a reinforcement learning agent (e.g., the RL agent 254) that applies a policy model to a predicted network traffic load associated with a workload to assign the workload to a first processing core (702) of a plurality of processing cores of the computing device. A virtual router (e.g., the virtual router 221) executing on the first processing core can process network traffic for the workload (704) based on the assignment of the workload to the first processing core of step 702.
[0165] The techniques described herein, including any of the foregoing portions, can be implemented in hardware, software, firmware, or any combination thereof. The various features described as modules, units, or components can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices or other hardware devices. In some examples, the various features of the electronic circuit can be implemented as one or more integrated circuit devices, e.g., an integrated circuit chip or a chipset.
[0166] If implemented in hardware, the present disclosure can relate to a device, e.g., a processor or an integrated circuit device, e.g., an integrated circuit chip or a chipset. Alternatively or additionally, if implemented in software or firmware, the techniques can be at least partially implemented by a computer-readable data storage medium including instructions that, when executed, cause a processor to perform one or more of the above-described methods. For example, the computer-readable data storage medium can store such instructions for execution by the processor.
[0167] A computer-readable medium can form part of a computer program product, which can include packaging material. The computer-readable medium can include computer data storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. In some examples, a manufactured product can include one or more computer-readable storage media.
[0168] In some examples, the computer-readable storage medium can include a non-transitory medium. The term "non-transitory" can indicate that the storage medium is not embodied in a carrier wave or a propagated signal. In certain examples, the non-transitory storage medium can store data that can change over time (e.g., in RAM or a cache).
[0169] The code or instructions can be software and / or firmware executed by a processing circuit that includes one or more processors, such as one or more digital signal processors (DSPs), general microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described in this disclosure can also be provided in software modules or hardware modules.
Claims
1. A computer network method, comprising: applying, by the reinforcement learning agent, a policy model to a predicted network traffic load associated with the workload to allocate the workload to a first processing core of a plurality of processing cores of a computing device; as well as Network traffic for the workload is processed using the first processing core by the virtual router based on assigning the workload to the first processing core.
2. The computer network method according to claim 1, wherein: The policy model is trained using one or more of historical allocation data, historical throughput data, usage of the plurality of processing cores, types of the plurality of processing cores, requirements of the workload, types of the workload, and profiles associated with the workload.
3. The computer network method according to claim 1, wherein: The policy model includes a neural network.
4. The computer network method according to claim 1, wherein: The reinforcement learning agent is executed by one of the computing device and a controller for a virtualized computing infrastructure including the computing device.
5. The computer network method according to any one of claims 1 to 4, further comprising calculating the predicted network traffic load by the following steps: Historical data associated with a plurality of workloads and the plurality of processing cores is obtained, wherein: The plurality of workloads includes the workload; and The predicted network traffic load is determined based on the historical data.
6. The computer network method according to any one of claims 1 to 4, further comprising: collecting state data associated with the plurality of processing cores and the workload; calculating a reward signal for allocating the workload to the first processing core based on a reward function and the state data; as well as Based on the reward signal, the policy model is updated.
7. The computer network method according to claim 6, wherein: The status data includes one or more of utilization of the plurality of processing cores, throughput of each of the plurality of processing cores, a number of queues associated with each of the plurality of processing cores, tail packet loss, jitter, and packet delay.
8. The computer network method according to claim 6, wherein: The reward function is defined to do one or more of: minimize a delay associated with an average time for the first processing core to process the network traffic for the workload, minimize a number of idle processing cores of the plurality of processing cores, minimize a number of packet drops, and maximize a total throughput of the plurality of processing cores.
9. The computer network method according to claim 6, wherein: Updating the policy model includes applying a reinforcement learning algorithm including a policy gradient method.
10. The computer network method according to any one of claims 1 to 4, further comprising: Determining, based on the real-time metric, to allocate the workload to a second processing core among the plurality of processing cores; as well as selecting to assign the workload to the first processing core instead of assigning the workload to the second processing core, Wherein, the network traffic for the workload is processed using the first processing core based on the selection.
11. The computer network method according to any one of claims 1 to 4, further comprising: Allocating, by the virtual router, a queue to the first processing core; enqueuing the network traffic for the workload to the queue based on assigning network traffic processing for the workload to the first processing core; as well as Prior to processing the network traffic for the workload, the network traffic for the workload is obtained based on the queue.
12. A computing system comprising a processing circuit capable of accessing a storage device, the processing circuit being configured to: applying, by the reinforcement learning agent, the policy model to a predicted network traffic load associated with the workload to allocate the workload to a first processing core of a plurality of processing cores of a computing device; and Network traffic for the workload is processed using the first processing core by the virtual router based on assigning the workload to the first processing core.
13. The computing system of claim 12, wherein: The processing circuit is further configured to: collecting state data associated with the plurality of processing cores and the workload; calculating a reward signal for allocating the workload to the first processing core based on a reward function and the state data; as well as Based on the reward signal, the policy model is updated.
14. The computing system of claim 13, wherein: The status data includes one or more of utilization of the plurality of processing cores, throughput of each of the plurality of processing cores, a number of queues associated with each of the plurality of processing cores, tail packet loss, jitter, and packet delay.
15. The computing system of claim 13, wherein: The reward function is defined as minimizing a delay associated with an average time for the first processing core to process the network traffic for the workload, minimizing a number of idle processing cores of the plurality of processing cores, minimizing a number of packet drops, and maximizing a total throughput of the plurality of processing cores.
16. The computing system of claim 13, wherein: To update the policy model, the processing circuit is configured to apply a reinforcement learning algorithm including a policy gradient method.
17. A computing system according to any one of claims 12 to 16, wherein: The processing circuit is further configured to: obtaining historical data associated with a plurality of workloads and the plurality of processing cores, wherein the plurality of workloads includes the workload; and Based on the historical data, the predicted network traffic load is determined by a machine learning prediction method.
18. The computing system of any one of claims 12 to 16, wherein: The processing circuit is further configured to: Determining, based on the real-time metric, to allocate the workload to a second processing core among the plurality of processing cores; as well as selecting to assign the workload to the first processing core instead of assigning the workload to the second processing core, Wherein, the network traffic for the workload is processed using the first processing core based on the selection.
19. The computing system of any one of claims 12 to 16, wherein: The processing circuit is further configured to: Allocating, by the virtual router, a queue to the first processing core; enqueuing the network traffic for the workload to the queue based on assigning network traffic processing for the workload to the first processing core; as well as Prior to processing the network traffic for the workload, the network traffic for the workload is obtained based on the queue.
20. A computer-readable storage medium encoded with instructions for causing one or more programmable processors to be configured to perform the computer network method according to any one of claims 1 to 11, or to be configured as a computing system according to any one of claims 12 to 19.
Citation Information
Patent Citations
Data interfaces with isolation for containers deployed to compute nodes
US20220278927A1
Tunneled packet aggregation for virtual networks
US9571394B1