Load balancing method, network card, and storage medium
By setting up a redirection module and multiple redirection queues in the network card, combining the mapping relationship between the data flow and the redirection queue in the target direction, the mapping relationship periodically adjusts the mapping relationship to achieve load balancing between processor cores, solving the problem of load imbalance of processor cores after the virtual switch is offloaded to the network card, and improving the performance and scalability of load balancing.
Patent Information
- Application Number
- PCT/CN2024/120125
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-21
- Filing Date
- 2024-09-20
- Publication Date
- 2025-05-30
AI Technical Summary
In a cloud computing environment, when a virtual switch is offloaded to a network card, how to achieve load balancing between multiple processor cores to avoid resource exhaustion and packet loss from processor cores with higher loads.
By setting up a redirection module and multiple redirect queues in the network card, combining the mapping relationship between the data flow and the redirect queue in the target direction, the mapping relationship is periodically adjusted to achieve load balancing between processor cores.
It realizes load balancing between multiple processor cores in the network card, reduces the number of queues that processor cores need to poll on the software forwarding plane, reduces overhead, and improves the performance and scalability of load balancing.
Smart Images

Figure CN2024120125_30052025_PF_FP_ABST
Abstract
Description
Load balancing method, network card and storage medium Technical Field
[0001] The embodiments of the present application relate to the field of cloud computing technology, and in particular to a load balancing method, a network card, and a storage medium. Background Art
[0002] With the development of cloud computing and virtualization technologies, physical servers deployed in the cloud can be virtualized into multiple virtual machines (VMs) through virtualization technology, thereby improving resource utilization of the physical servers. To cope with the ever-increasing network bandwidth and to support virtualization functions at a lower cost, virtualization functions such as network virtualization can be offloaded to the network interface cards (NICs) (NICs) that communicate with the physical servers. For example, the virtual switch (vSwitch) running on the physical server can be offloaded to the NIC, achieving high-performance packet forwarding.
[0003] When the virtual switch is offloaded to the network card, the network card is provided with multiple processor cores (such as CPU cores) for implementing the virtual switch. In this context, how to achieve load balancing among the multiple processor cores has become a technical problem that those skilled in the art urgently need to solve.
[0004] Summary of the Invention
[0005] In view of this, embodiments of the present application provide a load balancing method, a network card, and a storage medium to achieve load balancing among multiple processor cores in a network card and improve the performance of load balancing.
[0006] To achieve the above objectives, the embodiments of the present application provide the following technical solutions.
[0007] In a first aspect, an embodiment of the present application provides a load balancing method, applied to a network card, the method comprising:
[0008] Acquire a data packet, wherein the data packet comes from a network card queue in a target direction;
[0009] Determine a summary value of the data packet, where the summary value of the data packet represents a data flow corresponding to the data packet;
[0010] Determining, based on a mapping relationship between data flows and redirection queues in a target direction, a redirection queue in the target direction to which the digest value of the data packet is mapped, wherein the data flow in the mapping relationship is represented by the digest value of the data packet corresponding to the data flow;
[0011] The data packet is allocated to the redirection queue of the determined target direction so that the data packet is processed by the processor core corresponding to the redirection queue of the determined target direction; wherein the multiple processor cores of the network card implement a virtual switch, and one processor core corresponds to one redirection queue in the target direction to transmit the data packet from the network card queue in the target direction, and the mapping relationship is periodically adjusted with the goal of balancing the load of the data packets in the target direction processed by each processor core.
[0012] In a second aspect, an embodiment of the present application provides a load balancing method, applied to a network card, the method comprising:
[0013] Determining a total data packet load of each processor core in a current cycle, wherein the total data packet load of a processor core in the current cycle is composed of data packet loads in each target direction of the processor core in the current cycle;
[0014] Determining, based on the total packet load of each processor core in the current cycle, whether an adjustment condition for adjusting a mapping relationship is met in the current cycle, the mapping relationship comprising: a mapping relationship between a data flow and a redirection queue in a target direction, wherein the data flow in the mapping relationship is represented by a summary value of a packet corresponding to the data flow; wherein the multiple processor cores of the network card implement a virtual switch, and each processor core corresponds to a redirection queue in the target direction to transmit packets from the network card queue in the target direction;
[0015] If the current cycle meets the adjustment condition for adjusting the mapping relationship, the mapping relationship between the data flow and the redirection queue in the target direction is adjusted with the goal of balancing the load of the data packets in the target direction processed by each processor core.
[0016] In a third aspect, an embodiment of the present application provides a network card, comprising: a hardware acceleration engine and an on-chip processor; the hardware acceleration engine is provided with a redirection module, multiple downlink redirection queues, and multiple uplink redirection queues; the on-chip processor is provided with multiple processor cores, the multiple processor cores implement a virtual switch, and the virtual switch is provided with a load adjustment module;
[0017] The redirection module is connected between multiple downlink redirection queues and virtual network card queues, and between multiple uplink redirection queues and physical network card queues; one processor core corresponds to one downlink redirection queue to transmit data packets from the virtual network card queue, and one processor core corresponds to one uplink redirection queue to transmit data packets from the physical network card queue;
[0018] The redirection module is configured to execute the load balancing method as described in the first aspect above; the load adjustment module is configured to execute the load balancing method as described in the second aspect above.
[0019] In a fourth aspect, an embodiment of the present application provides a storage medium, which stores one or more computer-executable instructions. When the one or more computer-executable instructions are executed, the load balancing method described in the first aspect above is implemented, or the load balancing method described in the second aspect above is implemented.
[0020] In an embodiment of the present application, a processor core corresponds to a redirection queue in the target direction to transmit data packets from the network card queue in the target direction; at the same time, through the mapping relationship between the data flow and the redirection queue in the target direction, the embodiment of the present application can allocate the redirection queue in the target direction for the data packets from the network card queue in the target direction, so that the data packets are transmitted to the corresponding processor core for processing through the allocated redirection queue in the target direction, wherein the data flow in the mapping relationship is represented by the summary value of the data packet corresponding to the data flow. It can be seen that the embodiment of the present application can achieve data flow-level data packet redirection capability, and combined with the method of corresponding a redirection queue in the target direction for a processor core, the processor core only needs to poll the corresponding redirection queue in the target direction without polling the network card queue in the target direction, reducing the number of queues that the processor core needs to poll on the software forwarding side, reducing overhead, and being able to adapt to the expansion of the number of virtual machines, thereby improving scalability.
[0021] At the same time, the mapping relationship between the data flow and the redirection queue in the target direction is adjusted periodically, and the goal of adjusting the mapping relationship is to make the data packet load in the target direction processed by each processor core tend to be balanced. Therefore, the embodiment of the present application allocates the redirection queue in the target direction to the data flow through the mapping relationship between the data flow and the redirection queue in the target direction, and then transmits it to the corresponding processor core for processing, which can achieve the purpose of load balancing between multiple processor cores in the network card.
[0022] It can be seen that the embodiment of the present application realizes load balancing among multiple processor cores in the network card, and improves the refinement of load balancing (load balancing adjustment at the data flow level), reduces overhead, and improves scalability, thereby improving the performance of load balancing. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0024] Figure 1 is an example diagram of the relationship between a network card and a physical server.
[0025] Figure 2 is an example diagram of the connection between a virtual machine and a virtual switch.
[0026] FIG3 is an example diagram of a network card provided in an embodiment of the present application.
[0027] FIG4 is a flow chart of a load balancing method provided in an embodiment of the present application.
[0028] FIG5 is an example diagram of a downlink mapping table provided in an embodiment of the present application.
[0029] FIG6 is an example diagram of an uplink mapping table provided in an embodiment of the present application.
[0030] FIG7 is another flow chart of the load balancing method provided in an embodiment of the present application.
[0031] FIG8 is a flowchart of adjusting the mapping relationship provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0033] A virtual switch is a network switch that can be configured to manage network traffic between virtual machines and between virtual machines and the physical network. A primary function of a virtual switch is packet forwarding. For example, a virtual switch can forward packets based on forwarding rules to ensure they are routed correctly. Packets can be categorized as those sent by virtual machines and those sent by the physical network.
[0034] Virtual switches are typically implemented by software running on physical servers. For example, a general-purpose processor (such as a CPU) in a physical server can implement a virtual switch through software functionality. However, with the continuous development of cloud computing and network technologies, users' demands for network performance are becoming increasingly stringent. The packet forwarding performance of virtual switches running on physical servers can no longer meet these requirements. Therefore, to achieve high-performance packet forwarding for virtual machines, the virtual switches running on physical servers can be offloaded to network adapters (such as SmartNICs) to improve packet forwarding performance.
[0035] It should be noted that in cloud computing applications, a network card (such as a SmartNIC) is a network adapter that communicates with physical servers. Taking a SmartNIC as an example, a SmartNIC combines network connectivity with intelligent features to enhance cloud computing network performance, management capabilities, and security. For example, in addition to performing the network transmission functions of a standard NIC, a SmartNIC can also enhance application performance through a built-in programmable and configurable hardware acceleration engine. In one application example, a network card (such as a SmartNIC) can be configured for a cloud computing virtualization environment, such as a cloud computing platform or data center, to support the network needs of virtual machines.
[0036] Offloading, short for hardware offload, is the process of offloading specific tasks from a general-purpose processor (such as a CPU) to a dedicated hardware device for processing, thereby accelerating their execution. Accordingly, offloading a virtual switch running on a physical server to a network interface card (such as a SmartNIC) can be considered as using the network interface card (such as a SmartNIC) as a dedicated hardware device to implement the virtual switch functionality. As a result, the general-purpose processor of the physical server no longer implements the virtual switch functionality, and the network interface card (such as a SmartNIC) replaces the physical processor to implement the virtual switch functionality.
[0037] For ease of understanding, when the virtual switch is unloaded to the network card (such as a smart network card), Figure 1 exemplifies the relationship between the network card and the physical server. As shown in Figure 1, the physical server 110 runs multiple virtual machines 111 to 11n (n is the number of virtual machines, depending on the actual situation) through virtualization technology to meet the computing needs of cloud computing; the virtual switch originally running on the physical server 110 is unloaded to the network card 120 (network card 120 is such as a smart network card), and the network card 120 carries the function of the virtual switch through the hardware acceleration engine 121 and the on-chip processor 122.
[0038] The network card 120 may include a hardware acceleration engine 121 and an on-chip processor 122. When the virtual switch is offloaded to the network card, the hardware acceleration engine 121 is responsible for the hardware forwarding portion of packet forwarding for the virtual switch. For example, the hardware acceleration engine 121 may configure certain forwarding rules to accelerate packet forwarding through hardware acceleration. In one example, the hardware acceleration engine 121 may configure a forwarding rule table that records packet forwarding rules, thereby accelerating packet forwarding at the hardware level and improving packet forwarding performance.
[0039] The hardware acceleration engine 121 can be, for example, FPGA (Field Programmable Gate Array) hardware. Through the programmability of FPGA hardware, the network card (such as a smart network card) can adapt to changing network requirements and security challenges, thereby improving the flexibility of the network card (such as a smart network card) for actual application scenarios.
[0040] On-chip processor 122 is a processor, such as a CPU, provided in a network card (e.g., a smart network card). When the virtual switch is offloaded to the network card (e.g., a smart network card), the on-chip processor 122 implements the virtual switch functionality in software. For example, the on-chip processor 122 can carry out the virtual switch functionality by running software, and is the software forwarding portion of the virtual switch responsible for packet forwarding. In other words, the complete virtual switch functionality can be implemented on the on-chip processor of the network card (e.g., a smart network card) and implemented by running software on the on-chip processor.
[0041] It should be noted that the network card 120 can communicate with the physical server 110. For example, the network card 120 can communicate with the physical server 110 through a PCIE (PerIP address herald Component Interconnect Express, a high-speed serial computer expansion bus standard) bus. The network card 120 can also communicate with the physical network so that the physical server 110 can access the physical network. For example, the network card 120 can be connected to the physical devices of the physical network (physical devices of the physical network such as network switches, routers, firewalls, etc.) through communication connection methods such as the PCIE bus to achieve access to the physical network. It should be noted that one side of the network card (such as an intelligent network card) is connected to the physical server and the other side is connected to the physical network to meet the external communication needs of the physical server, and the specific form and type of the physical device of the physical network to which the network card (such as an intelligent network card) is connected can be determined according to the actual application situation.
[0042] When the virtual switch is offloaded to the network card, the virtual machines running on the physical server can connect to the virtual switch through the virtual network card queue, allowing the virtual machines and the virtual switch to interact through the virtual network card queue. At the same time, on the network card side, the network card's on-chip processor can implement the virtual switch through multiple processor cores (such as multiple CPU cores), so the virtual switch implemented by multiple processor cores can interact with the virtual machines through the virtual network card queue to forward the virtual machine's data packets. In other words, the virtual machine is connected to the multiple processor cores of the network card's on-chip processor (the multiple processor cores that implement the virtual switch function) through the virtual network card queue, thereby achieving data packet forwarding.
[0043] For ease of understanding, FIG2 exemplarily shows an example diagram of the connection between a virtual machine and a virtual switch. As shown in FIG1 and FIG2 , a physical server runs multiple virtual machines 111 to 11n, and a plurality of processor cores 211 to 21m (m is the number of processor cores, depending on the actual situation) are set in the on-chip processor of the network card to realize a virtual switch. Each virtual machine is connected to the processor core through the corresponding virtual network card queue to realize the forwarding of the virtual machine's data packets through the processor core. The virtual network card queue corresponding to a virtual machine can be one or more, and there is no limit. At the same time, each processor core can be connected to the corresponding physical network card queue to connect to the physical network, thereby processing data packets transmitted from the physical network interface.
[0044] It should be noted that the virtual network card queue is the network card queue of the virtual machine, which is configured to send data packets from the virtual machine to the virtual switch (such as the processor core that implements the virtual switch). Each virtual machine can have one or more virtual network card queues, so that when the virtual machine has multiple virtual network card queues, the virtual machine can send data packets in parallel. The physical network card queue is the hardware queue of the network card (such as the smart network card), which is configured to process data packets of the physical network interface. For example, the processor core can process data packets of the physical network interface through the corresponding physical network card queue. The physical network card queue allows the network card (such as the smart network card) to process multiple network connections and data packets simultaneously, thereby improving the performance of the network card (such as the smart network card).
[0045] The combined use of virtual NIC queues and physical NIC queues can improve network performance, especially in virtualized environments. For example, a virtual machine can communicate with multiple physical NIC queues on a network adapter (such as a Smart NIC) through its virtual NIC queue, thereby enabling communication between the virtual machine and the physical network and effectively utilizing network resources.
[0046] As can be seen, on the NIC (e.g., SmartNIC) side, the on-chip processor uses multiple processor cores to forward packets. However, the number of processor cores configured to forward packets is limited, so the processor cores need to use certain mapping relationships to achieve packet forwarding. For example, the processor cores need to use the mapping relationship between data flows and virtual NIC queues, and the mapping relationship between processor cores and virtual NIC queues, to distribute the data flow packets to the processor core mapped to the corresponding virtual NIC queue for forwarding.
[0047] That is to say, by establishing a mapping relationship between data streams and virtual network card queues, the data packets of the same data stream are routed to the virtual network card queue mapped to the data stream, and then based on the established mapping relationship between the virtual network card queue and the processor core, the data packets can be forwarded through the processor core mapped to the virtual network card queue to which the data packets need to be routed, thereby achieving the forwarding of data packets of the same data stream with the same processing logic. It should be noted that a data stream is a collection of data packets. If data packet forwarding is processed at the flow level, it means that the processing logic of data packets with the same five-tuple is the same. The five-tuple can be the source IP address, destination IP address, source port number, destination port number and protocol number of the data packet.
[0048] Forwarding packets by mapping data flows to virtual network interface card (VNIC) queues, and by mapping processor cores to VNIC queues, can lead to load imbalances across multiple processor cores. For example, multiple VNIC queues or data flows with high loads may ultimately be mapped to the same processor core, resulting in a high packet load on that core and a high risk of packet loss due to resource exhaustion. This can also leave idle processor cores unused.
[0049] To solve the problem of packet loss in processor cores with higher loads caused by unbalanced loads between the processor cores, the inventors of this application considered fine-grained division and distribution of the load of each processor core, and considered the following two approaches.
[0050] The first method is to adjust the mapping relationship between data streams and virtual network card queues to achieve load balancing across processor cores. For example, the mapped virtual network card queues are dynamically adjusted for multiple data streams, ensuring that the number of data streams required to be processed by multiple processor cores is balanced. In other words, based on the mapping relationship between virtual network card queues and processor cores, and the changes in the number of data streams mapped by each virtual network card queue, the virtual network card queues mapped by the data streams are adjusted so that the number of data streams mapped by the virtual network card queues mapped to different processor cores is close to the same, thereby balancing the number of data streams required to be processed by each processor core.
[0051] The first method mentioned above requires the processor core to calculate the data flow when dynamically changing the mapping relationship between the data flow and the virtual network card queue, and to evenly distribute multiple data flows to different processor cores through communication between the processor cores, so that multiple data flows can be distributed to different processor cores in a balanced manner; the above process involves a large overhead of multiple processor cores, resulting in a large overhead of load balancing.
[0052] The second method is to dynamically adjust the mapping relationship between virtual network card queues and processor cores to achieve load balancing across the processor cores. This method suffers from a low level of load balancing refinement, i.e., the granularity of load balancing is relatively coarse. For example, if a virtual network card queue has a heavy data flow load, then even if the processor core mapped to the virtual network card queue is adjusted, the processor core ultimately mapped to the virtual network card queue will need to handle the heavy data flow. This may cause the processor core ultimately mapped to the virtual network card queue to run out of resources and cause packet loss, rendering load balancing ineffective.
[0053] In addition, adjusting the mapping relationship between the virtual network card queue and the processor core is an invasive behavior, which can easily lead to some potential problems, especially the inability to correctly forward data packets. That is, due to the complexity of the hardware and operating system, if the mapping relationship between the virtual network card queue and the processor core is arbitrarily adjusted, a higher probability of failure may occur, such as potential problems such as virtual network card queue stagnation or processor core inability to receive packets from the virtual network card queue. Among them, virtual network card queue stagnation means that the virtual network card queue has stopped working normally and can no longer receive or send data packets.
[0054] Furthermore, with the increasing density of virtual machine deployments, the first and second approaches are also difficult to scale. For example, with the development of cloud computing, the density of virtual machines deployed on physical servers has increased, and accordingly, the number of virtual network interface card (VNIC) queues has also increased significantly. If the mapping relationship between VNIC queues and processor cores, or the mapping relationship between VNIC queues and data flows, needs to be adjusted, the processor cores will need to frequently switch between the large number of VNIC queues, resulting in additional overhead. Furthermore, with large-scale deployments of virtual machines, statistics and scheduling of VNIC queues are also difficult to implement.
[0055] It can be seen that the above load balancing method has the problems of high overhead, low degree of refinement, high probability of failure, and difficulty in scalability. Therefore, the performance of the load balancing method needs to be improved.
[0056] Based on this, an embodiment of the present application provides an improved load balancing solution to achieve load balancing between multiple processor cores in a network card (such as an intelligent network card), and to improve the performance of load balancing. As an optional implementation, Figure 3 exemplarily shows an example diagram of a network card provided by an embodiment of the present application. In combination with Figures 1, 2 and 3, in an embodiment of the present application, a redirection module 310, multiple downlink redirection queues 321 to 32m (m is the number of downlink redirection queues, which is consistent with the number of multiple processor cores), and multiple uplink redirection queues 331 to 33m (m is the number of uplink redirection queues, which is consistent with the number of multiple processor cores) are provided in the hardware acceleration engine 121 of the network card (such as an intelligent network card). The multiple processor cores 211 to 21m set in the on-chip processor 122 of the network card (such as an intelligent network card) can realize a virtual switch, and an embodiment of the present application is provided with a load adjustment module 340 in the virtual switch.
[0057] Among them, the redirection module 310, the downstream redirection queues 321 to 32m, and the upstream redirection queues 331 to 33m can be implemented by a hardware acceleration engine, for example, the redirection module 310, the downstream redirection queues 321 to 32m, and the upstream redirection queues 331 to 33m can be implemented through the programmable capabilities of FPGA hardware.
[0058] On the one hand, the redirection module is connected between the virtual network card queue and the downstream redirection queue, and can allocate data packets received from the virtual network card queue to the downstream redirection queue according to the data packet allocation strategy provided in the embodiment of the present application; on the other hand, the redirection module is connected between the physical network card queue and the upstream redirection queue, and can allocate data packets received from the physical network card queue to the upstream redirection queue according to the data packet allocation strategy provided in the embodiment of the present application. In other words, the redirection module can receive data packets from the physical network card queue and the virtual network card queue, and based on the data packet allocation strategy provided in the embodiment of the present application, allocate the data packets of the virtual network card queue to the downstream redirection queue, and then transmit them to the processor core corresponding to the downstream redirection queue; at the same time, based on the data packet allocation strategy provided in the embodiment of the present application, allocate the data packets of the physical network card queue to the upstream redirection queue, and then transmit them to the processor core corresponding to the upstream redirection queue.
[0059] It should be noted that the downstream direction refers to the direction in which the virtual machine sends data. In other words, the downstream direction is the direction in which the NIC (e.g., SmartNIC) receives packets from the virtual NIC queue. Accordingly, the downstream redirection queue allocates packets received by the NIC (e.g., SmartNIC) from the virtual NIC queue, i.e., packets sent by the virtual machine. The upstream direction refers to the direction in which the physical network sends data. In other words, the upstream direction is the direction in which the NIC (e.g., SmartNIC) receives packets from the physical NIC queue. Accordingly, the upstream redirection queue allocates packets received by the NIC (e.g., SmartNIC) from the physical NIC queue, i.e., packets sent by the physical network.
[0060] In an embodiment of the present application, the number of downstream redirection queues is equal to the number of multiple processor cores, and the number of upstream redirection queues is equal to the number of multiple processor cores, so that each processor core can correspond to one downstream redirection queue and one upstream redirection queue, and then from the perspective of data packet forwarding, each processor core only needs to forward and process the data packets of the corresponding downstream redirection queue and upstream redirection queue, that is, each processor core only needs to poll the corresponding pair of downstream redirection queues and upstream redirection queues, without the need to poll other queues, so that when a downstream redirection queue and an upstream redirection queue corresponding to the binding processor core are bound, even if the number of virtual network card queues increases, the data packets processed by the processor core are provided by the corresponding downstream redirection queue and upstream redirection queue, which can be easily applied and expanded when the number of virtual network card queues is large.
[0061] In one example, the downstream redirection queue and the upstream redirection queue can be two groups of PCI (Peripheral Component Interconnect) devices implemented by FPGA hardware, that is, m downstream redirection queues 321 to 32m are a group of PCI devices implemented by FPGA hardware, and m upstream redirection queues 331 to 33m are another group of PCI devices implemented by FPGA hardware, and m is set to the number of processor cores.
[0062] The load adjustment module 340 may be a software function module implemented in the virtual switch, and may periodically adjust the data packet allocation policy of the redirection module 310 for allocating data packets to the downlink redirection queue and the uplink redirection queue.
[0063] To facilitate understanding of the packet distribution strategy provided in the embodiment of the present application, the optional method flow of the load balancing method provided in the embodiment of the present application is introduced below. As an optional implementation, Figure 4 exemplarily shows an optional flow chart of the load balancing method provided in the embodiment of the present application. This method flow can be applied to a network card (such as a smart network card) and executed by a hardware acceleration engine (such as FPGA hardware) in the network card. In an optional implementation, the redirection module implemented by the FPGA hardware can execute the method flow. Referring to Figure 4, this method flow may include the following steps.
[0064] In step S410, a data packet is obtained, wherein the data packet comes from a network card queue in a target direction.
[0065] In an embodiment of the present application, the redirection module can obtain a data packet from the network card queue in the target direction. As an optional implementation, the target direction can be divided into a downstream direction and an upstream direction. If the target direction is the downstream direction, the network card queue in the target direction is a virtual network card queue, and it is necessary to perform data packet forwarding processing in the downstream direction; accordingly, the redirection module can obtain a data packet in the virtual network card queue, which is a data packet sent by the virtual machine and needs to be forwarded by the network card (such as an intelligent network card). If the target direction is the upstream direction, the network card queue in the target direction is a physical network card queue, and it is necessary to perform data packet forwarding processing in the upstream direction; accordingly, the redirection module can obtain a data packet in the physical network card queue, which is a data packet sent by the physical network and needs to be forwarded by the network card (such as an intelligent network card).
[0066] In step S411, a summary value of the data packet is determined, where the summary value of the data packet represents a data flow corresponding to the data packet.
[0067] As an optional implementation, the summary value of a data packet can be a value of a set length, which is used to represent the data stream to which the data packet belongs. That is, the data packet is transmitted in a streaming form, and after any data packet of the data stream is transmitted to the network card queue in the target direction, the redirection module can obtain the data packet from the network card queue in the target direction, and determine the summary value of the data packet, so as to represent the data stream corresponding to the data packet through the summary value of the data packet. For example, after any data packet in the data stream sent by the virtual machine is transmitted to the virtual network card queue, the redirection module can obtain the data packet from the virtual network card queue and determine the summary value of the data packet, so as to represent the data stream sent by the virtual machine to which the data packet corresponds through the summary value of the data packet. For another example, after any data packet in the data stream sent by the physical network is transmitted to the physical network card queue, the redirection module can obtain the data packet from the physical network card queue and determine the summary value of the data packet, so as to represent the data stream sent by the physical network to which the data packet corresponds through the summary value of the data packet.
[0068] As an optional implementation, the redirection module can use a digest algorithm to determine the digest value of the data packet based on the information representing the data flow in the data packet, so that the digest value of the data packet can represent the data flow corresponding to the data packet. In one implementation example, the five-tuple information of the data packet (source IP address, destination IP address, source port number, destination port number, and protocol number) can be used as the information representing the data flow in the data packet. In other words, data packets of the same data flow have the same five-tuple information, so the redirection module can use a digest algorithm to determine the digest value of the data packet based on the five-tuple information of the data packet, so that the digest value of the data packet can represent the data flow corresponding to the data packet.
[0069] In one example, the digest value of the data packet can be a hash value, and accordingly, the digest algorithm can be a hash algorithm; thus, the redirection module can use the hash algorithm based on the five-tuple information of the data packet to determine the hash value of the data packet, so as to represent the data flow corresponding to the data packet through the hash value of the data packet. In a further example, the hash value can be a value of a specific bit length calculated based on the five-tuple information of the data packet, such as an 8-bit hash value.
[0070] As an optional implementation, the five-tuple information of the data packet can be carried in the header information of the data packet, so that after the redirection module obtains the data packet in the target direction, it can extract the five-tuple information (such as source IP address, destination IP address, source port number, destination port number and protocol number) from the header information of the data packet, thereby determining the hash value corresponding to the five-tuple information of the data packet.
[0071] In step S412, the target redirection queue to which the digest value of the data packet is mapped is determined based on the mapping relationship between the data flow and the target redirection queue, wherein the data flow in the mapping relationship is represented by the digest value of the data packet corresponding to the data flow.
[0072] In an embodiment of the present application, multiple processor cores of an on-chip processor of a network interface card (e.g., a smart network interface card) implement a virtual switch, and each processor core corresponds to a redirection queue in the target direction; that is, a processor core in the target direction obtains data packets from the network interface card queue in the target direction through a corresponding redirection queue. Thus, data packets from the network interface card queue in the target direction can be assigned to the redirection queue in the target direction by a redirection module, and then forwarded by the processor core corresponding to the assigned redirection queue in the target direction.
[0073] As an optional implementation, combined with the above description, if the target direction is downlink, the redirection queue for the target direction is the downlink redirection queue. Thus, one processor core in the downlink direction corresponds to one downlink redirection queue, and one processor core in the downlink direction retrieves packets from the downlink virtual network card queue through the corresponding downlink redirection queue. In other words, packets from the virtual network card queue can be assigned to the downlink redirection queue by the redirection module, and then forwarded by the processor core corresponding to the assigned downlink redirection queue.
[0074] If the target direction is upstream, the redirection queue for the target direction is the upstream redirection queue. Thus, one processor core in the upstream direction corresponds to one upstream redirection queue. A processor core in the upstream direction retrieves data packets from the upstream physical network card queue through the corresponding upstream redirection queue. In other words, data packets from the physical network card queue can be assigned to the upstream redirection queue by the redirection module, and then forwarded by the processor core corresponding to the assigned upstream redirection queue.
[0075] As an optional implementation, when the redirection module assigns a redirection queue of a target direction to a data packet, it can use the summary value of the data packet to determine the redirection queue of the corresponding mapped target direction based on the mapping relationship between the data flow and the redirection queue of the target direction, thereby assigning the data packet to the redirection queue of the target direction corresponding to the summary value.
[0076] In a further optional implementation, based on the target direction being divided into a downstream direction and an upstream direction (correspondingly, the redirection queue for the target direction is divided into a downstream redirection queue and an upstream redirection queue), the mapping relationship between the data flow and the redirection queue for the target direction can be divided into: a downstream mapping relationship between the data flow and the downstream redirection queue, and an upstream mapping relationship between the data flow and the upstream redirection queue. In other words, when the data flow in the mapping relationship is represented by the summary value of the data packet corresponding to the data flow, the downstream mapping relationship records the relationship between the summary value of the data flow and the corresponding mapped downstream redirection queue, and the upstream mapping relationship records the relationship between the summary value of the data flow and the corresponding mapped upstream redirection queue.
[0077] Furthermore, based on a processor core corresponding to a redirection queue in the target direction, the mapping relationship represents the processor core to which the data packets of the data flow are assigned (i.e., the processor core corresponding to the redirection queue in the target direction to which the data packets are assigned). Therefore, based on the mapping relationship, it is possible to determine the processor core to which the data packets of the data flow are assigned, and complete the assignment of the data packets to the processor core. For example, the downlink mapping relationship corresponds to: the downlink redirection queue to which the data packets of the data flow from the virtual network card queue are assigned, i.e., the processor core to which the data packets of the data flow from the virtual network card queue are assigned (the processor core and the downlink redirection queue have a one-to-one correspondence). For another example, the uplink mapping relationship corresponds to: the uplink redirection queue to which the data packets of the data flow from the physical network card queue are assigned, i.e., the processor core to which the data packets of the data flow from the physical network card queue are assigned (the processor core and the uplink redirection queue have a one-to-one correspondence).
[0078] As an optional implementation, the mapping relationship between data flows and target redirection queues can be in the form of a table, referred to as a mapping table. For example, the mapping table can record the digest value representing the data flow (i.e., the digest value of the data packet corresponding to the data flow) and the redirection queue corresponding to the mapped target direction. In one example, the mapping table can have multiple entries, each of which records the digest value of a data flow and a redirection queue corresponding to the mapped target direction. For example, the index of an entry is the digest value of a data flow, and the content of an entry is the identifier of a redirection queue corresponding to the mapped target direction. Thus, when the redirection module determines the target redirection queue corresponding to the digest value of a data packet based on the mapping relationship between the data flow and the target redirection queue, it can query the multiple entries of the mapping table based on the digest value of the data packet, thereby querying the entry with the digest value of the data packet as the index, and then confirming the content of the query entry to obtain the identifier of the redirection queue corresponding to the mapped target direction, thereby determining the redirection queue corresponding to the target direction mapped by the digest value of the data packet. In one example, if the digest value of the data packet is in the form of a hash value, the mapping table can also be called a rehash table.
[0079] As an optional implementation, based on the target direction being divided into a downstream direction and an upstream direction, the mapping table can be divided into a downstream mapping table and an upstream mapping table. The downstream mapping table can record the mapping relationship between the data flow and the downstream redirection queue in a tabular form; in an example, the downstream mapping table can have multiple entries, and one entry records the summary value of a data flow and a corresponding mapped downstream redirection queue; for example, the index of an entry in the downstream mapping table is the summary value of a data flow, and the content of an entry is the identifier of a corresponding mapped downstream redirection queue. For ease of understanding, taking 256 entries as an example, FIG5 exemplifies an example diagram of a downstream mapping table (also called a downstream rehash table) for reference. Thus, when the redirection module determines the downstream redirection queue to which the summary value of a data packet corresponds, it can query multiple entries in the downstream mapping table based on the summary value of the data packet, thereby querying the entry with the summary value of the data packet as the index, and then confirming the content of the queried entry to obtain the identifier of the corresponding mapped downstream redirection queue.
[0080] The uplink mapping table can record the mapping relationship between data streams and uplink redirection queues in a tabular form; in an example, the uplink mapping table can have multiple entries, and one entry records the summary value of a data stream and a corresponding mapped uplink redirection queue; for example, the index of an entry in the uplink mapping table is the summary value of a data stream, and the content of an entry is the identifier of a corresponding mapped uplink redirection queue. For ease of understanding, taking 256 entries as an example, FIG6 exemplifies an example diagram of an uplink mapping table (also called an uplink rehash table) for reference. Thus, when the redirection module determines the uplink redirection queue to which the summary value of a data packet corresponds, it can query multiple entries in the uplink mapping table based on the summary value of the data packet, thereby querying the entry with the summary value of the data packet as the index, and then confirming the content of the queried entry to obtain the identifier of the corresponding mapped uplink redirection queue.
[0081] In step S413 , the data packet is allocated to the redirection queue of the determined target direction, so that the data packet is processed by the processor core corresponding to the redirection queue of the determined target direction.
[0082] In which, multiple processor cores of a network card (such as a smart network card) implement a virtual switch, and one processor core corresponds to a redirection queue in the target direction to transmit data packets from the network card queue in the target direction. In addition, the mapping relationship is periodically adjusted with the goal of balancing the load of data packets in the target direction processed by each processor core.
[0083] After determining the redirection queue of the target direction to which the data packet is assigned based on the mapping relationship, the redirection module can assign the data packet to the redirection queue of the determined target direction, so that the redirection queue of the target direction to which the data packet is assigned can transmit the data packet to the corresponding processor core for forwarding processing, that is, the data packet can be forwarded and processed by the processor core corresponding to the redirection queue of the assigned target direction.
[0084] As an optional implementation, if the target direction is the downlink direction, then after determining the downlink redirection queue to which the summary value of the data packet is mapped based on the downlink mapping relationship, the embodiment of the present application can allocate the data packet to the determined downlink redirection queue so that the data packet can be forwarded by the processor core corresponding to the determined downlink redirection queue. If the target direction is the uplink direction, then after determining the uplink redirection queue to which the summary value of the data packet is mapped based on the uplink mapping relationship, the embodiment of the present application can allocate the data packet to the determined uplink redirection queue so that the data packet can be forwarded by the processor core corresponding to the determined uplink redirection queue.
[0085] As an optional implementation, the processor core may forward a data packet using forwarding rules after receiving the data packet. For example, the processor core may query multiple forwarding configurations of the virtual switch to generate forwarding rules, and then use the forwarding rules to forward the data packet. In another example, the processor core may forward the data packet based on the generated forwarding rules.
[0086] In a further optional implementation, in order to achieve hardware-accelerated forwarding of data packets, the processor core can send forwarding rules to the hardware acceleration engine, so that when the hardware acceleration engine finds a forwarding rule that matches the data packet, it can perform hardware-accelerated forwarding of the data packet according to the forwarding rule. When the hardware acceleration engine does not find a forwarding rule that matches the data packet, the data packet is passed to the processor core, and the processor core generates a forwarding rule and uses the forwarding rule to perform data packet forwarding processing.
[0087] In the embodiments of the present application, the mapping relationship between data flows and target redirection queues (e.g., the downstream mapping relationship and the upstream mapping relationship) is not fixed, but is periodically adjusted to ensure that after the target redirection queues are assigned to the data packets based on the mapping relationship, the target data packet load processed by each processor core tends to be balanced. As an optional implementation, a load adjustment module provided in the virtual switch can implement periodic adjustment of the mapping relationship. For relevant details, please refer to the corresponding description below and will not be repeated here.
[0088] In an embodiment of the present application, a processor core corresponds to a redirection queue in the target direction to transmit data packets from the network card queue in the target direction; at the same time, through the mapping relationship between the data flow and the redirection queue in the target direction, the embodiment of the present application can allocate the redirection queue in the target direction for the data packets from the network card queue in the target direction, so that the data packets are transmitted to the corresponding processor core for processing through the allocated redirection queue in the target direction, wherein the data flow in the mapping relationship is represented by the summary value of the data packet corresponding to the data flow. It can be seen that the embodiment of the present application can achieve data flow-level data packet redirection capability, and combined with the method of corresponding a redirection queue in the target direction for a processor core, the processor core only needs to poll the corresponding redirection queue in the target direction without polling the network card queue in the target direction, reducing the number of queues that the processor core needs to poll on the software forwarding side, reducing overhead, and being able to adapt to the expansion of the number of virtual machines, thereby improving scalability.
[0089] At the same time, the mapping relationship between the data flow and the redirection queue in the target direction is periodically adjusted, and the goal of adjusting the mapping relationship is to make the data packet load in the target direction processed by each processor core tend to be balanced. Therefore, the embodiment of the present application allocates the redirection queue in the target direction to the data flow through the mapping relationship between the data flow and the redirection queue in the target direction, and then transmits it to the corresponding processor core for processing, which can achieve the purpose of load balancing between multiple processor cores in the network card (such as the smart network card).
[0090] It can be seen that the embodiment of the present application realizes load balancing among multiple processor cores in a network card (such as a smart network card), and improves the refinement of load balancing (load balancing adjustment at the data flow level), reduces overhead, and improves scalability, thereby improving the performance of load balancing.
[0091] As an optional implementation, the mapping relationship between the data flow and the redirection queue of the target direction (such as the downlink mapping relationship and the uplink mapping relationship) can be periodically adjusted by the on-chip processor of the network card (such as the smart network card). For example, in the embodiment of the present application, a load adjustment module is set in the virtual switch implemented by multiple processor cores of the on-chip processor, so that the mapping relationship is periodically adjusted by the load adjustment module. Based on this, in order to facilitate the load adjustment module to periodically adjust the mapping relationship, in an optional implementation, before allocating the data packet to the redirection queue of the determined target direction, the redirection module can write the summary value (such as a hash value) of the data packet into the reserved field in the header of the data packet, and then allocate the data packet with the summary value written in the reserved field in the header to the redirection queue of the determined target direction, so that the load adjustment module can count the data packet load corresponding to each data flow in each cycle (that is, the number of data packets to be processed corresponding to each data flow), thereby performing periodic adjustment of the mapping relationship.
[0092] As an optional implementation, to facilitate understanding of the optional implementation method of adjusting the mapping relationship in the embodiment of the present application, Figure 7 exemplarily shows another optional flowchart of the load balancing method provided in the embodiment of the present application. This method flow can be applied to a network card (e.g., a smart network card) and executed by an on-chip processor in the network card (e.g., a smart network card). In an optional implementation, the load adjustment module in the virtual switch implemented by multiple processor cores of the on-chip processor can execute this method flow. Referring to Figure 7, this method flow can include the following steps.
[0093] In step S710 , the total data packet load of each processor core in the current cycle is determined. The total data packet load of a processor core in the current cycle is composed of the data packet loads of each target direction of the processor core in the current cycle.
[0094] As an optional implementation, the load adjustment module can be periodically awakened to execute the process shown in Figure 7 to achieve periodic adjustment of the mapping relationship. By periodically waking up the load adjustment module, the load adjustment module can periodically determine the total packet load of each processor core in the current cycle, that is, the number of packets processed by each processor core in the current cycle. For example, the load adjustment module is awakened at the end of the current cycle, so that for each processor core in the on-chip processor, the load adjustment module can separately determine the packet load of each processor core in each target direction in the current cycle. For any processor core, the packet load of each processor core in each target direction in the current cycle is added together to obtain the total packet load of the processor core in the current cycle. For example, the total packet load of a processor core in the current cycle can be the sum of the packet load of the processor core in the downstream direction and the upstream direction in the current cycle. In an optional implementation, a cycle can be regarded as a time period of a set time length (for example, 5 minutes). The time length corresponding to a cycle can be determined according to actual conditions, and the embodiments of the present application do not limit it. In addition, the embodiments of the present application can also set the time lengths of each cycle to be the same, or different cycles can be set to have different time lengths. The time lengths corresponding to different cycles can be determined according to actual conditions, and the embodiments of the present application do not limit it.
[0095] As an optional implementation, the data packet load of the processor core can be regarded as the number of data packets forwarded and processed by the processor core. That is, for any processor core, the embodiment of the present application can periodically determine the number of data packets forwarded and processed by the processor core in the current cycle, involving the number of data packets in the downlink direction, the number of data packets in the uplink direction, and the total number of data packets (the sum of the number of data packets in the downlink direction and the number of data packets in the uplink direction forwarded and processed by the processor core in the current cycle, i.e., the total data packet load of the processor core in the current cycle).
[0096] In step S711 , the utilization rate of each processor core in the current cycle is determined according to the total load of data packets of each processor core in the current cycle.
[0097] As an optional implementation, for any processor core, the embodiment of the present application can determine the utilization rate of the processor core in the current cycle based on the total packet load of the processor core in the current cycle and the preset packet load upper limit of the processor core in the current cycle. For example, the embodiment of the present application can set the packet load upper limit corresponding to the utilization rate reaching 100%. That is to say, if the number of packets processed and forwarded by the processor core in one cycle reaches the packet load upper limit, it is considered that the utilization rate of the processor core has reached 100%, and the processor core is fully used. Furthermore, the embodiment of the present application can divide the total packet load of the processor core in the current cycle by the packet load upper limit of the processor core in the current cycle to determine the utilization rate of the processor core in the current cycle.
[0098] As an optional implementation, for any processor core, an embodiment of the present application can determine the utilization rate of the processor core in the current cycle based on the effective usage time corresponding to the total load of data packets of the processor core in the current cycle, and the time length corresponding to the cycle. For example, the utilization rate of a processor core in the current cycle can be determined based on the proportion of time corresponding to the total load of data packets processed by the processor core in the current cycle. In one example, for any processor core, an embodiment of the present application can determine the effective usage time corresponding to the total load of data packets processed by the processor core in the current cycle, and then divide the effective usage time by the time length of the cycle to obtain the utilization rate corresponding to the total load of data packets processed by the processor core in the current cycle. Among them, the effective usage time corresponding to the total load of data packets processed by the processor core in the current cycle can be regarded as the total usage time of the processor core in forwarding and processing data packets in the downlink and uplink directions in the current cycle.
[0099] In step S712, based on the utilization rate of each processor core in the current cycle, it is determined whether the current cycle meets the adjustment condition for adjusting the mapping relationship. If not, step S713 is executed; if so, step S714 is executed.
[0100] In an embodiment of the present application, the mapping relationship is divided into a downlink mapping relationship between a data flow and a downlink redirection queue, and an uplink mapping relationship between a data flow and an uplink redirection queue; wherein the summary value of the data packet represents the data flow corresponding to the data packet.
[0101] After determining the utilization rate of each processor core in the current cycle, the embodiment of the present application needs to judge whether the current cycle has reached the adjustment condition for adjusting the mapping relationship based on the utilization rate of each processor core in the current cycle. For example, the embodiment of the present application can set the adjustment condition required for adjusting the mapping relationship in advance, so that after the load adjustment module determines the utilization rate of each processor core in the current cycle, it can judge whether the current cycle has reached the adjustment condition for adjusting the mapping relationship based on the utilization rate of each processor core in the current cycle. If the adjustment condition for adjusting the mapping relationship is not reached, the embodiment of the present application can execute step S713, end the process, and wait for the load adjustment module to be awakened in the next cycle. If the adjustment condition for adjusting the mapping relationship is reached, the embodiment of the present application can execute step S714 to adjust the downlink mapping relationship between the data stream and the downlink redirection queue, and the uplink mapping relationship between the data stream and the uplink redirection queue, respectively.
[0102] As an optional implementation, an adjustment condition for adjusting the mapping relationship may be defined as follows: the number of processor cores whose utilization exceeds a utilization threshold is greater than a set number, and the utilization difference between the processor core with the highest utilization and the processor core with the lowest utilization exceeds a preset difference. It should be noted that the processor core with the highest utilization may be considered to be the processor core with the highest packet load among the multiple processor cores, and the processor core with the lowest utilization may be considered to be the processor core with the lowest packet load among the multiple processor cores.
[0103] In an example, if the load adjustment module determines, based on the usage rates of each processor core in the current cycle, that the number of processor cores whose usage rates exceed the usage rate threshold is greater than the set number, and the usage rate difference between the processor core with the highest usage rate and the processor core with the lowest usage rate exceeds the preset difference value, then it can be confirmed that the adjustment conditions for adjusting the mapping relationship are met, and step S714 can be executed to adjust the downlink mapping relationship and the uplink mapping relationship respectively; if, based on the usage rates of each processor core in the current cycle, it is determined that the number of processor cores whose usage rates exceed the usage rate threshold is not greater than the set number, and / or the usage rate difference between the processor core with the highest usage rate and the processor core with the lowest usage rate does not exceed the preset difference value, then it can be confirmed that the adjustment conditions for adjusting the mapping relationship are not met, and step S713 can be executed to end the process and wait for the load adjustment module to be awakened in the next cycle.
[0104] In one example, taking the on-chip processor as an example, the four processor cores can implement a virtual switch, and a load adjustment module is provided in the virtual switch. Thus, if the load adjustment module determines in the current cycle that the number of processor cores with a utilization rate exceeding 90% (an example of a utilization rate threshold) is greater than 1 (1 can be an example of a set number), and the utilization rate difference between the processor core with the highest utilization rate (i.e., the processor core with the highest data packet load) and the processor core with the lowest utilization rate (i.e., the processor core with the lowest data packet load) among the four processor cores reaches 30% (an example of a preset difference), it can be confirmed that the adjustment conditions for adjusting the mapping relationship are met; if the load adjustment module determines in the current cycle that the number of processor cores with a utilization rate exceeding 90% is not greater than 1, and / or the utilization rate difference between the processor core with the highest utilization rate and the processor core with the lowest utilization rate among the four processor cores does not reach 30%, it can be confirmed that the adjustment conditions for adjusting the mapping relationship are not met.
[0105] It should be noted that step S711 and step S712 can be optional implementation methods for determining whether the current cycle meets the adjustment conditions for adjusting the mapping relationship based on the total data packet load of each processor core in the current cycle. The embodiment of the present application can also set other forms of adjustment conditions, and make the form of the adjustment conditions associated with the total data packet load of each processor core in the current cycle, and is not limited to the form of the adjustment conditions described above.
[0106] In step S713, the process ends.
[0107] When the adjustment condition for adjusting the mapping relationship is not met in the current cycle, the embodiment of the present application can end the process and wait for the load adjustment module to be awakened in the next cycle.
[0108] In step S714 , with the goal of balancing the load of data packets in the target direction processed by each processor core, the mapping relationship between the data flow and the redirection queue in the target direction is adjusted.
[0109] When the adjustment conditions for adjusting the mapping relationship are met in the current cycle, the embodiment of the present application can adjust the mapping relationship between the data flow and the redirection queue in each target direction, so that the data packet load in the target direction processed by each processor core tends to be balanced. For example, the embodiment of the present application can adjust the downlink mapping relationship in the downlink direction so that the data packet load in the downlink direction processed by each processor core tends to be balanced; and in the uplink direction, adjust the uplink mapping relationship so that the data packet load in the uplink direction processed by each processor core tends to be balanced.
[0110] As an optional implementation, FIG8 exemplarily shows an optional flowchart for adjusting the mapping relationship provided in an embodiment of the present application. As shown in FIG8 , the method flow may include the following steps.
[0111] In step S810, for any target direction, the average utilization rate of the target direction is determined based on the utilization rate of each processor core in the target direction of the current cycle; wherein, any processor core whose utilization rate of the target direction of the current cycle is higher than the average utilization rate is a first-category processor core, and any processor core whose utilization rate of the target direction of the current cycle is lower than the average utilization rate is a second-category processor core.
[0112] In the process of adjusting the mapping relationship between the data flow and the redirection queue in the target direction, the embodiment of the present application adjusts the mapping relationship between the data flow and the redirection queue in the target direction. Therefore, for the target direction (divided into the downstream direction and the upstream direction), the embodiment of the present application can determine the utilization rate of each processor core in the target direction of the current cycle. For example, for any processor core, the packet load of the target direction processed by the processor core in the current cycle is divided by the packet load upper limit of the processor core in the current cycle to obtain the utilization rate of the processor core in the target direction of the current cycle. For another example, for any processor core, the effective usage time corresponding to the packet load of the target direction processed by the processor core in the current cycle is divided by the time length of the current cycle to obtain the utilization rate of the processor core in the target direction of the current cycle.
[0113] Taking the downlink as the target direction, the downlink utilization rate of each processor core in the current cycle can be determined. For example, for any processor core, the downlink packet load processed by the processor core in the current cycle is divided by the upper limit of the packet load of the processor core in the current cycle to obtain the downlink utilization rate of the processor core in the current cycle. For another example, for any processor core, the effective usage time corresponding to the downlink packet load processed by the processor core in the current cycle is divided by the duration of the current cycle to obtain the downlink utilization rate of the processor core in the current cycle.
[0114] Taking the uplink direction as an example, the uplink utilization rate of each processor core in the current cycle can be determined. For example, for any processor core, the uplink packet load processed by the processor core in the current cycle is divided by the upper limit of the packet load of the processor core in the current cycle to obtain the uplink utilization rate of the processor core in the current cycle. For another example, for any processor core, the effective usage time corresponding to the uplink packet load processed by the processor core in the current cycle is divided by the duration of the current cycle to obtain the uplink utilization rate of the processor core in the current cycle.
[0115] After confirming the usage rate of each processor core in the target direction of the current cycle, in an optional implementation, the embodiment of the present application can add the usage rates of each processor core in the target direction of the current cycle, and divide the sum by the number of processor cores to obtain the average usage rate in the target direction. In one example, assuming that the number of processor cores is m, and the usage rate of the i-th processor core in the target direction of the current cycle is C i (i belongs to m), the average usage rate of the target direction is set to C avg , then C avg It can be expressed as: C avg= Sum(C i ) / m.
[0116] After determining the average utilization rate of multiple processor cores in the target direction of the current cycle, the embodiment of the present application can identify the first type of processor core and the second type of processor core based on whether the utilization rate of the processor core in the target direction of the current cycle is higher than the average utilization rate, that is, the first type of processor core can be regarded as a processor core whose utilization rate in the target direction of the current cycle is higher than the average utilization rate, and the second type of processor core can be regarded as a processor core whose utilization rate in the target direction of the current cycle is lower than the average utilization rate. Since the utilization rate of the first type of processor core in the target direction of the current cycle is higher, and the utilization rate of the second type of processor core in the target direction of the current cycle is lower, the embodiment of the present application can adjust the mapping relationship between the data flow and the redirection queue of the target direction to divert part of the data flow of the first type of processor core in the target direction to the second type of processor core for processing, so that the utilization rate of the first type of processor core and the second type of processor core in the target direction can approach the average utilization rate, so as to achieve the purpose of load balancing among multiple processor cores in the target direction.
[0117] It should be noted that, since the target direction is divided into a downlink direction and an uplink direction, the embodiment of the present application needs to determine the first type of processor core and the second type of processor core in the downlink direction and the uplink direction respectively. Moreover, since the packet load of the same processor core in the downlink direction and the uplink direction may be different (corresponding to the different utilization rates of the same processor in the downlink direction and the uplink direction), the same processor core may become the first type of processor core in the downlink direction and the second type of processor core in the uplink direction. In other words, the first type of processor core and the second type of processor core in the downlink direction and the uplink direction need to be determined separately, and may be different.
[0118] For example, after confirming the downlink utilization rate of each processor core in the current cycle, the embodiment of the present application can add the downlink utilization rate of each processor core in the current cycle, divide the addition result by the number of processor cores, and thus obtain the average utilization rate in the downlink direction; if the downlink utilization rate of the processor core in the current cycle is higher than the average utilization rate in the downlink direction, then the processor core is a first-type processor core in the downlink direction, and data flow diversion in the downlink direction is required; if the downlink utilization rate of the processor core in the current cycle is lower than the average utilization rate in the downlink direction, then the processor core is a second-type processor core in the downlink direction.
[0119] Similarly, after confirming the uplink usage rate of each processor core in the current cycle, the embodiment of the present application can add the uplink usage rate of each processor core in the current cycle, divide the addition result by the number of processor cores, and thus obtain the average usage rate in the uplink direction; if the uplink usage rate of the processor core in the current cycle is higher than the average usage rate in the uplink direction, then the processor core is a first-type processor core in the uplink direction, and data flow diversion in the uplink direction is required; if the uplink usage rate of the processor core in the current cycle is lower than the average usage rate in the uplink direction, then the processor core is a second-type processor core in the uplink direction.
[0120] In step S811, for any first-type processor core, the usage rate corresponding to each data flow in the target direction processed by the first-type processor core in the current cycle is determined according to the summary value of each data packet in the target direction processed by the first-type processor core in the current cycle.
[0121] In an optional implementation, for any first-class processor core whose utilization rate in the target direction of the current cycle is higher than the average utilization rate in the target direction, an embodiment of the present application can determine the utilization rate corresponding to each data flow in the target direction processed by the first-class processor core in the current cycle. Since the data flow is represented by the summary value of the data packet corresponding to the data flow, the embodiment of the present application can determine the utilization rate corresponding to each data flow in the target direction processed by the first-class processor core in the current cycle based on the summary value of each data packet in the target direction processed by the first-class processor core in the current cycle. In an implementation example, an embodiment of the present application can collect the summary value of each data packet in the target direction processed by the first-class processor core in the current cycle, thereby attributing data packets with the same summary value to the same data flow, obtaining the number of data packets in each data flow in the target direction processed by the first-class processor core in the current cycle, and then determining the utilization rate corresponding to each data flow in the target direction processed by the first-class processor core in the current cycle based on the number of data packets in each data flow in the target direction processed by the first-class processor core in the current cycle. For example, the number of data packets in each data flow in the target direction processed by the first type processor core in the current cycle is divided by the data packet load upper limit of the first type processor core in the current cycle to obtain the utilization rate corresponding to each data flow in the target direction processed by the first type processor core in the current cycle. For another example, the effective usage time corresponding to the number of data packets in each data flow in the target direction processed by the first type processor core in the current cycle is divided by the duration of the current cycle to obtain the utilization rate corresponding to each data flow in the target direction processed by the first type processor core in the current cycle.
[0122] For example, taking the target direction as the downlink direction, for any first-class processor core whose downlink utilization rate in the current cycle is higher than the average downlink utilization rate, embodiments of the present application can determine the utilization rate corresponding to each downlink data flow processed by the first-class processor core in the current cycle. Since a data flow is represented by the summary value of the data packet corresponding to the data flow, embodiments of the present application can determine the utilization rate corresponding to each downlink data flow processed by the first-class processor core in the current cycle based on the summary value of each downlink data packet processed by the first-class processor core in the current cycle.
[0123] For another example, taking the target direction as the uplink direction, for any first-class processor core whose uplink utilization rate in the current cycle is higher than the average uplink utilization rate, the embodiment of the present application can determine the utilization rate corresponding to each uplink data flow processed by the first-class processor core in the current cycle. Since the data flow is represented by the summary value of the data packet corresponding to the data flow, the embodiment of the present application can determine the utilization rate corresponding to each uplink data flow processed by the first-class processor core in the current cycle based on the summary value of each uplink data packet processed by the first-class processor core in the current cycle.
[0124] In step S812, for any first-class processor core, the diverted data flow of the first-class processor core in the target direction is determined based on the usage rate corresponding to each data flow in the target direction processed by the first-class processor core in the current cycle, wherein after the diverted data flow of the first-class processor core in the target direction is diverted, the usage rate of the first-class processor core in the target direction tends to the average usage rate of the target direction.
[0125] For any first-class processor core, the diverted data flow of the first-class processor core in the target direction can be regarded as the data flow that needs to be diverted from the first-class processor core to the second-class processor core in the target direction. In order to make the utilization rate of the first-class processor core in the target direction approach the average utilization rate of the target direction after the first-class processor core diverts the diverted data flow in the target direction to the second-class processor core, the embodiment of the present application can determine the diverted data flow that needs to be diverted from the first-class processor core in the target direction based on the utilization rate corresponding to each data flow in the target direction processed by the first-class processor core in the current cycle.
[0126] As an optional implementation, in order to avoid the situation where a single diverted data flow with a higher load in the target direction is diverted to the second type of processor core, and the second type of processor core exceeds the average utilization rate in the target direction, the embodiment of the present application can filter out data flows with non-zero utilization rates from the data flows in the target direction processed by the first type of processor core in the current cycle according to the utilization rates corresponding to the various data flows in the target direction processed by the first type of processor core in the current cycle (the utilization rate corresponding to the data flow is zero, indicating that the data flow is zero load on the first type of processor core and there is no diversion demand); thus, for the filtered data flows, the utilization rates of the data flows are accumulated in order of utilization from low to high until the accumulated utilization rate approaches the difference between the utilization rate of the first type of processor core in the target direction of the current cycle and the average utilization rate, and then the utilization rate accumulation is stopped; for example, when the accumulated utilization rate is closest to the difference, it is still a certain percentage away from the difference, which can also be regarded as the accumulated utilization rate approaching the difference, and the utilization rate accumulation can be stopped. The certain percentage of difference can be determined according to actual conditions, such as 20%.
[0127] Furthermore, the data flow corresponding to the accumulated utilization rate is used as the diverted data flow of the first type of processor core in the target direction, so that the utilization rate corresponding to the diverted data flow of the first type of processor core in the target direction approaches the difference between the utilization rate of the first type of processor core in the target direction of the current cycle and the average utilization rate, so as to ensure that after the diverted data flow of the target direction of the first type of processor core is diverted, the utilization rate of the first type of processor core in the target direction of the current cycle tends to the average utilization rate.
[0128] For example, taking the target direction as the downlink direction, an embodiment of the present application can filter data flows with non-zero usage rates based on the usage rates corresponding to each data flow in the downlink direction processed by the first type of processor core in the current cycle; thus, for the filtered data flows, the usage rates of the data flows are accumulated in order of usage rate from low to high until the accumulated usage rate approaches the difference between the downlink usage rate of the first type of processor core in the current cycle and the average usage rate (the average usage rate in the downlink direction), and the accumulation of usage rates is stopped; further, the data flow corresponding to the accumulated usage rate is used as the diverted data flow of the first type of processor core in the downlink direction, so that the usage rate corresponding to the diverted data flow of the first type of processor core in the downlink direction approaches the difference between the downlink usage rate of the first type of processor core in the current cycle and the average usage rate, so as to ensure that after the diverted data flow in the downlink direction of the first type of processor core is diverted, the usage rate of the first type of processor core in the downlink direction of the current cycle approaches the average usage rate (the average usage rate in the downlink direction).
[0129] For example, taking the target direction as the upstream direction, an embodiment of the present application can filter data streams with non-zero usage rates based on the usage rates corresponding to each data stream in the upstream direction processed by the first type of processor core in the current cycle; thus, for the filtered data streams, the usage rates of the data streams are accumulated in order of usage rate from low to high until the accumulated usage rate approaches the difference between the upstream usage rate of the first type of processor core in the current cycle and the average usage rate (the average usage rate in the upstream direction), and the accumulation of usage rates is stopped; further, the data stream corresponding to the accumulated usage rate is used as the diverted data stream of the first type of processor core in the upstream direction, so that the usage rate corresponding to the diverted data stream of the first type of processor core in the upstream direction approaches the difference between the upstream usage rate of the first type of processor core in the current cycle and the average usage rate, so as to ensure that after the diverted data stream in the upstream direction of the first type of processor core is diverted, the usage rate of the first type of processor core in the upstream direction of the current cycle tends to the average usage rate (the average usage rate in the upstream direction).
[0130] In step S813, the mapping relationship between the diverted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the first type of processor core in the target direction is released.
[0131] After confirming the diverted data flow of the first type of processor core in the target direction, the embodiment of the present application can release the mapping relationship between the diverted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the first type of processor core in the target direction, so that the diverted data flow of the first type of processor core in the target direction will no longer be allocated to the redirection queue corresponding to the first type of processor core in the target direction, that is, it will no longer be allocated to the first type of processor core for processing.
[0132] As an optional implementation, the embodiment of the present application can clear the relevant table entries in the mapping table (such as clearing the contents of the relevant table entries) to achieve the release of the mapping relationship between the diverted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the first type of processor core in the target direction. For example, the embodiment of the present application can query the corresponding table entry in the mapping table based on the summary value of the diverted data flow of the first type of processor core in the target direction and the identifier of the redirection queue corresponding to the first type of processor core in the target direction, thereby clearing the table entry queried in the mapping table to achieve the release of the mapping relationship between the diverted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the first type of processor core in the target direction; wherein the mapping table can have multiple table entries, and one table entry records the summary value of a data flow and the identifier of a redirection queue corresponding to the mapped target direction.
[0133] For example, taking the target direction as the downlink direction, the embodiment of the present application can release the downlink mapping relationship between the diverted data flow of the first type of processor core in the downlink direction and the downlink redirection queue corresponding to the first type of processor core, so that the diverted data flow of the first type of processor core in the downlink direction is no longer subsequently allocated to the downlink redirection queue corresponding to the first type of processor core. In one example, the embodiment of the present application can query the corresponding table entry in the downlink mapping table based on the summary value of the diverted data flow of the first type of processor core in the downlink direction and the identifier of the downlink redirection queue corresponding to the first type of processor core, thereby clearing the table entry queried in the downlink mapping table; wherein the downlink mapping table can have multiple table entries, and one table entry records the summary value of a data flow and the identifier of a corresponding mapped downlink redirection queue.
[0134] For example, taking the target direction as the uplink direction, the embodiment of the present application can release the uplink mapping relationship between the diverted data flow of the first type of processor core in the uplink direction and the uplink redirection queue corresponding to the first type of processor core, so that the diverted data flow of the first type of processor core in the uplink direction will no longer be allocated to the uplink redirection queue corresponding to the first type of processor core. In one example, the embodiment of the present application can query the corresponding table entry in the uplink mapping table based on the summary value of the diverted data flow of the first type of processor core in the uplink direction and the identifier of the uplink redirection queue corresponding to the first type of processor core, thereby clearing the table entry queried in the uplink mapping table; wherein, the uplink mapping table can have multiple table entries, and one table entry records the summary value of a data flow and the identifier of a corresponding mapped uplink redirection queue.
[0135] In step S814, the diverted data flow of the first type processor core in the target direction is diverted to at least one second type processor core, so that the utilization rate of the second type processor core to which the diverted data flow is diverted in the target direction approaches the average utilization rate in the target direction.
[0136] In step S815, a mapping relationship is added between the shunted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the shunted data flow of the second type of processor core in the target direction.
[0137] After confirming the diverted data flow of the first type of processor core in the target direction, the embodiment of the present application can divert the diverted data flow of the first type of processor core in the target direction to one or more second type of processor cores, and make the utilization rate of the second type of processor core diverted by the diverted data flow in the target direction tend to the average utilization rate of the target direction, and not more than the average utilization rate of the target direction. For example, taking the target direction as the downlink direction as an example, the diverted data flow of the first type of processor core in the downlink direction can be diverted to one or more second type of processor cores, and make the utilization rate of the second type of processor core diverted by the diverted data flow in the downlink direction tend to the average utilization rate of the downlink direction, and not more than the average utilization rate of the downlink direction. For example, taking the target direction as the uplink direction as an example, the diverted data flow of the first type of processor core in the uplink direction can be diverted to one or more second type of processor cores, and make the utilization rate of the second type of processor core diverted by the diverted data flow in the uplink direction tend to the average utilization rate of the uplink direction, and not more than the average utilization rate of the uplink direction.
[0138] For the diverted data flow of the first type of processor core in the target direction, and the second type of processor core to which the diverted data flow is diverted, the embodiment of the present application can add a mapping relationship between the diverted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the diverted second type of processor core in the target direction, so that the diverted data flow of the first type of processor core in the target direction can be subsequently allocated to the redirection queue corresponding to the second type of processor core in the target direction, that is, diverted to the second type of processor core for processing.
[0139] As an optional implementation, the embodiment of the present application can add relevant table entries in the mapping table to realize the mapping relationship between the newly added diverted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the diverted second type of processor core in the target direction. For example, the embodiment of the present application can add a corresponding table entry in the mapping table based on the summary value of the diverted data flow of the first type of processor core in the target direction and the identifier of the redirection queue corresponding to the diverted second type of processor core in the target direction, thereby recording through the newly added table entry: the mapping relationship between the summary value of the diverted data flow of the first type of processor core in the target direction and the identifier of the redirection queue corresponding to the diverted second type of processor core in the target direction.
[0140] For example, taking the target direction as the downlink direction, the embodiment of the present application can add relevant table entries in the mapping table to realize the mapping relationship between the newly added diverted data flow of the first type of processor core in the downlink direction and the downlink redirection queue corresponding to the diverted second type of processor core. For example, the embodiment of the present application can add a corresponding table entry in the mapping table based on the summary value of the diverted data flow of the first type of processor core in the uplink direction and the identifier of the downlink redirection queue corresponding to the diverted second type of processor core, thereby recording through the newly added table entry: the mapping relationship between the summary value of the diverted data flow of the first type of processor core in the downlink direction and the identifier of the downlink redirection queue corresponding to the diverted second type of processor core.
[0141] For example, taking the target direction as the uplink direction as an example, the embodiment of the present application can realize the mapping relationship between the newly added diverted data flow of the first type of processor core in the uplink direction and the uplink redirection queue corresponding to the diverted second type of processor core by adding relevant table entries in the mapping table. For example, the embodiment of the present application can add corresponding table entries in the mapping table based on the summary value of the diverted data flow of the first type of processor core in the uplink direction and the identifier of the uplink redirection queue corresponding to the diverted second type of processor core, thereby recording through the newly added table entries: the mapping relationship between the summary value of the diverted data flow of the first type of processor core in the uplink direction and the identifier of the uplink redirection queue corresponding to the diverted second type of processor core.
[0142] It can be seen that based on the process shown in Figure 8, the embodiment of the present application can adjust the mapping relationship in the downstream direction and the upstream direction respectively, so that the processor core with a higher utilization rate in the downstream direction can divert part of the data stream to the processor core with a lower utilization rate for processing, and the processor core with a higher utilization rate in the upstream direction can divert part of the data stream to the processor core with a lower utilization rate for processing, thereby realizing that multiple processor cores can adjust the load balancing in the downstream and upstream directions.
[0143] The embodiment of the present application utilizes the programmable hardware on the network card (such as the smart network card) so that a processor core corresponds to a redirection queue in the target direction to transmit data packets from the network card queue in the target direction. For example, a processor core corresponds to a downstream redirection queue in the downstream direction to transmit data packets from the virtual network card queue in the downstream direction, and a processor core corresponds to an upstream redirection queue in the upstream direction to transmit data packets from the virtual network card queue in the upstream direction. At the same time, through the mapping relationship between the data flow and the redirection queue in the target direction, the redirection queue in the target direction is assigned to the data packet from the network card queue in the target direction, and the data flow in the mapping relationship is represented by the summary value of the data packet corresponding to the data flow. Therefore, the embodiment of the present application can realize the data flow level data packet redirection capability, thereby combining the method of one processor core corresponding to a redirection queue in the target direction. The processor core only needs to poll the corresponding redirection queue in the target direction without polling the network card queue in the target direction, reducing the number of queues that the processor core needs to poll on the software forwarding side, reducing overhead, and being able to adapt to the expansion of the number of virtual machines, thereby improving scalability.
[0144] Furthermore, the load adjustment module set in the virtual switch implemented by multiple processor cores can periodically monitor the data packet load in the target direction processed by each processor core, and thus periodically adjust the mapping relationship (for example, periodically overwrite the mapping table in the programmable hardware) so that the data packet load in the target direction processed by each processor core can be balanced, thereby realizing load balancing among multiple processor cores in a network card (for example, a smart network card).
[0145] Specifically, the number of downlink redirection queues and uplink redirection queues of the processor core is kept consistent with the number of processor cores, thereby reducing the overhead of the processor core polling queue; and, data packet allocation scheduling at the data flow granularity is implemented, and the processor core carrying the data flow is switched at the data flow level, thereby avoiding the problem of high load on a single queue, which causes packet loss due to resource exhaustion of the processor core bound to the queue; and, there is no need to adjust the binding relationship between the virtual network card queue and the processor core, only the mapping relationship between the data flow and the redirection queue (downlink redirection queue and uplink redirection queue) corresponding to the processor core in the target direction on the network card (such as the smart network card) hardware is modified. Therefore, there will be no potential problems such as stagnation of the virtual network card queue due to modification of the binding relationship between the virtual network card queue and the processor core, or the processor core being unable to receive packets from the virtual network card queue, thereby improving the robustness of the system; and, since the overhead of the processor core is for the corresponding downlink redirection queue and uplink redirection queue, the overhead does not increase with the increase of virtual machine density (as the virtual machine density increases, the number of virtual network card queues increases accordingly), which can support high-density deployment scenarios of virtual machines and improve scalability.
[0146] It can be seen that the embodiments of the present application can improve the refinement of load balancing (i.e., the granularity of load balancing is finer), reduce overhead, and support scalability, thereby improving the performance of load balancing. In other words, the embodiments of the present application not only reduce the overhead of polling queues of processor cores of network cards (such as smart network cards) in high-density deployment scenarios of virtual machines, but also make the load balancing adjustment method more fine-grained and lightweight, thereby improving the performance of load balancing when implementing load balancing among multiple processor cores in network cards (such as smart network cards).
[0147] The present application also provides a network card (e.g., a smart network card), as shown in FIG3 . The network card (e.g., a smart network card) may include a hardware acceleration engine and an on-chip processor; the hardware acceleration engine is provided with a redirection module, multiple downlink redirection queues, and multiple uplink redirection queues; the on-chip processor is provided with multiple processor cores, the multiple processor cores implement a virtual switch, and the virtual switch is provided with a load adjustment module;
[0148] The redirection module is connected between multiple downstream redirection queues and virtual network card queues, and between multiple upstream redirection queues and physical network card queues. One processor core corresponds to one downstream redirection queue to transmit data packets from the virtual network card queue, and one processor core corresponds to one upstream redirection queue to transmit data packets from the physical network card queue.
[0149] In an embodiment of the present application, the redirection module is configured to execute the load balancing method performed by the redirection module as provided in the embodiment of the present application; the load adjustment module is configured to execute the load balancing method performed by the load adjustment module as provided in the embodiment of the present application.
[0150] The present application also provides a storage medium storing one or more computer-executable instructions. When executed, the one or more computer-executable instructions implement the load balancing method performed by the redirection module as provided in the present application, or the load balancing method performed by the load adjustment module as provided in the present application. As an optional implementation, the one or more computer-executable instructions may be instructions in programmable software of the FPGA hardware of a network card (e.g., a smart network card), or instructions in software executed by a processor core in an on-chip processor of the network card (e.g., a smart network card).
[0151] The above describes multiple embodiment schemes provided by the embodiments of the present application. The various optional methods introduced in each embodiment scheme can be combined and cross-referenced with each other without conflict, thereby extending a variety of possible embodiment schemes, which can all be considered as embodiment schemes disclosed and open in the embodiments of the present application.
[0152] Although the embodiments of the present application are disclosed above, the present application is not limited thereto. Any person skilled in the art may make various changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims.
Claims
1. A load balancing method, applied to a network card, comprising: Acquire a data packet, wherein the data packet comes from a network card queue in a target direction; Determine a summary value of the data packet, wherein the summary value of the data packet represents a data flow corresponding to the data packet; Determine the redirection queue of the target direction to which the summary value of the data packet is mapped according to the mapping relationship between the data flow and the redirection queue of the target direction, wherein the data flow in the mapping relationship is represented by the summary value of the data packet corresponding to the data flow; The data packet is allocated to the redirection queue of the determined target direction so that the data packet is processed by the processor core corresponding to the redirection queue of the determined target direction; wherein the multiple processor cores of the network card implement a virtual switch, and one processor core corresponds to a redirection queue in the target direction to transmit the data packet from the network card queue in the target direction, and the mapping relationship is periodically adjusted with the goal of balancing the load of the data packets in the target direction processed by each processor core.
2. The method according to claim 1, wherein: The step of determining the redirection queue of the target direction corresponding to the summary value of the data packet according to the mapping relationship between the data flow and the redirection queue of the target direction includes: Based on the summary value of the data packet, searching for an entry with the summary value of the data packet as an index in multiple entries of a mapping table; Determine the content of the queried table entry, and obtain the identifier of the redirection queue of the target direction mapped to the summary value of the data packet; wherein the mapping table records the mapping relationship between the data flow and the redirection queue of the target direction through multiple table entries, the index of an entry is the summary value of the data packet corresponding to a data flow, and the content of an entry is the identifier of a redirection queue corresponding to the mapped target direction.
3. The method according to claim 1 or 2, wherein: The target direction is divided into a downlink direction and an uplink direction of the network card, the downlink direction is the direction in which the network card receives data packets from the virtual network card queue, and the uplink direction is the direction in which the network card receives data packets from the physical network card queue; The redirection queue in the target direction is divided into a downlink redirection queue and an uplink redirection queue. One processor core corresponds to one downlink redirection queue to transmit data packets from the virtual network card queue, and one processor core corresponds to one uplink redirection queue to transmit data packets from the physical network card queue.
4. The method according to claim 3, wherein: The target direction is a downlink direction, the data packet comes from a virtual network card queue, and the mapping relationship between the data flow and the redirection queue of the target direction includes: a downlink mapping relationship between the data flow and the downlink redirection queue; wherein the downlink mapping relationship is recorded in a downlink mapping table, and the downlink mapping table records the mapping relationship between the data flow and the downlink redirection queue through multiple table entries, the index of one table entry is a summary value of a data packet corresponding to one data flow, and the content of one table entry is an identifier of a corresponding mapped downlink redirection queue; Alternatively, the target direction is the uplink direction, the data packet comes from the physical network card queue, and the mapping relationship between the data flow and the redirection queue of the target direction includes: an uplink mapping relationship between the data flow and the uplink redirection queue; The mapping relationship is recorded in the uplink mapping table. The uplink mapping table records the mapping relationship between the data flow and the uplink redirection queue through multiple table entries. The index of an entry is the summary value of the data packet corresponding to a data flow, and the content of an entry is the identifier of a corresponding mapped uplink redirection queue.
5. A load balancing method, applied to a network card, the method comprising: Determine the total data packet load of each processor core in the current cycle, wherein the total data packet load of a processor core in the current cycle is composed of the data packet loads of each target direction of the processor core in the current cycle; According to the total data packet load of each processor core in the current cycle, it is judged whether the current cycle reaches the adjustment condition of adjusting the mapping relationship, wherein the mapping relationship includes: a mapping relationship between a data flow and a redirection queue in a target direction, wherein the data flow in the mapping relationship is represented by a summary value of a data packet corresponding to the data flow; wherein the multiple processor cores of the network card implement a virtual switch, and one processor core corresponds to one redirection queue in the target direction to transmit data packets from the network card queue in the target direction; If the current cycle meets the adjustment condition for adjusting the mapping relationship, the mapping relationship between the data flow and the redirection queue in the target direction is adjusted with the goal of balancing the load of the data packets in the target direction processed by each processor core.
6. The method according to claim 5, wherein: The determining, based on the total data packet load of each processor core in the current cycle, whether the current cycle meets the adjustment condition for adjusting the mapping relationship includes: Determine the utilization rate of each processor core in the current cycle according to the data packet load of each processor core in the current cycle; According to the utilization rate of each processor core in the current cycle, determining whether the current cycle meets the adjustment condition for adjusting the mapping relationship; The adjustment conditions include: the number of processor cores whose utilization rate exceeds the utilization rate threshold is greater than a set number, and the utilization rate difference between the processor core with the highest utilization rate and the processor core with the lowest utilization rate exceeds a preset difference.
7. The method according to claim 5, wherein: The step of adjusting the mapping relationship between the data flow and the redirection queue in the target direction with the goal of balancing the load of the data packets in the target direction processed by each processor core includes: For any target direction, the average utilization rate of the target direction is determined according to the utilization rate of each processor core in the target direction of the current cycle; wherein any processor core whose utilization rate of the target direction of the current cycle is higher than the average utilization rate is a first-category processor core, and any processor core whose utilization rate of the target direction of the current cycle is lower than the average utilization rate is a second-category processor core; Adjust the mapping relationship between the data flow and the redirection queue in the target direction to divert part of the data flow of the first type of processor core in the target direction to the second type of processor core for processing, so that the utilization rate of the first type of processor core and the second type of processor core in the target direction tends to the average utilization rate.
8. The method according to claim 7, wherein: The step of adjusting the mapping relationship between the data flow and the redirection queue in the target direction to divert part of the data flow of the first type of processor core in the target direction to the second type of processor core for processing, so that the utilization rates of the first type of processor core and the second type of processor core in the target direction tend to an average utilization rate includes: For any first-class processor core, according to the usage rate corresponding to each data flow in the target direction processed by the first-class processor core in the current cycle, the shunted data flow of the first-class processor core in the target direction is determined, wherein after the shunted data flow of the first-class processor core in the target direction is shunted, the usage rate of the first-class processor core in the target direction tends to the average usage rate of the target direction; Release the mapping relationship between the diverted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the first type of processor core in the target direction; And, diverting the diverted data flow of the first type of processor core in the target direction to at least one second type of processor core, so that the utilization rate of the second type of processor core to which the diverted data flow is diverted in the target direction tends to the average utilization rate in the target direction; A mapping relationship between the shunted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the shunted second type of processor core in the target direction is added.
9. The method according to claim 8, wherein: The mapping relationship between the release of the diverted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the first type of processor core in the target direction includes: Based on the summary value of the diverted data flow of the first type of processor core in the target direction and the identifier of the redirection queue corresponding to the first type of processor core in the target direction, query the corresponding table entry in the mapping table; Clear the table entries found in the mapping table; wherein the mapping table records the mapping relationship between the data flow and the redirection queue of the target direction through multiple table entries, the index of an entry is the summary value of the data packet corresponding to a data flow, and the content of an entry is the identifier of a redirection queue corresponding to the mapped target direction; The mapping relationship between the newly added shunted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the shunted second type of processor core in the target direction includes: Based on the summary value of the shunted data flow of the first type of processor core in the target direction and the identifier of the redirection queue corresponding to the shunted second type of processor core in the target direction, a corresponding table entry is added to the mapping table.
10. The method according to claim 8, wherein: The determining, for any first-class processor core, according to the usage rate corresponding to each data flow in the target direction processed by the first-class processor core in the current cycle, of the diverted data flow of the first-class processor core in the target direction comprises: For any first-category processor core, according to the usage rates corresponding to the data flows in the target direction processed by the first-category processor core in the current cycle, filter the data flows with non-zero usage rates; For the filtered data streams, the usage rates of the data streams are accumulated in order of usage rate from low to high, until the accumulated usage rate approaches the difference between the usage rate of the first type of processor core in the target direction of the current cycle and the average usage rate, and then the usage rate accumulation is stopped; The data flow corresponding to the accumulated usage rate is used as the diversion data flow of the first type of processor core in the target direction.
11. The method according to any one of claims 5 to 10, wherein: The target direction is divided into a downlink direction and an uplink direction of the network card, the downlink direction is the direction in which the network card receives data packets from the virtual network card queue, and the uplink direction is the direction in which the network card receives data packets from the physical network card queue; The redirection queue in the target direction is divided into a downlink redirection queue and an uplink redirection queue. One processor core corresponds to one downlink redirection queue to transmit data packets from the virtual network card queue, and one processor core corresponds to one uplink redirection queue to transmit data packets from the physical network card queue.
12. The method according to claim 11, wherein: The mapping relationship between the data flow and the redirection queue in the target direction is divided into a downlink mapping relationship between the data flow and the downlink redirection queue, and an uplink mapping relationship between the data flow and the uplink redirection queue; The downlink mapping relationship is recorded in a downlink mapping table, which records the mapping relationship between the data flow and the downlink redirection queue through multiple table entries, the index of an entry is a summary value of a data packet corresponding to a data flow, and the content of an entry is an identifier of a corresponding mapped downlink redirection queue; The uplink mapping relationship is recorded in an uplink mapping table, which records the mapping relationship between data streams and uplink redirection queues through multiple table entries. The index of an entry is a summary value of a data packet corresponding to a data stream, and the content of an entry is an identifier of a correspondingly mapped uplink redirection queue.
13. A network card, comprising: Hardware acceleration engine and on-chip processor; The hardware acceleration engine is provided with a redirection module, a plurality of downlink redirection queues, and a plurality of uplink redirection queues; The on-chip processor is provided with a plurality of processor cores, the plurality of processor cores implement a virtual switch, and the virtual switch is provided with a load adjustment module; The redirection module is connected between multiple downstream redirection queues and virtual network card queues, and between multiple upstream redirection queues and physical network card queues; one processor core corresponds to one downstream redirection queue to transmit data packets from the virtual network card queue, and one processor core corresponds to one upstream redirection queue to transmit data packets from the physical network card queue; The redirection module is configured to execute the load balancing method according to any one of claims 1 to 4; The load adjustment module is configured to execute the load balancing method as described in any one of claims 5-12.
14. A storage medium storing one or more computer executable instructions, wherein when the one or more computer executable instructions are executed, the load balancing method according to any one of claims 1 to 4 or the load balancing method according to any one of claims 5 to 12 is implemented.
Citation Information
Patent Citations
Network load balancing method and system
CN106533978A
Message processing method and apparatus, and device
CN108632165A
Load balancing method and device and network card
CN114553780A
System and method for operation management of autonomous driving golf cart in wireless power transmission environment
KR102189723B1