Load balancing method, network card and storage medium

In the cloud computing environment, the processing core load balancing of the virtual switch after being offloaded to the network card is achieved, which solves the problem of load imbalance between multiple processor cores and improves load balancing performance and scalability.

CN120075163APending Publication Date: 2025-05-30HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311561019.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In a cloud computing environment, when a virtual switch is offloaded to a network card, how to achieve load balancing between multiple processor cores and improve load balancing performance.

Method used

By obtaining the summary value of the data packet, determining its corresponding data stream, and assigning the data packet to the corresponding redirect queue according to the mapping relationship between the data stream and the redirect queue in the target direction, and then transmitting it to the corresponding processor core for processing. At the same time, the mapping relationship is periodically adjusted to ensure that the load of each processor core tends to be balanced.

Benefits of technology

It realizes load balancing between multiple processor cores in the network card, improves the degree of refinement and scalability of load balancing, reduces overhead, and improves overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075163A_ABST
    Figure CN120075163A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a load balancing method, a network card and a storage medium, and the method comprises the steps: obtaining a data packet which is from a network card queue in a target direction; determining an abstract value of the data packet; according to a mapping relationship between the data stream and the redirection queue in the target direction, determining the redirection queue in the target direction correspondingly mapped by the digest value of the data packet, the data stream in the mapping relationship being represented by the digest value of the data packet corresponding to the data stream; distributing the data packet to a redirection queue in the determined target direction so as to process the data packet through a corresponding processor core; wherein one processor core of the network card corresponds to one redirection queue in the target direction so as to transmit data packets from the network card queue in the target direction, and the mapping relation is periodically adjusted with the goal that the load of the data packets, processed by the processor cores, in the target direction tends to be balanced. According to the embodiment of the invention, load balancing is realized, and the load balancing performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of cloud computing technology, and in particular, to a load balancing method, a network card, and a storage medium. Background Art

[0002] With the development of cloud computing and virtualization technologies, physical servers deployed in the cloud can virtualize multiple virtual machines (VMs) through virtualization technology to improve the resource utilization rate of physical servers. To cope with the continuously increasing network bandwidth and support virtualization functions at a relatively low cost, virtualization functions such as network virtualization can be offloaded to a network card (such as a smart network card) that communicates with the physical server. For example, a virtual switch (vSwitch) running on a physical server can be offloaded to the network card to achieve high-performance forwarding of data packets.

[0003] In the case of offloading the virtual switch to the network card, multiple processor cores (such as CPU cores) for implementing the virtual switch are provided in the network card. In this context, how to achieve load balancing among multiple processor cores has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0004] In view of this, embodiments of the present application provide a load balancing method, a network card, and a storage medium to achieve load balancing among multiple processor cores in the network card and improve the performance of load balancing.

[0005] To achieve the above object, embodiments of the present application provide the following technical solutions.

[0006] In a first aspect, embodiments of the present application provide a load balancing method applied to a network card. The method includes:

[0007] Obtain a data packet, where the data packet comes from a network card queue in a target direction;

[0008] Determine a digest value of the data packet, where the digest value of the data packet represents the data stream corresponding to the data packet;

[0009] According to the mapping relationship between the data stream and a redirection queue in the target direction, determine the redirection queue in the target direction mapped by the digest value of the data packet, where the data stream in the mapping relationship is represented by the digest value of the data packet corresponding to the data stream;

[0010] Allocate the data packet to the redirection queue in the determined target direction, so that the data packet is processed by the processor core corresponding to the redirection queue in the determined target direction; wherein, multiple processor cores of the network card implement a virtual switch, and one processor core corresponds to one redirection queue in the target direction to transmit data packets from the network card queue in the target direction, and the mapping relationship is periodically adjusted with the aim of balancing the data packet loads in the target directions processed by each processor core.

[0011] In a second aspect, an embodiment of the present application provides a load balancing method applied to a network card. The method includes:

[0012] Determine the total data packet load of each processor core in the current cycle, where the total data packet load of a processor core in the current cycle consists of the data packet loads in each target direction of the processor core in the current cycle;

[0013] According to the total data packet loads of each processor core in the current cycle, determine whether the adjustment condition for adjusting the mapping relationship is met in the current cycle. The mapping relationship includes: the mapping relationship between the data flow and the redirection queue in the target direction, and the data flow in the mapping relationship is represented by the digest value of the data packet corresponding to the data flow; wherein, multiple processor cores of the network card implement a virtual switch, and one processor core corresponds to one redirection queue in the target direction to transmit data packets from the network card queue in the target direction;

[0014] If the adjustment condition for adjusting the mapping relationship is met in the current cycle, then adjust the mapping relationship between the data flow and the redirection queue in the target direction with the aim of balancing the data packet loads in the target directions processed by each processor core.

[0015] In a third aspect, an embodiment of the present application provides a network card, including: a hardware acceleration engine and an on-chip processor; the hardware acceleration engine is provided with a redirection module, multiple downstream redirection queues, and multiple upstream redirection queues; the on-chip processor is provided with multiple processor cores, the multiple processor cores implement a virtual switch, and the virtual switch is provided with a load adjustment module;

[0016] Wherein, the redirection module is connected between the multiple downstream redirection queues and the virtual network card queue, and between the multiple upstream redirection queues and the physical network card queue; one processor core corresponds to one downstream redirection queue to transmit data packets from the virtual network card queue, and one processor core corresponds to one upstream redirection queue to transmit data packets from the physical network card queue;

[0017] The redirection module is configured to execute the load balancing method described in the first aspect above; the load adjustment module is configured to execute the load balancing method described in the second aspect above.

[0018] In a fourth aspect, an embodiment of the present application provides a storage medium storing one or more computer-executable instructions, which, when executed, implement the load balancing method as described in the first aspect above, or the load balancing method as described in the second aspect above.

[0019] In an embodiment of the present application, a processor core corresponds to a redirection queue in a target direction to transmit data packets from a network card queue in the target direction; meanwhile, through the mapping relationship between the data stream and the redirection queue in the target direction, the embodiment of the present application can allocate the redirection queue in the target direction to the data packets from the network card queue in the target direction, so that the data packets are transmitted to the corresponding processor core for processing through the allocated redirection queue in the target direction, where the data stream in the mapping relationship is represented by the digest value of the data packet corresponding to the data stream. It can be seen that the embodiment of the present application can implement the data packet redirection ability at the data stream level, and combined with the method that a processor core corresponds to a redirection queue in a target direction, the processor core only needs to poll the corresponding redirection queue in the target direction, instead of polling the network card queue in the target direction, reducing the number of queues that the processor core needs to poll in the software forwarding plane, reducing the overhead, being able to adapt to the number expansion situation of virtual machines, and thus improving the scalability.

[0020] Meanwhile, the mapping relationship between the data stream and the redirection queue in the target direction is adjusted periodically, and the goal of adjusting the mapping relationship is to make the data packet loads in the target directions processed by each processor core tend to be balanced. Therefore, the embodiment of the present application allocates the redirection queue in the target direction to the data stream through the mapping relationship between the data stream and the redirection queue in the target direction, and then transmits it to the corresponding processor core for processing, which can achieve the purpose of load balancing among multiple processor cores in the network card.

[0021] It can be seen that the embodiment of the present application realizes the load balancing among multiple processor cores in the network card, improves the refinement degree of load balancing (adjusting load balancing at the data stream level), reduces the overhead, and improves the scalability, thus improving the performance of load balancing. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0023] Figure 1It is an example diagram of the relationship between a network card and a physical server.

[0024] Figure 2 It is an example diagram of the connection between a virtual machine and a virtual switch.

[0025] Figure 3 It is an example diagram of the network card provided by the embodiment of the present application.

[0026] Figure 4 It is a flowchart of the load balancing method provided by the embodiment of the present application.

[0027] Figure 5 It is an example diagram of the downlink mapping table provided by the embodiment of the present application.

[0028] Figure 6 It is an example diagram of the uplink mapping table provided by the embodiment of the present application.

[0029] Figure 7 It is another flowchart of the load balancing method provided by the embodiment of the present application.

[0030] Figure 8 It is a flowchart of adjusting the mapping relationship provided by the embodiment of the present application. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0032] As a network switch, a virtual switch can be used to manage the network traffic between virtual machines and between virtual machines and the physical network. A main function of the virtual switch is to be responsible for forwarding data packets; for example, the virtual switch can forward data packets based on forwarding rules to ensure that the data packets can be correctly routed; among them, the data packets can be divided into data packets sent by virtual machines and data packets sent by the physical network.

[0033] The virtual switch is usually implemented by software running on a physical server. For example, a general-purpose processor (such as a CPU) in the physical server can implement the virtual switch through software functions. However, with the continuous development of cloud computing and network technologies, users have higher and higher requirements for network performance, and the data packet forwarding performance of the virtual switch running on the physical server can no longer meet the performance requirements; therefore, in order to achieve high-performance forwarding of data packets of virtual machines, the virtual switch running on the physical server can be offloaded to a network card (such as a smart network card) to improve the data packet forwarding performance.

[0034] It should be noted that in the application of cloud computing, a network card (such as a smart network card) is a network adapter that communicates with a physical server. Taking the smart network card as an example, the smart network card can combine network connection functions with some intelligent functions to improve the network performance, management performance, and security of cloud computing. For example, in addition to being able to complete the network transmission functions of a standard network card, the smart network card can also use a built-in programmable and configurable hardware acceleration engine to improve application performance. In an application example, a network card (such as a smart network card) can be used in the virtualization environment of cloud computing, such as a cloud computing platform and a data center, to support the network requirements of virtual machines.

[0035] Offload is short for Hardware Offload, which is the process of offloading specific tasks from a general-purpose processor (such as a CPU) to a dedicated hardware device for processing, and can accelerate the execution of specific tasks. Correspondingly, offloading the virtual switch running on a physical server to a network card (such as a smart network card) can be regarded as using the network card (such as a smart network card) as a dedicated hardware device to implement the virtual switch function, so that the general-purpose processor of the physical processor no longer implements the virtual switch function, but is replaced by the network card (such as a smart network card) to implement the virtual switch function.

[0036] For ease of understanding, in the case of offloading the virtual switch to a network card (such as a smart network card), Figure 1 an exemplary relationship diagram of the network card and the physical server is shown, as Figure 1 shown, the physical server 110 runs multiple virtual machines 111 to 11n (n is the number of virtual machines, depending on the actual situation) through virtualization technology to meet the computing requirements of cloud computing; the virtual switch originally running on the physical server 110 is offloaded to the network card 120 (the network card 120 is, for example, a smart network card), and the network card 120 bears the function of the virtual switch through the hardware acceleration engine 121 and the on-chip processor 122.

[0037] The network card 120 can be a network card with a hardware acceleration engine 121 and an on-chip processor 122. In the case of offloading the virtual switch to the network card, the hardware acceleration engine 121 is responsible for the hardware forwarding part of the packet forwarding of the virtual switch. For example, the hardware acceleration engine 121 can set some forwarding rules to accelerate the packet forwarding in a hardware acceleration manner. In an example, the hardware acceleration engine 121 can set a forwarding rule table, and the forwarding rule table can be used to record the packet forwarding rules, so as to achieve accelerated forwarding of packets at the hardware level and improve the packet forwarding performance.

[0038] The hardware acceleration engine 121 can be, for example, FPGA (Field Programmable Gate Array) hardware. Through the programmability of the FPGA hardware, the network card (such as a smart network card) can be made to adapt to the ever-changing network requirements and security challenges, thereby enhancing the flexibility of the network card (such as a smart network card) for actual application scenarios.

[0039] The on-chip processor 122 is a processor set in the network card (such as a smart network card), such as a CPU. In the case where the virtual switch is offloaded to the network card (such as a smart network card), the on-chip processor 122 implements the function of the virtual switch in software; for example, the on-chip processor 122 can carry the function of the virtual switch by running software, and it is the software forwarding part of the virtual switch responsible for packet forwarding. That is to say, the complete virtual switch function can be implemented on the on-chip processor of the network card (such as a smart network card) and realized by the on-chip processor running software.

[0040] It should be noted that the network card 120 can communicate with the physical server 110. For example, the network card 120 can communicate with the physical server 110 through the PCIE (Peripheral Component Interconnect Express) bus. The network card 120 can also communicate with the physical network so that the physical server 110 can access the physical network. For example, the network card 120 can be connected to the physical devices of the physical network (such as network switches, routers, firewalls, etc.) through communication connection methods such as the PCIE bus to achieve access to the physical network. It should be noted that one side of the network card (such as a smart network card) is connected to the physical server, and the other side is connected to the physical network to meet the external communication needs of the physical server, and the specific form and type of the physical devices of the physical network to which the network card (such as a smart network card) is connected can be determined according to the actual application situation.

[0041] In the case where the virtual switch is offloaded to the network card, the virtual machines running on the physical server can be connected to the virtual switch through the virtual network card queue, so that the virtual machines and the virtual switch can interact through the virtual network card queue. At the same time, on the network card side, the on-chip processor of the network card can implement the virtual switch through multiple processor cores (such as multiple CPU cores), so that the virtual switch implemented by the multiple processor cores can interact with the virtual machines through the virtual network card queue to achieve the forwarding of the virtual machine's packets. That is to say, the virtual machine is connected to the multiple processor cores (the multiple processor cores implementing the virtual switch function) of the on-chip processor of the network card through the virtual network card queue, thereby achieving packet forwarding.

[0042] For ease of understanding, Figure 2 An exemplary connection diagram of a virtual machine and a virtual switch is shown. In combination with Figure 1 and Figure 2 As shown, the physical server runs multiple virtual machines 111 to 11n. Multiple processor cores 211 to 21m (m is the number of processor cores, depending on the actual situation) are set in the on-chip processor of the network card to implement the virtual switch. Each virtual machine is connected to the processor core through a corresponding virtual network card queue to forward the data packets of the virtual machine through the processor core. The virtual network card queue corresponding to a virtual machine can be one or more, without limitation. At the same time, each processor core can be connected to a corresponding physical network card queue to interface with the physical network, so as to process the data packets transmitted from the physical network interface.

[0043] It should be noted that the virtual network card queue is the network card queue of the virtual machine, which is used to send data packets from the virtual machine to the virtual switch (such as the processor core that implements the virtual switch). Each virtual machine can have one or more virtual network card queues, so that when the virtual machine has multiple virtual network card queues, the virtual machine is allowed to send data packets in parallel. The physical network card queue is the hardware queue of the network card (such as a smart network card), which is used to process the data packets of the physical network interface. For example, the processor core can process the data packets of the physical network interface through the corresponding physical network card queue. The physical network card queue allows the network card (such as a smart network card) to process multiple network connections and data packets at the same time, thereby improving the performance of the network card (such as a smart network card).

[0044] The combined use of the virtual network card queue and the physical network card queue can improve network performance, especially in a virtualized environment; for example, the virtual machine can communicate with multiple physical network card queues on the network card (such as a smart network card) through the virtual network card queue of the virtual machine, so as to realize the communication between the virtual machine and the physical network, and thus effectively utilize network resources.

[0045] It can be seen that on the network card (such as a smart network card) side, the on-chip processor realizes the forwarding of data packets through multiple processor cores, but the number of processor cores used for forwarding data packets is limited. Therefore, the processor core needs to use a certain mapping relationship to realize the forwarding of data packets. For example, the processor core needs to use the mapping relationship between the data stream and the virtual network card queue and the mapping relationship between the processor core and the virtual network card queue to allocate the data packets of the data stream to the processor core mapped by the mapped virtual network card queue for forwarding.

[0046] That is to say, by establishing a mapping relationship between data streams and virtual network card queues, data packets of the same data stream are routed to the virtual network card queues mapped to the data stream, and then based on the established mapping relationship between the virtual network card queues and the processor cores, the data packets can be forwarded by the processor core mapped to the virtual network card queues to which the data packets need to be routed, thereby achieving forwarding of data packets of the same data stream with the same processing logic. It should be noted that a data stream is a collection of data packets, and if data packet forwarding is processed based on the flow level, it means that the processing logic of data packets with the same five-tuple is the same. The five-tuple can be the source IP address, destination IP address, source port number, destination port number and protocol number of the data packet.

[0047] Forwarding data packets through the mapping relationship between data streams and virtual network card queues, and the mapping relationship between processor cores and virtual network card queues, may cause load imbalance between multiple processor cores. For example, multiple virtual network card queues or multiple data streams with high loads may eventually be mapped to the same processor core, which results in a high data packet load on the processor core, which is prone to packet loss due to resource exhaustion, while at the same time, there may be other idle processor cores that are not effectively used.

[0048] In order to solve the problem of packet loss in processor cores with higher loads caused by unbalanced loads between the above-mentioned processor cores, the inventors of the present application considered fine-grained division and distribution of the load of each processor core, and considered the following two ideas.

[0049] The first method is to adjust the mapping relationship between data streams and virtual network card queues to achieve load balancing of processor cores. For example, dynamically adjust the mapped virtual network card queues for multiple data streams, and ensure that the adjustment result is that the number of data streams that need to be processed by multiple processor cores is balanced; that is, based on the mapping relationship between virtual network card queues and processor cores, and the changes in the number of data streams mapped by each virtual network card queue, adjust the virtual network card queues mapped by the data streams so that the number of data streams mapped by virtual network card queues mapped by different processor cores is the same, so that the number of data streams that need to be processed by each processor core is balanced.

[0050] The first method mentioned above requires the processor core to calculate the data flow when dynamically changing the mapping relationship between the data flow and the virtual network card queue, and to evenly distribute multiple data flows to different processor cores through communication between processor cores, so that multiple data flows can be distributed to different processor cores in a balanced manner; the above process involves a large overhead of multiple processor cores, resulting in a large overhead of load balancing.

[0051] Second, dynamically adjust the mapping relationship between the virtual network card queues and the processor cores to achieve load balancing of the processor cores. However, this method has the problem of low refinement degree of load balancing, that is, the load balancing granularity is relatively coarse. For example, if the data flow load of a virtual network card queue is large, even if the processor core mapped by the virtual network card queue is adjusted, it will still cause the processor core finally mapped by the virtual network card queue to need to process a large-load data flow, resulting in the problem that the processor core finally mapped by the virtual network card queue may run out of resources and lose packets, leading to the failure of load balancing.

[0052] In addition, adjusting the mapping relationship between the virtual network card queue and the processor core is an invasive behavior, which is likely to cause some potential problems, especially that the packet forwarding cannot be correctly executed. That is to say, due to the complexity of the hardware and the operating system, if the mapping relationship between the virtual network card queue and the processor core is adjusted randomly, a relatively high failure probability may occur, such as potential problems like the virtual network card queue stagnation or the processor core being unable to receive packets from the virtual network card queue. Among them, the virtual network card queue stagnation means that the virtual network card queue stops normal operation and cannot continue to receive or send packets.

[0053] Furthermore, with the high-density deployment of virtual machines, the above first and second methods are also difficult to scale. For example, due to the development of cloud computing, the density of virtual machines deployed on physical servers increases. Correspondingly, the number of virtual network card queues also increases significantly. If it is necessary to adjust the mapping relationship between the virtual network card queue and the processor core, or the mapping relationship between the virtual network card queue and the data flow, then the processor core needs to frequently switch the mapping relationship among a large number of virtual network card queues, resulting in additional overhead. At the same time, under the large-scale deployment of virtual machines, it is also relatively difficult to implement the statistics and scheduling of virtual machine network card queues.

[0054] It can be seen that the above load balancing methods have problems such as large overhead, low refinement degree, relatively high failure probability, and difficulty in scaling. Therefore, the performance of load balancing of the above methods needs to be improved.

[0055] Based on this, the embodiments of the present application provide an improved load balancing scheme to achieve load balancing among multiple processor cores in a network card (such as a smart network card), and can improve the performance of load balancing. As an optional implementation, Figure 3 Exemplarily shows an example diagram of the network card provided by the embodiments of the present application, combined with Figure 1 、 Figure 2 and Figure 3As shown in the figure, in the embodiment of the present application, a redirection module 310, multiple downstream redirection queues 321 to 32m (m is the number of downstream redirection queues, which is kept consistent with the number of multiple processor cores), and multiple upstream redirection queues 331 to 33m (m is the number of upstream redirection queues, which is kept consistent with the number of multiple processor cores) are provided in the hardware acceleration engine 121 of a network card (such as a smart network card). Multiple processor cores 211 to 21m provided in the on-chip processor 122 of the network card (such as a smart network card) can implement a virtual switch, and a load adjustment module 340 is provided in the virtual switch in the embodiment of the present application.

[0056] Among them, the redirection module 310, the downstream redirection queues 321 to 32m, and the upstream redirection queues 331 to 33m can be implemented by the hardware acceleration engine. For example, the redirection module 310, the downstream redirection queues 321 to 32m, and the upstream redirection queues 331 to 33m are implemented through the programmable capabilities of FPGA hardware.

[0057] On the one hand, the redirection module is connected between the virtual network card queue and the downstream redirection queue, and can, according to the packet distribution strategy provided by the embodiment of the present application, distribute the packets received from the virtual network card queue to the downstream redirection queue. On the other hand, the redirection module is connected between the physical network card queue and the upstream redirection queue, and can, according to the packet distribution strategy provided by the embodiment of the present application, distribute the packets received from the physical network card queue to the upstream redirection queue. That is to say, the redirection module can receive packets from the physical network card queue and the virtual network card queue, and based on the packet distribution strategy provided by the embodiment of the present application, distribute the packets of the virtual network card queue to the downstream redirection queue, and then transmit them to the processor core corresponding to the downstream redirection queue. At the same time, based on the packet distribution strategy provided by the embodiment of the present application, distribute the packets of the physical network card queue to the upstream redirection queue, and then transmit them to the processor core corresponding to the upstream redirection queue.

[0058] It should be noted that the downstream direction refers to the direction in which the virtual machine sends data. That is to say, the downstream direction is the direction in which the network card (such as a smart network card) receives packets from the virtual network card queue. Correspondingly, the downstream redirection queue distributes the packets received by the network card (such as a smart network card) from the virtual network card queue, that is, the packets sent by the virtual machine. The upstream direction refers to the direction in which the physical network sends data. That is to say, the upstream direction is the direction in which the network card (such as a smart network card) receives packets from the physical network card queue. Correspondingly, the upstream redirection queue distributes the packets received by the network card (such as a smart network card) from the physical network card queue, that is, the packets sent by the physical network.

[0059] In the embodiments of the present application, the number of downlink redirection queues is equal to the number of multiple processor cores, and the number of uplink redirection queues is equal to the number of multiple processor cores. Thus, each processor core can correspond to a downlink redirection queue and an uplink redirection queue. Further, from the perspective of the packet forwarding plane, each processor core only needs to forward the packets in the corresponding downlink redirection queue and uplink redirection queue. That is to say, each processor core only needs to poll the corresponding pair of downlink redirection queue and uplink redirection queue, without polling other queues. Therefore, in the case of binding a downlink redirection queue and an uplink redirection queue corresponding to a processor core, even if the number of virtual network card queues increases, the packets processed by the processor core are provided by the corresponding downlink redirection queue and uplink redirection queue, which is convenient for application and expansion in the case of a large number of virtual network card queues.

[0060] In one example, the downlink redirection queue and the uplink redirection queue can be two groups of PCI (Peripheral Component Interconnect) devices implemented by FPGA hardware. That is to say, m downlink redirection queues 321 to 32m are used as a group of PCI devices implemented by FPGA hardware, and m uplink redirection queues 331 to 33m are used as another group of PCI devices implemented by FPGA hardware, and m is set to the number of processor cores.

[0061] The load adjustment module 340 can be a software function module implemented in a virtual switch, and can periodically adjust the packet distribution strategy by which the redirection module 310 distributes packets to the downlink redirection queue and the uplink redirection queue.

[0062] For ease of understanding of the packet distribution strategy provided in the embodiments of the present application, the following introduces an optional method flow of the load balancing method provided in the embodiments of the present application. As an optional implementation, Figure 4 An optional flowchart of the load balancing method provided in the embodiments of the present application is exemplarily shown. This method flow can be applied to a network card (such as a smart network card) and is executed by a hardware acceleration engine (such as FPGA hardware) in the network card. In an optional implementation, the method flow can be executed by a redirection module implemented by FPGA hardware. Referring to Figure 4 , this method flow can include the following steps.

[0063] In step S410, a packet is obtained, where the packet comes from a network card queue in a target direction.

[0064] In an embodiment of the present application, the redirection module may obtain data packets from the network card queue in the target direction. As an alternative implementation, the target direction may be divided into a downstream direction and an upstream direction. If the target direction is the downstream direction, the network card queue in the target direction is a virtual network card queue, and data packet forwarding processing needs to be performed in the downstream direction; correspondingly, the redirection module may obtain the data packets in the virtual network card queue, and the data packets are data packets sent by a virtual machine that need to be forwarded by a network card (such as a smart network card). If the target direction is the upstream direction, the network card queue in the target direction is a physical network card queue, and data packet forwarding processing needs to be performed in the upstream direction; correspondingly, the redirection module may obtain the data packets in the physical network card queue, and the data packets are data packets sent by the physical network that need to be forwarded by a network card (such as a smart network card).

[0065] In step S411, a digest value of the data packet is determined, and the digest value of the data packet represents the data stream corresponding to the data packet.

[0066] As an alternative implementation, the digest value of the data packet may be a value of a set length, which is used to represent the data stream where the data packet is located. That is to say, the data packets are transmitted in a streaming form. After any data packet in the data stream is transmitted to the network card queue in the target direction, the redirection module may obtain the data packet from the network card queue in the target direction and determine the digest value of the data packet, so as to represent the data stream corresponding to the data packet through the digest value of the data packet. For example, after any data packet in the data stream sent by the virtual machine is transmitted to the virtual network card queue, the redirection module may obtain the data packet from the virtual network card queue and determine the digest value of the data packet, so as to represent the data stream sent by the virtual machine corresponding to the data packet through the digest value of the data packet. Another example is that after any data packet in the data stream sent by the physical network is transmitted to the physical network card queue, the redirection module may obtain the data packet from the physical network card queue and determine the digest value of the data packet, so as to represent the data stream sent by the physical network corresponding to the data packet through the digest value of the data packet.

[0067] As an alternative implementation, the redirection module may use a digest algorithm to determine the digest value of the data packet according to the information representing the data stream in the data packet, so that the digest value of the data packet can represent the data stream corresponding to the data packet. In an implementation example, the five-tuple information (source IP address, destination IP address, source port number, destination port number, and protocol number) of the data packet may be used as the information representing the data stream in the data packet. That is to say, data packets in the same data stream have the same five-tuple information. Therefore, the redirection module may use a digest algorithm to determine the digest value of the data packet according to the five-tuple information of the data packet, so that the digest value of the data packet can represent the data stream corresponding to the data packet.

[0068] In one example, the digest value of a data packet can be a hash value. Correspondingly, the digest algorithm can be a hash algorithm. Thus, the redirection module can use the hash algorithm to determine the hash value of the data packet based on the five-tuple information of the data packet, so as to represent the data stream corresponding to the data packet through the hash value of the data packet. In a further example, the hash value can be a numerical value with a specific bit length calculated based on the five-tuple information of the data packet, such as an 8-bit hash value.

[0069] As an alternative implementation, the five-tuple information of the data packet can be carried in the header information of the data packet. Thus, after the redirection module obtains the data packet in the target direction, it can extract the five-tuple information (such as source IP address, destination IP address, source port number, destination port number, and protocol number) from the header information of the data packet, so as to determine the hash value corresponding to the five-tuple information of the data packet.

[0070] In step S412, according to the mapping relationship between the data stream and the redirection queue in the target direction, determine the redirection queue in the target direction corresponding to the digest value of the data packet, where the data stream in the mapping relationship is represented by the digest value of the data packet corresponding to the data stream.

[0071] In the embodiments of the present application, multiple processor cores of the on-chip processor of the network card (such as a smart network card) implement a virtual switch, and one processor core corresponds to one redirection queue in the target direction; that is, one processor core obtains the data packets from the network card queue in the target direction through a corresponding redirection queue in the target direction. Thus, the data packets from the network card queue in the target direction can be allocated by the redirection module to the redirection queue in the target direction, and then the data packets are forwarded through the processor core corresponding to the allocated redirection queue in the target direction.

[0072] As an alternative implementation, in combination with the foregoing description, if the target direction is the downstream direction, the redirection queue in the target direction is the downstream redirection queue. Thus, one processor core corresponds to one downstream redirection queue in the downstream direction, and one processor core obtains the data packets from the virtual network card queue in the downstream direction through the corresponding downstream redirection queue. That is, the data packets from the virtual network card queue can be allocated by the redirection module to the downstream redirection queue, so that the data packets are forwarded through the processor core corresponding to the allocated downstream redirection queue.

[0073] If the target direction is the upward direction, the redirection queue for the target direction is the upward redirection queue. Thus, for one processor core, there is one upward redirection queue in the upward direction. A processor core obtains packets from the physical network card queue in the upward direction through the corresponding upward redirection queue. That is to say, the packets from the physical network card queue can be allocated by the redirection module to the upward redirection queue, and then the packets are forwarded through the processor core corresponding to the allocated upward redirection queue.

[0074] As an optional implementation, when the redirection module allocates the redirection queue for the target direction to a packet, it can use the digest value of the packet to determine the corresponding mapped redirection queue for the target direction based on the mapping relationship between the data stream and the redirection queue for the target direction, and thus allocate the packet to the redirection queue for the target direction corresponding to the mapped digest value.

[0075] In a further optional implementation, based on the target direction being divided into the downward direction and the upward direction (correspondingly, the redirection queue for the target direction is divided into the downward redirection queue and the upward redirection queue), the mapping relationship between the data stream and the redirection queue for the target direction can be divided into: the downward mapping relationship between the data stream and the downward redirection queue, and the upward mapping relationship between the data stream and the upward redirection queue. That is to say, when the data stream in the mapping relationship is represented by the digest value of the packet corresponding to the data stream, the downward mapping relationship records the relationship between the digest value of the data stream and the correspondingly mapped downward redirection queue, and the upward mapping relationship records the relationship between the digest value of the data stream and the correspondingly mapped upward redirection queue.

[0076] Furthermore, based on one processor core corresponding to one redirection queue in the target direction, the mapping relationship represents the processor core to which the packets of the data stream are allocated (that is, the processor core corresponding to the redirection queue for the target direction to which the packet is allocated). Therefore, based on the mapping relationship, it is possible to determine the processor core to which the packets of the data stream are allocated and complete the allocation of the packets to the processor core. For example, what the downward mapping relationship corresponds to is: the downward redirection queue to which the packets of the data stream from the virtual network card queue are allocated, that is, the processor core to which the packets of the data stream from the virtual network card queue are allocated (the processor core and the downward redirection queue are in one-to-one correspondence). Another example is that what the upward mapping relationship corresponds to is: the upward redirection queue to which the packets of the data stream from the physical network card queue are allocated, that is, the processor core to which the packets of the data stream from the physical network card queue are allocated (the processor core and the upward redirection queue are in one-to-one correspondence).

[0077] As an optional implementation, the mapping relationship between the data stream and the redirection queue in the target direction can be in the form of a table, called a mapping table. For example, the mapping table can record the digest value representing the data stream (i.e., the digest value of the data packet corresponding to the data stream) and the redirection queue in the corresponding mapped target direction. In one example, the mapping table can have multiple entries, and one entry records the digest value of a data stream and a redirection queue in the corresponding mapped target direction. For example, the index of an entry is the digest value of a data stream, and the content of an entry is the identifier of a redirection queue in the corresponding mapped target direction. Thus, when the redirection module determines the redirection queue in the target direction corresponding to the digest value of the data packet based on the mapping relationship between the data stream and the redirection queue in the target direction, it can query multiple entries of the mapping table based on the digest value of the data packet, thereby querying the entry with the digest value of the data packet as the index, and then confirming the content of the queried entry to obtain the identifier of the redirection queue in the corresponding mapped target direction, so as to implement determining the redirection queue in the target direction corresponding to the digest value of the data packet. In one example, if the digest value of the data packet is in the form of a hash value, the mapping table can also be called a rehash table.

[0078] As an optional implementation, based on the target direction being divided into a downlink direction and an uplink direction, the mapping table can be divided into a downlink mapping table and an uplink mapping table. Among them, the downlink mapping table can record the mapping relationship between the data stream and the downlink redirection queue in tabular form; in one example, the downlink mapping table can have multiple entries, and one entry records the digest value of a data stream and a downlink redirection queue in the corresponding mapped direction; for example, the index of an entry in the downlink mapping table is the digest value of a data stream, and the content of an entry is the identifier of a downlink redirection queue in the corresponding mapped direction. For ease of understanding, taking 256 entries as an example, Figure 5 An exemplary example diagram of the downlink mapping table (also called the downlink rehash table) is shown for reference. Thus, when the redirection module determines the downlink redirection queue corresponding to the digest value of the data packet, it can query multiple entries of the downlink mapping table based on the digest value of the data packet, thereby querying the entry with the digest value of the data packet as the index, and then confirming the content of the queried entry to obtain the identifier of the redirection queue in the corresponding mapped downlink direction.

[0079] The uplink mapping table can record the mapping relationship between the data stream and the uplink redirection queue in tabular form; in one example, the uplink mapping table can have multiple entries, and one entry records the digest value of a data stream and an uplink redirection queue in the corresponding mapped direction; for example, the index of an entry in the uplink mapping table is the digest value of a data stream, and the content of an entry is the identifier of an uplink redirection queue in the corresponding mapped direction. For ease of understanding, taking 256 entries as an example, Figure 6An exemplary diagram showing an example of an uplink mapping table (also known as an uplink rehash table) can be referred to. Thus, when the redirection module determines the uplink redirection queue corresponding to the mapping of the digest value of the data packet, it can query multiple entries of the uplink mapping table based on the digest value of the data packet, so as to query the entry indexed by the digest value of the data packet, and then confirm the content of the queried entry to obtain the identifier of the uplink direction redirection queue corresponding to the mapping.

[0080] In step S413, the data packet is assigned to the redirection queue in the determined target direction, so that the data packet can be processed by the processor core corresponding to the redirection queue in the determined target direction.

[0081] Among them, multiple processor cores of the network card (such as a smart network card) implement a virtual switch, and one processor core corresponds to a redirection queue in the target direction to transmit data packets from the network card queue in the target direction. Moreover, the mapping relationship is periodically adjusted with the goal of balancing the data packet load in the target direction processed by each processor core.

[0082] After determining the redirection queue in the target direction to which the data packet is assigned based on the mapping relationship, the redirection module can assign the data packet to the redirection queue in the determined target direction. Thus, the redirection queue in the target direction to which the data packet is assigned can transmit the data packet to the corresponding processor core for forwarding processing. That is to say, the data packet can be forwarded and processed by the processor core corresponding to the redirection queue in the assigned target direction.

[0083] As an alternative implementation, if the target direction is the downlink direction, after determining the downlink redirection queue corresponding to the mapping of the digest value of the data packet based on the downlink mapping relationship, the embodiments of the present application can assign the data packet to the determined downlink redirection queue so as to perform forwarding processing of the data packet by the processor core corresponding to the determined downlink redirection queue. If the target direction is the uplink direction, after determining the uplink redirection queue corresponding to the mapping of the digest value of the data packet based on the uplink mapping relationship, the embodiments of the present application can assign the data packet to the determined uplink redirection queue so as to perform forwarding processing of the data packet by the processor core corresponding to the determined uplink redirection queue.

[0084] As an alternative implementation, an optional way for the processor core to perform forwarding processing of the data packet can be: after obtaining the data packet, the processor core uses the forwarding rule to perform forwarding processing of the data packet. For example, the processor core can query multiple forwarding configuration information of the virtual switch, generate a forwarding rule, and use the forwarding rule to perform forwarding processing of the data packet. Another example is that the processor core can perform forwarding processing of the data packet based on the generated forwarding rule.

[0085] In a further optional implementation, to achieve hardware-accelerated forwarding of data packets, the processor core can send the forwarding rules to the hardware acceleration engine. Thus, when the hardware acceleration engine finds a forwarding rule matching the data packet, it can perform hardware-accelerated forwarding of the data packet according to the forwarding rule. When the hardware acceleration engine fails to find a forwarding rule matching the data packet, it passes the data packet to the processor core, which generates a forwarding rule and uses the forwarding rule to perform forwarding processing of the data packet.

[0086] In the embodiments of the present application, the mapping relationship between the data stream and the redirection queue in the target direction (e.g., the downstream mapping relationship and the upstream mapping relationship) is not fixed, but is adjusted periodically to ensure that after the redirection queue in the target direction is allocated to the data packet based on the mapping relationship, the load of the data packets in the target direction processed by each processor core tends to be balanced. As an optional implementation, the load adjustment module set in the virtual switch can be used to implement the periodic adjustment of the mapping relationship. For related content, reference can be made to the corresponding part in the following text, and details are not described here.

[0087] In the embodiments of the present application, a processor core corresponds to a redirection queue in the target direction to transmit data packets from the network card queue in the target direction. At the same time, through the mapping relationship between the data stream and the redirection queue in the target direction, the embodiments of the present application can allocate a redirection queue in the target direction to the data packets from the network card queue in the target direction, so that the data packets are transmitted to the corresponding processor core for processing through the allocated redirection queue in the target direction. Among them, the data stream in the mapping relationship is represented by the digest value of the data packet corresponding to the data stream. It can be seen that the embodiments of the present application can achieve the data packet redirection ability at the data stream level. By combining the method that a processor core corresponds to a redirection queue in the target direction, the processor core only needs to poll the corresponding redirection queue in the target direction, instead of polling the network card queue in the target direction, reducing the number of queues that the processor core needs to poll in the software forwarding plane, reducing the overhead, being able to adapt to the expansion of the number of virtual machines, and thus improving the scalability.

[0088] At the same time, the mapping relationship between the data stream and the redirection queue in the target direction is adjusted periodically, and the goal of adjusting the mapping relationship is to make the load of the data packets in the target direction processed by each processor core tend to be balanced. Therefore, through the mapping relationship between the data stream and the redirection queue in the target direction, the embodiments of the present application allocate a redirection queue in the target direction to the data stream, and then transmit it to the corresponding processor core for processing, which can achieve the purpose of load balancing among multiple processor cores in a network card (e.g., a smart network card).

[0089] It can be seen that the embodiments of the present application realize load balancing among multiple processor cores in a network card (such as a smart network card), and improve the refinement of load balancing (load balancing adjustment at the data flow level), reduce overhead, and improve scalability, thereby improving the performance of load balancing.

[0090] As an optional implementation, the mapping relationship between the data flow and the redirection queue of the target direction (such as the downlink mapping relationship and the uplink mapping relationship) can be periodically adjusted by the on-chip processor of the network card (such as the smart network card). For example, in the embodiment of the present application, a load adjustment module is set in the virtual switch implemented by multiple processor cores of the on-chip processor, so that the mapping relationship is periodically adjusted by the load adjustment module. Based on this, in order to facilitate the load adjustment module to periodically adjust the mapping relationship, in an optional implementation, before allocating the data packet to the redirection queue of the determined target direction, the redirection module can write the summary value (such as a hash value) of the data packet into the reserved field of the header of the data packet, and then allocate the data packet with the summary value written in the reserved field of the header to the redirection queue of the determined target direction, so that the load adjustment module can count the data packet load corresponding to each data flow in each cycle (that is, the number of data packets to be processed corresponding to each data flow), thereby performing periodic adjustment of the mapping relationship.

[0091] As an optional implementation, in order to facilitate understanding of the optional implementation method of adjusting the mapping relationship in the embodiment of the present application, Figure 7 Another optional flow chart of the load balancing method provided in the embodiment of the present application is exemplarily shown. The method flow can be applied to a network card (e.g., an intelligent network card) and executed by an on-chip processor in the network card (e.g., an intelligent network card). In an optional implementation, the load adjustment module in the virtual switch implemented by multiple processor cores of the on-chip processor executes the method flow. Figure 7 The method flow may include the following steps.

[0092] In step S710, the total data packet load of each processor core in the current cycle is determined. The total data packet load of a processor core in the current cycle is composed of the data packet loads of each target direction of the processor core in the current cycle.

[0093] As an optional implementation, the load adjustment module can be woken up periodically to perform Figure 7The process shown is to achieve a periodic adjustment of the mapping relationship. By periodically waking up the load adjustment module, the load adjustment module can periodically determine the total packet load of each processor core in the current cycle, that is, the number of packets processed by each processor core in the current cycle. For example, the load adjustment module is woken up at the end of the current cycle. Thus, for each processor core in the on-chip processor, the load adjustment module can separately determine the packet loads of each processor core in each target direction in the current cycle. Then, for any processor core, the packet loads of each processor core in each target direction in the current cycle are added together to obtain the total packet load of the processor core in the current cycle. For example, the total packet load of a processor core in the current cycle can be the sum of the packet loads in the downstream direction and the upstream direction of the processor core in the current cycle. In an alternative implementation, a cycle can be regarded as a time period with a set time length (such as 5 minutes), and the time length corresponding to a cycle can be determined according to the actual situation, which is not limited in the embodiments of the present application; in addition, in the embodiments of the present application, the time lengths of each cycle can also be set to be the same, or different cycles can be set to have different time lengths, and the situations of the time lengths corresponding to different cycles can be determined according to the actual situation, which is not limited in the embodiments of the present application.

[0094] As an alternative implementation, the packet load of a processor core can be regarded as the number of packets forwarded and processed by the processor core. That is to say, for any processor core, in the embodiments of the present application, the number of packets forwarded and processed by the processor core in the current cycle can be periodically determined, including the number of packets in the downstream direction and the number of packets in the upstream direction forwarded and processed by the processor core in the current cycle, as well as the total number of packets (the sum of the number of packets in the downstream direction and the upstream direction forwarded and processed by the processor core in the current cycle, that is, the total packet load of the processor core in the current cycle).

[0095] In step S711, according to the total packet load of each processor core in the current cycle, determine the utilization rate of each processor core in the current cycle.

[0096] As an alternative implementation, for any processor core, in the embodiments of the present application, the utilization rate of the processor core in the current cycle can be determined according to the total packet load of the processor core in the current cycle and the preset upper limit of the packet load of the processor core in the current cycle. For example, in the embodiments of the present application, the upper limit of the packet load corresponding to a utilization rate of 100% can be set. That is to say, if the number of packets processed and forwarded by a processor core in a cycle reaches the upper limit of the packet load, it is regarded that the utilization rate of the processor core reaches 100%, and the processor core is fully utilized. Furthermore, in the embodiments of the present application, the total packet load of the processor core in the current cycle can be divided by the upper limit of the packet load of the processor core in the current cycle to determine the utilization rate of the processor core in the current cycle.

[0097] As an optional implementation, for any processor core, embodiments of the present application can determine the utilization rate of the processor core in the current cycle according to the effective usage time corresponding to the total packet load of the processor core in the current cycle and the time length corresponding to the cycle. For example, the utilization rate of a processor core in the current cycle can be determined based on the time ratio of the total packet load processed by the processor core in the current cycle. In one example, for any processor core, embodiments of the present application can determine the effective usage time corresponding to the total packet load processed by the processor core in the current cycle, and then divide the effective usage time by the time length of the cycle to obtain the utilization rate of the processor core corresponding to the total packet load processed in the current cycle. Among them, the effective usage time corresponding to the total packet load processed by the processor core in the current cycle can be regarded as the total usage time of the processor core for forwarding and processing packets in the downlink and uplink directions in the current cycle.

[0098] In step S712, according to the utilization rates of each processor core in the current cycle, it is judged whether the current cycle reaches the adjustment condition for adjusting the mapping relationship. If not, step S713 is executed; if so, step S714 is executed.

[0099] In embodiments of the present application, the mapping relationship is divided into the downlink mapping relationship between the data stream and the downlink redirection queue, and the uplink mapping relationship between the data stream and the uplink redirection queue; among them, the digest value of the packet represents the data stream corresponding to the packet.

[0100] After determining the utilization rates of each processor core in the current cycle, embodiments of the present application need to judge whether the current cycle reaches the adjustment condition for adjusting the mapping relationship based on the utilization rates of each processor core in the current cycle. For example, embodiments of the present application can preset the adjustment conditions required for adjusting the mapping relationship in advance, so that after the load adjustment module determines the utilization rates of each processor core in the current cycle, it can judge whether the current cycle reaches the adjustment condition for adjusting the mapping relationship based on the utilization rates of each processor core in the current cycle. If the adjustment condition for adjusting the mapping relationship is not reached, embodiments of the present application can execute step S713 to end the process and wait for the load adjustment module to be awakened in the next cycle. If the adjustment condition for adjusting the mapping relationship is reached, embodiments of the present application can execute step S714 to adjust the downlink mapping relationship between the data stream and the downlink redirection queue and the uplink mapping relationship between the data stream and the uplink redirection queue respectively.

[0101] As an alternative implementation, the adjustment condition for adjusting the mapping relationship can be defined as: the number of processor cores with a utilization rate exceeding the utilization rate threshold is greater than a set number, and the difference in utilization rates between the processor core with the highest utilization rate and the processor core with the lowest utilization rate exceeds a preset difference. It should be noted that the processor core with the highest utilization rate can be regarded as the processor core with the highest packet load among multiple processor cores, and the processor core with the lowest utilization rate can be regarded as the processor core with the lowest packet load among multiple processor cores.

[0102] In an example, if the load adjustment module determines, based on the utilization rates of each processor core in the current cycle, that the number of processor cores with a utilization rate exceeding the utilization rate threshold is greater than a set number, and the difference in utilization rates between the processor core with the highest utilization rate and the processor core with the lowest utilization rate exceeds a preset difference, it can be confirmed that the adjustment condition for adjusting the mapping relationship is met, and step S714 can be executed to adjust the downlink mapping relationship and the uplink mapping relationship respectively; if, based on the utilization rates of each processor core in the current cycle, it is determined that the number of processor cores with a utilization rate exceeding the utilization rate threshold is not greater than a set number, and / or the difference in utilization rates between the processor core with the highest utilization rate and the processor core with the lowest utilization rate does not exceed a preset difference, it can be confirmed that the adjustment condition for adjusting the mapping relationship is not met, and step S713 can be executed to end the process and wait for the load adjustment module to be awakened in the next cycle.

[0103] In an example, taking the on-chip processor with 4 processor cores as an example, these 4 processor cores can implement a virtual switch, and a load adjustment module is set in the virtual switch. Thus, if the load adjustment module determines in the current cycle that the number of processor cores with a utilization rate exceeding 90% (an example of the utilization rate threshold) is greater than 1 (1 can be an example of the set number), and the difference in utilization rates between the processor core with the highest utilization rate (i.e., the processor core with the highest packet load) and the processor core with the lowest utilization rate (i.e., the processor core with the lowest packet load) among the 4 processor cores reaches 30% (an example of the preset difference), it can be confirmed that the adjustment condition for adjusting the mapping relationship is met; if the load adjustment module determines in the current cycle that the number of processor cores with a utilization rate exceeding 90% is not greater than 1, and / or the difference in utilization rates between the processor core with the highest utilization rate and the processor core with the lowest utilization rate among the 4 processor cores does not reach 30%, it can be confirmed that the adjustment condition for adjusting the mapping relationship is not met.

[0104] It should be noted that steps S711 and S712 can be alternative implementation methods for determining whether the adjustment condition for adjusting the mapping relationship is met based on the total packet load of each processor core in the current cycle. Embodiments of the present application can also set other forms of adjustment conditions, and make the form of the adjustment condition associated with the total packet load of each processor core in the current cycle, rather than being limited to the form of the adjustment condition described above.

[0105] In step S713, the process ends.

[0106] When the adjustment condition for adjusting the mapping relationship is not reached in the current cycle, the embodiment of the present application can end the process and wait for the load adjustment module to be woken up in the next cycle.

[0107] In step S714, with the goal of making the packet loads in the target directions processed by each processor core tend to be balanced, the mapping relationship between the data stream and the redirection queue in the target direction is adjusted.

[0108] When the adjustment condition for adjusting the mapping relationship is reached in the current cycle, the embodiment of the present application can adjust the mapping relationship between the data stream and the redirection queue in the target direction in each target direction, so that the packet loads in the target directions processed by each processor core tend to be balanced. For example, the embodiment of the present application can adjust the downlink mapping relationship in the downlink direction to make the packet loads in the downlink direction processed by each processor core tend to be balanced; and in the uplink direction, adjust the uplink mapping relationship to make the packet loads in the uplink direction processed by each processor core tend to be balanced.

[0109] As an optional implementation, Figure 8 Exemplarily shows an optional flowchart for adjusting the mapping relationship provided by the embodiment of the present application, as Figure 8 shown, the method flow may include the following steps.

[0110] In step S810, for any target direction, according to the usage rate of the target direction of each processor core in the current cycle, determine the average usage rate of the target direction; where any processor core with a usage rate of the target direction in the current cycle higher than the average usage rate is a first type of processor core, and any processor core with a usage rate of the target direction in the current cycle lower than the average usage rate is a second type of processor core.

[0111] In the process of adjusting the mapping relationship between the data stream and the redirection queue in the target direction, the embodiment of the present application adjusts the mapping relationship between the data stream and the redirection queue in the target direction. Therefore, for the target direction (divided into the downlink direction and the uplink direction), the embodiment of the present application can determine the usage rate of the target direction of each processor core in the current cycle. For example, for any processor core, divide the packet load in the target direction processed by the processor core in the current cycle by the packet load upper limit of the processor core in the current cycle to obtain the usage rate of the target direction of the processor core in the current cycle. Another example is that for any processor core, divide the effective usage time corresponding to the packet load in the target direction processed by the processor core in the current cycle by the time length of the current cycle to obtain the usage rate of the target direction of the processor core in the current cycle.

[0112] Taking the target direction as the downward direction as an example, the utilization rate of each processor core in the downward direction in the current cycle can be determined. For example, for any processor core, divide the packet load in the downward direction processed by the processor core in the current cycle by the packet load upper limit of the processor core in the current cycle to obtain the utilization rate of the processor core in the downward direction in the current cycle. Another example is that for any processor core, divide the effective usage time corresponding to the packet load in the downward direction processed by the processor core in the current cycle by the time length of the current cycle to obtain the utilization rate of the processor core in the downward direction in the current cycle.

[0113] Taking the target direction as the upward direction as an example, the utilization rate of each processor core in the upward direction in the current cycle can be determined. For example, for any processor core, divide the packet load in the upward direction processed by the processor core in the current cycle by the packet load upper limit of the processor core in the current cycle to obtain the utilization rate of the processor core in the upward direction in the current cycle. Another example is that for any processor core, divide the effective usage time corresponding to the packet load in the upward direction processed by the processor core in the current cycle by the time length of the current cycle to obtain the utilization rate of the processor core in the upward direction in the current cycle.

[0114] After confirming the utilization rate of the target direction of each processor core in the current cycle, in an optional implementation, the embodiments of the present application can add up the utilization rates of the target direction of each processor core in the current cycle, and divide the sum by the number of processor cores to obtain the average utilization rate of the target direction. In an example, assume that the number of processor cores is m, and the utilization rate of the target direction of the i-th processor core in the current cycle is set as C i (i belongs to m), and the average utilization rate of the target direction is set as C avg , then C avg can be expressed as: C avg= Sum(C i ) / m.

[0115] After determining the average utilization rate of the target directions of multiple processor cores in the current cycle, the embodiments of the present application can confirm the first type of processor cores and the second type of processor cores based on whether the utilization rate of the target direction of the processor core in the current cycle is higher than the average utilization rate. That is, the first type of processor cores can be regarded as the processor cores with a utilization rate of the target direction higher than the average utilization rate in the current cycle, and the second type of processor cores can be regarded as the processor cores with a utilization rate of the target direction lower than the average utilization rate in the current cycle. Since the utilization rate of the first type of processor cores in the target direction of the current cycle is high, while the utilization rate of the second type of processor cores in the target direction of the current cycle is low, the embodiments of the present application can adjust the mapping relationship between the data stream and the redirection queue of the target direction, and divert a part of the data stream of the first type of processor cores in the target direction to the second type of processor cores for processing, so that the utilization rates of the first type of processor cores and the second type of processor cores in the target direction can tend to the average utilization rate, so as to achieve the purpose of load balancing among multiple processor cores in the target direction.

[0116] It should be noted that since the target direction is divided into the downstream direction and the upstream direction, the embodiments of the present application need to determine the first type of processor cores and the second type of processor cores in the downstream direction and the upstream direction respectively. And since the packet loads of the same processor core in the downstream direction and the upstream direction may be different (corresponding to the utilization rates of the same processor in the downstream direction and the upstream direction may be different), the same processor core may be the first type of processor core in the downstream direction and the second type of processor core in the upstream direction. That is to say, the first type of processor cores and the second type of processor cores in the downstream direction and the upstream direction need to be determined separately and may be different.

[0117] For example, after confirming the utilization rate of the downstream direction of each processor core in the current cycle, the embodiments of the present application can add up the utilization rates of the downstream direction of each processor core in the current cycle, and divide the sum by the number of processor cores to obtain the average utilization rate of the downstream direction; if the utilization rate of the processor core in the current cycle in the downstream direction is higher than the average utilization rate of the downstream direction, the processor core is the first type of processor core in the downstream direction and needs to perform data stream diversion in the downstream direction; if the utilization rate of the processor core in the current cycle in the downstream direction is lower than the average utilization rate of the downstream direction, the processor core is the second type of processor core in the downstream direction.

[0118] Similarly, after confirming the upstream direction utilization rate of each processor core in the current cycle, the embodiments of the present application can add up the upstream direction utilization rates of each processor core in the current cycle, and divide the sum by the number of processor cores to obtain the average utilization rate in the upstream direction. If the upstream direction utilization rate of a processor core in the current cycle is higher than the average utilization rate in the upstream direction, the processor core is a first type of processor core in the upstream direction and needs to perform data stream diversion in the upstream direction. If the upstream direction utilization rate of a processor core in the current cycle is lower than the average utilization rate in the upstream direction, the processor core is a second type of processor core in the upstream direction.

[0119] In step S811, for any first type of processor core, according to the digest values of the data packets in the target direction processed by the first type of processor core in the current cycle, determine the utilization rates corresponding to the data streams in the target direction processed by the first type of processor core in the current cycle.

[0120] In an alternative implementation, for any first type of processor core whose utilization rate in the target direction in the current cycle is higher than the average utilization rate in the target direction, the embodiments of the present application can determine the utilization rates corresponding to the data streams in the target direction processed by the first type of processor core in the current cycle. Since the data stream is represented by the digest value of the data packet corresponding to the data stream, the embodiments of the present application can determine the utilization rates corresponding to the data streams in the target direction processed by the first type of processor core in the current cycle based on the digest values of the data packets in the target direction processed by the first type of processor core in the current cycle. In an implementation example, the embodiments of the present application can collect the digest values of the data packets in the target direction processed by the first type of processor core in the current cycle, so as to classify the data packets with the same digest value into the same data stream, obtain the number of data packets in each data stream in the target direction processed by the first type of processor core in the current cycle, and then determine the utilization rates corresponding to the data streams in the target direction processed by the first type of processor core in the current cycle based on the number of data packets in each data stream in the target direction processed by the first type of processor core in the current cycle. For example, divide the number of data packets in each data stream in the target direction processed by the first type of processor core in the current cycle by the data packet load upper limit of the first type of processor core in the current cycle to obtain the utilization rates corresponding to the data streams in the target direction processed by the first type of processor core in the current cycle. Another example is to divide the effective usage time corresponding to the number of data packets in each data stream in the target direction processed by the first type of processor core in the current cycle by the time length of the current cycle to obtain the utilization rates corresponding to the data streams in the target direction processed by the first type of processor core in the current cycle.

[0121] For example, taking the downlink direction as the target direction, for any first - type processor core whose downlink direction utilization rate in the current cycle is higher than the average utilization rate of the downlink direction, the embodiments of the present application can determine the utilization rates corresponding to the respective data streams in the downlink direction processed by the first - type processor core in the current cycle. Since the data stream is represented by the digest value of the data packet corresponding to the data stream, the embodiments of the present application can determine the utilization rates corresponding to the respective data streams in the downlink direction processed by the first - type processor core in the current cycle based on the digest values of the respective data packets in the downlink direction processed by the first - type processor core in the current cycle.

[0122] Also for example, taking the uplink direction as the target direction, for any first - type processor core whose uplink direction utilization rate in the current cycle is higher than the average utilization rate of the uplink direction, the embodiments of the present application can determine the utilization rates corresponding to the respective data streams in the uplink direction processed by the first - type processor core in the current cycle. Since the data stream is represented by the digest value of the data packet corresponding to the data stream, the embodiments of the present application can determine the utilization rates corresponding to the respective data streams in the uplink direction processed by the first - type processor core in the current cycle based on the digest values of the respective data packets in the uplink direction processed by the first - type processor core in the current cycle.

[0123] In step S812, for any first - type processor core, according to the utilization rates corresponding to the respective data streams in the target direction processed by the first - type processor core in the current cycle, determine the shunt data streams of the first - type processor core in the target direction. After the shunt data streams of the first - type processor core in the target direction are shunted, the utilization rate of the first - type processor core in the target direction tends to the average utilization rate of the target direction.

[0124] For any first - type processor core, the shunt data streams of the first - type processor core in the target direction can be regarded as the data streams that need to be shunted from the first - type processor core to the second - type processor core in the target direction. To make the utilization rate of the first - type processor core in the target direction tend to the average utilization rate of the target direction after the shunt data streams in the target direction of the first - type processor core are shunted to the second - type processor core, the embodiments of the present application can determine the shunt data streams that need to be shunted from the first - type processor core in the target direction based on the utilization rates corresponding to the respective data streams in the target direction processed by the first - type processor core in the current cycle.

[0125] As an optional implementation, to avoid the situation where a single shunted data stream with a high load in the target direction is shunted to the second type of processor core and the utilization rate of the second type of processor core in the target direction exceeds the average utilization rate, the embodiment of the present application can, according to the utilization rates corresponding to the data streams in the target direction processed by the first type of processor core in the current cycle, screen out the data streams with non-zero utilization rates from the data streams in the target direction processed by the first type of processor core in the current cycle (if the utilization rate corresponding to the data stream is zero, it means that the data stream has zero load on the first type of processor core and there is no shunting requirement); thus, for the screened data streams, in the order of increasing utilization rate, accumulate the utilization rates of the data streams until the accumulated utilization rate approaches the difference between the utilization rate of the first type of processor core in the target direction in the current cycle and the average utilization rate; for example, when the accumulated utilization rate is closest to the difference and there is still a certain proportion of gap from the difference, it can also be regarded as the accumulated utilization rate approaching the difference, and the accumulated utilization rate can be stopped. The certain proportion of gap can be determined according to the actual situation, such as 20% or the like.

[0126] Furthermore, the data stream corresponding to the accumulated utilization rate is used as the shunted data stream of the first type of processor core in the target direction, so that the utilization rate corresponding to the shunted data stream of the first type of processor core in the target direction approaches the difference between the utilization rate of the first type of processor core in the target direction in the current cycle and the average utilization rate, so as to ensure that after the shunted data stream of the first type of processor core in the target direction is shunted, the utilization rate of the first type of processor core in the target direction in the current cycle tends to the average utilization rate.

[0127] For example, taking the target direction as the downlink direction as an example, the embodiment of the present application can, according to the utilization rates corresponding to the data streams in the downlink direction processed by the first type of processor core in the current cycle, screen out the data streams with non-zero utilization rates; thus, for the screened data streams, in the order of increasing utilization rate, accumulate the utilization rates of the data streams until the accumulated utilization rate approaches the difference between the utilization rate of the first type of processor core in the downlink direction in the current cycle and the average utilization rate (the average utilization rate in the downlink direction), and then stop accumulating the utilization rate; furthermore, the data stream corresponding to the accumulated utilization rate is used as the shunted data stream of the first type of processor core in the downlink direction, so that the utilization rate corresponding to the shunted data stream of the first type of processor core in the downlink direction approaches the difference between the utilization rate of the first type of processor core in the downlink direction in the current cycle and the average utilization rate, so as to ensure that after the shunted data stream of the first type of processor core in the downlink direction is shunted, the utilization rate of the first type of processor core in the downlink direction in the current cycle tends to the average utilization rate (the average utilization rate in the downlink direction).

[0128] For example, taking the upward direction as the target direction, in the embodiment of the present application, the data streams with non-zero usage rates can be filtered according to the usage rates corresponding to the data streams in the upward direction processed by the first type of processor core in the current cycle. Then, for the filtered data streams, the usage rates of the data streams are accumulated in ascending order of the usage rates until the accumulated usage rate approaches the difference between the usage rate of the first type of processor core in the upward direction in the current cycle and the average usage rate (the average usage rate in the upward direction). At this time, the accumulation of the usage rate stops. Furthermore, the data stream corresponding to the accumulated usage rate is used as the diverted data stream of the first type of processor core in the upward direction, so that the usage rate corresponding to the diverted data stream of the first type of processor core in the upward direction approaches the difference between the usage rate of the first type of processor core in the upward direction in the current cycle and the average usage rate, ensuring that after the diverted data stream of the first type of processor core in the upward direction is diverted, the usage rate of the first type of processor core in the upward direction in the current cycle tends to the average usage rate (the average usage rate in the upward direction).

[0129] In step S813, the mapping relationship between the diverted data stream of the first type of processor core in the target direction and the redirection queue corresponding to the first type of processor core in the target direction is removed.

[0130] After confirming the diverted data stream of the first type of processor core in the target direction, in the embodiment of the present application, the mapping relationship between the diverted data stream of the first type of processor core in the target direction and the redirection queue corresponding to the first type of processor core in the target direction can be removed, so that the diverted data stream of the first type of processor core in the target direction will no longer be allocated to the redirection queue corresponding to the first type of processor core in the target direction in the future, that is, it will no longer be allocated to the first type of processor core for processing.

[0131] As an optional implementation, in the embodiment of the present application, the mapping relationship between the diverted data stream of the first type of processor core in the target direction and the redirection queue corresponding to the first type of processor core in the target direction can be removed by clearing the relevant table entries in the mapping table (such as clearing the content of the relevant table entries). For example, in the embodiment of the present application, based on the digest value of the diverted data stream of the first type of processor core in the target direction and the identifier of the redirection queue corresponding to the first type of processor core in the target direction, the corresponding table entry can be queried in the mapping table, and then the queried table entry in the mapping table can be cleared to remove the mapping relationship between the diverted data stream of the first type of processor core in the target direction and the redirection queue corresponding to the first type of processor core in the target direction. Among them, the mapping table can have multiple table entries, and one table entry records the digest value of a data stream and the identifier of a redirection queue corresponding to the target direction mapped to it.

[0132] For example, taking the downward direction as the target direction, in the embodiments of the present application, the shunt data stream of the first type of processor core in the downward direction can be released from the downward mapping relationship with the downward redirection queue corresponding to the first type of processor core, so that the shunt data stream of the first type of processor core in the downward direction will no longer be allocated to the downward redirection queue corresponding to the first type of processor core in the future. In one example, in the embodiments of the present application, based on the digest value of the shunt data stream of the first type of processor core in the downward direction and the identifier of the downward redirection queue corresponding to the first type of processor core, the corresponding entry can be queried in the downward mapping table, and then the entry queried in the downward mapping table can be cleared; wherein, the downward mapping table can have multiple entries, and one entry records the digest value of a data stream and the identifier of a downward redirection queue corresponding to the mapping.

[0133] For example, taking the upward direction as the target direction, in the embodiments of the present application, the shunt data stream of the first type of processor core in the upward direction can be released from the upward mapping relationship with the upward redirection queue corresponding to the first type of processor core, so that the shunt data stream of the first type of processor core in the upward direction will no longer be allocated to the upward redirection queue corresponding to the first type of processor core in the future. In one example, in the embodiments of the present application, based on the digest value of the shunt data stream of the first type of processor core in the upward direction and the identifier of the upward redirection queue corresponding to the first type of processor core, the corresponding entry can be queried in the upward mapping table, and then the entry queried in the upward mapping table can be cleared; wherein, the upward mapping table can have multiple entries, and one entry records the digest value of a data stream and the identifier of an upward redirection queue corresponding to the mapping.

[0134] In step S814, the shunt data stream of the first type of processor core in the target direction is shunted to at least one second type of processor core, so that the utilization rate of the second type of processor core to which the shunt data stream is shunted in the target direction tends to the average utilization rate in the target direction.

[0135] In step S815, a mapping relationship is newly added between the shunt data stream of the first type of processor core in the target direction and the redirection queue corresponding to the second type of processor core shunted in the target direction.

[0136] After confirming the shunted data stream of the first type of processor core in the target direction, the embodiment of the present application can shunt the shunted data stream of the first type of processor core in the target direction to one or more second type of processor cores, and make the utilization rate of the second type of processor cores shunted by the shunted data stream in the target direction tend to the average utilization rate in the target direction and not be greater than the average utilization rate in the target direction. For example, taking the target direction as the downlink direction as an example, the shunted data stream of the first type of processor core in the downlink direction can be shunted to one or more second type of processor cores, and make the utilization rate of the second type of processor cores shunted by the shunted data stream in the downlink direction tend to the average utilization rate in the downlink direction and not be greater than the average utilization rate in the downlink direction. For another example, taking the target direction as the uplink direction as an example, the shunted data stream of the first type of processor core in the uplink direction can be shunted to one or more second type of processor cores, and make the utilization rate of the second type of processor cores shunted by the shunted data stream in the uplink direction tend to the average utilization rate in the uplink direction and not be greater than the average utilization rate in the uplink direction.

[0137] For the shunted data stream of the first type of processor core in the target direction and the second type of processor cores shunted by the shunted data stream, the embodiment of the present application can add a mapping relationship between the shunted data stream of the first type of processor core in the target direction and the redirected queue corresponding to the second type of processor cores shunted by the shunted data stream in the target direction, so that the shunted data stream of the first type of processor core in the target direction can be subsequently assigned to the redirected queue corresponding to the second type of processor cores shunted by the shunted data stream in the target direction, that is, shunted to the second type of processor cores for processing.

[0138] As an optional implementation, the embodiment of the present application can add relevant entries in the mapping table to implement the mapping relationship between the shunted data stream of the first type of processor core in the target direction and the redirected queue corresponding to the second type of processor cores shunted by the shunted data stream in the target direction. For example, the embodiment of the present application can add corresponding entries in the mapping table based on the digest value of the shunted data stream of the first type of processor core in the target direction and the identifier of the redirected queue corresponding to the second type of processor cores shunted by the shunted data stream in the target direction, so as to record through the added entries: the mapping relationship between the digest value of the shunted data stream of the first type of processor core in the target direction and the identifier of the redirected queue corresponding to the second type of processor cores shunted by the shunted data stream in the target direction.

[0139] For example, taking the downlink direction as the target direction, the embodiments of the present application can add relevant entries to the mapping table to implement the mapping relationship between the shunted data stream of the newly added first type of processor core in the downlink direction and the downlink redirection queue corresponding to the shunted second type of processor core. For example, the embodiments of the present application can add corresponding entries to the mapping table based on the digest value of the shunted data stream of the first type of processor core in the uplink direction and the identifier of the downlink redirection queue corresponding to the shunted second type of processor core, so as to record through the newly added entries: the mapping relationship between the digest value of the shunted data stream of the first type of processor core in the downlink direction and the identifier of the downlink redirection queue corresponding to the shunted second type of processor core.

[0140] For example, taking the uplink direction as the target direction, the embodiments of the present application can add relevant entries to the mapping table to implement the mapping relationship between the shunted data stream of the newly added first type of processor core in the uplink direction and the uplink redirection queue corresponding to the shunted second type of processor core. For example, the embodiments of the present application can add corresponding entries to the mapping table based on the digest value of the shunted data stream of the first type of processor core in the uplink direction and the identifier of the uplink redirection queue corresponding to the shunted second type of processor core, so as to record through the newly added entries: the mapping relationship between the digest value of the shunted data stream of the first type of processor core in the uplink direction and the identifier of the uplink redirection queue corresponding to the shunted second type of processor core.

[0141] It can be seen that based on Figure 8 the process shown, the embodiments of the present application can adjust the mapping relationship in the downlink direction and the uplink direction respectively, so that the processor cores with a higher utilization rate in the downlink direction can shunt part of the data stream to the processor cores with a lower utilization rate for processing, and the processor cores with a higher utilization rate in the uplink direction can shunt part of the data stream to the processor cores with a lower utilization rate for processing, realizing the load balancing adjustment of multiple processor cores in the downlink direction and the uplink direction.

[0142] In the embodiments of the present application, by using the programmable hardware on a network card (such as a smart network card), one processor core corresponds to a redirection queue in the target direction to transmit data packets from the network card queue in the target direction. For example, one processor core corresponds to a downlink redirection queue in the downlink direction to transmit data packets from the virtual network card queue in the downlink direction, and one processor core corresponds to an uplink redirection queue in the uplink direction to transmit data packets from the virtual network card queue in the uplink direction. At the same time, through the mapping relationship between the data stream and the redirection queue in the target direction, the redirection queue in the target direction is allocated to the data packets from the network card queue in the target direction, and the data stream in the mapping relationship is represented by the digest value of the data packet corresponding to the data stream. Therefore, the embodiments of the present application can achieve the data packet redirection ability at the data stream level. Thus, combined with the method that one processor core corresponds to a redirection queue in the target direction, the processor core only needs to poll the corresponding redirection queue in the target direction, rather than polling the network card queue in the target direction, reducing the number of queues that the processor core needs to poll in the software forwarding plane, reducing the overhead, being able to adapt to the number expansion of virtual machines, and thus improving the scalability.

[0143] Furthermore, the load adjustment module set in the virtual switch implemented by multiple processor cores can periodically monitor the data packet load in the target direction processed by each processor core, and thus periodically adjust the mapping relationship (such as periodically overwriting the mapping table in the programmable hardware), so that the data packet load in the target direction processed by each processor core can tend to be balanced, achieving the load balance among multiple processor cores in the network card (such as a smart network card).

[0144] Specifically, the number of the downlink redirection queue and the uplink redirection queue of the processor core is kept consistent with the number of processor cores, thus reducing the overhead of the processor core polling the queue. And, the data packet allocation and scheduling at the data stream granularity are realized, switching the processor core carrying the data stream at the data stream level, thus avoiding the problem of high load of a single queue and packet loss caused by resource exhaustion of the processor core bound to the queue. And, there is no need to adjust the binding relationship between the virtual network card queue and the processor core, only modifying the mapping relationship between the data stream on the network card (such as a smart network card) hardware and the redirection queue (downlink redirection queue and uplink redirection queue) corresponding to the processor core in the target direction. Therefore, potential problems such as the stagnation of the virtual network card queue or the inability of the processor core to receive packets from the virtual network card queue caused by modifying the binding relationship between the virtual network card queue and the processor core will not occur, improving the system robustness. And, since the overhead of the processor core is for the corresponding downlink redirection queue and uplink redirection queue, the overhead does not increase with the increase in the virtual machine density (as the virtual machine density increases, the number of virtual network card queues increases accordingly), being able to support the high-density deployment scenario of virtual machines and improving the scalability.

[0145] It can be seen that the embodiments of the present application can improve the refinement degree of load balancing (i.e., the load balancing granularity is finer), reduce the overhead, and support scalability, thus improving the performance of load balancing. That is to say, the embodiments of the present application not only reduce the overhead of the polling queue of the processor core of the network card (such as a smart network card) in the scenario of high-density virtual machine deployment, but also make the load balancing adjustment means more fine-grained and lightweight. When realizing the load balancing among multiple processor cores in the network card (such as a smart network card), the performance of load balancing is improved.

[0146] The embodiments of the present application further provide a network card (such as a smart network card), combined with Figure 3 As shown, the network card (such as a smart network card) may include a hardware acceleration engine and an on-chip processor; the hardware acceleration engine is provided with a redirection module, a plurality of downstream redirection queues, and a plurality of upstream redirection queues; the on-chip processor is provided with a plurality of processor cores, and the plurality of processor cores implement a virtual switch, and the virtual switch is provided with a load adjustment module;

[0147] Among them, the redirection module is connected between the plurality of downstream redirection queues and the virtual network card queue, and between the plurality of upstream redirection queues and the physical network card queue. One processor core corresponds to one downstream redirection queue for transmitting data packets from the virtual network card queue, and one processor core corresponds to one upstream redirection queue for transmitting data packets from the physical network card queue;

[0148] In the embodiments of the present application, the redirection module is configured to execute the load balancing method executed by the redirection module provided by the embodiments of the present application; the load adjustment module is configured to execute the load balancing method executed by the load adjustment module provided by the embodiments of the present application.

[0149] The embodiments of the present application further provide a storage medium, which stores one or more computer-executable instructions. When the one or more computer-executable instructions are executed, the load balancing method executed by the redirection module provided by the embodiments of the present application, or the load balancing method executed by the load adjustment module provided by the embodiments of the present application is implemented. As an optional implementation, the one or more computer-executable instructions may be instructions in the programmable software of the FPGA hardware of the network card (such as a smart network card), or instructions in the software running on the processor core of the on-chip processor of the network card (such as a smart network card).

[0150] The above describes multiple embodiment solutions provided by the embodiments of the present application. The optional ways introduced in each embodiment solution can be combined and cross-referenced with each other without conflict, so as to extend a variety of possible embodiment solutions, all of which can be considered as the embodiment solutions disclosed and made public by the embodiments of the present application.

[0151] Although the embodiments of the present application are disclosed as above, the present application is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application shall be subject to the scope defined by the claims.

Claims

1. A load balancing method, wherein, applied to a network card, the method includes: Obtain a data packet, wherein the data packet comes from the network card queue in the target direction; Determine the digest value of the data packet, and the digest value of the data packet represents the data stream corresponding to the data packet; According to the mapping relationship between the data stream and the redirect queue in the target direction, determine the redirect queue in the target direction mapped by the digest value of the data packet, wherein the data stream in the mapping relationship is represented by the digest value of the data packet corresponding to the data stream; Allocate the data packet to the determined redirect queue in the target direction, so that the data packet is processed by the processor core corresponding to the determined redirect queue in the target direction; wherein, multiple processor cores of the network card implement a virtual switch, and one processor core corresponds to one redirect queue in the target direction to transmit data packets from the network card queue in the target direction, and the mapping relationship is periodically adjusted with the goal of making the data packet load processed by each processor core tend to be balanced.

2. The method according to claim 1, wherein, The step of determining the redirect queue in the target direction mapped by the digest value of the data packet according to the mapping relationship between the data stream and the redirect queue in the target direction includes: Based on the digest value of the data packet, query in multiple entries of the mapping table for the entry with the digest value of the data packet as the index; Determine the content of the queried entry to obtain the identifier of the redirect queue in the target direction mapped by the digest value of the data packet; wherein, the mapping table records the mapping relationship between the data stream and the redirect queue in the target direction through multiple entries, the index of one entry is the digest value of the data packet corresponding to a data stream, and the content of one entry is the identifier of a redirect queue in the target direction corresponding to the mapping.

3. The method according to claim 1 or 2, wherein, The target direction is divided into the downstream direction and the upstream direction of the network card. The downstream direction is the direction in which the network card receives data packets from the virtual network card queue, and the upstream direction is the direction in which the network card receives data packets from the physical network card queue; The redirect queues in the target direction are divided into a downstream redirect queue and an upstream redirect queue. One processor core corresponds to one downstream redirect queue to transmit data packets from the virtual network card queue, and one processor core corresponds to one upstream redirect queue to transmit data packets from the physical network card queue.

4. The method according to claim 3, wherein, When the target direction is the downstream direction, the data packet comes from the virtual network card queue, and the mapping relationship between the data stream and the redirect queue in the target direction includes: the downstream mapping relationship between the data stream and the downstream redirect queue; wherein, the downstream mapping relationship is recorded in the downstream mapping table, and the downstream mapping table records the mapping relationship between the data stream and the downstream redirect queue through multiple entries. The index of one entry is the digest value of the data packet corresponding to a data stream, and the content of one entry is the identifier of a downstream redirect queue corresponding to the mapping; Alternatively, the target direction is the upstream direction, the data packet comes from the physical network card queue, and the mapping relationship between the data stream and the redirection queue in the target direction includes: the upstream mapping relationship between the data stream and the upstream redirection queue; the upstream mapping relationship is recorded in the upstream mapping table, and the upstream mapping table records the mapping relationship between the data stream and the upstream redirection queue through multiple table entries. The index of one table entry is the digest value of the data packet corresponding to a data stream, and the content of one table entry is the identifier of an upstream redirection queue corresponding to the mapping.

5. A load balancing method, wherein, applied to a network card, the method includes: Determine the total packet load of each processor core in the current cycle. Among them, the total packet load of a processor core in the current cycle is composed of the packet loads of each target direction of the processor core in the current cycle; According to the total packet load of each processor core in the current cycle, determine whether the current cycle reaches the adjustment condition for adjusting the mapping relationship. The mapping relationship includes: the mapping relationship between the data stream and the redirection queue in the target direction, and the data stream in the mapping relationship is represented by the digest value of the data packet corresponding to the data stream; wherein, multiple processor cores of the network card implement a virtual switch, and a processor core corresponds to a redirection queue in the target direction to transmit data packets from the network card queue in the target direction; If the current cycle reaches the adjustment condition for adjusting the mapping relationship, then target to balance the packet loads of the target directions processed by each processor core, and adjust the mapping relationship between the data stream and the redirection queue in the target direction.

6. The method according to claim 5, wherein, The determining whether the current cycle reaches the adjustment condition for adjusting the mapping relationship according to the total packet load of each processor core in the current cycle includes: Determine the utilization rate of each processor core in the current cycle according to the packet load of each processor core in the current cycle; According to the utilization rate of each processor core in the current cycle, determine whether the current cycle reaches the adjustment condition for adjusting the mapping relationship; Among them, the adjustment condition includes: the number of processor cores with a utilization rate exceeding the utilization rate threshold is greater than the set number, and the utilization rate difference between the processor core with the highest utilization rate and the processor core with the lowest utilization rate exceeds the preset difference.

7. The method according to claim 5 or 6, wherein, The adjusting the mapping relationship between the data stream and the redirection queue in the target direction with the goal of balancing the packet loads of the target directions processed by each processor core includes: For any target direction, determine the average utilization rate of the target direction according to the utilization rate of the target direction of each processor core in the current cycle; among them, any processor core with a utilization rate of the target direction in the current cycle higher than the average utilization rate is a first type of processor core, and any processor core with a utilization rate of the target direction in the current cycle lower than the average utilization rate is a second type of processor core; Adjust the mapping relationship between the data flow and the redirection queue in the target direction to divert part of the data flow of the first type of processor core in the target direction to the second type of processor core for processing, so that the utilization rate of the first type of processor core and the second type of processor core in the target direction tends to the average utilization rate.

8. The method according to claim 7, in, The step of adjusting the mapping relationship between the data flow and the redirection queue in the target direction to divert part of the data flow of the first type of processor core in the target direction to the second type of processor core for processing, so that the utilization rates of the first type of processor core and the second type of processor core in the target direction tend to an average utilization rate includes: For any first-class processor core, according to the usage rate corresponding to each data flow in the target direction processed by the first-class processor core in the current cycle, the shunted data flow of the first-class processor core in the target direction is determined, wherein after the shunted data flow of the first-class processor core in the target direction is shunted, the usage rate of the first-class processor core in the target direction tends to the average usage rate of the target direction; Release the mapping relationship between the diverted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the first type of processor core in the target direction; And, diverting the diverted data flow of the first type of processor core in the target direction to at least one second type of processor core, so that the utilization rate of the second type of processor core to which the diverted data flow is diverted in the target direction tends to the average utilization rate in the target direction; A mapping relationship between the shunted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the shunted second type of processor core in the target direction is added.

9. The method according to claim 8, in, The mapping relationship between the release of the diverted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the first type of processor core in the target direction includes: Based on the summary value of the diverted data flow of the first type of processor core in the target direction and the identifier of the redirection queue corresponding to the first type of processor core in the target direction, query the corresponding table entry in the mapping table; Clear the table entries found in the mapping table; wherein the mapping table records the mapping relationship between the data flow and the redirection queue of the target direction through multiple table entries, the index of an entry is the summary value of the data packet corresponding to a data flow, and the content of an entry is the identifier of a redirection queue corresponding to the mapped target direction; The mapping relationship between the newly added shunted data flow of the first type of processor core in the target direction and the redirection queue corresponding to the shunted second type of processor core in the target direction includes: Based on the summary value of the shunted data flow of the first type of processor core in the target direction and the identifier of the redirection queue corresponding to the shunted second type of processor core in the target direction, a corresponding table entry is added to the mapping table.

10. The method according to claim 8, in, The determining, for any first-class processor core, according to the usage rate corresponding to each data flow in the target direction processed by the first-class processor core in the current cycle, of the diverted data flow of the first-class processor core in the target direction comprises: For any first - type processor core, according to the utilization rates corresponding to each data stream in the target direction processed by the first - type processor core in the current cycle, filter out the data streams with non - zero utilization rates; For the filtered data streams, in the order of increasing utilization rate, accumulate the utilization rates of the data streams until the accumulated utilization rate approaches the difference between the utilization rate of the target direction of the first - type processor core in the current cycle and the average utilization rate, and then stop accumulating the utilization rates; Use the data stream corresponding to the accumulated utilization rate as the split data stream of the first - type processor core in the target direction.

11. The method according to any one of claims 5 - 10, wherein, the target direction is divided into the downstream direction and the upstream direction of the network card. The downstream direction is the direction in which the network card receives data packets from the virtual network card queue, and the upstream direction is the direction in which the network card receives data packets from the physical network card queue; the redirection queues of the target direction are divided into a downstream redirection queue and an upstream redirection queue. One processor core corresponds to one downstream redirection queue to transmit data packets from the virtual network card queue, and one processor core corresponds to one upstream redirection queue to transmit data packets from the physical network card queue.

12. The method according to claim 11, wherein, the mapping relationship between the data stream and the redirection queue of the target direction is divided into a downstream mapping relationship between the data stream and the downstream redirection queue, and an upstream mapping relationship between the data stream and the upstream redirection queue; wherein, the downstream mapping relationship is recorded in the downstream mapping table. The downstream mapping table records the mapping relationship between the data stream and the downstream redirection queue through multiple table entries. The index of one table entry is the digest value of the data packet corresponding to a data stream, and the content of one table entry is the identifier of a downstream redirection queue corresponding to the mapping; the upstream mapping relationship is recorded in the upstream mapping table. The upstream mapping table records the mapping relationship between the data stream and the upstream redirection queue through multiple table entries. The index of one table entry is the digest value of the data packet corresponding to a data stream, and the content of one table entry is the identifier of an upstream redirection queue corresponding to the mapping.

13. A network card, wherein, comprises: a hardware acceleration engine and an on - chip processor; the hardware acceleration engine is provided with a redirection module, multiple downstream redirection queues, and multiple upstream redirection queues; the on - chip processor is provided with multiple processor cores, the multiple processor cores implement a virtual switch, and the virtual switch is provided with a load adjustment module; wherein, the redirection module is connected between the multiple downstream redirection queues and the virtual network card queue, and between the multiple upstream redirection queues and the physical network card queue; one processor core corresponds to one downstream redirection queue to transmit data packets from the virtual network card queue, and one processor core corresponds to one upstream redirection queue to transmit data packets from the physical network card queue; the redirection module is configured to execute the load - balancing method according to any one of claims 1 - 4; the load adjustment module is configured to execute the load - balancing method according to any one of claims 5 - 12.

14. A storage medium, wherein, The storage medium stores one or more computer-executable instructions, which, when executed, implement the load balancing method described in any one of claims 1-4, or the load balancing method described in any one of claims 5-12.

Citation Information

Cited By

  • Data processing method, storage medium, electronic equipment and program product

    CN120602416A