Network performance optimization method and optimization device

By obtaining and binding the interrupt request of the network card to the corresponding CPU core in the server system with multi-NUMA node architecture, the network performance problems caused by improper software and hardware matching are solved, and more efficient data processing and a more stable network environment are achieved.

CN120196437APending Publication Date: 2025-06-24LENOVO (BEIJING) LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510265385.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In server systems with multi-NUMA node architecture, improper software and hardware matching will affect the transmission and reception performance of service network messages, resulting in excessive network delays, long-tail problems, packet loss and retransmission, reducing service quality and user experience. The existing technology requires users to understand and customize hardware topology, which increases the threshold for use and technical difficulty.

Method used

By obtaining the NUMA node information loaded by the operating system kernel and the NUMA node information where the network card is located, we judge and bind the interrupt request of the network card to the CPU core of the NUMA node where the network card is located, ensuring that data processing is carried out within the local memory range and reducing cross-node access delay. Optional steps include modifying the bootloader to ensure that the kernel loads to the same NUMA node as the network card and starting the process and memory migration service.

Benefits of technology

Through software-level optimization, network performance problems caused by software and hardware mismatch are solved, cross-node data transmission delay is reduced, network processing efficiency and overall system performance are improved, network jitter and instability factors are reduced, and a more stable network environment and a better user experience is provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196437A_ABST
    Figure CN120196437A_ABST
Patent Text Reader

Abstract

The invention discloses a network performance optimization method and device, and is applied to a server system of a multi-non-uniform memory access node architecture, and the network performance optimization method comprises the steps: obtaining non-uniform memory access node information loaded by an operating system kernel and non-uniform memory access node information where a network card is located; if it is judged that the non-uniform memory access node loaded by the operating system kernel is the same as the non-uniform memory access node where the network card is located, checking the interrupt request binding condition of the network card; according to the interrupt request binding condition of the network card, if the interrupt request of the network card is not bound to the CPU core of the non-uniform memory access node where the network card is located, binding the interrupt request of the network card to the CPU core of the non-uniform memory access node where the network card is located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and particularly to a network performance optimization method and an optimization device. Background Art

[0002] In the actual operating environment of a server system with a multi-NUMA (Non-Uniform Memory Access) node architecture, users often plug network cards into the CPUs of any NUMA node. However, when the server system runs tasks sensitive to network latency, if the software and hardware are not properly matched, it will seriously affect the sending and receiving performance of service network packets. For example, it may cause excessive network latency, resulting in the long-tail problem of services, and packet loss and retransmission under high-pressure network loads, thereby reducing the overall quality of services and the user experience.

[0003] To solve this problem, a fixed hardware topology structure is often adopted in the prior art. Specifically, the network card is fixedly plugged into the CPU of a specific NUMA node and combined with the upper-layer container scheduling to avoid the problem of software and hardware mismatch. However, this method has certain limitations. It requires users to have in-depth and comprehensive understanding of their own software and hardware environment and requires standard customization of the hardware topology. This not only increases the usage threshold and technical difficulty, but also makes it difficult for some users without professional technical capabilities or with complex and changeable hardware environments to implement, and it is difficult to meet diverse actual application requirements. Summary of the Invention

[0004] An embodiment of this application provides a network performance optimization method, which is applied to a server system with a multi-non-uniform memory access architecture and includes:

[0005] Obtain the non-uniform memory access node information loaded by the operating system kernel and the non-uniform memory access node information where the network card is located;

[0006] If it is determined that the non-uniform memory access node loaded by the operating system kernel is the same as the non-uniform memory access node where the network card is located, check the binding situation of the interrupt request of the network card;

[0007] According to the binding situation of the interrupt request of the network card, if the interrupt request of the network card is not bound to the CPU core of the non-uniform memory access node where the network card is located, bind the interrupt request of the network card to the CPU core of the non-uniform memory access node where the network card is located.

[0008] Optionally, if the non-uniform memory access node loaded by the operating system kernel is different from the non-uniform memory access node where the network card is located, modify and update the boot loader so that the kernel is loaded onto the same non-uniform memory access node as the network card when the operating system restarts.

[0009] Optionally, if the interrupt request of the network card has been bound to the CPU core of the non-uniform memory access node where the network card is located, start the process migration service and the memory migration service to migrate the processes and memory that need to ensure network performance to the non-uniform memory access node where the network card is located.

[0010] Optionally, before binding the interrupt request of the network card to the CPU core of the non-uniform memory access node where the network card is located, it includes:

[0011] Obtain the information of the CPU cores in each non-uniform memory access node;

[0012] Turn off the service for optimizing interrupt request allocation.

[0013] Optionally, binding the interrupt request of the network card to the CPU core of the non-uniform memory access node where the network card is located includes:

[0014] Obtain the names of each network port in the network card;

[0015] Obtain the interrupt number of each network port and the interrupt request corresponding to each interrupt number according to the network port name;

[0016] Bind each interrupt request to the CPU core of the non-uniform memory access node where the network card is located.

[0017] Optionally, binding each interrupt request to the CPU core of the non-uniform memory access node where the network card is located includes:

[0018] Query the interrupt number of the current network port and its corresponding interrupt request;

[0019] If it is determined that the interrupt request corresponding to the interrupt number of the current network port is not bound to the CPU core of the non-uniform memory access node where the network card is located, then bind the interrupt request corresponding to the interrupt number of the current network port to the CPU core of the non-uniform memory access node where the network card is located;

[0020] Query the next network port until all network ports are queried.

[0021] Optionally, if it is determined that the interrupt request corresponding to the interrupt number of the current network port has been bound to the CPU core of the non-uniform memory access node where the network card is located, then continue to query the next network port until all network ports are queried.

[0022] An embodiment of the present application also provides a network performance optimization device, which is applied to a server system with a multi-non-uniform memory access architecture, including:

[0023] A first acquisition module configured to acquire the non-uniform memory access node information loaded by the operating system kernel and the non-uniform memory access node information where the network card is located;

[0024] A checking module, configured to check the binding situation of the interrupt request of a network card if it is determined that the non-uniform memory access node loaded by the operating system kernel is the same as the non-uniform memory access node where the network card is located;

[0025] A binding module, configured to bind the interrupt request of the network card to the CPU core of the non-uniform memory access node where the network card is located if the interrupt request of the network card is not bound to the CPU core of the non-uniform memory access node where the network card is located.

[0026] Optionally, the network performance optimization device further includes:

[0027] A modification and update module, configured to modify and update the boot loader if the non-uniform memory access node loaded by the operating system kernel is different from the non-uniform memory access node where the network card is located, so that the kernel is loaded to the same non-uniform memory access node as the network card when the operating system restarts.

[0028] An embodiment of the present application further provides an electronic device, including a memory and a processor, where an executable program is stored in the memory, and the processor executes the executable program to perform the steps of the method described in any one of the above. Description of the Drawings

[0029] Figure 1 It is a flowchart of the network performance optimization method according to an embodiment of the present application;

[0030] Figure 2 It is for Figure 1 the first embodiment flowchart of step S300 in an embodiment of the present application;

[0031] Figure 3 It is for Figure 1 the second embodiment flowchart of step S300 in an embodiment of the present application;

[0032] Figure 4 It is for Figure 3 the flowchart of step S350 in an embodiment of the present application;

[0033] Figure 5 It is a structural block diagram of the network performance optimization device according to an embodiment of the present application. Detailed Embodiments

[0034] Various solutions and features of the present application are described herein with reference to the drawings.

[0035] It should be understood that various modifications can be made to the embodiments applied herein. Therefore, the above description should not be regarded as a limitation, but only as an example of the embodiments. Those skilled in the art will think of other modifications within the scope and spirit of the present application.

[0036] The accompanying drawings, which are included in and form a part of this specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.

[0037] These and other features of the present application will become apparent from the following description of the preferred forms of the embodiments given by way of non-limiting example with reference to the accompanying drawings.

[0038] It should also be understood that although the present application has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present application.

[0039] When taken in conjunction with the accompanying drawings, the above and other aspects, features, and advantages of the present application will become more apparent in view of the following detailed description.

[0040] Specific embodiments of the present application will be described hereinafter with reference to the accompanying drawings; however, it should be understood that the embodiments claimed are merely examples of the present application and can be implemented in various ways. Well-known and / or repetitive functions and structures are not described in detail to avoid obscuring the present application with unnecessary or redundant details. Thus, the specific structural and functional details claimed herein are not intended to be limiting, but rather are merely a basis and representative basis for the claims to teach those skilled in the art to use the present application in substantially any suitable detailed structure in a variety of ways.

[0041] This specification may use the phrases "in one embodiment", "in another embodiment", "in yet another embodiment", or "in other embodiments", each of which may refer to one or more of the same or different embodiments according to the present application.

[0042] A server system with a multi-NUMA node architecture is a high-performance computer architecture that is widely used in modern data centers and high-performance computing fields.

[0043] A server system with a multi-NUMA node architecture contains multiple NUMA nodes. Each NUMA node typically includes one or more CPUs, local memory, and related I / O devices, etc., where the I / O devices include network cards, storage device controllers, etc. The CPUs within a NUMA node can directly and quickly access the local memory through the local bus or interconnection structure, while when accessing the memory of other NUMA nodes, since data transmission needs to be carried out through the interconnection structure between nodes, there will be a relatively high access latency compared to accessing the local memory. This non-uniformity of memory access is the origin of the name NUMA (Non-Uniform Memory Access).

[0044] For example, a server with 2 NUMA nodes, each NUMA node has 4 CPUs, and each node has 1 independent local memory. CPUs within the same NUMA node can directly and quickly access the memory in that node, while a CPU in one NUMA node needs to access the memory in another NUMA node through the inter-node interconnection structure.

[0045] In a server system with a multi-NUMA node architecture, in the prior art, network cards are usually associated with specific NUMA nodes through a fixed hardware topology. In this way, when the server system receives data from an external network through a network card, the data can be preferentially sent to the associated NUMA node for processing, reducing the transmission delay of data within the server. At the same time, the operating system will also reasonably allocate network processing tasks to the CPUs of the corresponding nodes according to the association relationship between the network card and the NUMA nodes. However, this way of associating network cards with NUMA nodes not only requires high knowledge reserves and technical capabilities of users, but also requires standard customization of the hardware topology, and cannot flexibly respond to changes in different hardware environments, and cannot meet diverse actual application requirements.

[0046] A network performance optimization method according to an embodiment of the present application coordinates the association relationship between the network card and each NUMA node from a software perspective, has no requirements for the user's knowledge reserves and professional capabilities, can adapt to different hardware environments, and reasonably allocates the network processing tasks of the network card to the CPU cores of the corresponding NUMA nodes, thereby optimizing the performance of the network for data interaction between the server system with a multi-NUMA node architecture and external devices.

[0047] The network performance optimization method of the present application will be described in detail below with reference to the accompanying drawings. Figure 1 is a flowchart of the network performance optimization method according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:

[0048] S100. Obtain the non-uniform memory access node information loaded by the operating system kernel and the non-uniform memory access node information where the network card is located.

[0049] The operating system kernel, as the core part of the operating system, undertakes the important responsibilities of managing system resources (such as memory, CPU time, and I / O devices, etc.) and providing basic services (such as process management, file system management, and network communication, etc.).

[0050] When the computer starts up, a series of hardware initialization operations are first performed, and then the operating system kernel is loaded into memory. Next, the CPU starts to read the kernel code from memory and execute this code in the order of instructions. During operation, the kernel schedules CPU resources to execute different tasks according to the system's requirements and the current state. For example, running user programs, processing requests from device drivers, etc. The operating system kernel runs on the CPU and utilizes the computing power of the CPU to complete various functions. The operating system kernel usually loads NUMA node information during the computer system startup phase to facilitate subsequent reasonable management and optimization of system resources.

[0051] S200. If it is determined that the non-uniform memory access node loaded by the operating system kernel is the same as the non-uniform memory access node where the network card is located, then check the binding situation of the interrupt requests of the network card.

[0052] In a server system with a multi-NUMA node architecture, when the NUMA node loaded by the operating system kernel is the same as the NUMA node where the network card is located, it means that the network card and some resources managed by the operating system kernel (such as local memory and CPU cores, etc.) are on the same NUMA node. Since the latency caused by cross-node access is reduced, when the NUMA node loaded by the operating system kernel is the same as the NUMA node where the network card is located, theoretically the efficiency of data transmission and processing will be higher, and it is also necessary to further judge the binding situation of the interrupt requests of the network card.

[0053] If the NUMA node loaded by the operating system kernel is inconsistent with the NUMA node where the network card is located, then when the operating system kernel processes the data received by the network card, cross-node memory access needs to be performed through the system bus, which will increase the latency of data transmission.

[0054] When the network card receives network data or other events that require CPU processing, it will send an interrupt request signal to the CPU to notify the CPU that there is a task to be processed. Interrupt request binding can be understood as in a multi-CPU core system, the operating system needs to decide which CPU core to allocate these interrupt requests to for processing.

[0055] Reasonable interrupt request binding can improve system performance. If the interrupt requests of the network card are bound to the CPU core that belongs to the same NUMA node as the network card, then during the data processing process, when the CPU accesses network card-related data (such as data in the receive buffer), it can utilize local memory and reduce the overhead of cross-node memory access. On the contrary, if the binding is unreasonable, for example, the network card interrupt requests are bound to the CPU cores of other NUMA nodes, even if the NUMA nodes where the network card and the kernel are located are the same, the performance may still be reduced because the CPU needs to cross nodes when accessing network card data.

[0056] S300. According to the binding situation of the interrupt requests of the network card, if the interrupt requests of the network card are not bound to the CPU cores of the non-uniform memory access node where the network card is located, then bind the interrupt requests of the network card to the CPU cores of the non-uniform memory access node where the network card is located.

[0057] In the NUMA architecture, each NUMA node has its own local memory and CPU cores. When the interrupt requests of the network card are bound to the CPU cores of the NUMA node where it is located, data transmission can be carried out within the local NUMA node, which can avoid the additional latency caused by accessing memory across NUMA nodes. Because accessing memory across NUMA nodes usually requires data transmission through the system bus, which will increase the time overhead of data transmission.

[0058] Binding the interrupt requests of the network card to the appropriate CPU cores can make the processing tasks of the network card more concentrated and efficient. The communication between the CPU cores and the network card within the same NUMA node is closer, and the data processing efficiency is higher, thus improving the performance of the entire system. For example, in the case of high network load, this binding can reduce the competition and conflicts between CPU cores and improve the processing speed of network data.

[0059] Through reasonable interrupt binding, the resources of each NUMA node can be better utilized. Avoid the situation where the CPU cores of some NUMA nodes are overloaded while the resources of other NUMA nodes are idle, and the balanced allocation and optimized utilization of system resources can be achieved.

[0060] Exemplarily, in the Linux system, the irqbalance (interrupt balancer) tool or manually modifying the relevant files in the proc file system can be used to achieve the binding of the network card interrupt requests to the CPU cores. Among them, the proc file system is a special virtual file system that is dynamically created during the operation of the operating system and is used to reflect the current state of the operating system.

[0061] Specifically, manually modifying the relevant files in the proc file system to achieve the binding of the network card interrupt requests to the CPU cores includes: The following are the general steps for manual binding:

[0062] Determine the interrupt number of the network card:

[0063] The interrupt number of the corresponding network card can be found by viewing the / proc / interrupts file.

[0064] Determine the NUMA node where the network card is located:

[0065] The numactl tool or viewing the relevant system information can be used to determine the NUMA node where the network card is located.

[0066] Determine the CPU core to which the interrupt is to be bound:

[0067] Determine the CPU core to which the interrupt is to be bound according to the NUMA node where the network card is located.

[0068] Bind the interrupt request:

[0069] Enter the / proc / irq / <interrupt number> / smp_affinity file, which stores the CPU core information to which the current interrupt request is bound. Modify the value of the / proc / irq / <interrupt number> / smp_affinity file to the mask value of the CPU core to which the interrupt is to be bound. For example, assuming there are multiple CPU cores in the system, if the interrupt is to be bound to CPU cores 0 and 1, then set the value of the / proc / irq / <interrupt number> / smp_affinity file to 0x03 (binary: 00000011). A text editing tool (e.g., the echo command) can be used to modify the value of this file. For example, echo 0x03 > / proc / irq / <interrupt number> / smp_affinity.

[0070] Of course, in the Windows system, the CPU affinity of the interrupt can also be set through tools such as the Device Manager.

[0071] In one embodiment, in the above step S200, if the NUMA node loaded by the operating system kernel is different from the NUMA node where the network card is located, then modify and update the boot loader so that when the operating system restarts, the kernel is loaded onto the same non-uniform memory access node as the network card.

[0072] In a multi-NUMA node architecture, the speed at which the CPU accesses the memory within the same NUMA node (i.e., local memory) is much faster than accessing the memory of other NUMA nodes (i.e., remote memory). If the memory operations involved in network card data processing span NUMA nodes, it will increase the data transfer latency and reduce the network I / O performance. Modifying and updating the boot loader to load the operating system kernel onto the same NUMA node as the network card can enable the data transfer between the network card and the memory to occur within the local NUMA node, reducing the cross-node data transfer overhead, thereby improving the overall performance of the system and the network I / O efficiency.

[0073] A network card is a hardware device for a computer to connect to a network, responsible for receiving and sending network data. When the network card receives data, it sends an interrupt request signal to the CPU core through the hardware interrupt mechanism. When the CPU core receives the interrupt request from the network card, if the interrupt response condition is met (e.g., the CPU core is not in an uninterruptible state), then the CPU core pauses the currently executing task and instead executes the interrupt handler.

[0074] The operating system kernel is the core part of the operating system. It runs on the CPU and manages various resources of the computer system, including the CPU itself. The instructions and tasks executed by the CPU core are all carried out under the scheduling and management of the kernel. The kernel is responsible for allocating tasks to the CPU, managing the life cycle of processes, handling various system calls, etc. At the same time, the kernel also provides various driver interfaces for interacting with hardware devices, including network card drivers.

[0075] The operating system kernel communicates with the network card through the network card driver. The network card driver is a part of the kernel and it implements the specific functions of interacting with the network card hardware. When the CPU core receives an interrupt request from the network card and executes the interrupt handler, it is actually executing the interrupt service routine in the network card driver. This routine runs in the context of the kernel and is responsible for processing the data received by the network card. For example, it parses the data, stores it in the appropriate memory location, and notifies the relevant upper-layer applications that new data has arrived. In addition, the kernel also sends commands to the network card through the network card driver, such as configuring the parameters of the network card, controlling the startup and stop of the network card, etc.

[0076] The collaborative working process of the network card, CPU core, and operating system kernel is as follows:

[0077] After the network card receives data, it sends an interrupt request to the CPU core. The CPU core pauses the current task and instead executes the interrupt service routine provided by the network card driver in the operating system kernel. In this process, the kernel coordinates and manages the entire data processing flow, including memory allocation, data storage, and process scheduling, etc., to ensure that the data received by the network card can be correctly processed and delivered to the corresponding application programs.

[0078] In one embodiment, in the above step S300, if the interrupt request of the network card is bound to the CPU core of the non-uniform memory access node where the network card is located, then the process migration service and the memory migration service are started to migrate the processes and memory that need to ensure network performance to the non-uniform memory access node where the network card is located.

[0079] When the network card interrupt request is bound to the CPU core of the NUMA node where the network card is located, it means that this CPU core processes the network data related to the network card. If the processes and memory related to this network data processing are also migrated to the same node at this time, it can minimize the latency of the CPU accessing the memory, reduce the overhead of cross-node data transmission, and improve the overall performance of the system. For example, in the scenario of processing a large number of network packets, fast memory access can make the packet processing more efficient and reduce network latency.

[0080] Migrate the process and memory to the NUMA node where the CPU core bound to the network card interrupt request is located, so that when the CPU processes network-related tasks, it can obtain the required data from the local cache. Since the CPU is more likely to hit the cache when accessing data in local memory, the number of times of reading data from the main memory is reduced, thereby improving the data processing speed. This is very important for network applications with high performance requirements (such as high-speed network servers, real-time data processing systems, etc.).

[0081] The process migration service can schedule processes to appropriate nodes according to the load conditions of the nodes, and the memory migration service can reasonably allocate memory resources to different NUMA nodes. Combining the binding of network card interrupt requests and CPU cores can better achieve the optimal utilization of resources. For example, when the CPU cores on a certain NUMA node are highly loaded due to network card interrupt processing and related process operations, the process migration service can migrate some non-critical processes to other nodes with lower loads. At the same time, the memory migration service can also adjust the memory layout to make the loads of each node more balanced, avoiding the situation where one node is overly busy while other nodes are idle.

[0082] Specifically, the operating system can migrate a process from one NUMA node to another through the function interface of process migration. During the migration process, the operating system saves the current state of the process, including information such as the process's address space, register values, and open file descriptors, and restores these states on the target node so that the process can continue to run.

[0083] The operating system can use page migration technology to implement memory migration. The operating system migrates memory data from the source NUMA node to the target NUMA node (i.e., the NUMA node where the CPU core bound to the network card interrupt request is located) page by page. During the migration process, the operating system records the state of each page to ensure the integrity and consistency of the data.

[0084] Exemplarily, through the process migration interface and the memory migration interface, services sensitive to network latency can be automatically migrated to the CPU core in the NUMA node loaded by the operating system kernel and the memory on the same NUMA node, thereby significantly reducing memory access latency, improving the response speed of services, reducing cross-node communication overhead, and improving system resource utilization.

[0085] In one embodiment, in the above step S300, as Figure 2 shown, before binding the interrupt request of the network card to the CPU core of the NUMA node where the network card is located, it includes:

[0086] S310. Obtain the information of the CPU cores in each NUMA to determine the CPU core to which the interrupt request of the network card is bound;

[0087] It can be understood that by obtaining the information of CPU cores in each NUMA node, it can be determined which CPU core in which NUMA node the interrupt request of the network card is currently bound to. If the interrupt request of the network card is not currently bound to the CPU core of the NUMA node where the network card is located, the interrupt request of the network card is rebound to the CPU core of the NUMA node where the network card is located, so that operations related to network card data processing can be carried out within the local memory range as much as possible, reducing the latency caused by cross-node memory access, thereby significantly improving the overall performance of the system. For example, when processing a large amount of network data, if the network card interrupt request is bound to the CPU core of a non-local node, more time will be spent on memory access during the data transmission process, while binding to the CPU core of the local node can effectively avoid this situation.

[0088] After obtaining the information of the CPU cores of the NUMA node where the network card is located, the network card interrupt request can be more reasonably allocated to the appropriate cores, avoiding the situation where some CPU cores are overloaded while other cores are underloaded. For example, when there are multiple network cards, if the interrupt binding is not performed according to the information of the CPU cores of the NUMA node, it may cause one or several CPU cores to simultaneously process the interrupt requests of multiple network cards, resulting in excessive load on these cores, while other cores are idle. Through reasonable binding, the CPU load can be more balanced and the overall efficiency of the system can be improved.

[0089] S320. Turn off the service for optimizing interrupt request allocation.

[0090] Before binding the interrupt request of the network card to the CPU core of the NUMA node where the network card is located, turning off the service for optimizing interrupt request allocation can avoid conflicts between the automatic allocation mechanism of this service and the binding operation, ensuring that the interrupt request of the network card is stably bound to the specified CPU core. Turning off the service for optimizing interrupt request allocation can make the binding operation of the interrupt request of the network card not affected by other automatic allocation mechanisms, thereby ensuring that the interrupt request of the network card is indeed bound to the CPU core of the NUMA node where the network card is located, so as to play the expected role of performance improvement, such as reducing memory access latency and improving data processing efficiency.

[0091] In one embodiment, as Figure 3 shown, binding the interrupt request of the network card to the CPU core of the NUMA node where the network card is located includes:

[0092] S330. Obtain the names of each network port in the network card;

[0093] S340. Obtain the interrupt numbers of each network port and the interrupt requests corresponding to each interrupt number according to the network port names;

[0094] S350. Bind each interrupt request to the CPU core of the non-uniform memory access node where the network card is located.

[0095] Specifically, in a server system with a multi-NUMA node architecture, network cards are distributed on different NUMA nodes. Network cards on different NUMA nodes are relatively independent in terms of system resource allocation. The CPU, memory, and other resources on each node give priority to serving the network cards on the local node to improve data processing efficiency. Network ports, as physical interfaces on the network card, can be used to connect different types of network devices.

[0096] The network port name is a string used to identify each network port in the server, and the network port name can be obtained through the ifconfig command. The network port interrupt number is the identification number used when the network interface sends an interrupt request to the CPU core. When the network port receives data or completes data transmission and other operations, the corresponding interrupt number will be triggered to notify the CPU core for processing. Different interrupt numbers can be assigned to the network ports of different NUMA nodes to achieve resource isolation and optimization. This can avoid interference between the network port interrupts of different nodes and improve the stability and performance of the system.

[0097] The types of interrupt requests can be receive data interrupt and send data interrupt, etc. When the network port receives a new data frame, a receive data interrupt will be triggered to notify the CPU core of the NUMA node where it is located to read and process the data; a send data interrupt will be generated after the data transmission is completed.

[0098] In this embodiment, as Figure 4 shown, binding each interrupt request to the CPU core of the non-uniform memory access node where the network card is located includes:

[0099] S3510. Query the interrupt number of the current network port and its corresponding interrupt request.

[0100] S3520. If it is determined that the interrupt request corresponding to the interrupt number of the current network port is not bound to the CPU core of the non-uniform memory access node where the network card is located, then bind the interrupt request corresponding to the interrupt number of the current network port to the CPU core of the non-uniform memory access node where the network card is located.

[0101] S3530. Query the next network port until all network ports have been traversed and queried.

[0102] Further, if it is determined that the interrupt request corresponding to the interrupt number of the current network port has been bound to the CPU core of the non-uniform memory access node where the network card is located, then continue to query the next network port until all network ports have been queried.

[0103] The network performance optimization method provided by this application can, through software-level optimization, effectively solve the problems caused by software-hardware mismatch without changing the hardware topology. Moreover, it does not require users to have professional technical capabilities, nor does it require a standard customized hardware topology structure. The implementation operation is simple and can meet diverse actual application requirements.

[0104] The network performance optimization method provided by this application can reduce network jitter and unstable factors caused by cross-node data transmission and resource competition, enhance the stability and reliability of the network, optimize network performance, provide a more stable network environment for the server system, and also provide a better service experience for users by ensuring that the network card and the operating system kernel are in the same NUMA node and binding the interrupt request of the network card to the CPU core of the NUMA node where the network card is located.

[0105] Based on the same inventive concept, as Figure 5 shown, the embodiment of this application also provides a network performance optimization device, which is applied to a server system with a multi-non-uniform memory access architecture and includes:

[0106] A first acquisition module configured to acquire the non-uniform memory access node information loaded by the operating system kernel and the non-uniform memory access node information where the network card is located.

[0107] An inspection module configured to check the binding situation of the interrupt request of the network card if it is determined that the non-uniform memory access node loaded by the operating system kernel is the same as the non-uniform memory access node where the network card is located.

[0108] A binding module configured to, according to the binding situation of the interrupt request of the network card, if the interrupt request of the network card is not bound to the CPU core of the non-uniform memory access node where the network card is located, bind the interrupt request of the network card to the CPU core of the non-uniform memory access node where the network card is located.

[0109] In the embodiment of this application, if the interrupt request of the network card has been bound to the CPU core of the non-uniform memory access node where the network card is located, the process migration service and the memory migration service are started to migrate the processes and memory that need to ensure service quality to the CPU of the non-uniform memory access node where the network card is located, so as to ensure that the network card, the operating system kernel, the interrupt request of the network card, the processes, and their required memory are all in the same non-uniform memory access node.

[0110] The network performance optimization device provided by the embodiment of this application further includes:

[0111] A modification and update module, configured to modify and update the boot loader if the non-uniform memory access (NUMA) node loaded by the operating system kernel is different from the NUMA node where the network card is located, so that when the operating system restarts, the kernel is loaded onto the NUMA node that is the same as the network card.

[0112] Furthermore, the network performance optimization device provided by the embodiment of the present application further includes:

[0113] A second acquisition module, configured to acquire information about CPU cores in each NUMA before binding the interrupt request of the network card to the CPU core of the NUMA node where the network card is located;

[0114] A service shutdown module, configured to shut down the service for optimizing the distribution of interrupt requests before binding the interrupt request of the network card to the CPU core of the NUMA node where the network card is located.

[0115] Optionally, the binding module is further configured to:

[0116] Acquire the names of each network interface in the network card;

[0117] Acquire the interrupt number of each network interface and the interrupt request corresponding to each interrupt number according to the network interface name;

[0118] Binding each interrupt request to the CPU core of the non-uniform memory access node where the network card is located includes:

[0119] Query the interrupt number of the current network interface and its corresponding interrupt request;

[0120] If it is determined that the interrupt request corresponding to the interrupt number of the current network interface is not bound to the CPU core of the non-uniform memory access node where the network card is located, then bind the interrupt request corresponding to the interrupt number of the current network interface to the CPU core of the non-uniform memory access node where the network card is located;

[0121] Query the next network interface until all network interfaces have been traversed and queried.

[0122] The embodiment of the present application also provides an electronic device, including a memory and a processor, where an executable program is stored in the memory, and the processor executes the executable program to perform the steps of the method described in any one of the above.

[0123] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements to the present application within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.

Claims

1. A network performance optimization method, applied to a server system with multiple non-uniform memory access node architectures, comprising: Obtain the non-uniform memory access node information loaded by the operating system kernel and the non-uniform memory access node information where the network card is located; If it is determined that the non-uniform memory access node loaded by the operating system kernel is the same as the non-uniform memory access node where the network card is located, then check the interrupt request binding status of the network card; According to the interrupt request binding status of the network card, if the interrupt request of the network card is not bound to the CPU core of the non-uniform memory access node where the network card is located, the interrupt request of the network card is bound to the CPU core of the non-uniform memory access node where the network card is located.

2. According to the method described in claim 1, if the non-uniform memory access node loaded by the operating system kernel is different from the non-uniform memory access node where the network card is located, the boot loader is modified and updated so that when the operating system is restarted, the kernel is loaded onto the same non-uniform memory access node as the network card.

3. According to the method described in claim 1, if the interrupt request of the network card has been bound to the CPU core of the non-uniform memory access node where the network card is located, the process migration service and the memory migration service are started to migrate the processes and memory that need to ensure network performance to the non-uniform memory access node where the network card is located.

4. The method according to claim 1, before binding the interrupt request of the network card to the CPU core of the non-uniform memory access node where the network card is located, comprises: Get the information of CPU cores in each non-uniform memory access node; Turns off the service for optimizing interrupt request distribution.

5. The method according to claim 4, wherein binding the interrupt request of the network card to the CPU core of the non-uniform memory access node where the network card is located comprises: Get the name of each network port in the network card; Obtain the interrupt number of each network port and the interrupt request corresponding to each interrupt number according to the network port name; Bind each interrupt request to the CPU core of the non-uniform memory access node where the network card is located.

6. The method according to claim 5, wherein binding each interrupt request to a CPU core of the non-uniform memory access node where the network card is located comprises: Query the interrupt number of the current network port and its corresponding interrupt request; If it is determined that the interrupt request corresponding to the interrupt number of the current network port is not bound to the CPU core of the non-uniform memory access node where the network card is located, then the interrupt request corresponding to the interrupt number of the current network port is bound to the CPU core of the non-uniform memory access node where the network card is located; Query the next network port until all network ports are queried.

7. According to the method of claim 6, if it is determined that the interrupt request corresponding to the interrupt number of the current network port has been bound to the CPU core of the non-uniform memory access node where the network card is located, continue to query the next network port until all network ports are queried.

8. A network performance optimization device, applied to a server system with multiple non-uniform memory access architectures, comprising: A first acquisition module is configured to acquire non-uniform memory access node information loaded by the operating system kernel and non-uniform memory access node information where the network card is located; A checking module configured to check the interrupt request binding status of the network card if it is determined that the non-uniform memory access node loaded by the operating system kernel is the same as the non-uniform memory access node where the network card is located; The binding module is configured to bind the interrupt request of the network card to the CPU core of the non-uniform memory access node where the network card is located if the interrupt request of the network card is not bound to the CPU core of the non-uniform memory access node where the network card is located.

9. The apparatus according to claim 8, further comprising: The modification and update module is configured to modify and update the boot loader if the non-uniform memory access node loaded by the operating system kernel is different from the non-uniform memory access node where the network card is located, so that when the operating system is restarted, the kernel is loaded to the same non-uniform memory access node as the network card.

10. An electronic device, comprising a memory and a processor, wherein an executable program is stored in the memory, and the processor executes the executable program to perform the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Method and device for improving execution efficiency of compute-intensive tasks and medium

    CN121166378A