A method, apparatus, device, medium, and product for network congestion handling

By determining the congestion path in the network cluster and selecting the light-load path based on real-time traffic information, establishing mapping relationships to schedule message transmission, the congestion problem in the network cluster is solved, and load balancing and stability improvement is achieved.

CN119484417BActive Publication Date: 2025-07-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510049265.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-07-22
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

Congestion caused by limited network resources in network clusters leads to degradation of network performance, increased latency and increased packet loss rate, and may even lead to network cluster crashes, which are difficult to effectively alleviate the existing technology.

Method used

By determining the first network path in which congestion occurs in the network cluster, obtaining its corresponding first transmission multitude, and deciding from the multiple network paths based on real-time traffic information, a mapping relationship between the first transmission multitude and the second transmission multitude is established, and the packet to be transmitted is scheduled to be transmitted to the second network path for transmission.

Benefits of technology

Load balancing is achieved, the transmission pressure of the first network path is alleviated, the stability of the network cluster is improved, and the network cluster is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119484417B_ABST
    Figure CN119484417B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a network congestion handling method, apparatus, device, medium, and product. The method includes: determining a first network path that is congested in a network cluster, and obtaining a first transmission tuple corresponding to the first network path; based on the real-time traffic information of multiple network paths in the network cluster, determining a second network path from the multiple network paths, and obtaining a second transmission tuple corresponding to the second network path; the transmission load of the second network path satisfies the light load condition; establishing a mapping relationship between the first transmission tuple and the second transmission tuple; when there is a first packet with the first transmission tuple to be transmitted in the first network path, scheduling the first packet to the second network path for transmission based on the mapping relationship. The embodiment of the present application is beneficial to improving the stability of the network cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, particularly to the field of data transmission technology, and specifically relates to a network congestion handling method, a network congestion handling device, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] A network cluster is a computing system in which multiple computers are connected together through a high-speed network and work collaboratively to complete specific tasks. Network clusters have characteristics such as high computing power, high storage capacity, high reliability, and flexibility, and are widely used in fields such as scientific computing, big data processing, cloud computing, and artificial intelligence. Network congestion refers to the phenomenon in a network cluster where, due to the limited nature of network resources (such as bandwidth, memory, etc.), when the network load exceeds its capacity, the network performance deteriorates. This phenomenon usually leads to an increase in network latency, an increase in the packet loss rate, and may even cause the network cluster to crash. How to alleviate network congestion and improve the stability of the network cluster is a technical problem urgently to be solved in the field of data transmission technology. Summary of the Invention

[0003] Embodiments of this application provide a network congestion handling method, device, equipment, medium, and product, which are beneficial to improving the stability of the network cluster.

[0004] On the one hand, embodiments of this application provide a network congestion handling method, which includes:

[0005] Determine a first network path with congestion in the network cluster, and obtain a first transmission tuple corresponding to the first network path;

[0006] Based on the real-time traffic information of multiple network paths in the network cluster, decide on a second network path from the multiple network paths, and obtain a second transmission tuple corresponding to the second network path; the transmission load of the second network path meets the light-load condition;

[0007] Establish a mapping relationship between the first transmission tuple and the second transmission tuple;

[0008] When there is a first packet with the first transmission tuple to be transmitted in the first network path, schedule the first packet to the second network path for transmission based on the mapping relationship.

[0009] On the one hand, embodiments of this application provide a network congestion handling device, which includes:

[0010] An obtaining unit, configured to determine a first network path with congestion in the network cluster, and obtain a first transmission tuple corresponding to the first network path;

[0011] A processing unit, configured to determine a second network path from multiple network paths based on real-time traffic information of the multiple network paths in a network cluster, and obtain a second transmission tuple corresponding to the second network path; the transmission load of the second network path meets the light-load condition;

[0012] The processing unit is further configured to establish a mapping relationship between the first transmission tuple and the second transmission tuple;

[0013] The processing unit is further configured to, when there is a first packet with the first transmission tuple to be transmitted in the first network path, schedule the first packet to the second network path for transmission based on the mapping relationship.

[0014] In one embodiment, any network path in the network cluster includes at least one network port; the processing unit determines the first network path congested in the network cluster, including:

[0015] When detecting a congestion event occurring at a first network port in the network cluster, determining the network path in the network cluster that includes the first network port as the first network path congested in the network cluster, where the first network port is any network port in the network cluster.

[0016] In one embodiment, the first network port is provided with a first packet buffer queue, and the first packet buffer queue is used to store packets to be transmitted by the first network port, and the packets have transmission tuples; the obtaining unit obtains the first transmission tuple corresponding to the first network path, including:

[0017] Obtaining the occurrence times of the transmission tuples of each packet in the first packet buffer queue in the first packet buffer queue;

[0018] Determining the transmission tuple with the most occurrence times as the first transmission tuple corresponding to the first network path.

[0019] In one embodiment, the first transmission tuple includes a first element for indicating a source port number; the processing unit determines a second network path from multiple network paths based on real-time traffic information of the multiple network paths in the network cluster, including:

[0020] Generating N alternative elements for the first element, where N is a positive integer;

[0021] Replacing the first element in the first transmission tuple with the N alternative elements to obtain N alternative transmission tuples;

[0022] Filtering out N alternative network paths from the network cluster according to the N alternative transmission tuples;

[0023] Based on the real-time traffic information of N alternative network paths, determine the second network path from the N alternative network paths.

[0024] In one embodiment, the processing unit generates N alternative elements for the first element, including:

[0025] Create N offset parameters for the first element;

[0026] Perform exclusive OR operations on the first element and the N offset parameters respectively to generate N alternative elements.

[0027] In one embodiment, the processing unit filters out N alternative network paths from the network cluster according to the N alternative transmission tuples, including:

[0028] Obtain the static topology structure of the network cluster;

[0029] Perform hash calculations on each of the N alternative transmission tuples to obtain the hash value corresponding to each alternative transmission tuple;

[0030] According to the hash value corresponding to each alternative transmission tuple, perform path selection in the static topology structure starting from the first computing node to obtain the network path corresponding to each alternative transmission tuple. The network cluster includes multiple computing nodes, and the first computing node is the computing node that creates the packet with the first transmission tuple;

[0031] Determine the network path corresponding to each alternative transmission tuple as the N alternative network paths in the network cluster.

[0032] In one embodiment, the obtaining unit obtains the second transmission tuple corresponding to the second network path, including:

[0033] Determine the alternative transmission tuple used to filter out the second network path among the N alternative transmission tuples as the second transmission tuple corresponding to the second network path.

[0034] In one embodiment, the network cluster includes multiple computing nodes, and a target communication library runs in each computing node. The packet with the first transmission tuple is created by the target communication library running in the first computing node; the processing unit establishes a mapping relationship between the first transmission tuple and the second transmission tuple, including:

[0035] Send a negotiation instruction to the target communication library running in the first computing node. The negotiation instruction is used to inquire whether the target communication library running in the first computing node supports the packet modification function;

[0036] Receive the negotiation response information returned by the target communication library running in the first computing node. When the negotiation response information indicates that the target communication library running in the first computing node supports the packet modification function, send a scheduling instruction to the target communication library running in the first computing node. The scheduling instruction is used to instruct the target communication library running in the first computing node to establish a mapping relationship between the first transmission tuple and the second transmission tuple.

[0037] In one embodiment, the processing unit is further configured to:

[0038] When the negotiation response information indicates that the target communication library running in the first computing node does not support the packet modification function, send a congestion notification message to the target communication library running in the first computing node. The congestion notification message is used to instruct the target communication library running in the first computing node to adjust the packet sending rate.

[0039] In one embodiment, the first computing node includes at least one network card port, and at least one communication library process is created in the target communication library running in the first computing node. One network card port corresponds to one communication library process. The processing unit sends a negotiation instruction to the target communication library running in the first computing node, including:

[0040] Determine the node address of the first computing node using the source address in the first packet;

[0041] Establish a communication connection between the node address of the first computing node and each network card port in the first computing node, and send a negotiation instruction to each network card port in the first computing node. When each communication library process detects that the corresponding network card port has received the negotiation instruction, if the target communication library running in the first computing node supports the packet modification function, generate negotiation response information including the address of the corresponding network card port.

[0042] In one embodiment, the processing unit receives the negotiation response information returned by the target communication library running in the first computing node. When the negotiation response information indicates that the target communication library supports the packet modification function, send a scheduling instruction to the target communication library running in the first computing node, including:

[0043] Receive the negotiation response information returned by each communication library process through the corresponding network card port. If the address in the first negotiation response information is the same as the source address in the first packet, determine the network card port that returns the first negotiation response information as the first network card port;

[0044] Send a scheduling instruction to the first network card port. When the first communication library process corresponding to the first network card port detects that the first network card port has received the scheduling instruction, the first communication library process is used to establish a mapping relationship between the first transmission tuple and the second transmission tuple.

[0045] In one embodiment, when there is a first packet with a first transmission tuple to be transmitted in the first network path, the processing unit schedules the first packet to be transmitted in the second network path based on the mapping relationship, including:

[0046] When the first communication library process creates a first packet with a first transmission tuple to be transmitted, call the first communication library process to obtain a second transmission tuple that has a mapping relationship with the first transmission tuple;

[0047] Call the first communication library process to modify the first transmission tuple in the first packet to the second transmission tuple to obtain a second packet;

[0048] Send the second packet through the first network card port and transmit the second packet in the second network path.

[0049] On the one hand, an embodiment of the present application provides a computer device, which includes:

[0050] A processor for loading and executing a computer program;

[0051] A computer-readable storage medium storing a computer program, which when executed by the processor, implements the above network congestion handling method.

[0052] On the one hand, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which is adapted to be loaded and executed by a processor to implement the above network congestion handling method.

[0053] On the one hand, an embodiment of the present application provides a computer program product, which includes a computer program, which when executed by a processor, implements the above network congestion handling method.

[0054] The network congestion handling solution provided by the embodiments of this application can determine the first network path where congestion occurs in a network cluster and obtain the first transmission tuple corresponding to the first network path; Exemplarily, the packets transmitted in the network cluster have transmission tuples, and the transmission tuples possessed by the packets can be used to determine the network paths through which the packets pass when transmitted in the network cluster. The first transmission tuple can be determined according to the transmission tuples possessed by the packets transmitted on the first network path. For example, the first transmission tuple can be the transmission tuple possessed by any packet transmitted on the first network path. Based on the real-time traffic information of multiple network paths in the network cluster, the second network path is determined from the multiple network paths, and the second transmission tuple corresponding to the second network path is obtained; The transmission load of the second network path meets the light load condition; Exemplarily, the packets with the second transmission tuple can be transmitted on the second network path. A mapping relationship is established between the first transmission tuple and the second transmission tuple. When there is a first packet with the first transmission tuple to be transmitted in the first network path, the first packet is scheduled to be transmitted in the second network path based on the mapping relationship; For example, the first transmission tuple in the first packet can be modified to the second transmission tuple based on the mapping relationship, so that the first packet can be scheduled to be transmitted on the second network path. It can be seen that the embodiments of this application can redistribute some of the traffic transmitted through the first network path to be transmitted through the second network path, relieve the transmission pressure of the first network path, facilitate load balancing, and improve the stability of the network cluster. Description of the Drawings

[0055] Figure 1 is a schematic diagram of the architecture of a network cluster provided by the embodiments of this application;

[0056] Figure 2 is a schematic diagram of a five-tuple provided by the embodiments of this application;

[0057] Figure 3 is a schematic flowchart of a hash path selection strategy provided by the embodiments of this application;

[0058] Figure 4 is a schematic flowchart of the execution process of an RDMA network card provided by the embodiments of this application;

[0059] Figure 5 is a schematic diagram of path selection provided by the embodiments of this application;

[0060] Figure 6 is a schematic diagram of network link congestion provided by the embodiments of this application;

[0061] Figure 7 is a schematic diagram of the architecture of a network congestion handling system provided by the embodiments of this application;

[0062] Figure 8 It is a schematic flowchart of a network congestion handling solution provided by an embodiment of the present application Figure 1 ;

[0063] Figure 9 It is a schematic flowchart of a network congestion handling method provided by an embodiment of the present application;

[0064] Figure 10 It is a schematic diagram of a port number selection strategy provided by an embodiment of the present application;

[0065] Figure 11 It is a schematic flowchart of the execution process in a negotiation phase provided by an embodiment of the present application;

[0066] Figure 12 It is a schematic flowchart of the execution process in a scheduling phase provided by an embodiment of the present application;

[0067] Figure 13 It is a schematic diagram of the definition of a message header format provided by an embodiment of the present application;

[0068] Figure 14 It is a schematic diagram of the definition of an ECHO message body format provided by an embodiment of the present application;

[0069] Figure 15 It is a schematic diagram of the definition of an ECHO - RESPONSE message body format provided by an embodiment of the present application;

[0070] Figure 16 It is a schematic diagram of the definition of an UPDATE message body format provided by an embodiment of the present application;

[0071] Figure 17 It is a schematic diagram of the definition of an UPDATE - RESPONSE message body format provided by an embodiment of the present application;

[0072] Figure 18 It is a schematic flowchart of a network congestion handling solution provided by an embodiment of the present application Figure 2 ;

[0073] Figure 19 It is a schematic flowchart of the execution process of a communication library process provided by an embodiment of the present application;

[0074] Figure 20a It is a schematic illustration of a test result provided by an embodiment of the present application Figure 1 ;

[0075] Figure 20b It is a schematic illustration of a test result provided by an embodiment of the present application Figure 2 ;

[0076] Figure 21It is a schematic structural diagram of a network congestion handling device provided by an embodiment of the present application;

[0077] Figure 22 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0078] To more clearly understand the technical solutions provided by the embodiments of the present application, key terms related to the embodiments of the present application are introduced here first:

[0079] I. Network cluster:

[0080] A network cluster may include multiple computing nodes. Each computing node can be a high-performance computer (such as a server), equipped with a powerful processor (such as a GPU (Graphics Processing Unit)), memory, and storage devices. These computing nodes are connected together through a high-speed network, aiming to perform computing tasks with high performance and efficiency through parallel processing and distributed computing. High-performance computing refers to aggregating the computing capabilities of multiple computers to provide more powerful computing performance than a single computer. Therefore, a network cluster can also be called a high-performance computing cluster. A computing task refers to a series of operations or instructions that need to be executed to complete a specific goal or solve a problem; for example, tasks of simulating complex mathematical models such as physical phenomena and chemical reactions in the field of scientific computing, tasks of predicting future weather changes by analyzing a large amount of meteorological data in meteorological forecasting, and model training tasks in the field of artificial intelligence. Parallel processing means using multiple computing nodes to execute multiple tasks at the same time. Distributed computing means decomposing a large computing task into multiple smaller subtasks and distributing these subtasks to multiple computing nodes for processing. For example, a task scheduling system can be deployed in a network cluster. A task scheduling system refers to a system used to manage and schedule computing tasks in a network cluster. When the network cluster is executing a computing task, the task scheduling system can be used to divide the computing task into multiple subtasks, distribute these multiple subtasks to each computing node in the network cluster, each computing node independently completes its own subtask, and passes the result of the subtask to other computing nodes through a high-speed network, and finally summarizes to obtain the result of the computing task, ensuring the efficient execution of the computing task and the reasonable utilization of resources.

[0081] A high-speed network refers to a network system with a relatively high data transmission rate (i.e., high bandwidth) and low latency (i.e., low delay). Network devices (also known as network nodes) are dedicated hardware devices that make up a high-speed network. Common network nodes include switches, routers, load balancers, etc. The types of network nodes are not limited in this application. A network node contains at least one network port. A network port is an interface through which a network node connects to other network nodes or computing nodes, and can be used to establish a communication channel for data transmission between other network nodes or computing nodes. Network ports can be divided into upstream ports and downstream ports. An upstream port is used to transmit data from one network node to another higher-level network node, and a downstream port is used to transmit data from one network node to another lower-level network node (or computing node).

[0082] In this application, elements represented in the singular are intended to mean "one or more", rather than "one and only one", unless otherwise specified. In this application, unless otherwise specified, "at least one" is intended to mean "one or more", and "multiple" is intended to mean "two or more".

[0083] At least one network interface card (NIC) can be deployed in a computing node. A network interface card, also known as a network adapter, is a key hardware component used to connect other computing nodes and other network nodes. A network interface card can contain at least one network interface card port. A network interface card can connect to other computing nodes and other network nodes through the network interface card port. That is to say, a network interface card port is an interface through which a computing node connects to other network nodes or other computing nodes. In a network configuration, a network interface card port of a network interface card can be connected to one or more network ports of a network node, and a network port of a network node can be connected to one or more network ports of another network node. Exemplarily, in this application, a network interface card port of a network interface card can be connected to a network port of a network node, and a network port of a network node can be connected to a network port of another network node.

[0084] Please refer to Figure 1 which is a schematic diagram of the architecture of a network cluster provided by an embodiment of this application. The network cluster includes a high-speed network and a computing node layer. The computing node layer includes multiple computing nodes, and these multiple computing nodes can access the high-speed network; a computing node can be a GPU server, and a GPU server refers to a server whose processor type is GPU type; the high-speed network usually adopts a CLOS network architecture (a non-blocking multi-stage switching structure), and usually adopts a two-layer fat tree topology (including an access layer and an aggregation layer) or a three-layer fat tree topology (including an access layer, an aggregation layer, and a core layer). Figure 1Taking the three - layer fat - tree topology in a high - speed network as an example for illustration. Among them:

[0085] The access layer is used to connect to computing nodes and includes a relatively large number of network nodes, which can be switches. The network nodes in the access layer can be represented by access (LA) nodes. Since the network nodes in the access layer are usually located at the top of the rack, they can also be represented by TOR (Top of Rack) nodes.

[0086] The aggregation layer is located between the core layer and the access layer and can connect to the network nodes of the access layer and the core layer. It is used to aggregate and forward the data flow of the access layer to the core layer. The network nodes included in the aggregation layer can also be switches. The number of network nodes in the aggregation layer is usually more than that in the core layer and less than that in the access layer. Among them, the network nodes in the aggregation layer can be represented by leaf nodes or LC nodes. Optionally, the aggregation layer can include a line card (LC), and the LC is a hardware component in the network node used to process line interfaces and data transmission.

[0087] The core layer is used to provide high - speed and non - blocking connections to support data transmission of the entire network cluster. The network nodes included in the core layer can be high - performance switches or routers, with high bandwidth and low latency. The number of network nodes in the core layer is small, but they have high port density and connection capabilities. Among them, the network nodes in the core layer can be represented by spine nodes, which are responsible for connecting each leaf node to ensure high - speed forwarding of data flow between leaf nodes and are used to provide high - performance data transmission and forwarding capabilities. Optionally, the core layer can include a switched port analyzer (SPAN) or can include a super LC (enhanced version of the line card).

[0088] The typical feature of the fat - tree topology is that there is no bandwidth convergence. That is to say, in an ideal situation, the total egress bandwidth of the computing node layer, the total egress bandwidth of the access layer, and the total egress bandwidth of the aggregation layer in the network cluster are the same. For example, when there are 4 network cards with a bandwidth of 100 Gbs (gigabits per second) in the computing node layer, in an ideal situation, the total egress bandwidth of the access layer and the total egress bandwidth of the aggregation layer should be 400 Gbps. The total egress bandwidth of the access layer can be equal to the sum of the bandwidths of all the uplink ports in the access layer, and the total egress bandwidth of the aggregation layer can be equal to the sum of the bandwidths of all the uplink ports in the aggregation layer.

[0089] This application does not limit the type of high-speed network in the network cluster. For example, due to different numbers of computing nodes in the network cluster, the number of levels of the fat tree topology may vary. For example, some network clusters with fewer computing nodes do not need to introduce a core layer, while network clusters with more computing nodes need to introduce a core layer. Another example is that in addition to using the Figure 1 non-blocking fat tree topology shown, a ring topology, a mesh topology, etc. can also be used.

[0090] II. Equal-Cost Multi-Path (ECMP) algorithm:

[0091] There are multiple equivalent paths between the sender (referring to the computing node that sends the packet) and the receiver (referring to the computing node that receives the packet) in the network cluster. An equivalent path refers to a path with the same cost, which can be calculated based on the number of hops during packet transmission. For example, Figure 1 the packet sent by computing node 1 in Figure 1 can reach computing node 2 through network path 1 in Figure 1 : ①-②-③-④-⑤-⑥, or can also reach computing node 2 through network path 2 in Figure 1 : ①-②-③-⑦-⑤-⑥. Since the number of hops of network path 1 and network path 2 is both 6, network path 1 and network path 2 can be considered equivalent paths. A packet is the basic unit of data transmission in the network cluster and can be obtained by encapsulating the data to be transmitted using a packet header. The packet header can include but is not limited to the following information: source address (used to indicate the address of the sender of the packet (such as an IP address)), destination address (used to indicate the address of the receiver that the packet is to reach (such as an IP address)), source port number (used to indicate the port of a specific service used by the sender to send the packet), destination port number (used to indicate the port of a specific service used by the receiver to receive the packet), the length of the packet, protocol number (used to indicate the type of protocol used by the transport layer in the computer network protocol stack).

[0092] The ECMP algorithm is a method for implementing network load balancing. It allows network nodes to identify and distinguish different data streams based on the transmission multi-tuples of the packets and evenly distribute these data streams to multiple equivalent paths for forwarding. That is to say, packets with the same transmission multi-tuples belong to the same data stream, and packets with different transmission multi-tuples belong to different data streams. Packets belonging to the same data stream are transmitted using the same network path, and packets belonging to different data streams are transmitted using different network paths. The transmission multi-tuple can be obtained by combining multiple fields included in the packet header in the packet. For example, the transmission multi-tuple can be a five-tuple. Please refer to Figure 2It is a schematic diagram of a five-tuple provided by an embodiment of the present application. The five-tuple may include: source address, source port number, destination address, destination port number, and protocol number.

[0093] Common path selection strategies of the ECMP algorithm include hashing, polling, and path weight-based, etc. The hashing path selection strategy is a path selection strategy widely used in the ECMP algorithm. Please refer to Figure 3 It is a schematic flowchart of a hashing path selection strategy provided by an embodiment of the present application. Its implementation process can be as follows: A hash function and a hash seed (a numerical value) can be stored in a network node. The hash functions and hash seeds stored in each network node can be the same or different. For example, network nodes at the same layer in a network cluster can use the same hash seed and hash function. When a packet arrives at a network node, the network node can perform a hash calculation on the transmission multi-tuple of the packet using the hash function and the hash seed to obtain a hash value, and convert the hash value into an offset. For example, take the first 8 bits of the hash value as the offset. This offset can be used to determine the node (network node or computing node) of the next-hop route of the packet. Exemplarily, a network node may include at least one output port. The network node can perform a modulo operation (i.e., remainder calculation) on the offset and the number of output ports of the network node to obtain a modulo value. This modulo value can identify an output port. For example, if the number of output ports of the network node is 30, the range of the modulo value is [0 - 29]. The identifier of output port 1 can be 0, the identifier of output port 2 can be 1, and so on. When the modulo value is 1, this modulo value can identify output port 2. In this way, the modulo value can indicate from which output port the packet can be forwarded to determine the node of the next-hop route of the packet. For example, the network node forwards the packet through network port 1, and network port 1 is connected to network card port 1 of the first GPU server. Then the packet will be sent to the first GPU server through network card port 1 of the first GPU server. Among them, when the network node is sending a packet to a higher-level network node, the at least one output port included in the network node refers to all the upstream ports of the network node; when the network node is sending a packet to a lower-level network node (or computing node), the at least one output port included in the network node refers to all the downstream ports of the network node. Thus, it can be seen that when the hashing algorithm (including the hash function and the hash seed) adopted by the network node is determined, the only factor that can affect the network path of packet transmission is the transmission multi-tuple of the packet.

[0094] III. RDMA Network Card:

[0095] RDMA (Remote Direct Memory Access) is a technology that bypasses the operating system kernel of a remote computer and directly accesses the data in its memory. Since it does not go through the operating system, it not only saves a large amount of computing resources, but also improves system throughput and reduces the network communication latency of the system. It is especially suitable for scenarios of a large number of parallel computations in a network cluster. Exemplarily, the network card in a computing node can be an RDMA network card. In one implementation, a target communication library can be deployed in the computing node, and the target communication library can work in coordination with the RDMA network card to achieve efficient data transmission and communication between computing nodes.

[0096] A communication library is a software library used to handle and manage communication-related tasks. It usually provides a series of functions and interfaces for implementing data sending, receiving, parsing, processing, and communicating with other devices or systems. The target communication library is a communication library that can achieve efficient data exchange and collaborative work in a network cluster. Exemplarily, the target communication library can be a collective communication library. A collective communication library is a type of communication library specifically used in parallel computing and distributed computing environments, aiming to efficiently implement communication operations between multiple processes or nodes. It provides a set of predefined communication operations, such as AllReduce (performing a reduction operation on the data on all computing nodes participating in the computing task and then broadcasting the result to all computing nodes), Broadcast (broadcasting the data on one computing node to all other computing nodes), Allgather (each computing node has a piece of data, and the data of each computing node is collected to form a complete data set), ReduceScatter (after reducing the data on one computing node according to a certain rule, the result is evenly distributed to each computing node), AlltoAll (each computing node has a set of data, and a set of data of each computing node is passed to other computing nodes), etc. The target communication library can be an open-source communication library. For example, NCCL (Nvidia Collective Communication Library) is a widely used collective communication library currently, or it can be a self-developed communication library.

[0097] Please refer to Figure 4It is a schematic diagram of the execution process of an RDMA network card provided by an embodiment of the present application. The RDMA network card includes a pair of work queues (Queue Pair, QP): a send queue (Send Queue, SQ) and a receive queue (Receive Queue, RQ). The target communication library can prepare the message to be sent and send the message sending instruction to the RDMA network card in the form of a work request (WR) through the RDMA communication primitive. The RDMA network card driver can convert the message sending instruction into a work queue element (Work Queue Element, WQE) and publish it to the send queue in the RDMA network card. For example Figure 4 Work queue element 1 and work queue element 2 in. Among them, WR is the basic unit of RDMA communication. By sending WR to the RDMA network card, the RDMA network card can be instructed to perform specific RDMA operations. The RDMA communication primitive refers to a set of basic operations defined in the RDMA technology. These operations can be sent to the RDMA network card in the form of WR to achieve various data transmissions and memory operations. The work queue is a data structure used in the RDMA technology to store work requests. WQE is a key data structure used in the RDMA technology to describe tasks, containing detailed information about the tasks, such as operation type, data buffer, target address, etc. The RDMA network card will obtain the work queue element from the send queue, obtain the data at the corresponding address in the memory according to the indication of the work queue element (which can be the message to be sent), and send the obtained data to the remote RDMA network card. Correspondingly, the target communication library can send the message receiving instruction to the RDMA network card in the form of a work request (work request, WR) through the RDMA communication primitive. The RDMA network card driver can convert the message receiving instruction into a work queue element (Work Queue Element, WQE) and publish it to the receive queue in the RDMA network card. For example Figure 4 Work queue element 3 in. The RDMA network card will obtain the work queue element from the receive queue and store the received data at the corresponding address in the memory according to the indication of the work queue element.

[0098] When the transmission tuple is a five-tuple, for the message created by the target communication library, its source address and destination address determine the receiving end and sending end of the message and cannot be changed arbitrarily; its destination port number and protocol number are usually determined by the transmission protocol adopted by the RDMA network card. Different transmission protocols can have different destination port numbers and protocol numbers; therefore, the configuration of the source port number by the target communication library becomes the only hash factor (referring to the element used for hash calculation) that affects the path selection of the message in the network cluster. For example, please refer to Figure 5This is a schematic diagram of path selection provided by an embodiment of the present application. After the source port number of a packet changes from source port number A to source port number B, the network path for transmitting the packet changes from access node 1 (network module 1) - leaf node 1 - access node 1 (network module 16) to access node 1 (network module 1) - leaf node 2 - access node 1 (network module 16). Each network module (Block) consists of a computing node (such as a GPU server) and a network node in the access layer (represented by an access node), and there is a full connection relationship between the computing node in each network module and the network node in the access layer (i.e., there is a connection between each computing node and each network node in the access layer). Additionally, there is no connection relationship between nodes in different network modules, and there is also a full connection relationship between the network nodes in the access layer of each network module and the network nodes in the aggregation layer (represented by leaf nodes).

[0099] IV. Data Center Quantized Congestion Notification:

[0100] Data Center Quantized Congestion Notification (DCQCN) is a communication protocol for congestion control in a network cluster. Its implementation principle is as follows: when a network node detects network congestion, it marks the packet with an Explicit Congestion Notification (ECN). When the receiving end of the packet receives a packet with an ECN mark, it generates a Congestion Notification Packet (CNP) and sends it back to the sending end. The sending end quantifies the congestion level in the network cluster based on the number and frequency of received CNP, and adjusts the packet sending rate accordingly. The greater the congestion level, the lower the sending rate, which can reduce the packet loss rate and latency.

[0101] V. INT (In-band Network Telemetry) Technology:

[0102] INT technology is a technology for collecting and reporting network status. It realizes real-time detection and feedback of network conditions by embedding telemetry information in packets. Telemetry Information refers to information related to network status filled in by intermediate forwarding devices during the transmission of packets, such as queue length, network port for forwarding packets out, etc. INT latency detection is a method of using INT technology to measure and monitor the latency of packet transmission in a network.

[0103] VI. sFlow (Sampled Flow) Technology:

[0104] The sFlow technology is a technology used to monitor the traffic forwarding status of network nodes. For example, packets can be collected at a predetermined sampling rate on the interfaces of specific network nodes, and these packets can be deeply analyzed, including content parsing, forwarding path information, etc. Then, the processed statistical data and the original packets are sent to a dedicated collector for centralized processing. At the same time, it also supports periodic statistics of port traffic, as well as statistics on the processor and memory usage of network nodes, so as to comprehensively understand the operating status of network nodes, with high real-time performance and efficiency.

[0105] Network clusters are usually used in fields such as scientific computing, weather forecasting, big data processing, cloud computing, and artificial intelligence to meet the needs of large-scale computing tasks. For example, in the field of artificial intelligence (AI) technology, a large model refers to a deep learning model with a large number of parameters, and the parameter scale can usually reach millions, tens of millions, or even hundreds of millions, with high expressive ability and prediction performance, and can handle complex tasks. Large models have extremely high requirements for computing resources and storage space. Network clusters can provide powerful computing capabilities and efficient data processing capabilities to meet the training needs of large models. With the rapid development of artificial intelligence technology, the training requirements of large models are constantly increasing. In order to adapt to the training needs of large models, network clusters are usually configured to have non-converging bandwidth (such as using a fat tree topology). However, due to reasons such as network link congestion and uneven hash load, a certain degree of congestion is formed in the network cluster, and the theoretically non-converging bandwidth cannot be fully utilized, directly affecting the communication duration of large model training, and further affecting the training efficiency and inference efficiency of large models. In addition, although data center quantization congestion notification can alleviate the congestion of network clusters to a certain extent, the reduction of the packet sending rate still inevitably causes a decline in network communication performance, directly affecting the time required for large model training.

[0106] Among them, a network link refers to a communication channel established between one node and another node. This communication channel can be implemented through a physical connection (i.e., connected through a physical port) or through a logical connection (i.e., connected through a logical port). In one implementation, a node can be a network node or a computing node. A network card port and a network port can be either a physical port or a logical port. Network link congestion refers to the situation where the amount of data transmitted on a network link is too large or overloaded, resulting in the network link being unable to process all packets in a timely manner, causing problems such as data transmission delay, packet loss, and packet retransmission. A network path refers to all the routes that a packet passes through from the sender to the receiver, consisting of a series of network links and network nodes. Hash load imbalance means that since each target communication library is only responsible for data communication on its corresponding computing node, the target communication library cannot perceive more accurate global information in the configuration of transmitting multi-tuples. There is a lack of planning in the network paths for different data streams, and there is a high possibility of large hash conflicts. For example, the target communication library often uses a random configuration method to determine the source port number for packets to be sent, resulting in multiple data streams sharing the same network link in the network cluster for transmission, thus causing network link congestion; please refer to Figure 6 is a schematic diagram of network link congestion provided by an embodiment of the present application. The network path for transmitting data stream 1 is: access node 1 (network module 1) - leaf node 1 - access node 1 (network module 16). The network path for transmitting data stream 2 is: access node 1 (network module 1) - leaf node 1 - access node 2 (network module 16). Upward congestion occurs at the network link between access node 1 (network module 1) and leaf node 1. The network link for transmitting data stream 3 is: access node 3 (network module 1) - leaf node 16 - access node 4 (network module 16). The network link for transmitting data stream 4 is: access node 4 (network module 1) - leaf node 16 - access node 4 (network module 16). Downward congestion occurs at the network link between leaf node 16 and access node 4 (network module 16).

[0107] The embodiment of the present application provides a network congestion handling solution, which can flexibly change the hash value of the data stream passing through the congested network link to quickly eliminate the network link congestion in the network cluster. The network congestion handling solution includes: determining a first network path that is congested in the network cluster, and obtaining a first transmission tuple corresponding to the first network path. Exemplarily, the packets transmitted in the network cluster have transmission tuples, and the transmission tuples possessed by the packets can be used to determine the network path that the packets pass through when transmitted in the network cluster. For example, each network node can adopt a hash path selection strategy to perform a hash calculation on the transmission tuples possessed by the packets, and then determine the node of the next-hop route of the packets. The first transmission tuple can be determined according to the transmission tuples possessed by the packets transmitted on the first network path. For example, the first transmission tuple can be the transmission tuple possessed by any packet transmitted on the first network path. Based on the real-time traffic information of multiple network paths in the network cluster, a second network path is determined from the multiple network paths. Among them, the real-time traffic information of multiple network paths in the network cluster can be obtained based on monitoring technologies such as the INT technology and the sFlow technology. The transmission load of the second network path meets the light-load condition, that is to say, the second network path has certain idle resources to achieve the fast transmission of the data stream. The second transmission tuple corresponding to the second network path can be obtained, and among them, the packets possessing the second transmission tuple can be transmitted on the second network path. A mapping relationship is established between the first transmission tuple and the second transmission tuple. When there is a first packet possessing the first transmission tuple to be transmitted in the first network path, the first packet is scheduled to be transmitted in the second network path based on the mapping relationship; for example, the first transmission tuple in the first packet can be modified to the second transmission tuple, so that the first packet can be scheduled to be transmitted on the second network path. Thus, it can be seen that in the embodiment of the present application, part of the traffic transmitted through the first network path can be redistributed to be transmitted on the second network path, alleviating the transmission pressure of the first network path, which is beneficial to achieving load balancing and improving the stability of the network cluster.

[0108] Please refer to Figure 7 which is a schematic structural diagram of a network congestion handling system provided by the embodiment of the present application. The network congestion handling system includes a network cluster and a management device of the network cluster. The network cluster may include multiple computing nodes ( Figure 7 taking 2 computing nodes as an example), and a high-speed network. The high-speed network may include multiple network nodes ( Figure 7Taking 4 network nodes as an example, data can be transmitted between computing nodes through a high-speed network. The management device can communicate with each node in the network cluster (including network nodes and computing nodes) through a communication network. The communication network can be a wired network or a wireless network, which is not limited in this application. For example, the management device can be indirectly connected to each node in the network cluster through a wireless access point, or the management device can be directly connected to each node in the network cluster through the Internet. The management device can be a terminal device such as a smart phone, a tablet computer, a smart wearable device, a smart voice interaction device, a smart home appliance, a personal computer, a vehicle-mounted terminal, etc. The management device can also be a server. For example, it can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0109] The management device may run a network controller and a service program. The network controller refers to an application program used to manage and control the operation of a network cluster. Exemplarily, the network controller may be a Software-Defined Networking (SDN) controller. The SDN controller is a core component in the SDN architecture and is responsible for managing and controlling the data flow in the network. By separating the network control logic from the underlying hardware devices, the SDN controller realizes the centralized management and control of the network, thereby improving the flexibility, programmability, and manageability of the network. The main functions of the SDN controller include but are not limited to: 1) Network abstraction: The SDN controller can communicate with the underlying network nodes through the southbound interface to collect network status information. The southbound interface is a key concept in SDN technology and refers to the interface between the SDN controller and the underlying network devices (such as switches, routers, etc.). The network status information may include but is not limited to: ① Device information, such as the model, configuration, and operating status of network nodes; ② Link status: Describing indicators such as the bandwidth utilization rate, delay, and packet loss rate of each network link in the network cluster; ③ Topology information: Describing information about all nodes and network links in the network cluster; ④ Traffic statistics: Describing the traffic conditions and traffic distribution of each data flow in the network cluster. 2) Flow table management: The SDN controller can generate and distribute flow table rules to network nodes, and the flow table rules can guide the network nodes to forward packets. 3) Topology management: The SDN controller can be responsible for discovering and managing the static topology structure, which may include information about all nodes and network links in the network cluster. 4) Resource allocation: The SDN controller can dynamically allocate network resources, such as bandwidth and network paths, according to the network status and application requirements. 5) Fault recovery: The SDN controller can detect network faults and automatically take measures for recovery, such as recalculating paths and redistributing flow table rules.

[0110] The service program is an application program used to monitor, manage, and optimize the performance of the network cluster. Please refer to Figure 8 It is a flow schematic diagram of a network congestion handling solution provided by an embodiment of this application. Figure 1, in this network congestion handling solution, the network controller can collect information of all nodes and network links in the network cluster to form a static topology structure of the network cluster. The service program can monitor whether there is ECN information generated by the DCQCN protocol (i.e., packets with ECN markings) in the network cluster. When it monitors that there is ECN information generated by the DCQCN protocol in the network cluster, or when it detects that the number of ECN information in the network cluster exceeds a preset information quantity (the preset information quantity can be set as needed), it can report a link congestion alarm to the network controller. The network controller can determine a first network path where congestion occurs in the network cluster based on the reported link congestion alarm. In one implementation, the link congestion alarm reported by the service program can include information about the network path where congestion occurs. In this way, the network controller can directly determine the first network path based on the information about the network path where congestion occurs. In another implementation, the network controller can check which network paths have increased traffic when congestion occurs, and then determine the first network path where congestion occurs in the network cluster. The network controller can obtain a first transmission tuple corresponding to the first network path. For example, the first transmission tuple can be the transmission tuple of the data stream with the largest data volume transmitted on the first network path. The network controller can implement the hash algorithm in the network node in software form to provide a hash software service for calculating hash values according to hash factors in the network controller. The network controller can use the hash software service for scheduling and routing. For example, it can use a set of alternative source port numbers and the quadruple in the first transmission tuple (i.e., destination port number, destination address, source address, protocol number) to form a set of alternative transmission tuples, use the hash software service to perform hash calculations on each alternative transmission tuple, and then use the hash value corresponding to each alternative transmission tuple to determine the alternative network path corresponding to each alternative transmission tuple. According to the real-time traffic information of these alternative network paths in the network cluster, it can decide a second network path from the network cluster and obtain a second transmission tuple corresponding to the second network path (such as the alternative transmission tuple corresponding to the second network path), where the transmission load of the second network path meets the light load condition. The network controller can issue a scheduling instruction to the first computing node through the custom communication library - controller communication protocol of this application. The first computing node is the computing node that creates packets with the first transmission tuple. The target communication library running in the first computing node can, in response to the scheduling instruction, establish a mapping relationship between the first transmission tuple and the second transmission tuple. In this way, when the target communication library running in the first computing node creates a first packet with the first transmission tuple to be transmitted, it can modify the first transmission tuple in the first packet to the second transmission tuple based on the mapping relationship, so that the first packet can be scheduled to be transmitted on the second network path.

[0111] It can be seen that the embodiment of the present application implements a scheduling system based on a network controller. The network controller can perform dynamic data exchange with the target communication library on the end side (i.e., the computing node side), switch the data stream transmitted on the congested network link to the lightly loaded network link, quickly restore the packet sending rate on the end side, avoid affecting the execution of computing tasks, and is beneficial to improving the stability of the network cluster.

[0112] The above system architecture is only an example and does not constitute a limitation on the system architecture of the technical solution provided by the embodiment of the present application. The technical solution of the present application can also be applied to other system architectures. For example, as known to those of ordinary skill in the art, with the evolution of the system architecture and the emergence of new business scenarios, the technical solution provided by the embodiment of the present application is also applicable to similar technical problems.

[0113] It should also be noted that in the embodiment of the present application, the collection and processing of relevant data should be strictly in accordance with the requirements of relevant laws and regulations. Obtaining personal information requires the informed consent of the personal subject (or having a legal basis for information acquisition), and subsequent data use and processing behaviors should be carried out within the scope authorized by laws and regulations and the personal information subject.

[0114] The network congestion handling solution provided by the present application will be described in detail below. Please refer to Figure 9 , Figure 9 which is a schematic flowchart of a network congestion handling method provided by the embodiment of the present application. The network congestion handling method can be executed by a management device (the network controller and service program therein). The network congestion handling method includes but is not limited to the following steps S11 to S14:

[0115] S11. Determine the first network path where congestion occurs in the network cluster, and obtain the first transmission tuple corresponding to the first network path.

[0116] A network cluster is generally a computing system composed of a large number of computing nodes and a series of network nodes interconnected with each other. In a network cluster, the computing nodes are connected to the network nodes through network cards (such as RDMA network cards). Exemplarily, a computing node may include at least one network card, and each network card may include at least one network card port, and the network card port may be an interface for the computing node to connect to the network nodes. The network nodes may include at least one network port, and the network port may be an interface for the network nodes to connect to other network nodes or computing nodes. There are multiple network paths in the network cluster, and the network paths can be used to indicate the order of the network nodes and network links passed by a packet during the transmission process from the sending end (i.e., the computing node that sends the packet) to the receiving end (i.e., the computing node that receives the packet). A network link refers to a communication channel established between two nodes. In one implementation, the network link may include the ports connected when the two nodes establish the communication channel. For example, a certain network link may indicate that there is a connection relationship between network card port 1 of computing node 1 and network port 1 of network node 1. In this application, a node may include one or both of a network node and a computing node, and a port may include one or both of a network port and a network card port.

[0117] In an embodiment of this application, in addition to including a source port number, a transmission tuple may further include one or more of a source address, a destination address, a destination port number, and a protocol number. That is to say, the transmission tuple is not limited to being a five-tuple, and may also be a two-tuple, a three-tuple, etc. that includes a source port number. For example: the transmission tuple may be a two-tuple composed of a source port number and a destination port number; another example: the transmission tuple may be a three-tuple composed of a source port number, a destination port number, and a protocol number.

[0118] In one implementation, any network path in a network cluster may include at least one network port. For example, any network path may represent: sending from network card port 1 of computing node 1 to network port 1 of network node 1, and then transmitting from network port 1 of network node 1 through network port 2 of network node 1 to network card port 1 of computing node 2. Here, computing node 1 is the sending end of the message, computing node 2 is the receiving end of the message, and network node 1 is the intermediate forwarding device of the message. When a congestion event occurs at the first network port in the network cluster, the network path in the network cluster that includes the first network port can be determined as the first congested network path in the network cluster. The first network port is any network port in the network cluster, and the number of the first network paths can be one or more. Exemplarily, each network port may be provided with its own message buffer queue, and each network port can use the provided message buffer queue as a temporary buffer area for the received (or to-be-transmitted) messages. The first network port may be provided with a first message buffer queue, and the first message buffer queue can be used to store the messages to be transmitted by the first network port. The congestion event at the first network port may include at least one of the following: the number of messages passing through the first network port per unit time exceeds the first quantity threshold (the first quantity threshold can be set as needed); the delay of the messages passing through the first network port (referring to the time experienced by the message from being sent by the sending end to being received by the receiving end and finally being processed by the receiving end) exceeds the first time threshold (the first time threshold can be set as needed); the number of messages contained in the first message buffer queue exceeds the second quantity threshold (the second quantity threshold can be set as needed); the message loss rate of the first network port exceeds the first loss rate threshold (the first loss rate threshold can be set as needed); the number of messages containing congestion marks (such as ECN marks) in the first message buffer queue exceeds the third quantity threshold (the third quantity threshold can be set as needed); there are messages containing congestion marks in the first message buffer queue.

[0119] In one implementation, obtaining the first transmission tuple corresponding to the first network path includes: obtaining the number of occurrences of the transmission tuples of each packet in the first packet buffer queue in the first packet buffer queue. For example, the first packet buffer queue contains Packet 1, Packet 2, and Packet 3. Packet 1 has transmission tuple 1, Packet 2 has transmission tuple 1, and Packet 3 has transmission tuple 2. The number of occurrences of transmission tuple 1 in the first packet buffer queue is 2, and the number of occurrences of transmission tuple 2 in the first packet buffer queue is 1. The transmission tuple with the most occurrences can be determined as the first transmission tuple corresponding to the first network path. That is to say, the present application can schedule the data stream with the largest number of packets in the first packet buffer queue to be transmitted on other network links. In a feasible implementation, the transmission tuples with the number of occurrences greater than a preset number (the preset number can be set as required) can also be determined as the first transmission tuples corresponding to the first network path.

[0120] In another implementation, a service program can be called to monitor the congestion metrics of each network link in the network cluster. The congestion metrics can include at least one of the following: throughput (the amount of data that can be transmitted per unit time), latency, packet loss rate (the proportion of lost packets), bandwidth utilization, the number of retransmitted packets due to timeout, whether there is an ECN count (i.e., whether to transmit packets containing ECN marks), etc. The congestion score of the corresponding network link can be evaluated through each congestion metric. For example, through a certain algorithm, each congestion metric can be weighted and calculated to obtain a comprehensive congestion score. Exemplarily, the larger the throughput, the larger the congestion score; the larger the latency, the larger the congestion score; the larger the packet loss rate, the larger the congestion score; the larger the bandwidth utilization, the larger the congestion score; the larger the number of packets, the larger the congestion score; the presence of an ECN count, the larger the congestion score; conversely, the smaller the throughput, the smaller the congestion score; the smaller the latency, the smaller the congestion score; the smaller the packet loss rate, the smaller the congestion score; the smaller the bandwidth utilization, the smaller the congestion score; the smaller the number of packets, the smaller the congestion score; the absence of an ECN count, the smaller the congestion score. In this way, the larger the congestion score, the greater the degree of congestion, and the smaller the congestion score, the smaller the degree of congestion. When the congestion score of the first network link exceeds the first preset score threshold (the first preset score threshold can be set as required), the network path including the first network link can be determined as the first network path with congestion in the network cluster. The first network link can be any network link in the network cluster. The transmission tuple of the data stream with the largest amount of data transmitted on the first network link can be obtained and determined as the first transmission tuple. For example, the packets transmitted on the first network link can be captured, and the data streams can be distinguished based on the transmission tuples of the captured packets, and then the data stream with the largest amount of data transmitted on the first network link can be determined.

[0121] In a feasible implementation, historical traffic transmitted on the first network link (or passing through the first network port) can be obtained. The historical traffic may include all packets transmitted on the first network link (or passing through the first network port) within a neighboring preset time (for example, taking the current time point as the cut-off time point, and determining the previous 24 hours of the cut-off time point as the preset time). Based on the transmission tuples of these packets, the data flow with the largest data volume can be determined, and the transmission tuple of the data flow with the largest data volume is determined as the first transmission tuple.

[0122] S12. Based on the real-time traffic information of multiple network paths in the network cluster, determine a second network path from the multiple network paths, and obtain the second transmission tuple corresponding to the second network path; the transmission load of the second network path meets the light load condition.

[0123] The real-time traffic information of a network path refers to the real-time status and statistical information of the traffic generated when packets are transmitted on this network path, and may include at least one of the following information (which can be called load metrics): the number of packets transmitted on the network path per unit time, the amount of data transmitted on the network path per unit time (which can be in bits or bytes), the data transmission speed on the network path (which can be in bits per second), whether there is network link congestion in the network path, the latency of the packets transmitted on the network path, and the packet loss rate of the network path. The real-time traffic information of a network path can be obtained based on monitoring technologies such as INT technology and sFlow technology. In one implementation, each load metric in the real-time traffic information of each network path can be used to evaluate the load score of the corresponding network path. For example, through a certain algorithm, each load metric in the real-time traffic information of each network path can be weighted and calculated to obtain a comprehensive load score. Exemplarily, the number of packets transmitted on the network path per unit time can be positively correlated with the load score, the amount of data transmitted on the network path per unit time can be positively correlated with the load score, the data transmission speed on the network path can be inversely correlated with the load score, the latency of the packets transmitted on the network path can be positively correlated with the load score, and the packet loss rate of the network path can be positively correlated with the load score. In addition, if there is network link congestion in the network path, the load score is larger; if there is no network link congestion in the network path, the load score is smaller. In this way, the smaller the load score of the network path, the better the transmission performance. In a feasible implementation, the load score of each network path can also be evaluated according to the traffic of each network path. The traffic of a network path refers to the amount of data transmitted on the network path. The larger the traffic of the network path, the larger the corresponding load score; the smaller the traffic of the network path, the smaller the corresponding load score. That the transmission load of the second network path meets the light load condition can mean that the load score of the second network path is less than the second preset score threshold (the second preset score threshold can be set as needed).

[0124] In one implementation, the first transmission tuple includes a first element for indicating a source port number; based on the real-time traffic information of multiple network paths in a network cluster, a second network path is determined from the multiple network paths, including: generating N alternative elements for the first element, where the alternative elements are used to indicate alternative source port numbers, and N is a positive integer. Exemplarily, N offset parameters can be created for the first element, and the first element is respectively exclusive-ORed with the N offset parameters to generate N alternative elements. When performing the exclusive-OR operation on the first element and the offset parameter, the first element and the offset parameter can be converted into binary forms, and then each bit of the two binary numbers is subjected to a bitwise exclusive-OR operation (that is, if the two numbers in the same position are the same, the value is 1, and if the two numbers in the same position are different, the value is 0), and then the binary number obtained by the bitwise exclusive-OR operation is converted into a decimal number to obtain the alternative element. The first element in the first transmission tuple can be replaced with the N alternative elements to obtain N alternative transmission tuples. According to the N alternative transmission tuples, N alternative network paths can be screened out from the network cluster, and based on the real-time traffic information of the N alternative network paths, a second network path is determined from the N alternative network paths. For example, the alternative network path with the smallest load fraction among the N alternative network paths can be selected as the second network path.

[0125] In one embodiment, creating N offset parameters for the first element includes: randomly generating a set of values (including multiple numbers) for the first element, and determining this set of values as N alternative elements. In another embodiment, creating N offset parameters for the first element includes: obtaining a second element other than the first element in the first transmission tuple, and the second element can include one or more of a source address, a destination address, a destination port number, and a protocol number. The hash algorithm of each network node in the network cluster can form a hash software service in a management device (the network controller therein) in software form. The management device (the network controller therein) can use the hash software service and the second element to calculate N offset parameters for the first element, and using these N offset parameters can ensure that all optional network paths in the network cluster can be covered as much as possible. For example, the first network node is a network node including a first network port. When the first network port is an uplink port (or a downlink port) of the first network node, all uplink ports (or downlink ports) without congestion can be determined from the first network node, ensuring that the N alternative transmission tuples generated using these N offset parameters can respectively select these uplink ports (or downlink ports) without congestion when performing path selection.

[0126] In a feasible embodiment, N alternative network paths are screened out from a network cluster according to N alternative transmission tuples, including: obtaining the static topology of the network cluster; the static topology is formed by collecting information of all nodes and network links in the network cluster, and can be used to describe each network path in the network cluster, that is, it can describe the information of the nodes included in each network path and the connection relationship between the nodes (which may involve the connection relationship between ports). Hash calculation can be performed on each of the N alternative transmission tuples to obtain the hash value corresponding to each alternative transmission tuple. If multiple sets of hash algorithms are adopted in the network cluster, for example, each layer of network nodes uses the same hash function but different hash seeds, then each alternative transmission tuple can correspond to multiple hash values. According to the hash value corresponding to each alternative transmission tuple, path selection can be performed in the static topology starting from the first computing node to obtain the network path corresponding to each alternative transmission tuple; the strategy selected during path selection can be the hash path selection strategy, that is, when reaching each network node, the corresponding hash value will be converted into an offset, and the offset is used to determine the next routing node. The network path corresponding to each alternative transmission tuple can be determined as the N alternative network paths in the network cluster.

[0127] In an embodiment, obtaining the second transmission tuple corresponding to the second network path includes: the alternative transmission tuple used to screen out the second network path among the N alternative transmission tuples can be determined as the second transmission tuple corresponding to the second network path.

[0128] For example, please refer to Figure 10 is a schematic diagram of a port number selection strategy provided by an embodiment of the present application. The source port number indicated by the first element in the first transmission tuple is 60051. A set of offset parameters (i.e., 3, 12, 15, 513, 2048, 2074, 2063, 2069) is calculated for 60051. The exclusive OR operation can be performed between 60051 and each offset parameter respectively to obtain alternative parameters (i.e., 60048, 60063, 60060, 59538, 58003, 57993, 58012, 57990). These alternative parameters are used to replace the first element in the first transmission tuple to obtain multiple alternative transmission tuples. Hash calculation is performed on each alternative transmission tuple respectively to obtain the hash value corresponding to each alternative transmission tuple. The hash values corresponding to these alternative transmission tuples can be used to determine the network port selected from the first network node during path selection. When it is found that the network path where the network port Eth200GE88 in the first network node is located is relatively idle, the network path where the network port Eth200GE88 is located can be determined as the second network path, and 57793 is determined as the new source port number.

[0129] S13. Establish a mapping relationship between the first transmission tuple and the second transmission tuple.

[0130] In one implementation, establishing a mapping relationship between the first transmission tuple and the second transmission tuple includes: establishing a mapping relationship between the first transmission tuple and a third element, where the third element refers to the field in the second transmission tuple used to indicate the source port number.

[0131] The network cluster may include multiple computing nodes, and a target communication library runs in each computing node. The target communication library running in each computing node can be used to create packets. In a feasible embodiment, establishing a mapping relationship between the first transmission tuple and the second transmission tuple includes: sending a negotiation instruction to the target communication library running in the first computing node. The negotiation instruction is used to query whether the target communication library running in the first computing node supports the packet modification function, and the packet modification function may refer to the ability to modify the source port number in the packet. The negotiation response information returned by the target communication library running in the first computing node can be received. When the negotiation response information indicates that the target communication library running in the first computing node supports the packet modification function, a scheduling instruction is sent to the target communication library running in the first computing node. The scheduling instruction is used to instruct the target communication library running in the first computing node to establish a mapping relationship between the first transmission tuple and the second transmission tuple. For example, a data structure (such as a hash table or an array) can be created to store this mapping relationship.

[0132] In a feasible implementation, the first computing node includes at least one network card port, and at least one communication library process (referring to the process created by the target communication library) is created by the target communication library running in the first computing node. One network card port corresponds to one communication library process, and the packets created by each communication library process will be sent out through its corresponding network card port. The source address in the transmission tuple of the packet can be the address (such as an IP address) of the network card port that sends the packet. Please refer to Figure 11 is a schematic diagram of the execution process in the negotiation phase provided by the embodiments of the present application. Sending a negotiation instruction to the target communication library running in the first computing node includes: determining the node address of the first computing node using the source address in the first packet. For example, the node address associated with the source address in the first packet can be found through an address query function (such as an IP (Internet Protocol) query function), and the computing node corresponding to the node address is determined as the first computing node. Based on the node address of the first computing node and each network card port in the first computing node (such as Figure 11Establish a communication connection with network card ports 5001, ..., 5008 in the first computing node. Exemplarily, this communication connection can be a TCP (Transmission Control Protocol) connection. After establishing the communication connection, a negotiation instruction can be sent to each network card port in the first computing node; when each communication library process detects that the corresponding network card port has received the negotiation instruction, if the target communication library running in the first computing node supports the packet modification function, negotiation response information containing the address of the corresponding network card port is generated.

[0133] In one embodiment, receive the negotiation response information returned by the target communication library running in the first computing node. When the negotiation response information indicates that the target communication library supports the packet modification function, send a scheduling instruction to the target communication library running in the first computing node, including: receive the negotiation response information returned by each communication library process through the corresponding network card port. If the address in the first negotiation response information is the same as the source address in the first packet, determine the network card port that returns the first negotiation response information as the first network card port. That is to say, it is necessary to determine which network card port sent the packet with the first transmission multi-tuple through the address in the negotiation response information. A scheduling instruction can be sent to the first network card port. When the first communication library process corresponding to the first network card port detects that the first network card port has received the scheduling instruction, the first communication library process can be used to establish a mapping relationship between the first transmission multi-tuple and the second transmission multi-tuple. Please refer to Figure 12 is a schematic diagram of the execution process in the scheduling stage provided by an embodiment of the present application. If network card port 5008 is determined to be the first network card port, a scheduling instruction is sent to network card port 5008, and the communication library process corresponding to network card port 5008 can return scheduling response information through network card port 5008. The scheduling response information is used to indicate whether the communication library process corresponding to network card port 5008 has successfully established a mapping relationship between the first transmission multi-tuple and the second transmission multi-tuple.

[0134] Exemplarily, a negotiation message can be carried in the negotiation instruction. The negotiation message (ECHO message) is used to query whether the target communication library running in the first computing node supports the message modification function. The negotiation response information can carry a negotiation response message (ECHO-RESPONSE message), and the negotiation response message is used to indicate whether the target communication library running in the first computing node supports the message modification function. A scheduling instruction carries a scheduling message (UPDATE message), and the scheduling message is used to instruct the first communication library process to establish a mapping relationship between the first transmission tuple and the second transmission tuple. The negotiation response information carries a scheduling response message (UPDATE-RESPONSE message), and the scheduling response message is used to indicate whether the first communication library process has successfully established a mapping relationship between the first transmission tuple and the second transmission tuple. The lengths of the headers of the negotiation message, the negotiation response message, the scheduling message, and the scheduling response message are fixed, and the basic TLV (Type-Length-Value) method can be used for definition, where Value represents the data in the header body. Please refer to Figure 13 is a schematic diagram of a header format definition provided by an embodiment of the present application. The header can include the following fields: Marker, Version, Length, and Type. The length of Marker is 8 octets (8 eight-bit bytes), and the value is all 1s, that is, 8 oxff. Version is a 1-octet unsigned integer, which can be used to represent the protocol version and perform negotiation. Length is a 2-octet unsigned integer, representing the length of the message (including the header and the message body), and the computing node can find the position of the Marker of the next message through Length. Type is a 1-octet unsigned integer, representing the message type. For example, when Type = 0x01, it represents an ECHO message; when Type = 0x02, it represents an ECHO-RESPONSE message; when Type = 0x03, it represents an UPDATE message; when Type = 0x04, it represents an UPDATE-RESPONSE message.

[0135] Please refer to Figure 14 is a schematic diagram of an ECHO message body format definition provided by an embodiment of the present application. The message body in the ECHO message can carry Version. Please refer to Figure 15 is a schematic diagram of an ECHO-RESPONSE message body format definition provided by an embodiment of the present application. The message body in the ECHO-RESPONSE message can carry Version, Type, Length, and the address of the network card port (Address V4). Please refer toFigure 16 This is a schematic diagram of the format definition of an UPDATE message body provided by an embodiment of the present application. The message body in the UPDATE message can carry: Update ID (message identifier), Update Type (modification flag), Update-records Length (length of the modification indication information), and Update-records (modification indication information). Among them, Update ID is a 4-octets unsigned long (32-bit long integer) and is used to uniquely identify the UPDATE message. Update Type is a 1-octet unsigned integer, and its value is 0x01, representing the source port number of the modified message. Update-records Length is a 2-octet unsigned integer, representing the length of the Update-records field. Update-records can be composed of one or more Update Record (modification indications). Exemplarily, an UpdateRecord can be composed of a source port number (Src Port), a destination port number (Dst Port), a source address (Src Address V4), a destination address (Dst Address V4), a protocol number (protocol), and a new source port number (New Src Port). That is to say, an Update Record can be composed of a first transmission tuple and a third element. Please refer to Figure 17 This is a schematic diagram of the format definition of an UPDATE-RESPONSE message body provided by an embodiment of the present application. The message body in the UPDATE-RESPONSE message can carry Update ID, Type, Length, and Response Code. The Response Code represents an error code and is used to feedback the modification result, that is, whether the mapping relationship between the first transmission tuple and the second transmission tuple is successfully established.

[0136] In one implementation, when the negotiation response information indicates that the target communication library running in the first computing node does not support the message modification function, a congestion notification message is sent to the target communication library running in the first computing node. The congestion notification message is used to instruct the target communication library running in the first computing node to adjust the message sending rate. For example, after the target communication library running in the first computing node receives the congestion notification message, it can reduce the message sending rate.

[0137] S14. When there is a first message with a first transmission tuple to be transmitted in the first network path, the first message is scheduled to be transmitted in the second network path based on the mapping relationship.

[0138] In one implementation, when the first communication library process creates a first packet with a first transmission tuple to be transmitted, it can call the first communication library process to obtain a second transmission tuple (or a third element) that has a mapping relationship with the first transmission tuple, and can call the first communication library process to modify the first transmission tuple in the first packet to the second transmission tuple (or modify the first element in the first packet to the third element) to obtain a second packet. Then, the second packet can be sent out through the first network card port, and each node in the network cluster can perform path selection based on the second transmission tuple in the second packet, and then transmit the second packet in the second network path.

[0139] In a feasible implementation, the network controller can use Policy-Based Routing (PBR) to replace the source port number modification on the end side, but it is necessary to write the mapping relationship between the first transmission tuple and the second transmission tuple (or the third tuple) into the configuration of the network node. Limited by the maximum number of table entries of the network node, and the later operation and maintenance and configuration management and maintenance are more complex than directly modifying the source port number on the end side.

[0140] It can be seen that the embodiment of the present application adopts the method of end-network collaboration to link the network controller in the network with the target communication library on the end side, realizing dynamic traffic scheduling in the scenarios of network congestion and hash conflict. And by adopting a lightweight source port number modification method, the traffic on the congested link can be redistributed to a suitable path in the network, thereby eliminating the influence of the network hot spot link and avoiding the problems of local hot spots and uneven load caused by random hashing of network nodes, ensuring the efficiency of the execution of computing tasks, and having strong practicability.

[0141] In summary, please refer to Figure 18 is a flow diagram of a network congestion handling solution provided by an embodiment of the present application Figure 2 and its steps include the following S21-S28:

[0142] S21. Link congestion monitoring: The service program can monitor the link congestion in the network cluster based on network monitoring technologies such as DCQCN. When it detects that a certain network link is congested, it notifies the network controller to perform congestion scheduling.

[0143] S22: Congested data flow selection: When the network controller receives a congestion scheduling request, it analyzes the packets on the congested network link to determine the data flow with the largest amount of data transmitted on the network link, and determines the transmission tuple of the data flow as the first transmission tuple.

[0144] S23: Path selection: Based on the static topology structure of the network cluster and the real-time traffic information of each network path, path selection is performed. When performing path selection, the network path with the smallest load fraction can be selected as the scheduling target, that is, the most idle network path is selected.

[0145] S24: Source port decision: Based on the scheduling target obtained in S23, the corresponding offset parameter can be determined. Performing an exclusive OR operation on the offset parameter and the first element in the first transmission tuple can determine the new source port number (i.e., the third element).

[0146] S25: Port modification negotiation: Send an ECHO message to detect whether the message modification function is supported. If it is supported, execute S26; if not, execute 27.

[0147] S26: Source port modification: Cooperate with the target communication library to notify the corresponding communication library process to establish a mapping relationship between the first transmission tuple and the third element.

[0148] S27: Adopt a network-side scheduling scheme to relieve network congestion: For example, adjust the message sending rate of some computing nodes.

[0149] S28: Congestion relief verification: Check whether the traffic is scheduled to the newly selected network path, whether the traffic on the original congested link has decreased, and whether the network congestion has been relieved.

[0150] The whole application relies on the mutual cooperation between the controller software on the network side and the target communication library on the end side to achieve. It is a dynamic scheduling scheme for end-network collaboration. Moreover, the modification of the source port number on the end side is only valid within the life cycle of the current communication library process. When the communication library process is destroyed, the information related to the communication library process will be cleared, and this source port number modification can automatically become invalid without the need to revoke the link scheduling, and it will not affect the network cluster. At the same time, in the communication library process on the end side, each communication library process only needs to monitor fixed network card ports and implement the corresponding communication protocol, and perform corresponding processing after receiving ECHO messages and UPDATE messages, which will not affect the collective communication performance on the end side.

[0151] For example, please refer to Figure 19It is a schematic diagram of the execution process of a communication library process provided by an embodiment of the present application. Each communication library process can call a socket function (Socket function) to create a socket, and the Socket function will return a socket descriptor (fd). Each communication library process can bind the socket to the corresponding network card port through the returned socket descriptor, which can be achieved by calling a binding function (bind function). Each communication library process can call a monitoring function (Listen function) to wait for a request to establish a communication connection sent by the network controller to the corresponding network card port, and call a first receiving function (Accept function) to accept this request to establish a communication connection, and generate a new socket descriptor for communicating with the network controller. Then, it can call a second receiving function (recv function) to receive an ECHO message (negotiation message) or an UPDATE message (scheduling message) sent by the network controller to the corresponding network card port, and perform corresponding processing, that is, call a sending function (Send function) to return an ECHO-RESPONSE message (negotiation response message) or an UPDATE-RESPONSE message (scheduling response message) to the network controller. After the processing is completed, it can call a closing function (Close function) to close the communication connection established with the network controller, and the newly created socket descriptor is closed (that is, the system resources associated with this new socket descriptor are released), and it will not affect the collective communication performance of the peer side.

[0152] NCCL Tests is a tool for evaluating the communication performance of a network cluster. It can simulate a parallel computing scenario in a multi-GPU environment by executing a series of collective communication operations to measure the data transfer speed, latency, and overall communication efficiency in the network cluster. Please refer to Figure 20a It is a schematic diagram of a test result provided by an embodiment of the present application Figure 1 When NCCL Tests starts the test, the traffic of the Eth200GE20 port of a switch in the network cluster surges. By analyzing the traffic passing through the Eth200GE20 port, a data stream with a source port number of 60051 is selected for scheduling, and the source port number 60051 of this data stream is modified to 57993. After the scheduling is completed, the traffic of the Eth200GE20 port starts to decline. Since this data stream is scheduled to a new link of the Eth200GE88 port, the traffic of the Eth200GE88 port increases. Please refer to Figure 20b It is a schematic diagram of a test result provided by an embodiment of the present application Figure 2 When the source port number 60051 of this data stream is modified to 57993, the bandwidth of the Eth200GE20 port starts to decline. Since this data stream is scheduled to a new link of the Eth200GE88 port, the bandwidth of the Eth200GE88 port increases.

[0153] The method of the embodiment of the present application is elaborated in detail above. To facilitate better implementation of the above solution of the embodiment of the present application, correspondingly, the device of the embodiment of the present application is provided below.

[0154] Figure 21 It is a schematic structural diagram of a network congestion handling device provided by an embodiment of the present application; the network congestion handling device can be used to execute some or all of the steps in the foregoing method embodiment. Please refer to Figure 21 , the network congestion handling device includes the following units: an acquisition unit 31 and a processing unit 32.

[0155] The acquisition unit 31 is configured to determine a first network path with congestion in the network cluster and acquire a first transmission tuple corresponding to the first network path;

[0156] The processing unit 32 is configured to determine a second network path from multiple network paths based on the real-time traffic information of the multiple network paths in the network cluster, and acquire a second transmission tuple corresponding to the second network path; the transmission load of the second network path meets the light load condition;

[0157] The processing unit 32 is further configured to establish a mapping relationship between the first transmission tuple and the second transmission tuple;

[0158] The processing unit 32 is further configured to, when there is a first packet with the first transmission tuple to be transmitted in the first network path, schedule the first packet to the second network path for transmission based on the mapping relationship.

[0159] In one embodiment, each network path in the network cluster includes at least one network port; the processing unit 32 determines the first network path with congestion in the network cluster, including:

[0160] When detecting a congestion event on a first network port in the network cluster, determining the network path including the first network port in the network cluster as the first network path with congestion in the network cluster, and the first network port is any network port in the network cluster.

[0161] In one embodiment, the first network port is provided with a first packet buffer queue, and the first packet buffer queue is used to store the packets to be transmitted by the first network port, and the packets have transmission tuples; the acquisition unit 31 acquires the first transmission tuple corresponding to the first network path, including:

[0162] Acquiring the occurrence times of the transmission tuples of each packet in the first packet buffer queue in the first packet buffer queue;

[0163] Determining the transmission tuple with the most occurrence times as the first transmission tuple corresponding to the first network path.

[0164] In one embodiment, the first transmission tuple includes a first element for indicating a source port number; the processing unit 32 determines a second network path from multiple network paths in the network cluster based on real-time traffic information of the multiple network paths, including:

[0165] Generate N alternative elements for the first element, where N is a positive integer;

[0166] Replace the first element in the first transmission tuple with the N alternative elements to obtain N alternative transmission tuples;

[0167] Filter out N alternative network paths from the network cluster according to the N alternative transmission tuples;

[0168] Based on the real-time traffic information of the N alternative network paths, determine a second network path from the N alternative network paths.

[0169] In one embodiment, the processing unit 32 generates N alternative elements for the first element, including:

[0170] Create N offset parameters for the first element;

[0171] Perform exclusive OR operations on the first element and the N offset parameters respectively to generate N alternative elements.

[0172] In one embodiment, the processing unit 32 filters out N alternative network paths from the network cluster according to the N alternative transmission tuples, including:

[0173] Obtain the static topology structure of the network cluster;

[0174] Perform hash calculations on each of the N alternative transmission tuples to obtain a hash value corresponding to each alternative transmission tuple;

[0175] According to the hash value corresponding to each alternative transmission tuple, perform path selection in the static topology structure starting from the first computing node to obtain the network path corresponding to each alternative transmission tuple. The network cluster includes multiple computing nodes, and the first computing node is the computing node that creates a message with the first transmission tuple;

[0176] Determine the network path corresponding to each alternative transmission tuple as N alternative network paths in the network cluster.

[0177] In one embodiment, the acquisition unit 31 acquires a second transmission tuple corresponding to the second network path, including:

[0178] Determine the alternative transmission tuple used to filter out the second network path among the N alternative transmission tuples as the second transmission tuple corresponding to the second network path.

[0179] In one embodiment, a network cluster includes multiple computing nodes, and a target communication library runs in each computing node. The packet with the first transmission tuple is created by the target communication library running in the first computing node. The processing unit 32 establishes a mapping relationship between the first transmission tuple and the second transmission tuple, including:

[0180] Sending a negotiation instruction to the target communication library running in the first computing node. The negotiation instruction is used to inquire whether the target communication library running in the first computing node supports the packet modification function;

[0181] Receiving the negotiation response information returned by the target communication library running in the first computing node. When the negotiation response information indicates that the target communication library running in the first computing node supports the packet modification function, sending a scheduling instruction to the target communication library running in the first computing node. The scheduling instruction is used to instruct the target communication library running in the first computing node to establish a mapping relationship between the first transmission tuple and the second transmission tuple.

[0182] In one embodiment, the processing unit 32 is further configured to:

[0183] When the negotiation response information indicates that the target communication library running in the first computing node does not support the packet modification function, sending a congestion notification message to the target communication library running in the first computing node. The congestion notification message is used to instruct the target communication library running in the first computing node to adjust the packet sending rate.

[0184] In one embodiment, the first computing node includes at least one network card port, and at least one communication library process is created by the target communication library running in the first computing node. One network card port corresponds to one communication library process. The processing unit 32 sends a negotiation instruction to the target communication library running in the first computing node, including:

[0185] Determining the node address of the first computing node by using the source address in the first packet;

[0186] Establishing a communication connection between the node address of the first computing node and each network card port in the first computing node, and sending a negotiation instruction to each network card port in the first computing node. When each communication library process detects that the corresponding network card port receives the negotiation instruction, if the target communication library running in the first computing node supports the packet modification function, generating negotiation response information including the address of the corresponding network card port.

[0187] In one embodiment, when the processing unit 32 receives the negotiation response information returned by the target communication library running in the first computing node and the negotiation response information indicates that the target communication library supports the packet modification function, sending a scheduling instruction to the target communication library running in the first computing node, including:

[0188] Receive the negotiation response information returned by each communication library process through the corresponding network card port. If the address in the first negotiation response information is the same as the source address in the first message, determine the network card port that returns the first negotiation response information as the first network card port;

[0189] Send a scheduling instruction to the first network card port. When the first communication library process corresponding to the first network card port detects that the first network card port receives the scheduling instruction, the first communication library process is used to establish a mapping relationship between the first transmission tuple and the second transmission tuple.

[0190] In one embodiment, when there is a first message with a first transmission tuple to be transmitted in the first network path, the processing unit 32 schedules the first message to be transmitted in the second network path based on the mapping relationship, including:

[0191] When the first communication library process creates a first message with a first transmission tuple to be transmitted, call the first communication library process to obtain a second transmission tuple that has a mapping relationship with the first transmission tuple;

[0192] Call the first communication library process to modify the first transmission tuple in the first message to the second transmission tuple to obtain a second message;

[0193] Send the second message through the first network card port and transmit the second message in the second network path.

[0194] According to an embodiment of the present application, Figure 21 Each unit in the network congestion handling device shown can be separately or entirely combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units with functional division to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the network congestion handling device can also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units. According to another embodiment of the present application, it is possible to construct, for example, on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), a computer program (including program code) that can execute each step involved in the foregoing method, Figure 21The network congestion handling device shown in [description], and to implement the network congestion handling method of the embodiments of the present application. The computer program can be recorded on, for example, a computer-readable recording medium, and loaded into the above-mentioned computing device through the computer-readable recording medium and run therein.

[0195] The network congestion handling solution provided by the embodiments of the present application can determine a first network path where congestion occurs in a network cluster, and obtain a first transmission tuple corresponding to the first network path; Exemplarily, the packets transmitted in the network cluster have transmission tuples, and the transmission tuples possessed by the packets can be used to determine the network path through which the packets pass when transmitted in the network cluster. The first transmission tuple can be determined according to the transmission tuples possessed by the packets transmitted on the first network path. Based on the real-time traffic information of multiple network paths in the network cluster, a second network path is determined from the multiple network paths, and a second transmission tuple corresponding to the second network path is obtained; The transmission load of the second network path meets the light load condition; Exemplarily, the packets with the second transmission tuple can be transmitted on the second network path. A mapping relationship is established between the first transmission tuple and the second transmission tuple. When there is a first packet with the first transmission tuple to be transmitted in the first network path, the first packet is scheduled to be transmitted in the second network path based on the mapping relationship; For example, the first transmission tuple in the first packet can be modified to the second transmission tuple, so that the first packet can be scheduled to be transmitted on the second network path. Thus, it can be seen that the embodiments of the present application can reallocate some of the traffic transmitted through the first network path to be transmitted on the second network path, relieve the transmission pressure of the first network path, facilitate load balancing, and improve the stability of the network cluster.

[0196] Figure 22 is a schematic structural diagram of a computer device provided by the embodiments of the present application. Please refer to Figure 22 , The computer device includes a processor 41, a communication interface 42, and a computer-readable storage medium 43. Among them, the processor 41, the communication interface 42, and the computer-readable storage medium 43 can be connected through a bus or other means. Among them, the communication interface 42 is used to receive and send data. The computer-readable storage medium 43 can be stored in the memory of the computer device. The computer-readable storage medium 43 is used to store a computer program. The computer program includes program instructions, and the processor 41 is used to execute the program instructions stored in the computer-readable storage medium 43. The processor 41 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the computer device, and is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.

[0197] An embodiment of the present application further provides a computer-readable storage medium (Memory). A computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the computer device. And, in this storage space, one or more instructions suitable for being loaded and executed by the processor 41 are also stored, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0198] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the processor 41 loads and executes one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the above-mentioned embodiment of the network congestion handling method; in a specific implementation, one or more instructions in the computer-readable storage medium are loaded and executed by the processor 41 to perform the following steps:

[0199] Determine the first network path where congestion occurs in the network cluster, and obtain the first transmission tuple corresponding to the first network path;

[0200] Based on the real-time traffic information of multiple network paths in the network cluster, decide on a second network path from the multiple network paths, and obtain the second transmission tuple corresponding to the second network path; the transmission load of the second network path meets the light load condition;

[0201] Establish a mapping relationship between the first transmission tuple and the second transmission tuple;

[0202] When there is a first packet with the first transmission tuple to be transmitted in the first network path, schedule the first packet to the second network path for transmission based on the mapping relationship.

[0203] In one embodiment, at least one network port is included on any network path in the network cluster; the processor 41 determines the first network path where congestion occurs in the network cluster, including:

[0204] When it is detected that a congestion event occurs at the first network port in the network cluster, determine the network path in the network cluster that includes the first network port as the first network path where congestion occurs in the network cluster, and the first network port is any network port in the network cluster.

[0205] In one embodiment, a first message buffer queue is provided for a first network port. The first message buffer queue is used to store messages to be transmitted by the first network port, and the messages have transmission tuples; the processor 41 obtains a first transmission tuple corresponding to a first network path, including:

[0206] Obtaining the number of occurrences of the transmission tuples of each message in the first message buffer queue in the first message buffer queue;

[0207] Determining the transmission tuple with the most occurrences as the first transmission tuple corresponding to the first network path.

[0208] In one embodiment, the first transmission tuple includes a first element for indicating a source port number; the processor 41 determines a second network path from multiple network paths based on the real-time traffic information of the multiple network paths in the network cluster, including:

[0209] Generating N alternative elements for the first element, where N is a positive integer;

[0210] Replacing the first element in the first transmission tuple with the N alternative elements to obtain N alternative transmission tuples;

[0211] Filtering out N alternative network paths from the network cluster according to the N alternative transmission tuples;

[0212] Determining a second network path from the N alternative network paths based on the real-time traffic information of the N alternative network paths.

[0213] In one embodiment, the processor 41 generates N alternative elements for the first element, including:

[0214] Creating N offset parameters for the first element;

[0215] Performing exclusive OR operations on the first element and the N offset parameters respectively to generate N alternative elements.

[0216] In one embodiment, the processor 41 filters out N alternative network paths from the network cluster according to the N alternative transmission tuples, including:

[0217] Obtaining the static topology structure of the network cluster;

[0218] Performing hash calculations on each of the N alternative transmission tuples to obtain a hash value corresponding to each alternative transmission tuple;

[0219] Based on the hash values corresponding to each alternative transmission tuple, path selection is performed in the static topology starting from the first computing node to obtain the network paths corresponding to each alternative transmission tuple. The network cluster includes multiple computing nodes, and the first computing node is the computing node that creates the message with the first transmission tuple.

[0220] Determine the network paths corresponding to each alternative transmission tuple as N alternative network paths in the network cluster.

[0221] In one embodiment, the processor 41 obtains the second transmission tuple corresponding to the second network path, including:

[0222] Determine the alternative transmission tuple among the N alternative transmission tuples that is used to filter out the second network path as the second transmission tuple corresponding to the second network path.

[0223] In one embodiment, the network cluster includes multiple computing nodes, and a target communication library runs in each computing node. The message with the first transmission tuple is created by the target communication library running in the first computing node; the processor 41 establishes a mapping relationship between the first transmission tuple and the second transmission tuple, including:

[0224] Send a negotiation instruction to the target communication library running in the first computing node. The negotiation instruction is used to inquire whether the target communication library running in the first computing node supports the message modification function;

[0225] Receive the negotiation response information returned by the target communication library running in the first computing node. When the negotiation response information indicates that the target communication library running in the first computing node supports the message modification function, send a scheduling instruction to the target communication library running in the first computing node. The scheduling instruction is used to instruct the target communication library running in the first computing node to establish a mapping relationship between the first transmission tuple and the second transmission tuple.

[0226] In one embodiment, the processor 41 is further configured to:

[0227] When the negotiation response information indicates that the target communication library running in the first computing node does not support the message modification function, send a congestion notification message to the target communication library running in the first computing node. The congestion notification message is used to instruct the target communication library running in the first computing node to adjust the message sending rate.

[0228] In one embodiment, the first computing node includes at least one network card port, and the target communication library running in the first computing node creates at least one communication library process. One network card port corresponds to one communication library process; the processor 41 sends a negotiation instruction to the target communication library running in the first computing node, including:

[0229] Determine the node address of the first computing node using the source address in the first message;

[0230] Establish a communication connection with each network card port in the first computing node based on the node address of the first computing node, and send a negotiation instruction to each network card port in the first computing node; when each communication library process detects that the corresponding network card port has received the negotiation instruction, if the target communication library running in the first computing node supports the message modification function, generate negotiation response information including the address of the corresponding network card port.

[0231] In one embodiment, the processor 41 receives the negotiation response information returned by the target communication library running in the first computing node. When the negotiation response information indicates that the target communication library supports the message modification function, send a scheduling instruction to the target communication library running in the first computing node, including:

[0232] Receive the negotiation response information returned by each communication library process through the corresponding network card port. If the address in the first negotiation response information is the same as the source address in the first message, determine the network card port that returns the first negotiation response information as the first network card port;

[0233] Send a scheduling instruction to the first network card port. When the first communication library process corresponding to the first network card port detects that the first network card port has received the scheduling instruction, the first communication library process is used to establish a mapping relationship between the first transmission tuple and the second transmission tuple.

[0234] In one embodiment, when there is a first message with a first transmission tuple to be transmitted in the first network path, the processor 41 schedules the first message to be transmitted in the second network path based on the mapping relationship, including:

[0235] When the first communication library process creates a first message with a first transmission tuple to be transmitted, call the first communication library process to obtain a second transmission tuple that has a mapping relationship with the first transmission tuple;

[0236] Call the first communication library process to modify the first transmission tuple in the first message to the second transmission tuple to obtain a second message;

[0237] Send the second message through the first network card port and transmit the second message in the second network path.

[0238] Based on the same inventive concept, the principle of solving problems and the beneficial effects of the computer device provided in the embodiments of the present application are similar to the principle of solving problems and the beneficial effects of the network congestion handling method in the method embodiments of the present application. The principle and beneficial effects of the method implementation can be referred to. For the sake of brevity, they will not be described here again.

[0239] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.

[0240] The embodiments of the present application also provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned network congestion handling method.

[0241] The above description is only the specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A network congestion handling method, characterized in that, The method includes: Determine a first network path with congestion in a network cluster, and obtain a first transmission tuple corresponding to the first network path; at least a first computing node is included in the network cluster, a target communication library runs in the first computing node, the target communication library running in the first computing node supports a message modification function, the first computing node includes at least one network card port, at least one communication library process is created by the target communication library running in the first computing node, and one network card port corresponds to one communication library process; Based on the real-time traffic information of multiple network paths in the network cluster, determine a second network path from the multiple network paths, and obtain a second transmission tuple corresponding to the second network path; the transmission load of the second network path meets a light load condition; Establish a mapping relationship between the first transmission tuple and the second transmission tuple, where the mapping relationship is established by a first communication library process corresponding to a first network card port in the first computing node, the first network card port is the network card port in the first computing node for sending a message with the first transmission tuple, and the first communication library process is used to create a message with the first transmission tuple; When the first communication library process creates a first message with the first transmission tuple to be transmitted, call the first communication library process to obtain the second transmission tuple that has a mapping relationship with the first transmission tuple, and modify the first transmission tuple in the first message to the second transmission tuple to obtain a second message; Call the first network card port to send the second message, and transmit the second message in the second network path; When the first communication library process is destroyed, perform an invalidation process on the modification performed by the first communication library process on the first message.

2. The method according to claim 1, wherein At least one network port is included on any network path in the network cluster; determining the first network path with congestion in the network cluster includes: When detecting a congestion event occurring at a first network port in the network cluster, determine the network path in the network cluster that includes the first network port as the first network path with congestion in the network cluster, and the first network port is any network port in the network cluster.

3. The method according to claim 2, wherein The first network port is provided with a first message buffer queue, and the first message buffer queue is used to store messages to be transmitted by the first network port, and the messages have transmission tuples; obtaining the first transmission tuple corresponding to the first network path includes: Obtain the occurrence times of the transmission tuples included in each message in the first message buffer queue in the first message buffer queue; Determine the transmission tuple with the most occurrence times as the first transmission tuple corresponding to the first network path.

4. The method according to any one of claims 1-3, characterized in that, The first transmission tuple includes a first element for indicating a source port number; based on the real-time traffic information of multiple network paths in the network cluster, determining a second network path from the multiple network paths includes: Generate N alternative elements for the first element, where N is a positive integer; Replacing the first element in the first transmission multi - tuple with the N alternative elements to obtain N alternative transmission multi - tuples; Filtering out N alternative network paths from the network cluster according to the N alternative transmission multi - tuples; Based on the real - time traffic information of the N alternative network paths, determining a second network path from the N alternative network paths.

5. The method according to claim 4, wherein Generating N alternative elements for the first element, including: Creating N offset parameters for the first element; Performing exclusive - OR operations on the first element and the N offset parameters respectively to generate N alternative elements.

6. The method according to claim 4, wherein Filtering out N alternative network paths from the network cluster according to the N alternative transmission multi - tuples, including: Obtaining the static topology structure of the network cluster; Performing hash calculation on each alternative transmission multi - tuple in the N alternative transmission multi - tuples to obtain the hash value corresponding to each alternative transmission multi - tuple; According to the hash value corresponding to each alternative transmission multi - tuple, selecting a path in the static topology structure starting from the first computing node to obtain the network path corresponding to each alternative transmission multi - tuple. The network cluster includes multiple computing nodes, and the first computing node is the computing node that creates a message with the first transmission multi - tuple; Determining the network paths corresponding to each alternative transmission multi - tuple as N alternative network paths in the network cluster.

7. The method according to claim 4, wherein Obtaining the second transmission multi - tuple corresponding to the second network path, including: Determining the alternative transmission multi - tuple used to filter out the second network path among the N alternative transmission multi - tuples as the second transmission multi - tuple corresponding to the second network path.

8. The method according to claim 1, wherein The network cluster includes multiple computing nodes, and a target communication library runs in each computing node. The message with the first transmission multi - tuple is created by the target communication library running in the first computing node; Establishing a mapping relationship between the first transmission multi - tuple and the second transmission multi - tuple, including: Sending a negotiation instruction to the target communication library running in the first computing node. The negotiation instruction is used to inquire whether the target communication library running in the first computing node supports the message modification function; Receiving the negotiation response information returned by the target communication library running in the first computing node. When the negotiation response information indicates that the target communication library running in the first computing node supports the message modification function, sending a scheduling instruction to the target communication library running in the first computing node. The scheduling instruction is used to instruct the target communication library running in the first computing node to establish a mapping relationship between the first transmission multi - tuple and the second transmission multi - tuple.

9. The method according to claim 8, wherein The method further includes: When the negotiation response information indicates that the target communication library running in the first computing node does not support the message modification function, sending a congestion notification message to the target communication library running in the first computing node. The congestion notification message is used to instruct the target communication library running in the first computing node to adjust the message sending rate.

10. The method according to claim 8, wherein Sending the negotiation instruction to the target communication library running in the first computing node, including: Determine the node address of the first computing node using the source address in the first message; Based on the node address of the first computing node, establish a communication connection with each network card port in the first computing node, and send a negotiation instruction to each network card port in the first computing node; when each communication library process detects that the corresponding network card port has received the negotiation instruction, if the target communication library running in the first computing node supports the message modification function, generate negotiation response information including the address of the corresponding network card port.

11. The method according to claim 10, wherein Receive the negotiation response information returned by the target communication library running in the first computing node, and when the negotiation response information indicates that the target communication library supports the message modification function, send a scheduling instruction to the target communication library running in the first computing node, including: Receive the negotiation response information returned by each communication library process through the corresponding network card port. If the address in the first negotiation response information is the same as the source address in the first message, determine the network card port that returns the first negotiation response information as the first network card port; Send a scheduling instruction to the first network card port. When the first communication library process corresponding to the first network card port detects that the first network card port has received the scheduling instruction, the first communication library process is used to establish a mapping relationship between the first transmission multi-tuple and the second transmission multi-tuple.

12. A network congestion handling device, characterized in that, The device includes: An acquisition unit, configured to determine a first network path with congestion in the network cluster, and acquire a first transmission multi-tuple corresponding to the first network path; the network cluster at least includes a first computing node, a target communication library runs in the first computing node, the target communication library running in the first computing node supports the message modification function, the first computing node includes at least one network card port, at least one communication library process is created by the target communication library running in the first computing node, and one network card port corresponds to one communication library process; A processing unit, configured to make a decision on a second network path from the multiple network paths based on the real-time traffic information of the multiple network paths in the network cluster, and acquire a second transmission multi-tuple corresponding to the second network path; the transmission load of the second network path meets the light load condition; The processing unit is further configured to establish a mapping relationship between the first transmission multi-tuple and the second transmission multi-tuple, and the mapping relationship is established by the first communication library process corresponding to the first network card port in the first computing node. The first network card port is the network card port in the first computing node used to send a message with the first transmission multi-tuple, and the first communication library process is used to create a message with the first transmission multi-tuple; The processing unit is further configured to, when the first communication library process creates a first message to be transmitted with the first transmission multi-tuple, call the first communication library process to obtain the second transmission multi-tuple having a mapping relationship with the first transmission multi-tuple, and modify the first transmission multi-tuple in the first message to the second transmission multi-tuple to obtain a second message; call the first network card port to send the second message and transmit the second message in the second network path; when the first communication library process is destroyed, perform invalidation processing on the modification performed by the first communication library process on the first message.

13. A computer device, characterized in that, including: a processor adapted to execute a computer program; a computer-readable storage medium storing a computer program, which when executed by the processor, implements the network congestion handling method according to any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the network congestion handling method according to any one of claims 1-11.

15. A computer program product, characterized in that, The computer program product includes a computer program, which when executed by the processor, implements the network congestion handling method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Congestion recovery method and device based on fast rerouting, equipment and medium

    CN117173834A

  • Method, device and equipment for processing network congestion, network system and storage medium

    CN117459460A