Packet processing method and system
By implementing the message processing method on the switching node of the distributed system, determining the congestion state of the output port and selectively processing the message, the problem of inefficient collective communication in the prior art is solved, and the AI training performance is improved.
Patent Information
- Application Number
- PCT/CN2024/141034
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-05
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
When performing collective communication, existing distributed systems have problems such as network topology dependence, high protocol processing complexity, large maintenance overhead, and unsuitable ring networks, resulting in low efficiency of collective communication and affecting AI training performance.
By implementing the message processing method on the switching node, it is determined whether the output port is in a congested state, and selective processing is performed based on the judgment results, including direct forwarding or standardized operations, reducing the length of the queue to be forwarded and alleviating network congestion.
It improves the efficiency of collective communication, reduces the number of messages transmitted within the network, alleviates network congestion, and offloads the computing pressure of computing devices, improving the overall performance of AI training.
Smart Images

Figure CN2024141034_26062025_PF_FP_ABST
Abstract
Description
Message processing method and system
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent applications filed with the China Patent Office on December 21, 2023, with application number 202311777191.8 and public name “Message Processing Method and System” and filed with the China Patent Office on August 5, 2024, with application number 202411070108.8 and public name “Message Processing Method and System”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to a distributed system, and specifically to a message processing method, system, electronic device, and medium suitable for artificial intelligence model training. Background Art
[0004] A trend in AI (Artificial Intelligence) applications is the increasing size of AI models and the increasing volume of AI training data. For example, the model parameter size of ChatGPT and GPT-3 is 175 billion, and the text training data is 570GB. Traditional single computing devices are no longer able to support such large model training. Distributed systems consisting of multiple computing devices and switching nodes are usually used to implement parallel computing training to improve overall AI training performance. Using this distributed parallel computing training method, computing devices use switching nodes to meet various communication requirements, such as transferring training data, updating model parameters, etc. The communication between multiple computing devices in the distributed system used for parallel computing training is different from point-to-point communication. There are often a large number of one-to-many or many-to-many communication requirements between computing devices, which is usually called collective communication. Among them, Reduce / Allreduce is a commonly used collective communication in parallel computing training, and improving Reduce / Allreduce performance can significantly improve the overall performance of AI training.
[0005] Some existing solutions propose on-network computing, which offloads collective communication operations to switching nodes (e.g., switch networks), thereby reducing the amount of data transmitted over the network and shortening message delivery operations. For example, SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) is a network offload technology for collective communication. As shown in Figure 1, before performing collective communication, SHARP constructs a tree-like logical network topology based on the physical network topology, consisting of aggregation nodes (typically network devices) and terminal nodes (typically computing devices). Terminal nodes send data to their respective aggregation nodes. After collecting data from their child nodes, the aggregation nodes perform on-network computations and send the results to the aggregation nodes in the next level up. The aggregation nodes in the next level, after collecting data from their child nodes, perform on-network computations and send the results to the aggregation nodes in the next level up. This cycle repeats until the root aggregation node completes its on-network computations and sends the results to all its direct child nodes. The direct child nodes, upon receiving the results, send the results to the direct child nodes in the next level up. This cycle repeats until the final result is sent to the terminal node. However, the disadvantages of this type of solution are:
[0006] 1) It is dependent on the physical network topology and must be able to form a tree-shaped aggregation network topology. Ring networks are not suitable for SHARP deployment.
[0007] 2) The introduction of the additional SHARP protocol increases the complexity of protocol processing;
[0008] 3) A logical communication tree needs to be established in advance, which introduces additional overhead;
[0009] 4) The logical communication tree needs to be maintained. For example, if a node fails, the logical communication tree needs to be re-established, which introduces additional maintenance overhead and increases the maintenance complexity of the entire network.
[0010] Therefore, it is necessary to provide a method or system that can optimize the efficiency of collective communication, which can be used in distributed systems to improve the overall performance of AI training. Summary of the Invention
[0011] Based on the above situation, the main purpose of this application is to provide a method, system, electronic device and medium for message processing.
[0012] To achieve the above objectives, the technical solutions adopted in this application are as follows:
[0013] A first aspect of the present application provides a message processing method for a distributed system, the method comprising the following steps:
[0014] The switching node of the distributed system receives a message from a computing device in the distributed system, and determines an output port of the switching node to which the message needs to be forwarded;
[0015] The switching node determines whether the output port is in a congested state, and processes the message according to the congestion state result obtained by the determination, specifically:
[0016] When the queue to be forwarded of the output port is empty, the message is directly forwarded from the output port;
[0017] When the queue to be forwarded of the output port is not empty, the message is analyzed, a processing strategy for the message is determined based on the analysis result, and the message is selectively processed according to the processing strategy.
[0018] Preferably, determining the output port of the switching node to which the message needs to be forwarded comprises the following steps:
[0019] The output port is determined according to a routing table of the switching node.
[0020] Preferably, analyzing the message, determining a processing strategy for the message based on the analysis result, and selectively processing the message according to the processing strategy comprises the following steps:
[0021] If the message is a collective communication message, the message is parsed and processed according to the parsing results, including:
[0022] According to the parsing result, determine whether the message is a message that needs to be regulated.
[0023] If so, perform the corresponding protocol operation according to the message.
[0024] If not, storing the message in the cache of the switching node and moving the message into the queue to be forwarded of the output port;
[0025] If the message is not a collective communication message, the message is stored in the cache of the switching node, and the message is moved into the queue to be forwarded of the output port.
[0026] Preferably, analyzing the message comprises the following steps:
[0027] Determine whether the message is a collective communication message based on a preset message identification condition.
[0028] Preferably, if the message is a collective communication message, parsing the message and processing the message according to the parsing result includes the following steps:
[0029] The protocol data unit in the message is parsed, and it is determined whether the message is a message that needs to be standardized based on the preset collective communication operation type field obtained by the parsing.
[0030] Preferably, performing corresponding protocol operations according to the message includes the following steps:
[0031] Extracting first data of a preset collective communication identification field from the protocol data unit of the message;
[0032] According to the first data, a query is performed in a preset table in the distributed system with the preset set communication identification field as the primary key,
[0033] If no record containing the first data is found in the preset table, the message is stored in the cache of the switching node and moved into the queue to be forwarded, and a record corresponding to the message is added to the preset table, and the record includes address information of the message in the cache;
[0034] If a record containing the first data is found in the preset table, the second data used for the reduction calculation is extracted from other messages corresponding to the record stored in the cache, the second data is subjected to corresponding reduction calculation with the first data used for the reduction calculation of the message, and the second data of other messages corresponding to the record in the cache are updated to the result obtained by the reduction calculation.
[0035] Preferably, the method further comprises the following steps:
[0036] When the switching node schedules a message from the cache according to the queue to be forwarded, if the message matches a record in the preset table, the switching node deletes the matched record from the preset table after forwarding the message.
[0037] Preferably, the preset collective communication identification field includes a group identification number field, and the group identification number field is globally unique and is used to group and identify the computing devices.
[0038] Preferably, the preset collective communication identification field further includes a message tag field for marking the protocol operation of the collective communication.
[0039] Preferably, the reduction operation includes finding the maximum value, finding the minimum value, finding the sum, finding the product, and finding the average value.
[0040] A second aspect of the present application provides a message processing system for a switching node of a distributed system, the system comprising:
[0041] a receiving unit, configured to receive a message from a computing device in the distributed system and determine an output port of the switching node to which the message is to be forwarded;
[0042] The processing unit is configured to determine whether the output port is in a congested state and process the message according to the congestion state result obtained by the determination, specifically:
[0043] When the queue to be forwarded of the output port is empty, the message is directly forwarded from the output port;
[0044] When the queue to be forwarded of the output port is not empty, the message is analyzed, a processing strategy for the message is determined based on the analysis result, and the message is selectively processed according to the processing strategy.
[0045] Preferably, the receiving unit comprises a first module, configured to determine the output port according to a routing table of the switching node.
[0046] Preferably, the processing unit includes a second module for determining whether the message is a collective communication message,
[0047] If the message is a collective communication message, the message is parsed and processed according to the parsing results, including:
[0048] According to the parsing result, determine whether the message is a message that needs to be regulated.
[0049] If so, perform the corresponding protocol operation according to the message.
[0050] If not, storing the message in the cache of the switching node and moving the message into the queue to be forwarded of the output port;
[0051] If the message is not a collective communication message, the message is stored in the cache of the switching node, and the message is moved into the queue to be forwarded of the output port.
[0052] Preferably, the second module is used to determine whether the message is a collective communication message according to a preset message identification condition.
[0053] Preferably, the second module is used to parse the protocol data unit in the collective communication message, and determine whether the message is a message that needs to be standardized based on the preset collective communication operation type field obtained by the parsing.
[0054] Preferably, the processing unit includes a third module, configured to extract first data of a preset aggregate communication identification field from the protocol data unit of the message;
[0055] and used to query a preset table in the distributed system with the preset set communication identification field as the primary key according to the first data,
[0056] If no record containing the first data is found in the preset table, the message is stored in the cache of the switching node and moved into the queue to be forwarded, and a record corresponding to the message is added to the preset table, and the record includes address information of the message in the cache;
[0057] If a record containing the first data is found in the preset table, the second data used for the reduction calculation is extracted from other messages corresponding to the record in the cache, the second data is subjected to corresponding reduction calculation with the first data used for the reduction calculation of the message, and the second data of other messages corresponding to the record in the cache are updated to the result obtained by the reduction calculation.
[0058] Preferably, the system further comprises:
[0059] The scheduling unit is used to schedule the messages in the queue to be forwarded in sequence, and if the message matches the record in the preset table, delete the matched record from the preset table after forwarding the message.
[0060] Preferably, the preset collective communication identification field includes a group identification number field, and the group identification number field is globally unique and is used to group and identify the computing devices.
[0061] Preferably, the preset collective communication identification field further includes a message tag field for marking the protocol operation of the collective communication.
[0062] Preferably, it is characterized in that the reduction operation includes finding the maximum value, finding the minimum value, finding the sum, finding the product, and finding the average value.
[0063] A third aspect of the present invention provides a distributed system, including several switching nodes and several computing devices, wherein the switching nodes can implement the method described in the first aspect.
[0064] Preferably, the plurality of computing devices are grouped according to a preset group identification number, and the group identification number is globally unique.
[0065] A fourth aspect of the present application provides an electronic device, comprising: a processor; and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method described in the first aspect can be implemented.
[0066] A fifth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is configured to be run to implement the method described in the first aspect above.
[0067] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] FIG1 is a schematic diagram of a logical network topology of SHARP in the prior art;
[0069] FIG2 is a flow chart of a preferred embodiment of the message processing method of the present application;
[0070] 3a to 3c are schematic diagrams of message forwarding of computing devices and switching nodes in the distributed system of the present application;
[0071] FIG4 is an example of the format of the collective communication message of the present application;
[0072] FIG5 is a schematic diagram of a preset table and cache of the distributed system of the present application;
[0073] FIG6 is a module diagram of a preferred embodiment of the message processing system of the present application. DETAILED DESCRIPTION
[0074] In order to further illustrate the technical means and effects adopted by this application to achieve the intended application purpose, the following, in combination with the accompanying drawings and preferred embodiments, describes in detail the specific implementation methods, methods, steps, features and effects of the methods, systems, electronic devices and computer-readable storage media proposed in this application.
[0075] Referring to FIG2 , a message processing method is used in a distributed system, and the method includes the following steps:
[0076] S100, a switching node of the distributed system receives a message from a computing device in the distributed system, and determines an output port of the switching node to which the message needs to be forwarded;
[0077] S200: The switching node determines whether the output port is in a congested state, and processes the message according to the congestion state result obtained by the determination, specifically:
[0078] S210, when the to-be-forwarded queue of the output port is empty, directly forwarding the message from the output port;
[0079] S220 , when the to-be-forwarded queue of the output port is not empty, analyzing the message, determining a processing strategy for the message based on the analysis result, and selectively processing the message according to the processing strategy.
[0080] A distributed system is typically composed of multiple computing devices that communicate over a network and coordinate their work to complete a common task. For example, a distributed system for AI model training includes multiple computing devices and multiple switching nodes. Computing devices, such as AI accelerators, are primarily used to complete computing tasks, while switching nodes, such as switching chips, switches, routing chips, and routers, are primarily used for communication between multiple computing devices.
[0081] Regarding step S210, in this embodiment, the message is parsed and processed accordingly (for example, protocol operation) according to the parsing result only when the output port is in a congested state (the queue to be forwarded is not empty) to alleviate the network congestion. When there is no message queued for processing at the output port, the message is forwarded directly without introducing additional waiting delay. Even if the message is a protocol operation, the message forwarding is completed as much as possible under the condition of smooth network, which helps to improve the problem of communication network delay.
[0082] Regarding collective communication, in distributed systems, there is often a large demand for collective communication between various computing devices. When conducting distributed AI model training, some relatively low-level message communication behaviors can usually be defined based on MPI (Message Passing Interface), such as Reduce, Allreduce, Scatter, Gather, Allgather, Broadcast, etc. Among them, Reduce, which is a reduction or reduction calculation, is a general term for a series of operations, generally including maximum value (MAX), minimum value (MIN), sum (SUM), product (PROD), average value (MEAN), etc.; Allreduce applies the same reduce operation to all node processes, which can be completed through reduce+broadcast operations on a single node.
[0083] As a further improvement to the above embodiment, in step S220, analyzing the message, determining a processing strategy for the message based on the analysis result, and selectively processing the message according to the processing strategy includes the following steps:
[0084] If the message is a collective communication message, the message is parsed and processed according to the parsing results, including:
[0085] According to the parsing result, determine whether the message is a message that needs to be regulated.
[0086] If so, perform the corresponding protocol operation according to the message.
[0087] If not, storing the message in the cache of the switching node and moving the message into the queue to be forwarded of the output port;
[0088] If the message is not a collective communication message, the message is stored in the cache of the switching node, and the message is moved into the queue to be forwarded of the output port.
[0089] For example, as shown in Figure 3a, AI accelerators 0-4 are connected as computing devices via a switch. When performing protocol reduction operations, AI accelerators 1-4 need to encapsulate their data into messages and transmit them to AI accelerator 0, which may cause congestion on the switch. As shown in Figure 3b, three messages will appear in the forwarding queue of the switch's output port, waiting to be processed. In this embodiment, based on the congestion state of the switching node, the collective communication messages are parsed to identify the messages that need to be reduced. The switching node performs protocol reduction calculations and updates the corresponding message data in the cache with the obtained protocol calculation results, thereby converting multiple related messages into a single message to be forwarded. As shown in Figure 3c, messages 2-4 originally in the forwarding queue are converted into message 2' after protocol reduction calculations by the switching node. It can be seen that performing on-network calculations on messages by the switching node can reduce the length of the forwarding queue and effectively alleviate network congestion.
[0090] Furthermore, after the on-line computational processing by the switching node, AI accelerator 0 no longer needs to process messages 2 to 4. It only needs to perform the corresponding reduction operation after receiving message 2', which reduces the computational burden on AI accelerator 0. Generally speaking, if there are N computing devices in a system, existing reduction calculation methods require N-1 or Log2(N) calculations. However, by adopting the method of this embodiment, the optimal number of reduction calculations can be reduced to 1.
[0091] Through the above steps, the message processing method of the present application is based on judging the congestion state of the switching node of the distributed system, and performing corresponding processing on the messages to be forwarded according to the judgment result, wherein, when a certain output port is in a congested state, the message is analyzed and the processing strategy is determined, and the message is selectively processed according to the processing strategy, for example, a protocol operation is performed on the collective communication message that needs to be reduced, so that when there are multiple messages queued for processing at the same output port of the switching node, the message that meets the protocol operation requirements is reduced. After the corresponding protocol calculation, multiple related messages can become one message, thereby reducing the length of the queue to be forwarded at the output port, which can effectively reduce the number of messages transmitted in the network and alleviate network congestion. At the same time, the switching node can also unload part of the computing pressure of the computing device, which helps to optimize the efficiency of the distributed system. In addition, there is no restriction on the specific transport layer protocol, and there is no need to add additional network communication protocols. There is no restriction on the network type or topology of the specific application, and the applicability is better.
[0092] As a further improvement to the above embodiment, in step S220, analyzing the message includes the following steps:
[0093] Determine whether the message is a collective communication message based on a preset message identification condition.
[0094] Specifically, collective communication is an upper-layer application that can be carried using network protocols such as UDP / IP, RDMA, and InfiniBand. Taking the use of UDP / IP protocol to carry collective communication as an example, compared to some well-known applications, such as the well-known port numbers 20 and 21 of FTP, and the well-known port number 80 of HTTP, there is no well-known port number when collective communication is carried using UDP / IP protocol. Instead, it depends on the specific implementation of the collective communication library on the computing node. Moreover, because the computing cluster network used for distributed systems is a relatively closed network, an arbitrary UDP port can be set to represent collective communication, usually using a non-well-known UDP port number, but of course a well-known port number can also be selected. Therefore, in a distributed system, message identification conditions can be set to identify collective communication messages. For example,
[0095] You can also consider adding a protection point where the conditions for identifying collective communication messages are configurable. For example, for a UDP / IP network, the identification conditions can be set as follows: "IP_Header.protocol == 17 (UDP), UDP_Header.destination_port == xxx", where xxx is the preset UDP port number used for collective communication.
[0096] Through the above steps, it is possible to quickly analyze and determine whether the message transmitted in the distributed system is a collective communication message, and perform corresponding processing based on the analysis results, which helps improve the efficiency of network message forwarding.
[0097] As an optional embodiment, in step S100, determining the output port of the switching node that is required to forward the message includes the following steps:
[0098] S110: Determine the output port according to the routing table of the switching node.
[0099] Specifically, after receiving a message from the computing device, the switching node obtains the forwarding path of the message by searching the routing table, and thereby determines the corresponding output port on the switching node.
[0100] Through the above steps, the switching node determines the output port for forwarding the message based on the received message, and subsequently determines the congestion status of the output port, and determines whether the message needs to be parsed and whether a protocol operation needs to be performed based on the congestion status structure obtained.
[0101] As an optional embodiment, in step S220, if the message is a collective communication message, parsing the message and processing the message according to the parsing result includes the following steps:
[0102] The protocol data unit in the message is parsed, and it is determined whether the message is a message that needs to be standardized based on the preset collective communication operation type field obtained by the parsing.
[0103] Specifically, collective communication messages are generally encapsulated using protocol data units (PDUs) for network transmission. Protocol data units usually provide reserved fields for defining the operation type of collective communication messages. Therefore, based on the preset collective communication operation type field, it can be determined whether the message is a reduction operation (reduce).
[0104] Through the above steps, it is possible to quickly determine whether a collective communication message needs to be subjected to protocol operation, without adding additional overhead to the existing network communication protocol, avoiding increasing the complexity of protocol processing, and facilitating overall network maintenance.
[0105] As an optional embodiment, in step S220, performing corresponding protocol operations according to the message includes the following steps:
[0106] Extracting first data of a preset collective communication identification field from the protocol data unit of the message;
[0107] According to the first data, a query is performed in a preset table in the distributed system with the preset set communication identification field as the primary key,
[0108] If no record containing the first data is found in the preset table, the message is stored in the cache of the switching node and moved into the queue to be forwarded, and a record corresponding to the message is added to the preset table, and the record includes address information of the message in the cache;
[0109] If a record containing the first data is found in the preset table, the second data used for the reduction calculation is extracted from other messages corresponding to the record stored in the cache, the second data is subjected to corresponding reduction calculation with the first data used for the reduction calculation of the message, and the second data of other messages corresponding to the record in the cache are updated to the result obtained by the reduction calculation.
[0110] Regarding the preset collective communication identification field, for example, referring to FIG4 , the message header (Header) and data (Data) are parsed from the protocol data unit (PDU) of the collective communication message, and the data content of the Group ID field and the TAG field in the Header are read as the first data. Among them, the standard currently used for collective communication is MPI, and collective communication libraries such as NCCL, MPICH, Open MPI, and Gloo are all specific implementations based on the MPI standard. The MPI standard defines a programming interface for collective communication and defines a series of collective operations, such as MPI_Bcast, MPI_Scatter, MPI_Reduce, MPI_Gather, MPI_Allreduce, and MPI_Allgather. In this embodiment, high-performance optimization of the MPI_Reduce operation can be achieved, including MPI_Allreduce implemented using MPI_Reduce+MPI_Bcast.
[0111] Regarding the preset table, it is used to record messages that require protocol operations, with the preset collective communication identification field as the primary key (Key). For example, referring to Figure 5, the preset table can form the Group ID field and the TAG field into a key. Each record in the table corresponds to a message in the cache. The data part of the record can store a pointer, and the pointer points to the address of the message in the cache.
[0112] It should be noted that the preset collective communication identification field may include but is not limited to the Group ID field and the TAG field. Currently, most collective communications are serial, and the Group ID field and the TAG field are used to maintain the preset table for on-line calculations. Generally, the reduce operations of two previous and subsequent time points will not be merged together. However, with the development of upper-level business and performance requirements, if there are multiple parallel collective communications, it is necessary to introduce more field information as the primary key of the preset table to distinguish multiple collective communications.
[0113] Specifically, when a received message needs to be reduced, it is possible to check in the preset table whether there is a record related to the first data of the message: if not, a cache is allocated to save the message, and the message enters the queue to be forwarded and waits, and a new record is added to the preset table accordingly. The record corresponds to the message and includes the address information of the message in the cache; if so, the second data of the corresponding message used for the reduction calculation is extracted from the cache according to the address information recorded in the record. After the first data and the second data are reduced, the second data of the corresponding message in the cache is updated to the result obtained by the reduction calculation.
[0114] Generally speaking, the protocol data unit (PDU) of an aggregate communication message includes a counter field, such as a Counter field. The value of this field is used to indicate that the PDU is the result of a reduction calculation of several data points in other PDUs. This field is updated accordingly after the reduction operation, for example, by adding 1 to the original value, or adding N to the original value, where N is the value of the Counter field in the other PDUs that the PDU is reducing. For example, in the calculation of the average value (MEAN), the first PDU and the second PDU need to be calculated. The first PDU contains the first data and the first Counter, and the second PDU contains the second data and the second Counter. Then the average value = (first data x first Counter + second data x second Counter) / (first Counter + second Counter).
[0115] Through the above steps, the switching node can perform protocol calculation on the qualified collective communication messages, and the results of the protocol calculation can be updated in the corresponding messages in the cache, so that multiple messages that need to be forwarded become one, thereby reducing the length of the queue to be forwarded at the output port, effectively reducing the number of messages transmitted in the network, alleviating network congestion, and at the same time unloading the computing pressure of the computing device.
[0116] As an optional embodiment, the message processing method further includes the following steps:
[0117] S300 , when the switching node schedules a message from the cache according to the queue to be forwarded, if the message matches a record in the preset table, the switching node deletes the matching record from the preset table after forwarding the message.
[0118] Generally speaking, when the output port of a switching node is performing egress scheduling, if it determines that the forwarding queue of a certain output port is not empty, that is, there are messages waiting to be scheduled, the message at the head of the queue is read from the cache and forwarded. After the forwarding is completed, the cache is recycled, and then the next message is scheduled.
[0119] Therefore, regarding the preset table for recording protocol operation messages in step S220, if the message forwarded by the switching node corresponds to a record in the preset table, the corresponding record needs to be deleted from the preset table after the message forwarding is completed.
[0120] Through the above steps, the forwarded messages can be deleted in the preset table in time, and the message records in the preset table that need to be operated by protocol can be updated in real time, which is conducive to the maintenance of data tables and the normal operation of the system, and ensures the accuracy of online calculations.
[0121] As a further improvement of the above embodiment, the preset collective communication identification field includes a group identification number field, and the group identification number field is globally unique and is used to group the computing devices.
[0122] Specifically, the group identification number field can adopt a predefined field in the collective communication message. For example, the group identification number field can be the Group ID field of the MPI standard. The group identification number field (for example, the Group ID field) needs to remain globally unique in a distributed system to implement effective group identification and management of all computing devices (for example, AI accelerators). It should be noted that the globally unique group identification number field is allowed to change, but must remain globally unique within a predefined time range.
[0123] As a further improvement of the above embodiment, the preset collective communication identification field further includes a message tag field for marking the protocol operation of the collective communication.
[0124] Specifically, the message tag field may adopt a predefined field in the aggregate communication message. For example, the message tag field may be a TAG field of the MPI standard.
[0125] Regarding the TAG field, generally speaking, in point-to-point communication, for example, A and B conduct point-to-point communication, and A needs to receive a bunch of data from B. In order to ensure that the data received from B has a place to be stored, A can reserve a location in the memory in advance and associate a TAG. When A sends a request to B, it will carry this TAG, and then B will also carry this TAG when returning data to A. After receiving the data, A can store the data according to the reserved location based on this TAG. In multi-point collective communication, the TAG field can be used to mark a collective communication operation. For example, there are four nodes A / B / C / D. A initiates a Reduce operation and first allocates a piece of memory locally to receive data and associates a TAG with it. Then, A sends a Reduce request to B / C / D, carrying this TAG. When B / C / D returns their respective data to A, they also carry this TAG. When A needs to initiate another Reduce operation, it can allocate additional memory and associate another TAG. Therefore, A sends another Reduce request to B / C / D, carrying this other TAG. When B / C / D returns their respective data to A, they also carry this other TAG. In this way, the two Reduce operations can be determined based on the Group ID field + TAG field.
[0126] In addition, the preset collective communication identification field supports the extension field in the collective communication PDU, which is used to enhance the existing Reduce operation or add a new type of Reduce operation to meet various network communication transmission and business development needs.
[0127] Through the above steps, multiple computing devices can be effectively grouped, so that the switching node can perform corresponding protocol operations on the messages transmitted by each computing device based on the preset collective communication identification field, thereby ensuring the accuracy and timeliness of the switching node when performing on-line computing.
[0128] As an optional embodiment, the reduction operation includes finding the maximum value, finding the minimum value, finding the sum, finding the product, and finding the average value.
[0129] Specifically, the reduction operation supports all Reduce operations defined by the collective communication specification, including but not limited to maximum value (MAX), minimum value (MIN), sum (SUM), product (PROD), average value (MEAN), etc.
[0130] Therefore, by utilizing the on-line computing power of the switching node, multiple related messages can be turned into one message after the corresponding protocol calculation, thereby reducing the length of the queue to be forwarded at the output port, effectively reducing the number of messages transmitted within the network, alleviating network congestion, and at the same time unloading the computing pressure of the computing equipment.
[0131] Referring to FIG6 , a message processing system is used in a switching node of a distributed system. The system includes:
[0132] a receiving unit, configured to receive a message from a computing device in the distributed system and determine an output port of the switching node to which the message is to be forwarded;
[0133] The processing unit is configured to determine whether the output port is in a congested state and process the message according to the congestion state result obtained by the determination, specifically:
[0134] When the queue to be forwarded of the output port is empty, the message is directly forwarded from the output port;
[0135] When the queue to be forwarded of the output port is not empty, the message is analyzed, a processing strategy for the message is determined based on the analysis result, and the message is selectively processed according to the processing strategy.
[0136] Therefore, by arranging the receiving unit and processing unit on the switching node of the distributed system, when the output port of the switching node is congested, the message is analyzed and the processing strategy is determined, and the message is selectively processed according to the processing strategy. For example, the corresponding protocol operation is performed on the message based on the analysis of the message, and the message that does not need to be simplified enters the cache queue and waits for processing, thereby realizing on-line computing at the switching node, effectively reducing the number of messages transmitted in the network, alleviating network congestion, and at the same time unloading the computing pressure of the computing device.
[0137] As an optional embodiment, the receiving unit includes a first module, configured to determine the output port according to a routing table of the switching node.
[0138] Therefore, after receiving a message from the computing device, the switching node obtains the forwarding path of the message by searching the routing table, and thereby determines the corresponding output port on the switching node.
[0139] As a further improvement of the above embodiment, the processing unit includes a second module for determining whether the message is a collective communication message,
[0140] If the message is a collective communication message, the message is parsed and processed according to the parsing results, including:
[0141] According to the parsing result, determine whether the message is a message that needs to be regulated.
[0142] If so, perform the corresponding protocol operation according to the message.
[0143] If not, storing the message in the cache of the switching node and moving the message into the queue to be forwarded of the output port;
[0144] If the message is not a collective communication message, the message is stored in the cache of the switching node, and the message is moved into the queue to be forwarded of the output port.
[0145] As a result, when there are multiple messages waiting to be processed in a queue at the same output port, the switching node can perform protocol calculations on the messages that meet the protocol operation requirements. After the corresponding protocol calculations, multiple related messages can become one message, thereby reducing the length of the queue to be forwarded at the output port, effectively reducing the number of messages transmitted within the network, alleviating network congestion, and at the same time offloading the computing pressure of the computing device. There is no need to add additional network communication protocols, there is no restriction on the network type of specific application, and it has better applicability.
[0146] As a further improvement of the above embodiment, the second module is used to determine whether the message is a collective communication message according to a preset message identification condition.
[0147] As a result, the switching node can quickly analyze and determine whether the message transmitted in the distributed system is a collective communication message, and perform corresponding processing based on the analysis results, which helps improve the efficiency of network message forwarding.
[0148] As an optional embodiment, the second module is used to parse the protocol data unit in the collective communication message, and determine whether the message needs to be standardized based on the preset collective communication operation type field obtained by the parsing.
[0149] As a result, the switching node can quickly determine whether a protocol operation is required for the collective communication message, without adding additional overhead to the existing network communication protocol, avoiding increasing the complexity of protocol processing, and facilitating overall network maintenance.
[0150] As an optional embodiment, the processing unit includes a third module, configured to extract first data of a preset aggregate communication identification field from the protocol data unit of the message;
[0151] and used to query a preset table in the distributed system with the preset set communication identification field as the primary key according to the first data,
[0152] If no record containing the first data is found in the preset table, the message is stored in the cache of the switching node and moved into the queue to be forwarded, and a record corresponding to the message is added to the preset table, and the record includes address information of the message in the cache;
[0153] If a record containing the first data is found in the preset table, the second data used for the reduction calculation is extracted from other messages corresponding to the record in the cache, the second data is subjected to corresponding reduction calculation with the first data used for the reduction calculation of the message, and the second data of other messages corresponding to the record in the cache are updated to the result obtained by the reduction calculation.
[0154] Therefore, the switching node performs protocol calculation on the collective communication messages that meet the conditions, and updates the results of the protocol calculation in the corresponding messages in the cache, so that multiple messages that need to be forwarded become one, thereby reducing the length of the queue to be forwarded at the output port, which can effectively reduce the number of messages transmitted in the network, alleviate network congestion, and at the same time unload the computing pressure of the computing equipment.
[0155] As an optional embodiment, the system further includes:
[0156] The scheduling unit is used to schedule the messages in the queue to be forwarded in sequence, and if the message matches the record in the preset table, delete the matched record from the preset table after forwarding the message.
[0157] In this way, forwarded messages can be deleted in the preset table in time, and message records in the preset table that need to be operated by protocol can be updated in real time, which is beneficial to the maintenance of data tables and the normal operation of the system, and ensures the accuracy of online calculations.
[0158] As an optional embodiment, the preset collective communication identification field includes a group identification number field, and the group identification number field is globally unique and is used to group the computing devices.
[0159] As a further improvement of the above embodiment, the preset collective communication identification field further includes a message tag field for marking the protocol operation of the collective communication.
[0160] The group identification number field and the message tag field may respectively adopt predefined fields in the collective communication message. For example, the message tag field may be the Group ID field and the TAG field of the MPI standard.
[0161] In this way, multiple computing devices can be effectively grouped, and multiple protocol operations of collective communications can be identified, so that the switching node can perform corresponding protocol operations on the messages transmitted by each computing device based on the preset collective communication identification field, thereby ensuring the accuracy and timeliness of the switching node when performing on-line calculations.
[0162] As an optional embodiment, the reduction operation includes finding the maximum value, finding the minimum value, finding the sum, finding the product, and finding the average value.
[0163] Therefore, by utilizing the on-line computing power of the switching node, multiple related messages can be turned into one message after the corresponding protocol calculation, thereby reducing the length of the queue to be forwarded at the output port, effectively reducing the number of messages transmitted within the network, alleviating network congestion, and at the same time unloading the computing pressure of the computing equipment.
[0164] The present application also provides a distributed system, including several switching nodes and several computing devices, wherein the switching nodes can implement the message processing method described in the above embodiment.
[0165] As an optional embodiment, the plurality of computing devices are grouped according to a preset group identification number, and the group identification number is globally unique.
[0166] Specifically, for example, the group identification number field can be the Group ID field of the MPI standard, which ensures global uniqueness. The group identification number field must remain globally unique in a distributed system to enable effective group identification and management of all computing devices. It should be noted that the globally unique group identification number field is allowed to change but must remain globally unique within a predefined timeframe.
[0167] The present application also provides an electronic device, comprising: a processor; and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the message processing method as described in the above embodiment can be implemented.
[0168] The present application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is used to be run to implement the message processing method as described in the above embodiment.
[0169] The distributed system, electronic device and computer-readable storage medium of the present application can reduce the number of messages transmitted within the network and alleviate network congestion by adopting the above-mentioned method, while at the same time unloading the computing pressure of the computing device.
[0170] The above description is merely a preferred embodiment of the present application and does not constitute any form of limitation to the present application. Although the present application has been disclosed as a preferred embodiment, it is not intended to limit the present application. Any technician familiar with the present profession can make slight changes or modifications to equivalent embodiments of the technical contents disclosed above without departing from the scope of the technical solution of the present application. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.
Claims
1. A message processing method for a distributed system, characterized in that: The method comprises the following steps: The switching node of the distributed system receives a message from a computing device in the distributed system, and determines an output port of the switching node to which the message needs to be forwarded; The switching node determines whether the output port is in a congested state, and processes the message according to the congested state result obtained by the determination, specifically: When the queue to be forwarded of the output port is empty, the message is directly forwarded from the output port; When the queue to be forwarded of the output port is not empty, the message is analyzed, a processing strategy for the message is determined based on the analysis result, and the message is selectively processed according to the processing strategy.
2. The method according to claim 1, characterized in that The step of determining the output port of the switching node to which the message needs to be forwarded comprises the following steps: The output port is determined according to a routing table of the switching node.
3. The method according to claim 1, characterized in that The step of analyzing the message, determining a processing strategy for the message based on the analysis result, and selectively processing the message according to the processing strategy comprises the following steps: If the message is a collective communication message, the message is parsed and processed according to the parsing result, including: According to the parsing result, determine whether the message is a message that needs to be regulated. If so, perform the corresponding protocol operation according to the message. If not, storing the message in the cache of the switching node and moving the message into the queue to be forwarded of the output port; If the message is not a collective communication message, the message is stored in the cache of the switching node, and the message is moved into the queue to be forwarded of the output port.
4. The method according to claim 3, characterized in that The analysis of the message comprises the following steps: Determine whether the message is a collective communication message according to a preset message identification condition.
5. The method according to claim 3, characterized in that If the message is a collective communication message, the message is parsed, and the message is processed according to the parsing result, including the following steps: The protocol data unit in the message is parsed, and it is determined whether the message is a message that needs to be standardized according to the preset collective communication operation type field obtained by the parsing.
6. The method according to claim 3, characterized in that The corresponding protocol operation is performed according to the message, comprising the following steps: Extracting first data of a preset collective communication identification field from the protocol data unit of the message; According to the first data, a query is performed in a preset table in the distributed system with the preset set communication identification field as the primary key, If no record containing the first data is found in the preset table, the message is stored in the cache of the switching node and the message is moved into the queue to be forwarded, and a record corresponding to the message is added to the preset table, and the record includes address information of the message in the cache; If a record containing the first data is found in the preset table, second data used for reduction calculation is extracted from other messages corresponding to the record stored in the cache, and the second data is subjected to corresponding reduction calculation with the first data used for reduction calculation of the message, and the second data of other messages corresponding to the record in the cache are updated to the result obtained by the reduction calculation.
7. The method according to claim 6, characterized in that The following steps are also included: When the switching node schedules a message from the cache according to the queue to be forwarded, if the message matches a record in the preset table, the matching record is deleted from the preset table after the message is forwarded.
8. The method according to claim 6, characterized in that The preset set communication identification field includes a group identification number field, and the group identification number field is globally unique and is used to group the computing devices.
9. The method according to claim 8, characterized in that The preset collective communication identification field also includes a message tag field, which is used to mark the protocol operation of the collective communication.
10. The method according to claim 3, characterized in that The reduction operations include finding the maximum value, finding the minimum value, finding the sum, finding the product, and finding the average value.
11. A message processing system, used in a switching node of a distributed system, characterized in that: The system comprises: A receiving unit, configured to receive a message from a computing device in the distributed system, and determine an output port of the switching node to which the message needs to be forwarded; The processing unit is used to determine whether the output port is in a congested state, and process the message according to the congested state result obtained by the determination, specifically: When the queue to be forwarded of the output port is empty, the message is directly forwarded from the output port; When the queue to be forwarded of the output port is not empty, the message is analyzed, a processing strategy for the message is determined based on the analysis result, and the message is selectively processed according to the processing strategy.
12. The system according to claim 11, characterized in that The receiving unit includes a first module, which is used to determine the output port according to the routing table of the switching node.
13. The system according to claim 11, characterized in that The processing unit includes a second module, which is used to determine whether the message is a collective communication message. If the message is a collective communication message, the message is parsed and processed according to the parsing result, including: According to the parsing result, determine whether the message is a message that needs to be regulated. If so, perform the corresponding protocol operation according to the message. If not, storing the message in the cache of the switching node and moving the message into the queue to be forwarded of the output port; If the message is not a collective communication message, the message is stored in the cache of the switching node, and the message is moved into the queue to be forwarded of the output port.
14. The system of claim 13, wherein: The second module is used to determine whether the message is a collective communication message according to a preset message identification condition.
15. The system of claim 13, wherein: The second module is used to parse the protocol data unit in the collective communication message, and determine whether the message is a message that needs to be reduced based on the preset collective communication operation type field obtained by the parsing.
16. The system of claim 13, wherein: The processing unit includes a third module, which is used to extract the first data of the preset collective communication identification field from the protocol data unit of the message; and used to query a preset table in the distributed system with the preset set communication identification field as a primary key according to the first data, If no record containing the first data is found in the preset table, the message is stored in the cache of the switching node and the message is moved into the queue to be forwarded, and a record corresponding to the message is added to the preset table, and the record includes address information of the message in the cache; If a record containing the first data is found in the preset table, second data used for reduction calculation is extracted from other messages corresponding to the record in the cache, and corresponding reduction calculation is performed on the second data and the first data used for reduction calculation of the message, and the second data of other messages corresponding to the record in the cache are updated to the results obtained by the reduction calculation.
17. The system of claim 16, wherein: The system further comprises: The scheduling unit is used to schedule the messages in the queue to be forwarded in sequence, and if the message matches the record in the preset table, delete the matched record from the preset table after forwarding the message.
18. The system of claim 16, wherein: The preset set communication identification field includes a group identification number field, and the group identification number field is globally unique and is used to group the computing devices.
19. The system of claim 18, wherein: The preset collective communication identification field also includes a message tag field, which is used to mark the protocol operation of the collective communication.
20. The system of claim 13, wherein: The reduction operations include finding the maximum value, finding the minimum value, finding the sum, finding the product, and finding the average value.
21. A distributed system, characterized in that: The method comprises several switching nodes and several computing devices, wherein the switching nodes can implement the method according to any one of claims 1 to 10.
22. The system of claim 21, wherein: The computing devices are grouped according to a preset group identification number, and the group identification number is globally unique.
23. An electronic device, characterized in that: include: processor; as well as A memory having a computer program stored thereon, wherein the computer program, when executed by the processor, can implement the method according to any one of claims 1 to 10.
24. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program is used to run to implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Message processing method and equipment
CN102325092A
Forwarding queue flushing method and system of switch chip
CN114221916A
Data stream transmission control method based on congestion feedback in data center lossless network
CN114938350A
Fast message forwarding method, network device, storage medium and computer program product
CN115967687A
Method for congestion control in a network
WO2018225039A1