Message processing method and system
By analyzing congestion at switch nodes and performing selective processing of collective communication messages, the method optimizes network efficiency and reduces computational load in distributed AI training systems, addressing inefficiencies in existing protocols like SHARP.
Patent Information
- Application Number
- CN202411070108.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-05
- Publication Date
- 2025-07-15
AI Technical Summary
The existing distributed systems have problems in collective communications, such as dependence on physical network topology, high protocol complexity, large maintenance overhead and network congestion, resulting in limited improvement in AI training performance.
By judging the congestion status of the output port on the switching node, selective processing of the collective communication messages, including parsing and regulation operations, reducing the number of messages transmitted in the network, alleviating network congestion and unloading the computing pressure of computing devices.
Effectively reduce the number of messages transmitted within the network, alleviate network congestion, optimize the efficiency of distributed systems, is suitable for various network types and topology, and reduces the computing pressure of computing devices.
Smart Images

Figure CN120321183A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a distributed system, and in particular, to a method, a system, an electronic device, and a medium for message processing applicable to artificial intelligence model training. Background Art
[0002] A trend in AI (Artificial Intelligence) applications is that the scale of AI models is getting larger and the AI training data is getting larger. For example, for ChatGPT and GPT-3, the scale of model parameters is 175 billion, and the text-based training data is 570 GB. However, traditional single computing devices can no longer support such large-scale model training. Usually, a distributed system including multiple computing devices and switching nodes is adopted to achieve parallel computing training to improve the overall AI training performance. With such a distributed parallel computing training method, the computing devices satisfy various communication requirements, such as transmitting training data and updating model parameters, through the switching nodes. The communication between multiple computing devices in the distributed system for parallel computing training is different from point-to-point communication. There are often a large number of one-to-many or many-to-many communication requirements between computing devices, which is usually called collective communication. Among them, Reduce / Allreduce is a commonly used collective communication in parallel computing training, and the improvement of Reduce / Allreduce performance can significantly improve the overall performance of AI training.
[0003] Some existing solutions propose in-network computing, which offloads collective communication operations to switching nodes (such as switch networks), thereby reducing the amount of data transmitted over the network and shortening the message passing operation time. For example, SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) is a collective communication network offloading technology. See Figure 1, before SHARP performs collective communication, it is necessary to construct a tree-shaped logical network topology based on the physical network topology, which is divided into aggregation nodes (usually network devices) and terminal nodes (usually computing devices). The terminal nodes send data to their affiliated aggregation nodes; after the aggregation nodes collect the data of their child nodes, they perform in-network computing and send the computing results to the upper-layer aggregation nodes; after the upper-layer aggregation nodes collect the data of their child nodes, they perform in-network computing and send the computing results to the upper-layer aggregation nodes; and so on, until the root aggregation node; after the root aggregation node completes in-network computing, it sends the computing results to all its direct child nodes; after the direct child nodes receive the computing results, they send them to the direct child nodes of the next layer; and so on, until the final results are sent to the terminal nodes. However, the disadvantages of such solutions are as follows:
[0004] 1) It depends on the physical network topology. It must be a network topology that can form a tree-shaped aggregation, and a ring network is not suitable for deploying SHARP;
[0005] 2) Introduce an additional SHARP protocol, increasing the complexity of protocol processing;
[0006] 3) It is necessary to establish a logical communication tree in advance, introducing additional overhead;
[0007] 4) It is necessary to maintain the logical communication tree. For example, when a node fails, it is necessary to re-establish the logical communication tree, that is, introduce additional maintenance overhead, and at the same time, it also increases the maintenance complexity of the entire network.
[0008] Therefore, it is necessary to provide a method or system that can optimize the efficiency of collective communication, which can be used in distributed systems to improve the overall performance of AI training. Summary of the Invention
[0009] Based on the above situation, the main object of the present invention is to provide a method, system, electronic device and medium for message processing. Based on the congestion state of the output port of the switching node, corresponding reduction operations are performed on the collective communication messages that need to be reduced, which helps to reduce the length of the forwarding queue of the output port, can effectively reduce the number of messages transmitted in the network, relieve network congestion, and at the same time can offload the computing pressure of the computing device.
[0010] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0011] The first aspect of the present invention provides a method for message processing, which is used in a distributed system. The method includes the following steps:
[0012] The switching node of the distributed system receives a message from a computing device in the distributed system and determines the output port of the switching node that needs to forward the message;
[0013] The switching node determines whether the output port is in a congested state, and processes the packet according to the obtained congestion state result. Specifically:
[0014] When the forwarding queue of the output port is empty, the packet is directly forwarded from the output port;
[0015] When the forwarding queue of the output port is not empty, the packet is analyzed, a processing strategy for the packet is determined based on the analysis result, and the packet is selectively processed according to the processing strategy.
[0016] Preferably, determining the output port of the switching node that needs to forward the packet includes the following steps:
[0017] Determine the output port according to the routing table of the switching node.
[0018] Preferably, analyzing the packet, determining a processing strategy for the packet based on the analysis result, and selectively processing the packet according to the processing strategy includes the following steps:
[0019] If the packet is a collective communication packet, the packet is parsed and processed according to the parsing result, including:
[0020] Judge whether the packet is a packet that needs to be reduced according to the parsing result,
[0021] If so, perform corresponding reduction operations according to the packet,
[0022] If not, store the packet in the cache of the switching node and move the packet into the forwarding queue of the output port;
[0023] If the packet is not a collective communication packet, store the packet in the cache of the switching node and move the packet into the forwarding queue of the output port.
[0024] Preferably, analyzing the packet includes the following steps:
[0025] Determine whether the packet is a collective communication packet according to preset packet recognition conditions.
[0026] Preferably, if the packet is a collective communication packet, the packet is parsed and processed according to the parsing result, including the following steps:
[0027] Parse the protocol data unit in the packet, and determine whether the packet is a packet that needs to be reduced according to the preset collective communication operation type field obtained by parsing.
[0028] Preferably, performing corresponding protocol operations according to the message includes the following steps:
[0029] Extract the first data of the preset set communication identification field from the protocol data unit of the message;
[0030] Query according to the first data in the preset table with the preset set communication identification field as the primary key in the distributed system,
[0031] If no record containing the first data is found in the preset table, store the message in the cache of the switching node, move the message into the to-be-forwarded queue, and add a record corresponding to the message in the preset table, and the record includes the address information of the message in the cache;
[0032] If a record containing the first data is found in the preset table, extract the second data for protocol calculation from other messages corresponding to the record stored in the cache, perform corresponding protocol calculation on the second data and the first data for protocol calculation of the message, and update the second data of other messages corresponding to the record in the cache to the result obtained by the protocol calculation.
[0033] Preferably, it further includes the following steps:
[0034] When the switching node schedules messages from the cache according to the to-be-forwarded queue, if the message matches the record in the preset table, delete the matching record from the preset table after forwarding the message.
[0035] Preferably, the preset set communication identification field includes a group identification number field, and the group identification number field is globally unique and is used for grouping and identifying the computing devices.
[0036] Preferably, the preset set communication identification field further includes a message tag field, which is used to mark the protocol operation of the collective communication.
[0037] Preferably, the protocol operations include finding the maximum value, finding the minimum value, summing, multiplying, and finding the average value.
[0038] The second aspect of the present invention provides a message processing system for a switching node of a distributed system, and the system includes:
[0039] A receiving unit, configured to receive a message from a computing device in the distributed system and determine an output port of the switching node that needs to forward the message;
[0040] A processing unit, configured to determine whether the output port is congested and process the message according to the determined congestion status result, specifically:
[0041] When the forwarding queue of the output port is empty, forward the packet directly from the output port;
[0042] When the forwarding queue of the output port is not empty, analyze the packet, determine the processing strategy for the packet based on the analysis result, and selectively process the packet according to the processing strategy.
[0043] Preferably, the receiving unit includes a first module for determining the output port according to the routing table of the switching node.
[0044] Preferably, the processing unit includes a second module for determining whether the packet is a collective communication packet.
[0045] If the packet is a collective communication packet, parse the packet and process the packet according to the parsing result, including:
[0046] Judge whether the packet is a packet that needs to be reduced according to the parsing result.
[0047] If so, perform corresponding reduction operations according to the packet.
[0048] If not, store the packet in the cache of the switching node and move the packet into the forwarding queue of the output port.
[0049] If the packet is not a collective communication packet, store the packet in the cache of the switching node and move the packet into the forwarding queue of the output port.
[0050] Preferably, the second module is used to determine whether the packet is a collective communication packet according to preset packet identification conditions.
[0051] Preferably, the second module is used to parse the protocol data unit in the collective communication packet and determine whether the packet is a packet that needs to be reduced according to the preset collective communication operation type field obtained by parsing.
[0052] Preferably, the processing unit includes a third module for extracting first data of a preset collective communication identification field from the protocol data unit of the packet.
[0053] And it is used to query in a preset table of the distributed system with the preset collective communication identification field as the primary key according to the first data.
[0054] If no record containing the first data is found in the preset table, store the packet in the cache of the switching node, move the packet into the to-be-forwarded queue, add a record corresponding to the packet in the preset table, and the record includes the address information of the packet in the cache;
[0055] If a record containing the first data is found in the preset table, extract second data for reduction calculation from other packets corresponding to the record in the cache, perform corresponding reduction calculation on the second data of the packet and the first data of the packet for reduction calculation, and update the second data of other packets corresponding to the record in the cache to the result obtained by the reduction calculation.
[0056] Preferably, the system further includes:
[0057] A scheduling unit, configured to schedule the packets in the to-be-forwarded queue in sequence. If the packet matches the record in the preset table, delete the matched record from the preset table after forwarding the packet.
[0058] Preferably, the preset set communication identification field includes a group identification number field, and the group identification number field is globally unique and is used for group identification of the computing devices.
[0059] Preferably, the preset set communication identification field further includes a message tag field, which is used to mark the reduction operation of the collective communication.
[0060] Preferably, the reduction operation includes finding the maximum value, finding the minimum value, summing, multiplying, and finding the average value.
[0061] A third aspect of the present invention provides a distributed system, including a plurality of switching nodes and a plurality of computing devices, and the switching nodes can implement the method described in the first aspect above.
[0062] Preferably, the plurality of computing devices are grouped according to a preset group identification number, and the group identification number is globally unique.
[0063] A fourth aspect of the present invention provides an electronic device, including: a processor; and a memory, on which a computer program is stored, and when the computer program is executed by the processor, it can implement the method described in the first aspect above.
[0064] A fifth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to run to implement the method described in the first aspect above.
[0065] Compared with the prior art, the present invention has obvious advantages and beneficial effects, and it has at least the following advantages:
[0066] The message processing method of the present invention is based on judging the congestion state of the switching node of the distributed system, and performing corresponding processing on the message to be forwarded according to the judgment result, wherein when a certain output port is in a congested state, the message is analyzed and a processing strategy is determined, and the message is selectively processed according to the processing strategy, for example, a protocol operation is performed on the collective communication message that needs to be reduced, so that when there are multiple messages queued for processing at the same output port of the switching node, the message that meets the protocol operation requirements is reduced. After the corresponding protocol calculation, multiple related messages can become one message, thereby reducing the length of the queue to be forwarded at the output port, and can effectively reduce the number of messages transmitted in the network, alleviate network congestion, and at the same time, the switching node can also unload part of the computing pressure of the computing device, which helps to optimize the efficiency of the distributed system. In addition, there is no restriction on the specific transport layer protocol, and there is no need to add additional network communication protocols. There is no restriction on the network type or topology of the specific application, and the applicability is better.
[0067] The message processing system of the present invention analyzes the message and determines the processing strategy when the output port of the switching node is congested, through the receiving unit and the processing unit arranged on the switching node of the distributed system, and selectively processes the message according to the processing strategy, for example, performing corresponding protocol operations on the message based on the analysis of the message, and the message that does not need to be reduced enters the cache queue to wait for processing, thereby realizing on-network computing at the switching node, effectively reducing the number of messages transmitted in the network, alleviating network congestion, and at the same time unloading the computing pressure of the computing device.
[0068] The distributed system, electronic device and computer-readable storage medium of the present invention can reduce the number of messages transmitted in the network and alleviate network congestion by adopting the above method, while unloading the computing pressure of the computing device. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 A schematic diagram of a logical network topology of SHARP in the prior art;
[0070] Figure 2 A flowchart of a preferred implementation of the message processing method of the present invention;
[0071] 3a to 3c are schematic diagrams of message forwarding of computing devices and switching nodes of the distributed system of the present invention;
[0072] Figure 4 An example of the format of the collective communication message of the present invention;
[0073] Figure 5Schematic diagram of a preset table and cache of the distributed system of the present invention;
[0074] Figure 6 Module schematic diagram of a preferred embodiment of the message processing system of the present invention. Detailed implementation manner
[0075] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following describes in detail the specific implementation manners, methods, steps, features and effects of the methods, systems, electronic devices and computer-readable storage media proposed according to the present invention in conjunction with the accompanying drawings and preferred embodiments.
[0076] See Figure 2 , a method for message processing, used in a distributed system, the method includes the following steps:
[0077] S100, the switching node of the distributed system receives a message from a computing device in the distributed system, and determines the output port of the switching node that needs to forward the message;
[0078] S200, the switching node determines whether the output port is congested, and processes the message according to the determined congestion status result, specifically:
[0079] S210, when the forwarding queue of the output port is empty, directly forward the message from the output port;
[0080] S220, when the forwarding queue of the output port is not empty, analyze the message, determine a processing strategy for the message based on the analysis result, and selectively process the message according to the processing strategy.
[0081] Among them, for a distributed system, it is usually a system composed of a group of multiple computing devices that communicate through a network and coordinate to complete a common task. For example, a distributed system for AI model training, which includes multiple computing devices and multiple switching nodes. Computing devices are mainly used to complete computing tasks. For example, they can be AI accelerators, while switching nodes are mainly used for communication between multiple computing devices. For example, they can be switching chips, switches, routing chips, routers, etc.
[0082] Regarding step S210, in this embodiment, the message is parsed and corresponding processing (such as protocol operation) is performed only when the output port is in a congested state (the waiting forwarding queue is not empty) to relieve the network congestion. When there is no message waiting to be processed in the output port, the message is directly forwarded without introducing additional waiting delay. Even if the message is a protocol operation, the message forwarding is preferably completed as much as possible when the network is unobstructed, which helps to improve the problem of communication network delay.
[0083] Regarding collective communication, in a distributed system, there are often a large number of collective communication requirements among various computing devices. When performing distributed AI model training, some relatively low-level message communication behaviors can usually be defined based on MPI (Message Passing Interface), such as Reduce, Allreduce, Scatter, Gather, Allgather, Broadcast, etc. Among them, Reduce is reduction or reduction calculation, which is a collective term for a series of arithmetic operations, generally including finding the maximum value (MAX), minimum value (MIN), sum (SUM), product (PROD), average value (MEAN), etc.; Allreduce applies the same reduce operation to all node processes and can be completed through the reduce + broadcast operation on a single node.
[0084] As a further improvement of the above embodiment, in step S220, the analyzing the message, determining a processing strategy for the message based on the analysis result, and selectively processing the message according to the processing strategy include the following steps:
[0085] If the message is a collective communication message, the message is parsed and processed according to the parsing result, including:
[0086] Judging whether the message is a message that needs to be reduced according to the parsing result,
[0087] If so, perform corresponding reduction operations according to the message,
[0088] If not, store the message in the cache of the switching node and move the message into the waiting forwarding queue of the output port;
[0089] If the message is not a collective communication message, store the message in the cache of the switching node and move the message into the waiting forwarding queue of the output port.
[0090] For example, referring to Figure 3a, AI accelerators 0 to 4 are connected through a switch as computing devices. When performing a reduction operation, AI accelerators 1 to 4 need to encapsulate their respective data into packets and transmit them to AI accelerator 0, which may cause congestion on the switch. Referring to Figure 3b, there will be 3 packets queuing in the forwarding queue of the output port of the switch waiting to be processed. In this embodiment, based on the congestion state of the switching node, by parsing the collective communication packets, the packets that need to be reduced are identified, and the switching node performs reduction calculations and updates the corresponding packet data in the cache with the obtained reduction calculation results, so that multiple related packets can be converted into a packet that needs to be forwarded. Referring to Figure 3c, packets 2 to 4 in the original forwarding queue are converted into packet 2' by the switching node after reduction calculations. It can be seen that by performing in-network calculations on the packets by the switching node, the length of the forwarding queue can be reduced, effectively alleviating network congestion.
[0091] In addition, after the in-network calculation processing by the switching node, AI accelerator 0 also does not need to perform calculation processing on packets 2 to 4, and only needs to perform corresponding reduction operations after receiving packet 2', so that the calculation pressure on AI accelerator 0 is unloaded. Generally speaking, if there are N computing devices in the system, according to the existing reduction calculation method, N - 1 times or Log2(N) times of calculations are required, while after adopting the method of this embodiment, the optimal number of reduction calculations can be reduced to 1 time.
[0092] Through the above steps, when there are multiple packets queuing in the same output port of the switching node waiting to be processed, reduction calculations are performed on the packets that meet the requirements of the reduction operation. After the corresponding reduction calculations, multiple related packets can become one packet, thereby reducing the length of the forwarding queue at the output port, effectively reducing the number of packets transmitted in the network, alleviating network congestion, while also unloading the calculation pressure of the computing device, and there is no need to add an additional network communication protocol, and there is no restriction on the network type of the specific application, so the applicability is better.
[0093] As a further improvement of the above embodiment, in step S220, the analysis of the packet includes the following steps:
[0094] Determine whether the packet is a collective communication packet according to the preset packet identification condition.
[0095] Specifically, collective communication is an upper-layer application that can be carried using network protocols such as UDP / IP, RDMA, and InfiniBand. Taking the case where collective communication is carried using the UDP / IP protocol as an example, compared to some well-known applications, such as the well-known port numbers 20 and 21 for FTP and the well-known port number 80 for HTTP, collective communication does not have a well-known port number when carried using the UDP / IP protocol. Instead, it depends on the specific implementation of the collective communication library on the computing nodes. Moreover, since the computing cluster network for the distributed system is a relatively closed network, an arbitrary UDP port can be set to represent collective communication. Usually, a non-well-known UDP port number is used. Of course, a well-known port number can also be selected. Thus, in the distributed system, message recognition conditions can be set to identify collective communication messages. For example,
[0096] It is also possible to consider adding a protection point, and the conditions for identifying collective communication messages can be configured. For example: For the UDP / IP network, the recognition conditions can be set as follows: "IP_Header.protocol == 17 (UDP), UDP_Header.destination_port == xxx", where xxx is the preset UDP port number for collective communication.
[0097] Through the above steps, it is possible to quickly analyze and determine whether the message transmitted in the distributed system is a collective communication message, and perform corresponding processing according to the analysis result, which helps improve the network message forwarding efficiency.
[0098] As an optional embodiment, in step S100, determining the output port of the switching node that needs to forward the message includes the following steps:
[0099] S110, determining the output port according to the routing table of the switching node.
[0100] Specifically, after receiving the message from the computing device, the switching node obtains the forwarding path of the message by looking up the routing table, and thereby determines the corresponding output port on the switching node.
[0101] Through the above steps, the switching node determines the output port for forwarding the message according to the received message, which is used to subsequently determine the congestion status of the output port, and determine whether to parse the message and whether to perform a reduction operation according to the determined congestion status structure.
[0102] As an optional embodiment, in step S220, if the message is a collective communication message, then parse the message and process the message according to the parsing result, including the following steps:
[0103] The protocol data unit in the message is parsed, and it is determined whether the message is a message that needs to be standardized according to the preset collective communication operation type field obtained by the parsing.
[0104] Specifically, collective communication messages are generally encapsulated in protocol data units (PDUs) for network transmission. The protocol data units usually provide reserved fields for defining the operation type of the collective communication messages. Thus, whether the message is a reduction operation (reduce) can be determined based on the preset collective communication operation type field.
[0105] Through the above steps, it is possible to quickly determine whether a protocol operation is required for the collective communication message, without adding additional overhead to the existing network communication protocol, avoiding increasing the complexity of protocol processing, and facilitating overall network maintenance.
[0106] As an optional embodiment, in step S220, performing a corresponding protocol operation according to the message includes the following steps:
[0107] Extracting first data of a preset collective communication identification field from the protocol data unit of the message;
[0108] According to the first data, a query is performed in a preset table in the distributed system with the preset set communication identification field as the primary key,
[0109] If no record containing the first data is found in the preset table, the message is stored in the cache of the switching node and the message is moved into the queue to be forwarded, and a record corresponding to the message is added to the preset table, and the record includes address information of the message in the cache;
[0110] If a record containing the first data is found in the preset table, second data used for reduction calculation is extracted from other messages corresponding to the record stored in the cache, and the second data is subjected to corresponding reduction calculation with the first data used for reduction calculation of the message, and the second data of other messages corresponding to the record in the cache are updated to the result obtained by the reduction calculation.
[0111] For the preset aggregate communication identification field, see, for example, Figure 4, the header (Header) and data (Data) are parsed from the protocol data unit (PDU) of the collective communication message, and the data contents of the Group ID field and the TAG field in the Header are read as the first data. Currently, the standard for collective communication is MPI, and collective communication libraries such as NCCL, MPICH, OpenMPI, and Gloo are all specific implementations based on the MPI standard. The MPI standard defines the programming interface for collective communication and defines a series of collective operations. For example, MPI_Bcast, MPI_Scatter, MPI_Reduce, MPI_Gather, MPI_Allreduce, MPI_Allgather, etc. In this embodiment, high-performance optimization of the MPI_Reduce operation can be achieved, including MPI_Allreduce implemented using MPI_Reduce + MPI_Bcast.
[0112] Regarding the preset table, it is used to record the messages that need to be reduced. Using the preset collective communication identification field as the primary key (Key). For example, see Figure 5 , the preset table can form a Key with the Group ID field and the TAG field. Each record in the table corresponds to a message in the cache, and the data part of the record can store a pointer that points to the address of the message in the cache.
[0113] It should be noted that the preset collective communication identification field may include, but is not limited to, the Group ID field and the TAG field. Currently, most collective communications are in serial mode, and the Group ID field and the TAG field are used to maintain the preset table for in-network computing. Generally, the reduce operations at two consecutive time points will not be combined. However, with the development of upper-layer services and performance requirements, if there are multiple parallel collective communications, more field information needs to be introduced as the primary key of the preset table to distinguish multiple collective communications.
[0114] Specifically, when a received message needs to be reduced, it can be queried in the preset table whether there is a record related to the first data of this message: if not, allocate a cache to save this message, and this message enters the waiting-forward queue to queue and wait, and a new record is added accordingly in the preset table. This record corresponds to this message and includes the address information of this message in the cache; if so, extract the second data for reduction calculation of the corresponding message from the cache according to the address information recorded in the record. After the first data and the second data are reduced, the second data of the corresponding message in the cache is updated to the result obtained from the reduction calculation.
[0115] Generally, the protocol data unit (PDU) of a collective communication message includes a count field, such as a Counter field, whose value is used to indicate that the PDU is the result of a reduction calculation of several data of other PDUs. This field is updated accordingly after the reduction operation. For example, add 1 to the original value, or add N to the original value, where N is the value of the Counter field in other PDUs reduced by this PDU. For example, in the reduction calculation of finding the average value (MEAN), it is necessary to calculate the first PDU and the second PDU. The first PDU contains the first data and the first Counter, and the second PDU contains the second data and the second Counter. Then
[0116] Average value = (First data × First Counter + Second data × Second Counter) / (First Counter + Second Counter).
[0117] Through the above steps, the switching node can perform reduction calculations on the eligible collective communication messages, update the results of the reduction calculations in the corresponding messages in the cache, so that multiple messages to be forwarded become one, thereby reducing the length of the forwarding queue of the output port, effectively reducing the number of messages transmitted in the network, alleviating network congestion, and at the same time unloading the computing pressure of the computing device.
[0118] As an optional embodiment, the method for message processing further includes the following steps:
[0119] S300, when the switching node schedules messages from the cache according to the forwarding queue, if the message matches the record in the preset table, delete the matching record from the preset table after forwarding the message.
[0120] Generally speaking, when the output port of the switching node performs egress scheduling, if it is determined that the forwarding queue of a certain output port is not empty, that is, there are messages queuing waiting for scheduling, then read the message at the head of the queue from the cache and forward it. After the forwarding is completed, recycle the cache, and then perform the scheduling process of the next message.
[0121] Therefore, regarding the preset table for recording reduction operation messages in step S220, if the message forwarded by the switching node corresponds to the record in the preset table, the corresponding record needs to be deleted from the preset table after the message is forwarded.
[0122] Through the above steps, the forwarded messages can be deleted from the preset table in a timely manner, and the message records that need to perform reduction operations in the preset table can be updated in real time, which is beneficial to the maintenance of the data table and the normal operation of the system, and ensures the accuracy of in-network computing.
[0123] As a further improvement of the above embodiment, the preset set communication identification field includes a group identification number field, and the group identification number field is globally unique and is used to group the computing devices.
[0124] Specifically, the group identification number field may adopt a predefined field in the collective communication message, for example, the group identification number field may be the Group ID field of the MPI standard. The group identification number field (for example, the Group ID field) needs to remain globally unique in a distributed system to implement effective group identification and management of all computing devices (for example, AI accelerators). It should be noted that the globally unique group identification number field is allowed to change, but must remain globally unique within a predefined time range.
[0125] As a further improvement of the above embodiment, the preset collective communication identification field also includes a message tag field for marking the protocol operation of the collective communication.
[0126] Specifically, the message tag field may adopt a predefined field in the aggregate communication message, for example, the message tag field may be a TAG field of the MPI standard.
[0127] Regarding the TAG field, generally speaking, in point-to-point communication, for example, A and B conduct point-to-point communication, and A needs to receive a bunch of data from B. In order to ensure that the data received from B has a place to be stored, A can reserve a location in the memory in advance and associate a TAG. When A sends a request to B, it will carry this TAG, and then when B returns data to A, it will also carry this TAG. After receiving the data, A can store the data according to the reserved location based on this TAG. In multi-point collective communication, the TAG field can be used to mark a collective communication operation. For example, there are four nodes A / B / C / D. A initiates a Reduce operation and first allocates a piece of memory locally to receive data, and associates a TAG with it. Then, A sends a Reduce request to B / C / D, and the request carries this TAG. When B / C / D returns their respective data to A, they also carry this TAG. When A needs to initiate another Reduce operation, it can allocate another memory and associate another TAG. Therefore, A sends another Reduce request to B / C / D, and the request carries this other TAG. When B / C / D returns their respective data to A, they also carry this other TAG. In this way, the two Reduce operations can be determined based on the Group ID field + TAG field.
[0128] In addition, the preset collective communication identification field supports the extension field in the collective communication PDU, which is used to enhance the existing Reduce operation or add a new type of Reduce operation to meet various network communication transmission and business development needs.
[0129] Through the above steps, multiple computing devices can be effectively grouped, so that the switching node can perform corresponding protocol operations on the messages transmitted by each computing device based on the preset collective communication identification field, thereby ensuring the accuracy and timeliness of the switching node when performing on-line computing.
[0130] As an optional embodiment, the reduction operation includes finding the maximum value, finding the minimum value, finding the sum, finding the product, and finding the average value.
[0131] Specifically, the reduction operation supports all Reduce operations defined by the collective communication specification, including but not limited to maximum value (MAX), minimum value (MIN), sum (SUM), product (PROD), average value (MEAN), etc.
[0132] Therefore, by utilizing the on-line computing power of the switching node, multiple related messages can become one message after corresponding protocol calculation, thereby reducing the length of the queue to be forwarded at the output port, effectively reducing the number of messages transmitted in the network, alleviating network congestion, and at the same time unloading the computing pressure of the computing equipment.
[0133] See also Figure 6 , a message processing system, used for a switching node of a distributed system, the system comprising:
[0134] A receiving unit, configured to receive a message from a computing device in the distributed system, and determine an output port of the switching node to which the message needs to be forwarded;
[0135] The processing unit is used to determine whether the output port is in a congested state, and process the message according to the congested state result obtained by the determination, specifically:
[0136] When the queue to be forwarded of the output port is empty, the message is directly forwarded from the output port;
[0137] When the queue to be forwarded of the output port is not empty, the message is analyzed, a processing strategy for the message is determined based on the analysis result, and the message is selectively processed according to the processing strategy.
[0138] Therefore, by arranging the receiving unit and the processing unit on the switching node of the distributed system, when the output port of the switching node is congested, the message is analyzed and the processing strategy is determined, and the message is selectively processed according to the processing strategy. For example, the message is subjected to corresponding protocol operations based on the analysis of the message, and the message that does not need to be subjected to the protocol enters the cache queue and waits for processing, thereby realizing on-network computing at the switching node, effectively reducing the number of messages transmitted in the network, alleviating network congestion, and at the same time unloading the computing pressure of the computing device.
[0139] As an alternative embodiment, the receiving unit includes a first module configured to determine the output port according to the routing table of the switching node.
[0140] Thereby, after receiving a packet from the computing device, the switching node obtains the forwarding path of the packet by looking up the routing table, and thereby determines the corresponding output port on the switching node.
[0141] As a further improvement of the above embodiment, the processing unit includes a second module configured to determine whether the packet is a collective communication packet.
[0142] If the packet is a collective communication packet, then the packet is parsed, and the packet is processed according to the parsing result, including:
[0143] Determine whether the packet is a packet that needs to be reduced according to the parsing result.
[0144] If so, perform a corresponding reduction operation according to the packet.
[0145] If not, store the packet in the cache of the switching node and move the packet into the forwarding queue of the output port.
[0146] If the packet is not a collective communication packet, then store the packet in the cache of the switching node, and move the packet into the forwarding queue of the output port.
[0147] Thereby, when there are multiple packets queuing for processing at the same output port, the switching node can perform reduction calculations on the packets that meet the requirements of the reduction operation. After corresponding reduction calculations, multiple related packets can become one packet, thereby reducing the length of the forwarding queue of the output port, effectively reducing the number of packets transmitted in the network, alleviating network congestion, while also unloading the computing pressure of the computing device, and not requiring additional network communication protocols, and having no restrictions on the network type of specific applications, with better applicability.
[0148] As a further improvement of the above embodiment, the second module is configured to determine whether the packet is a collective communication packet according to preset packet identification conditions.
[0149] Thereby, the switching node can quickly analyze and determine whether the packet transmitted in the distributed system is a collective communication packet, and perform corresponding processing according to the analysis result, which helps to improve the network packet forwarding efficiency.
[0150] As an alternative embodiment, the second module is configured to parse the protocol data unit in the collective communication packet, and determine whether the packet needs to be reduced according to the preset collective communication operation type field obtained by parsing.
[0151] Thus, the switching node can quickly determine whether a collective communication message needs to be reduced, without adding extra overhead to the existing network communication protocol, avoiding increasing the complexity of protocol processing, and being beneficial to the overall network maintenance.
[0152] As an optional embodiment, the processing unit includes a third module configured to extract first data of a preset collective communication identification field from a protocol data unit of the message;
[0153] and is configured to query in a preset table of the distributed system with the preset collective communication identification field as the primary key according to the first data,
[0154] if no record containing the first data is found in the preset table, store the message in the cache of the switching node, move the message into the to-be-forwarded queue, and add a record corresponding to the message in the preset table, and the record includes the address information of the message in the cache;
[0155] if a record containing the first data is found in the preset table, extract second data for reduction calculation from other messages corresponding to the record in the cache, perform corresponding reduction calculation on the second data of the other messages corresponding to the record in the cache and the first data of the message for reduction calculation, and update the second data of the other messages corresponding to the record in the cache to the result obtained by the reduction calculation.
[0156] Thus, the switching node performs reduction calculation on collective communication messages that meet the conditions, updates the result obtained by the reduction calculation in the corresponding messages in the cache, so that multiple messages to be forwarded become one, thereby reducing the length of the to-be-forwarded queue of the output port, effectively reducing the number of messages transmitted in the network, alleviating network congestion, and at the same time unloading the calculation pressure of the computing device.
[0157] As an optional embodiment, the system further includes:
[0158] a scheduling unit configured to sequentially schedule the messages in the to-be-forwarded queue, and if the message matches a record in the preset table, delete the record it matches from the preset table after forwarding the message.
[0159] Thus, the messages that have been forwarded can be deleted from the preset table in a timely manner, and the message records that need to be reduced in the preset table are updated in real time, which is beneficial to the maintenance of the data table and the normal operation of the system, and ensures the accuracy of in-network calculation.
[0160] As an optional embodiment, the preset collective communication identification field includes a group identification number field, and the group identification number field is globally unique and is used for grouping and identifying the computing devices.
[0161] As a further improvement of the above embodiments, the preset set communication identification field further includes a message tag field for marking the protocol operations of the set communication.
[0162] Among them, the group identification number field and the message tag field can respectively adopt the fields predefined in the set communication message. For example, the message tag field can be the Group ID field and the TAG field of the MPI standard.
[0163] Thus, it is possible to effectively group multiple computing devices and also identify multiple protocol operations of the set communication, enabling the switching node to perform corresponding protocol operations on the messages transmitted by each computing device based on the preset set communication identification field, thereby ensuring the accuracy and timeliness of the switching node during in-network computing.
[0164] As an optional embodiment, the protocol operations include finding the maximum value, finding the minimum value, summing, multiplying, and finding the average value.
[0165] Thus, by utilizing the in-network computing ability of the switching node, after corresponding protocol calculations, multiple associated messages can become one message, thereby reducing the length of the forwarding queue of the output port, effectively reducing the number of messages transmitted in the network, alleviating network congestion, and at the same time unloading the computing pressure of the computing device.
[0166] The present invention also provides a distributed system, including a plurality of switching nodes and a plurality of computing devices, and the switching nodes can implement the method for message processing described in the above embodiments.
[0167] As an optional embodiment, the plurality of computing devices are grouped according to a preset group identification number, and the group identification number is globally unique.
[0168] Specifically, for example, the group identification number field can be the Group ID field of the MPI standard, and the GroupID field under this standard can ensure global uniqueness. The group identification number field needs to maintain global uniqueness in the distributed system for effective grouping identification and management of all computing devices. It should be noted that the globally unique group identification number field is allowed to change, but it must maintain global uniqueness within a predefined time range.
[0169] The present invention also provides an electronic device, including: a processor; and a memory, on which a computer program is stored, and when the computer program is executed by the processor, it can implement the method for message processing described in the above embodiments.
[0170] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is used to run to implement the message processing method as described in the above embodiments.
[0171] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A method for message processing, which is used in a distributed system, characterized in that, The method includes the following steps: The switching node of the distributed system receives a packet from a computing device in the distributed system and determines the output port of the switching node that needs to forward the packet; The switching node determines whether the output port is congested and processes the packet according to the determined congestion status result. Specifically: When the forwarding queue of the output port is empty, the packet is directly forwarded from the output port; When the forwarding queue of the output port is not empty, the packet is analyzed, a processing policy for the packet is determined based on the analysis result, and the packet is selectively processed according to the processing policy.
2. The method according to claim 1, characterized in that, Determining the output port of the switching node that needs to forward the packet includes the following steps: Determine the output port according to the routing table of the switching node.
3. The method according to claim 1, wherein Analyzing the packet, determining a processing policy for the packet based on the analysis result, and selectively processing the packet according to the processing policy includes the following steps: If the packet is a collective communication packet, the packet is parsed and processed according to the parsing result, which includes: Determine whether the packet is a packet that needs to be reduced according to the parsing result; If so, perform corresponding reduction operations according to the packet; If not, store the packet in the cache of the switching node and move the packet into the forwarding queue of the output port; If the packet is not a collective communication packet, store the packet in the cache of the switching node and move the packet into the forwarding queue of the output port.
4. The method according to claim 3, wherein Analyzing the packet includes the following steps: Determine whether the packet is a collective communication packet according to preset packet recognition conditions.
5. The method according to claim 3, characterized in that If the packet is a collective communication packet, parsing the packet and processing the packet according to the parsing result includes the following steps: Parse the protocol data unit in the packet, and determine whether the packet is a packet that needs to be reduced according to the preset collective communication operation type field obtained by parsing.
6. The method according to claim 3, wherein Performing corresponding reduction operations according to the packet includes the following steps: Extract the first data of the preset collective communication identification field from the protocol data unit of the packet; Query in a preset table of the distributed system with the preset collective communication identification field as the primary key according to the first data; If no record containing the first data is found in the preset table, store the packet in the cache of the switching node, move the packet into the forwarding queue, and add a record corresponding to the packet to the preset table, and the record includes the address information of the packet in the cache; If a record containing the first data is found in the preset table, extract the second data for reduction calculation from other packets corresponding to the record stored in the cache, perform corresponding reduction calculation on the second data of the other packets corresponding to the record in the cache and the first data of the packet for reduction calculation, and update the second data of the other packets corresponding to the record in the cache to the result obtained by the reduction calculation.
7. The method according to claim 6, wherein It also includes the following steps: When the switching node schedules a message from the cache according to the queue to be forwarded, if the message matches a record in the preset table, the matching record is deleted from the preset table after the message is forwarded.
8. The method according to claim 6, wherein The preset set communication identification field includes a group identification number field, and the group identification number field is globally unique and is used to group the computing devices.
9. The method according to claim 8, wherein The preset collective communication identification field also includes a message tag field, which is used to mark the protocol operation of the collective communication.
10. The method according to claim 3, wherein The reduction operations include finding the maximum value, finding the minimum value, finding the sum, finding the product, and finding the average value.
11. A message processing system for a switching node of a distributed system, characterized in that, The system comprises: A receiving unit, configured to receive a message from a computing device in the distributed system, and determine an output port of the switching node to which the message needs to be forwarded; The processing unit is used to determine whether the output port is in a congested state, and process the message according to the congested state result obtained by the determination, specifically: When the queue to be forwarded of the output port is empty, the message is directly forwarded from the output port; When the queue to be forwarded of the output port is not empty, the message is analyzed, a processing strategy for the message is determined based on the analysis result, and the message is selectively processed according to the processing strategy.
12. The system according to claim 11, wherein The receiving unit includes a first module, which is used to determine the output port according to the routing table of the switching node.
13. The system according to claim 11, wherein, The processing unit includes a second module, which is used to determine whether the message is a collective communication message. If the message is a collective communication message, the message is parsed and processed according to the parsing result, including: According to the parsing result, determine whether the message is a message that needs to be regulated. If so, perform the corresponding protocol operation according to the message. If not, storing the message in the cache of the switching node and moving the message into the queue to be forwarded of the output port; If the message is not a collective communication message, the message is stored in the cache of the switching node, and the message is moved into the queue to be forwarded of the output port.
14. The system according to claim 13, wherein The second module is used to determine whether the message is a collective communication message according to a preset message identification condition.
15. The system according to claim 13, wherein The second module is used to parse the protocol data unit in the collective communication message, and determine whether the message is a message that needs to be reduced based on the preset collective communication operation type field obtained by the parsing.
16. The system according to claim 13, wherein, The processing unit includes a third module, which is used to extract the first data of the preset collective communication identification field from the protocol data unit of the message; and used to query a preset table in the distributed system with the preset set communication identification field as a primary key according to the first data, If no record containing the first data is found in the preset table, the message is stored in the cache of the switching node and the message is moved into the queue to be forwarded, and a record corresponding to the message is added to the preset table, and the record includes address information of the message in the cache; If a record containing the first data is found in the preset table, the second data for reduction calculation is extracted from other messages in the cache corresponding to this record. The second data is subjected to corresponding reduction calculation with the first data for reduction calculation of this message, and the second data of other messages in the cache corresponding to this record is updated to the result obtained from the reduction calculation.
17. The system according to claim 16, wherein The system further includes: A scheduling unit configured to sequentially schedule the messages in the to-be-forwarded queue. If a message matches a record in the preset table, the record it matches is deleted from the preset table after the message is forwarded.
18. The system according to claim 16, wherein The preset set communication identification field includes a group identification number field, which is globally unique and used for group identification of the computing devices.
19. The system according to claim 18, wherein The preset set communication identification field further includes a message tag field, which is used to mark the reduction operation of the set communication.
20. The system according to claim 13, wherein The reduction operations include finding the maximum value, finding the minimum value, summing, multiplying, and finding the average value.
21. A distributed system, characterized in that, It includes a plurality of switching nodes and a plurality of computing devices, and the switching nodes can implement the method according to any one of claims 1-10.
22. The system according to claim 21, wherein The plurality of computing devices are grouped according to a preset group identification number, and the group identification number is globally unique.
23. An electronic device, characterized in that, It includes: A processor; And A memory, on which a computer program is stored. When the computer program is executed by the processor, it can implement the method according to any one of claims 1 to 10.
24. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is used to run to implement the method according to any one of claims 1 to 10.