Data communication method and communication equipment

By adopting a data communication method combining reliable unicast and unreliable multicast in distributed clusters, the bottleneck problem of data exchange between nodes is solved, and efficient data synchronization and task processing is achieved.

CN120416264APending Publication Date: 2025-08-01LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510550315.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In distributed clusters, data communication between nodes becomes a bottleneck in performance optimization, and it is difficult for the prior art to achieve fast and efficient data exchange.

Method used

Reliable unicast method is used to receive beacon information and data information, determine the expected waiting time, and then receive data information through unreliable multicast method, combined with the unreliable multicast mechanism of multi-tree interleaving, to ensure the efficiency and reliability of data communication.

Benefits of technology

It realizes fast and efficient data communication between nodes in a distributed cluster, improves network resource utilization and task processing efficiency, and avoids waste of network resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416264A_ABST
    Figure CN120416264A_ABST
Patent Text Reader

Abstract

The invention discloses a data communication method and communication equipment, and relates to the technical field of artificial intelligence and communication, the communication equipment comprises first communication equipment and second communication equipment, the first communication equipment comprises a first transceiver and a first processor, and the second communication equipment comprises a second transceiver and a second processor; the first processor is configured to receive beacon information sent by at least one second communication device in a reliable unicast mode in a first stage, and receive first data information sent by the at least one second communication device in the reliable unicast mode; determining expected waiting time according to the receiving time of the beacon information and the receiving time of each piece of first data information; in the second stage, beacon information sent by at least one second communication device is received in a reliable unicast mode; receiving second data information sent by at least one second communication device through an unreliable multicast mode based on the expected waiting time; wherein the first data information and the second data information comprise related data in a model training or reasoning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and communication technologies, and particularly to a data communication method and a communication device. Background Art

[0002] With the rise of large models in recent years, AI (Artificial Intelligence) models have been continuously growing. The huge amount of computation and data makes it no longer practical to execute tasks such as model training or inference in a single-card (such as a single graphics card) or even a single-machine environment. The distributed processing method has emerged. People use a multi-machine computing cluster to execute tasks such as AI model training or inference. This can increase the parallelism of computation, improve the task execution efficiency, and at the same time relieve the data storage and processing pressure of a single card / single machine.

[0003] In the above processing method, frequent data exchanges are required between the nodes in the distributed cluster to achieve behavior synchronization. Thus, communication has become a bottleneck in optimizing the performance of the distributed cluster. How to perform fast and efficient communication between the nodes of the distributed cluster has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0004] To this end, the present application discloses the following technical solutions:

[0005] A first communication device, comprising:

[0006] A first transceiver; and

[0007] A first processor, which is coupled to the first transceiver, wherein the first processor is configured to:

[0008] In a first stage, receive beacon information sent by at least one second communication device through a reliable unicast method, and receive first data information sent by at least one second communication device through a reliable unicast method;

[0009] Determine an expected waiting time according to the reception time of the beacon information and the reception time of each of the first data information; the expected waiting time represents the waiting time required to collect all the data information currently synchronized by each of the second communication devices after the beacon information arrives;

[0010] In a second stage, receive beacon information sent by at least one second communication device through a reliable unicast method; based on the expected waiting time, receive second data information sent by at least one second communication device through an unreliable multicast method;

[0011] Wherein, the first data information and the second data information include relevant data in the model training or inference process.

[0012] Optionally, the first processor is further configured to:

[0013] Based on each received first data information, perform data processing corresponding to the to-be-processed task in the first stage;

[0014] Based on each received second data information, perform data processing corresponding to the to-be-processed task in the second stage;

[0015] Wherein, the to-be-processed task is a model training task or a model inference task, and the first communication device and each of the second communication devices are used for distributed processing of the to-be-processed task.

[0016] Optionally, when the first processor receives beacon information sent by at least one second communication device, it is further configured to:

[0017] Receive beacon information sent by the beacon source device;

[0018] Wherein, the beacon source device is a device selected as the beacon source among each of the second communication devices.

[0019] Optionally, when the first processor receives second data information sent by at least one second communication device, it is further configured to:

[0020] Receive different target data blocks sent by each second communication device through different unreliable multicast paths represented by different multicast trees;

[0021] Wherein, each target data block sent by each second communication device includes multiple data blocks obtained by splitting the second data information to be sent by the second communication device; each node on the multicast tree corresponds to a switch and the network interface of the corresponding communication device, and each multicast tree is used to represent the unreliable multicast path between the corresponding communication devices.

[0022] Optionally, when the first processor performs data processing corresponding to the to-be-processed task in the second stage based on each received second data information, it is further configured to:

[0023] If all the second data information sent by each of the second communication devices is collected within the expected waiting time, perform data processing corresponding to the to-be-processed task in the second stage based on all the second data information;

[0024] If all the second data information sent by each of the second communication devices is not collected within the expected waiting time, perform a complement processing on the missing second data information, and perform data processing corresponding to the to-be-processed task in the second stage based on the complemented second data information.

[0025] Optionally, the first processor is further configured to:

[0026] If all the second data information sent by each of the second communication devices is collected within the expected waiting time, update the expected waiting time based on the time taken to collect all the second data information.

[0027] A second communication device, comprising:

[0028] A second transceiver; and

[0029] A second processor coupled to the second transceiver, wherein the second processor is configured to:

[0030] In a first phase, send beacon information to at least one first communication device by reliable unicast, and send first data information to at least one first communication device by reliable unicast, so that each of the first communication devices determines an expected waiting time; the expected waiting time represents the waiting time required to collect all the currently synchronized data information after the beacon information arrives;

[0031] In a second phase, send beacon information to at least one first communication device by reliable unicast, and send second data information to at least one first communication device by unreliable multicast, so that the at least one communication device receives each of the second data information for synchronization based on the expected waiting time;

[0032] Wherein, the first data information and the second data information include relevant data in the model training or inference process.

[0033] Optionally, when the second processor sends beacon information to at least one first communication device, it is further configured to:

[0034] If the second communication device is a beacon source device, the second processor sends beacon information to at least one first communication device;

[0035] Wherein, the beacon source device is a communication device selected as the beacon source.

[0036] Optionally, when the second processor sends second data information to at least one first communication device, it is further configured to:

[0037] Send each target data block to each of the first communication devices through different unreliable multicast paths represented by corresponding different multicast trees;

[0038] Among them, each of the target data blocks includes a plurality of data blocks obtained by splitting the second data information to be sent by the second communication device; each node on the multicast tree corresponds to a network interface of a switch and a corresponding communication device, and each multicast tree is used to represent an unreliable multicast path between corresponding communication devices.

[0039] A switch, comprising:

[0040] A third transceiver; and

[0041] A third processor, which is coupled to the third transceiver, wherein the third processor is configured to:

[0042] Receive a target data block sent by a second communication device; the target data block includes corresponding data blocks obtained by splitting the second data information to be sent by the second communication device, and a multicast address;

[0043] Determine a target multicast tree corresponding to the multicast address;

[0044] Transmit the target data block to a corresponding first communication device through the unreliable multicast path represented by the target multicast tree;

[0045] Among them, each node on the multicast tree corresponds to a network interface of a switch and a corresponding communication device, and each multicast tree is used to represent an unreliable multicast path between corresponding communication devices; a communication device pair formed by corresponding two communication devices corresponds to a plurality of multicast trees, and the plurality of multicast trees corresponding to the communication device pair formed by two communication devices are used to represent a plurality of unreliable multicast paths between the two communication devices, and the plurality of unreliable multicast paths are respectively used to transmit different target data blocks.

[0046] A data communication method applied to a first communication device, comprising:

[0047] In a first stage, receive beacon information sent by at least one second communication device through a reliable unicast method, and receive first data information sent by at least one second communication device through a reliable unicast method;

[0048] Determine an expected waiting time according to the reception time of the beacon information and the reception time of each of the first data information; the expected waiting time represents the waiting time required to collect all the data information currently synchronized by each of the second communication devices after the beacon information arrives;

[0049] In a second stage, receive beacon information sent by at least one second communication device through a reliable unicast method; based on the expected waiting time, receive second data information sent by at least one second communication device through an unreliable multicast method;

[0050] Among them, the first data information and the second data information include relevant data in the model training or inference process.

[0051] A data communication method applied to a second communication device includes:

[0052] In the first stage, beacon information is sent to at least one first communication device by means of reliable unicast, and first data information is sent to at least one first communication device by means of reliable unicast, so that each of the first communication devices determines an expected waiting time; the expected waiting time represents the waiting time required to collect all the currently synchronized data information after the beacon information arrives.

[0053] In the second stage, beacon information is sent to at least one first communication device by means of reliable unicast, and second data information is sent to at least one first communication device by means of unreliable multicast, so that the at least one communication device receives each of the synchronized second data information based on the expected waiting time.

[0054] Among them, the first data information and the second data information include relevant data in the model training or inference process.

[0055] A data communication method applied to a switch includes:

[0056] Receiving a target data block sent by a second communication device; the target data block includes corresponding data blocks obtained by splitting the second data information to be sent by the second communication device, and a multicast address.

[0057] Determining a target multicast tree corresponding to the multicast address.

[0058] Transmitting the target data block to the corresponding first communication device through the unreliable multicast path represented by the target multicast tree.

[0059] Among them, each node on the multicast tree corresponds to the network interface of the switch and the corresponding communication device respectively, and each multicast tree is used to represent the unreliable multicast path between the corresponding communication devices; a communication device pair formed by two corresponding communication devices corresponds to multiple multicast trees, and the multiple multicast trees corresponding to the communication device pair formed by two communication devices are used to represent multiple unreliable multicast paths between the two communication devices, and the multiple unreliable multicast paths are respectively used to transmit different target data blocks. Description of the Drawings

[0060] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments or related technologies. Obviously, the accompanying drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided accompanying drawings.

[0061] Figure 1 is the composition structure diagram of the first communication device provided by the present application;

[0062] Figure 2 is the schematic diagram of distributed model training in a distributed cluster provided by the present application;

[0063] Figure 3 is the multicast schematic diagram provided by the present application;

[0064] Figure 4(a) - Figure 4(d) is the schematic diagram of different multicast trees provided by the present application;

[0065] Figure 5 is the composition structure diagram of the second communication device provided by the present application;

[0066] Figure 6 is the composition structure diagram of the switch provided by the present application;

[0067] Figure 7 is the schematic diagram of realizing lossy gradient synchronization based on RDMA unreliable multicast provided by the present application;

[0068] Figure 8 is the complete distributed cluster communication flow chart based on unreliable multicast provided by the present application;

[0069] Figure 9 is the data communication method flow chart applied to the first communication device provided by the present application;

[0070] Figure 10 is the data communication method flow chart applied to the second communication device provided by the present application;

[0071] Figure 11 is the data communication method flow chart applied to the switch provided by the present application. Specific Embodiments

[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0073] Embodiments of the present application provide a data communication method and a communication device, which are used for scenarios where data exchange is required between nodes in a distributed cluster to synchronize behaviors among nodes, etc., and realize fast and efficient data communication between nodes in the distributed cluster.

[0074] The communication device provided by the embodiments of the present application includes a first communication device and a second communication device. Among them, the first communication device and the second communication device are respectively a receiving-side device for receiving data and a sending-side device for sending data in distributed cluster communication, and specifically may be, but are not limited to, devices such as servers in the distributed cluster that have data sending and receiving requirements. In the distributed cluster, the same communication device can be used as a receiving-side device, a sending-side device, or can also be used as both a receiving-side device and a sending-side device according to the current actual data sending and receiving requirements, which is not limited and depends on the actual data sending and receiving requirements of each communication device in the distributed cluster.

[0075] See Figure 1 , which shows the composition structure diagram of the first communication device. The first communication device provided by the embodiments of the present application includes: a first transceiver 11 and a first processor 12.

[0076] The first transceiver is used for data sending and receiving, such as sending or receiving data synchronized in the distributed cluster, etc.

[0077] The first transceiver may specifically include a first transmitter for sending data and a first receiver for receiving data.

[0078] The first processor is coupled to the first transceiver.

[0079] Among them, the first processor is configured to:

[0080] In the first stage, receive beacon information sent by at least one second communication device through reliable unicast, and receive first data information sent by at least one second communication device through reliable unicast;

[0081] Determine an expected waiting time according to the receiving time of the beacon information and the receiving time of each piece of the first data information; the expected waiting time represents the waiting time required to collect all the data information currently synchronized by each second communication device after the beacon information arrives;

[0082] In the second stage, receive beacon information sent by at least one second communication device through reliable unicast; based on the expected waiting time, receive second data information sent by at least one second communication device through unreliable multicast;

[0083] In the embodiments of the present application, each communication device in the distributed cluster is used for distributed processing of tasks to be processed. Optionally, the task to be processed may be a model training task or a model inference task.

[0084] The first data information and the second data information may correspondingly include relevant data in the model training or inference process.

[0085] Each communication device in the distributed cluster needs to complete the distributed processing of the task to be processed through data exchange with each other. Taking the model training task as an example (a similar situation exists in the model inference scenario), refer to Figure 2 , which shows a schematic diagram of distributed model training between GPUs (Graphics Processing Unit) of different communication devices (such as servers) in a distributed cluster. Among them, each server's GPU in the distributed cluster maintains a copy of the model parameters. The training data is divided into N parts (assuming there are N GPUs in the cluster participating in the model training task), which are respectively used for forward and backward calculations on different GPUs to determine the model gradient data through forward and backward calculations. After obtaining the local gradient data, the GPUs in each server use relevant communication primitives such as the All-Reduce collective communication primitive to complete the global gradient synchronization of the entire cluster, so as to update the model parameters on the GPUs of each server based on the synchronized global gradient data.

[0086] The implementation method provided by the current related technology is to use RC (Reliable Connection) unicast and RDMA (Remote Direct Memory Access) transmission technologies to achieve cross-server communication, such as using RDMA RC unicast and GPU Direct RDMA transmission technologies to achieve cross-server GPU communication.

[0087] Based on the differences between the communication sender and receiver, the communication modes are divided into four categories: one-to-one, one-to-many, many-to-one, and many-to-many. The one-to-one communication mode is also known as point-to-point communication (Point-to-Point, P2P), while the one-to-many, many-to-one, and many-to-many communication modes involving multiple communication nodes belong to the category of collective communication (Collective Communication, CC). Among them, the collective communication of one-to-many and many-to-many (the many-to-many communication can be split into multiple one-to-many communications) conforms to the multicast / groupcast (Multicast) transmission method, which has the advantages of efficient, flexible, and fast communication, and can significantly reduce the use of network bandwidth and improve the utilization rate of network resources. Refer to Figure 3The multicast schematic diagram shown, especially the multicast accelerated by switch hardware, has a more efficient passing process. Based on this, in the embodiments of this application, the unicast method is no longer adopted in distributed cluster communication, but the multicast method is adopted for data communication between different communication devices.

[0088] In addition, the RDMA technology meets the increasingly stringent communication requirements of tasks such as distributed AI, such as high throughput, microsecond-level latency, and low CPU (Central Processing Unit) overhead, etc. Currently, it has become the de facto network technology standard for distributed clusters. The reliable connection (RC) mode of RDMA supports complete reliable transmission functions, so it is widely adopted. However, there is a natural mismatch between the multicast method and the semantics of the RDMA RC mode, which prevents the use of multicast in the RDMA network.

[0089] Specifically, first, the multicast data packet is sent in a one-to-many manner, and it is necessary to copy one data stream into multiple data streams in the switch, which violates the one-to-one connection semantics of RDMA RC. Second, RDMA requires feedback from the receiving end to perform functions such as reliability guarantee, retransmission, and congestion control. However, the RDMA RC feedback processing logic is designed for a single feedback stream, and multiple feedback streams in multicast will confuse the sending end and reduce the overall transmission performance.

[0090] To solve the above problems, at the same time, considering the characteristics of high tolerance for distributed data loss widely existing in the field of AI communication (data communication related to AI model training or inference tasks), the embodiments of this application utilize the insensitivity to data loss in AI communication and propose an unreliable multicast mechanism based on RDMA UD for AI communication. Based on this mechanism, this application specifically adopts the RDMA unreliable multicast method to perform data communication related to tasks such as model training or inference between different communication devices in the distributed cluster.

[0091] In the embodiments of this application, the data communication between communication devices in the distributed cluster is divided into two stages, namely the first stage and the second stage described above.

[0092] Among them, the first stage is essentially the WarmUp stage (preheating stage), which is used to determine the expected waiting time by performing beacon information and data information communication based on the reliable unicast method between different communication devices in the distributed cluster. The expected waiting time represents the waiting time required to collect all the currently synchronized data information in the distributed cluster after the beacon information arrives. The second stage is essentially the formal communication stage based on the WarmUp stage, which is used to perform data transmission between different communication devices in the distributed cluster by the unreliable multicast method based on the expected waiting time.

[0093] For the first communication device acting as the receiving side, in the first stage, specifically, its first processor can receive beacon information sent by at least one second communication device through reliable unicast, and receive first data information sent by at least one second communication device through reliable unicast. On this basis, according to the reception time of the beacon information and the reception time of each first data information, the expected waiting time is determined. In the second stage, the first processor can receive beacon information sent by at least one second communication device through reliable unicast; and based on the expected waiting time, receive second data information sent by at least one second communication device through unreliable multicast.

[0094] Among them, the reliable unicast method can specifically be the RDMA RC method, and the unreliable multicast method can specifically be the RDMA UD method.

[0095] In practical applications, corresponding communication devices can be selected from the distributed cluster as beacon source devices to use the beacon source devices to send beacon information. The selection method of the beacon source device can be, but is not limited to, the random selection method. For the first communication device on the receiving side, the beacon source device can specifically be the device selected as the beacon source among the second communication devices other than the first communication device in the distributed cluster.

[0096] After that, in the first and second stages of distributed communication, the beacon source device can be used to send beacon information to each receiving side device along with the data communication tasks between communication devices, so that the receiving side device, such as the first communication device, can perform corresponding processing based on the received beacon information.

[0097] Optionally, the beacon information is high-priority RC beacon information without carrying any data.

[0098] Correspondingly, in the first and second stages, the first processor in the first communication device, after receiving the beacon information sent by at least one second communication device, is further configured to: receive the beacon information sent by the beacon source device.

[0099] Among them, in the first stage, in addition to receiving the beacon information sent by the beacon source device based on reliable unicast, the first processor also receives the first data information sent by each second communication device participating in the distributed task processing in the distributed cluster based on reliable unicast. On this basis, optionally, the time difference between the reception time of each first data information and the reception time of the beacon message can be further calculated, and the maximum difference is used as the expected waiting time, which can be specifically expressed as:

[0100] T = max{T datai -T beacon}.

[0101] where, T represents the expected waiting time; T datai represents the receiving moment of the first communication device for the i-th first data information, that is, the moment when the i-th first data information arrives at the first communication device, where 1 ≤ i ≤ n, both i and n are integers, and n represents the number of first data information; T beacon represents the receiving moment of the first communication device for the beacon information in the first stage, that is, the moment when the beacon information arrives at the first communication device.

[0102] However, it is not limited to this. In other embodiments, in the first stage, the first processor of the first communication device may also, after collecting all the first data information, calculate the time difference between the collection moment of all the first data information (that is, the moment when the last first data information arrives at the first communication device) and the receiving moment of the beacon information, and use this difference as the expected waiting time. It is easy to understand that the values of the expected waiting time calculated by the above two methods are the same.

[0103] In practical applications, in the first stage, optionally, multiple rounds of global data (first data information) synchronization, such as multiple rounds of global gradient data synchronization, can be used to calculate multiple reference expected waiting times respectively, and the multiple reference expected waiting times can be comprehensively calculated to obtain the final expected waiting time. For example, by calculating the average value or median of the multiple reference expected waiting times, and using the average value or median as the final expected waiting time to further improve the reliability / credibility of the expected waiting time.

[0104] After that, in the second stage, the first processor in the first communication device first receives the beacon information sent by the beacon source device based on the reliable unicast method, and within the expected waiting time after receiving the beacon information, further receives the second data information sent by each second communication device in the distributed cluster based on the unreliable multicast method. That is to say, in the second stage, after receiving the beacon information, the first processor can time the expected waiting time, for example, start an interval timer with a duration of the expected waiting time to achieve timing, etc., and during the timing period, receive the second data information in an unreliable multicast manner, and stop receiving data after the timing ends. That is, even if all the second data information of each second communication device has not been collected after the timing ends, the receiving of the second data information will not continue.

[0105] The first data information and the second data information include relevant data in the model training or inference process. Exemplarily, if the task to be processed is a model training task, the first data information and the second data information may respectively include gradient data to be synchronized between different communication devices in the corresponding stage, so as to obtain all the gradient data of the current stage on each communication device based on the gradient data synchronization between communication devices in the distributed cluster, and then update the model parameters based on the obtained gradient data. If the task to be processed is a model inference task, the first data information and the second data information may respectively include token data to be synchronized between different communication devices in the corresponding stage, so as to obtain all the token data of the current stage on each communication device based on the token data synchronization between communication devices in the distributed cluster, and then execute the corresponding inference task based on the obtained token data.

[0106] The token of a model refers to the smallest unit after splitting the data to be processed in AI models such as large models. When an AI model processes text, it cannot directly process the words in human natural language. Therefore, the text needs to be split into individual tokens so that the model can understand and process it. In natural language processing (NLP), a token can refer to the basic processing unit after splitting the text, which can be a phrase, a word, or a character, etc.

[0107] In summary, by dividing the communication process in the distributed cluster into the above two stages and designing information transmission for each stage, the embodiment of the present application can estimate the approximate time required for communication devices in the distributed cluster to collect all the data currently synchronized in the cluster in the first stage, that is, the expected waiting time. At the same time, by synchronizing the information (beacon information, first data information) for time estimation based on the reliable unicast method in the first stage, the reliability / credibility of the estimated expected waiting time is ensured. In the second stage, based on the expected waiting time estimated in the first stage, the second data information is received based on the unreliable multicast method, which not only ensures that the second data information synchronized by each device in the cluster can be collected as much as possible based on the unreliable multicast method, but also avoids excessive time waiting. Therefore, based on the unreliable multicast mechanism proposed in the embodiment of the present application, fast and efficient communication can be realized between nodes in the distributed cluster, while avoiding network resource waste and improving network resource utilization.

[0108] In an optional embodiment, the first processor in the first communication device is further configured to:

[0109] Execute the data processing corresponding to the first stage of the task to be processed based on the received first data information; execute the data processing corresponding to the second stage of the task to be processed based on the received second data information.

[0110] In the first stage, since the first processor receives the first data information based on the reliable unicast method, the first processor can collect the first data information sent by each second communication device in the cluster. Correspondingly, it can directly perform the data processing corresponding to the task to be processed in the first stage based on the collected first data information.

[0111] Taking the task to be processed as a model training task as an example, the first processor of the first communication device can directly update the model parameters based on the gradient data collected in the first stage, so as to iteratively train the AI model by updating the model parameters based on the gradient data.

[0112] In the second stage, since the first processor receives the second data information based on the unreliable multicast method, and the unreliable multicast method has the characteristic of unreliable communication, there are two situations where the first processor either collects or does not collect all the second data information sent by each second communication device within the expected waiting time, depending on the actual communication quality.

[0113] Among them, if all the second data information sent by each second communication device is collected within the expected waiting time, the first processor can perform the data processing corresponding to the task to be processed in the second stage based on all the second data information. For example, for the model training task, update the model parameters based on all the gradient data collected in the second stage; or for the model inference task, perform the corresponding model inference based on all the token data collected in the second stage.

[0114] On the contrary, if all the second data information sent by each second communication device is not collected within the expected waiting time, the first processor can first perform a complementary processing on the missing second data information.

[0115] In an optional implementation manner, a preset value can be used to complement the missing second data information. For example, directly use the value 0 to complement the missing gradient data, but not limited to this. In other optional implementation manners, an estimation algorithm based on the historical state can also be used to estimate and complement the missing second data information. In implementation, the complementary manner of the missing second data information can be determined according to the actual requirements.

[0116] After completing the complementary processing of the missing second data information to obtain the complemented second data information, the data processing corresponding to the task to be processed in the second stage can be further performed based on the complemented second data information. For example, in the model training task, update the model parameters based on the currently complemented gradient data (including the non-missing gradient data and the data such as the value 0 supplemented for the missing gradient data).

[0117] In this embodiment, based on each piece of first data information received, data processing corresponding to the first stage of the task to be processed is performed, and based on each piece of second data information received, data processing corresponding to the second stage of the task to be processed is performed. When all the second data information sent by each second communication device is not collected in the second stage, the missing second data information is first complemented, realizing distributed data synchronization and task processing based on an unreliable multicast mechanism. Compared with the related art that uses RDMA RC unicast for distributed communication and task processing, the embodiment of the present application can achieve fast and efficient communication between nodes of a distributed cluster, improve the efficiency of data communication and task processing, avoid waste of network resources, and improve the utilization rate of network resources.

[0118] In an alternative embodiment, the first processor in the first communication device is further configured to: if all the second data information sent by each second communication device is collected within the expected waiting time, update the expected waiting time based on the time taken to collect all the second data information.

[0119] Optionally, in practical applications, it can be specifically determined whether the actual time taken to collect all the second data information after receiving the beacon information is less than the expected waiting time. If it is less, it indicates that the time taken to collect all the second data information after receiving the beacon information does not require waiting for the duration of the expected waiting time, so the expected waiting time can be updated to the actual time taken to collect all the second data information after receiving the beacon information.

[0120] Conversely, if the actual time taken to collect all the second data information after receiving the beacon information is equal to the expected waiting time, there is no need to update the expected waiting time.

[0121] If the actual time taken to collect all the second data information after receiving the beacon information is greater than the expected waiting time, in order to avoid excessive time waiting and save network resources, the expected waiting time is also not updated.

[0122] In this embodiment, by updating the expected waiting time based on the actual time taken to collect all the second data information currently synchronized in the distributed cluster, and specifically updating the expected waiting time to the actual time when the actual time is lower than the expected waiting time, and not updating the expected waiting time when the actual time is not lower than the expected waiting time, the expected waiting time is optimized and adjusted according to the actual network condition to make it more consistent with the actual network condition, thereby further improving the communication efficiency and distributed task processing efficiency based on the unreliable multicast mechanism in the present application, avoiding waste of network resources, and improving the utilization rate of network resources.

[0123] In an optional embodiment, the first processor in the first communication device is further configured to:

[0124] Different target data blocks sent by each second communication device through different unreliable multicast paths represented by different multicast trees are received.

[0125] Among them, each target data block sent by each second communication device includes multiple data blocks obtained by dividing the second data information to be sent by the second communication device; each node on the multicast tree corresponds to the network interface (such as a network card) of the switch and the corresponding communication device, and each multicast tree is used to represent an unreliable multicast path between the corresponding communication devices.

[0126] In distributed clusters, such as AI distributed clusters, RDMA UD packets may be lost or experience significant delays during transmission. These issues are primarily caused by load imbalance within the distributed cluster network or by the disconnection or degradation of specific transmission links, such as frequent optical module failures in a live network. Live networks refer to the currently operational, production-ready network environment used by users, typically as opposed to development and testing environments. When using UD multicast for collective communication, congestion or disconnection of specific links can cause specific receivers in the multicast group to continuously not receive data, impacting global data synchronization across the entire cluster.

[0127] To alleviate this problem, this embodiment proposes an unreliable multicast mechanism based on multi-tree interleaving.

[0128] In an unreliable multicast mechanism based on multi-tree interleaving, optionally, within a distributed cluster, a communication device pair formed by any two different communication devices participating in distributed task processing may correspond to multiple multicast trees, and the multiple multicast trees corresponding to the communication device pair formed by the two communication devices represent multiple unreliable multicast paths between the two communication devices, and are used to support, when there is a need for data synchronization between the two communication devices, different data blocks of the data to be synchronized can be transmitted separately based on multiple different UD paths represented by the multiple multicast trees, so as to avoid as much as possible the continuous data loss or delay caused by a specific link represented by a single tree.

[0129] Based on this, when it is necessary to synchronize the second data information of each second communication device in a distributed cluster to the first communication device, the second data information to be sent by the second communication device can be first divided to obtain multiple target data blocks. On this basis, each target data block is sent to the first communication device through different UD paths represented by multiple multicast trees between the second communication device and the first communication device.

[0130] Correspondingly, the first processor of the first communication device can specifically receive different target data blocks sent by each second communication device through different unreliable multicast paths represented by different multicast trees. Subsequently, by integrating the received target data blocks of the same second communication device, the second data information from the same second communication device can be obtained, so as to further perform subsequent processing based on the second data information. For example, by integrating the gradient data blocks of the same second communication device, the gradient data from the same second communication device can be obtained, and then subsequent model parameter update processing can be performed based on the integrated gradient data.

[0131] See Figure 4(a) - Figure 4(d) The example diagram of the UD multicast scheme based on multi-tree interleaving shown in the figure. In this example, taking the server node pair (G0, G3) as an example, G0 is simultaneously connected to 4 multicast trees: namely, Tree0, Tree1, Tree2, and Tree3 shown in Figures 4(a)-4(d) in sequence. These 4 multicast trees represent 4 unreliable multicast paths between G0 and G3. Thus, when G0 needs to send data to G3, the data to be sent by G0 can be divided into multiple data blocks, such as divided into data blocks D0, D1, D2, and D3, and different data blocks are respectively sent to G3 using different unreliable multicast paths represented by Tree0, Tree1, Tree2, and Tree3.

[0132] In this embodiment, by proposing an unreliable multicast mechanism based on multi-tree interleaving, and using different unreliable multicast paths represented by multiple multicast trees to respectively transmit different data blocks of the data to be synchronized between communication devices, the problem of continuous data loss or delay caused by data transmission through a specific link represented by a single tree is avoided as much as possible. The problem that specific receivers in the multicast group cannot continuously receive data due to congestion or disconnection of a specific link in UD multicast communication, thereby affecting the global data synchronization of the entire cluster, is effectively solved, thus improving the reliability of unreliable multicast communication and distributed task processing based on this.

[0133] See Figure 5 shows the composition structure diagram of the second communication device. The second communication device provided in the embodiment of the present application includes: a second transceiver 51 and a second processor 52.

[0134] The second transceiver is used for data transceiver, such as sending or receiving data synchronized in a distributed cluster, etc.

[0135] The second transceiver may specifically include a second transmitter for sending data and a second receiver for receiving data.

[0136] The second processor is coupled to the second transceiver.

[0137] Wherein, the second processor is configured to:

[0138] In the first stage, beacon information is sent to at least one first communication device by reliable unicast, and first data information is sent to at least one first communication device by reliable unicast, so that each of the first communication devices can determine an expected waiting time; the expected waiting time represents the waiting time required to collect all the currently synchronized data information after the beacon information arrives.

[0139] In the second stage, beacon information is sent to at least one first communication device by reliable unicast, and second data information is sent to at least one first communication device by unreliable multicast, so that the at least one communication device can receive each of the synchronized second data information based on the expected waiting time.

[0140] Wherein, the first data information and the second data information include relevant data in the model training or inference process.

[0141] Matched with the communication processing of the first communication device on the receiving side, the communication processing of the second communication device on the sending side also includes two stages: the first stage and the second stage.

[0142] In the first stage, the second processor of the second communication device sends relevant information to the first communication device based on reliable unicast, for example, sends beacon information to the first communication device based on the RDMA RC mode, and sends first data information to the first communication device based on the RDMA RC, so that the first communication device can determine the expected waiting time based on the receiving moments of the beacon information and the first data information; in the second stage, the second processor of the second communication device further sends beacon information to the first communication device based on reliable unicast, and sends second data information to the first communication device based on unreliable multicast, for example, sends beacon information to the first communication device based on the RDMA RC mode, and sends second data information to the first communication device based on the RDMA UD mode, so that the first communication device can receive the second data information sent by each second communication device based on the beacon information and the expected waiting time in the second stage, thereby correspondingly realizing fast and efficient communication between the nodes of the distributed cluster based on unreliable multicast, while avoiding waste of network resources and improving network resource utilization.

[0143] In an optional implementation manner, when the second processor sends beacon information to at least one first communication device, it is further configured to: if the second communication device is a beacon source device, the second processor sends beacon information to at least one first communication device.

[0144] Wherein, the beacon source device is a communication device selected as the beacon source in the distributed cluster.

[0145] In an alternative embodiment, when the second processor sends second data information to at least one first communication device, it is further configured to: send each target data block to each of the first communication devices through different unreliable multicast paths represented by corresponding different multicast trees.

[0146] Wherein, each of the target data blocks includes a plurality of data blocks obtained by splitting the second data information to be sent by the second communication device; each node on the multicast tree corresponds to a switch and a network interface of a corresponding communication device, and each multicast tree is used to represent an unreliable multicast path between corresponding communication devices.

[0147] The second communication device provided in this embodiment matches the first communication device provided in the corresponding embodiment above, and is respectively used as the sending-side device and the receiving-side device to be synchronized with data in the distributed cluster. For a more detailed implementation process of the communication processing of the second communication device in two stages, reference can be made to the relevant description of the communication processing process of the first communication device in the above embodiment, which will not be elaborated here.

[0148] In summary, by dividing the communication process in the distributed cluster into the above two stages, the embodiment of the present application can estimate the approximate time required for the communication devices in the distributed cluster to collect all the data currently synchronized in the cluster in the first stage, that is, the expected waiting time. At the same time, by synchronizing the information (beacon information, first data information) for time estimation based on the reliable unicast method in the first stage, the reliability / credibility of the estimated expected waiting time is ensured. In the second stage, based on the expected waiting time estimated in the first stage, the second data information is communicated based on the unreliable multicast method, which not only ensures that the receiving-side device can collect the second data information synchronized by each device in the cluster as much as possible based on the unreliable multicast method, but also avoids excessive time waiting. Therefore, based on the unreliable multicast mechanism proposed in the embodiment of the present application, fast and efficient communication can be realized between the nodes of the distributed cluster, the efficiency of data synchronization and task processing is improved, and at the same time, network resource waste is avoided and the network resource utilization rate is improved.

[0149] The embodiment of the present application also provides a switch. Refer to Figure 6 , which shows the composition structure diagram of the switch, specifically including: a third transceiver 61 and a third processor 62.

[0150] The third transceiver is used for data transceiver, such as sending or receiving data synchronized in the distributed cluster, etc.

[0151] The third transceiver may specifically include a third transmitter for sending data and a third receiver for receiving data.

[0152] The third processor is coupled to the third transceiver.

[0153] Wherein, the third processor is configured to:

[0154] Receive a target data block sent by a second communication device; the target data block includes a corresponding data block obtained by splitting second data information to be sent by the second communication device, and a multicast address;

[0155] Determine a target multicast tree corresponding to the multicast address;

[0156] Transmit the target data block to a corresponding first communication device through an unreliable multicast path represented by the target multicast tree.

[0157] Wherein, each node on the multicast tree respectively corresponds to a network interface of a switch and a corresponding communication device, and each multicast tree is used to represent an unreliable multicast path between corresponding communication devices; a communication device pair formed by two communication devices corresponds to multiple multicast trees, and the multiple multicast trees corresponding to the communication device pair formed by two communication devices are used to represent multiple unreliable multicast paths between the two communication devices, and the multiple unreliable multicast paths are respectively used to transmit different target data blocks.

[0158] In addition to communication devices (such as servers) on the receiving side and the sending side, the distributed cluster further includes switches for data forwarding, and the number of switches is one or more, which is not limited and depends on the actual application scenario.

[0159] In this embodiment, optionally, the switch may first perform network topology awareness on the distributed cluster during the communication initialization phase of the distributed cluster, so as to obtain network topology information of the distributed cluster by combining network topology awareness and device information of the server. The network topology information may include, but is not limited to, servers and switches included in the distributed cluster, network cards deployed on the servers, GPUs mounted on the network cards, and connection relationships between devices such as servers and switches, etc. Then, based on the network topology information, multiple multicast trees corresponding to communication device pairs in the distributed cluster are constructed to form multiple unreliable multicast paths represented by the multiple multicast trees between different communication devices included in the communication device pair, so as to support that when there is a data synchronization requirement between different communication devices, different data blocks of the data to be synchronized can be transmitted respectively through multiple different UD paths represented by the corresponding multiple multicast trees, and the problem of continuous data loss or delay caused by data transmission through a specific link represented by a single tree is avoided as much as possible.

[0160] Afterwards, in the unreliable multicast communication phase of the distributed cluster, the second communication device can split the second data information to be sent into multiple data blocks, and bind a corresponding multicast address to each data block based on the actual data transmission requirements, forming respective target data blocks carrying the multicast address and data block information, and then send each target data block to a switch whose multicast address matches.

[0161] The switch can correspondingly receive the target data block sent by the second communication device, determine the target multicast tree corresponding to the multicast address in the target data block, and then transmit the target data block to the corresponding first communication device through the unreliable multicast path represented by the target multicast tree.

[0162] In this embodiment, by constructing corresponding multiple multicast trees for communication device pairs in the distributed cluster at the switch, and implementing unreliable multicast communication between different communications in the distributed cluster based on the unreliable multicast mechanism of multi-tree interleaving, it is possible to avoid as much as possible the continuous data loss or delay caused by a specific link represented by a single tree, effectively solving the problem that in UD multicast communication, due to congestion or disconnection of a specific link, specific receivers in the multicast group continuously cannot receive data, thus affecting the global data synchronization of the entire cluster, thereby improving the reliability of unreliable multicast communication and distributed task processing based on this.

[0163] The following takes the distributed processing of a model training task as an example to provide an application example of the present application.

[0164] In this example, the distributed cluster includes multiple servers for distributed model training and several switches, and a GPU for executing corresponding model training tasks based on gradient data synchronization is deployed on each server.

[0165] See Figure 7 , which is a schematic diagram of lossy gradient synchronization based on RDMA unreliable multicast in this example. Combining with Figure 8 the complete distributed cluster communication process based on unreliable multicast shown, the gradient synchronization based on RDMA UD multicast is divided into two stages: the WarmUp stage and the formal training stage.

[0166] I. WarmUp stage (the first stage)

[0167] In the RDMA UD multicast scheme of this example, the global gradient synchronization is completed using the RDMA RC unicast method in the WarmUp stage.

[0168] Among them, a corresponding server can be selected from the distributed cluster as the beacon source device, and the beacon source device is used to send high-priority RC beacon messages without carrying any data along with the communication task. Each receiving server calculates the time difference between the moment when the global gradient data set is complete and the moment when the beacon message arrives. And through multiple rounds of global gradient synchronization and the comprehensive operation of the time differences obtained in multiple rounds of global gradient synchronization, the expected waiting time T for all gradient data to be complete after the beacon message arrives in the current distributed network environment can be learned.

[0169] II. Formal training stage (the second stage)

[0170] After the formal training stage starts, the GPUs in each server use UD multicast to complete the synchronization of gradient data. Due to the unreliable characteristics of UD multicast, some multicast members may have data loss or data arriving with delay.

[0171] After each GPU receives the beacon message in this stage, it starts an interval timer with a duration of T. When the timer times out, it immediately collects the gradient data that has completed transmission. If all the gradient data for the current synchronization has not been collected, the gradient estimation algorithm can be used to estimate the gradient data that has not arrived in time to complete the missing gradient data. A simple estimation method is to directly use 0 to complete the missing gradient data to participate in the subsequent model parameter update based on the gradient data.

[0172] However, it is not limited to this. The missing gradient data can also be completed based on a more complex historical state-based estimation algorithm. This method requires more video memory and computing power, and at the same time will bring more accurate gradient estimation values. In implementation, the method for completing the missing gradient data can be determined based on actual requirements.

[0173] When the network transmission quality is good, if each data to be synchronized based on UD multicast can arrive at the destination receiver in time within the expected waiting time and with a duration less than the expected waiting time, the update of the expected waiting time T can also be triggered, so that the size of the expected waiting time T is more matched with the current network transmission quality, to further avoid wasting network resources and improve the utilization rate of network resources.

[0174] In practical applications, refer to Figure 8, the switch can also perform network topology awareness on the distributed cluster during the communication initialization phase of the distributed cluster, and complete the construction of the multicast tree based on the network topology information to support UD multicast communication based on multi-tree interleaving in the distributed cluster, and avoid as much as possible the continuous data loss or delay caused by specific links represented by a single tree, effectively solving the problem that specific receivers in the multicast group cannot continuously receive data due to congestion or disconnection of specific links in UD multicast communication, thereby affecting the global data synchronization of the entire cluster, and thus improving the reliability of unreliable multicast communication and distributed task processing based on this.

[0175] Corresponding to the above first communication device, an embodiment of the present application further provides a data communication method applied to the first communication device. Refer to Figure 9 , the data communication method includes the following processing steps:

[0176] Step 901, in the first stage, receive beacon information sent by at least one second communication device through reliable unicast, and receive first data information sent by at least one second communication device through reliable unicast;

[0177] Step 902, determine the expected waiting time according to the reception time of the beacon information and the reception time of each first data information; the expected waiting time represents the waiting time required to collect all the data information currently synchronized by each second communication device after the beacon information arrives;

[0178] Step 903, in the second stage, receive beacon information sent by at least one second communication device through reliable unicast; based on the expected waiting time, receive second data information sent by at least one second communication device through unreliable multicast.

[0179] Wherein, the first data information and the second data information include relevant data in the model training or inference process.

[0180] In an optional embodiment, the above method may further include:

[0181] Execute the data processing corresponding to the first stage of the task to be processed based on each received first data information;

[0182] Execute the data processing corresponding to the second stage of the task to be processed based on each received second data information;

[0183] Wherein, the task to be processed is a model training task or a model inference task, and the first communication device and each second communication device are used for distributed processing of the task to be processed.

[0184] In an alternative embodiment, receiving beacon information sent by at least one second communication device may include: receiving beacon information sent by a beacon source device;

[0185] wherein the beacon source device is a device selected as the beacon source among the second communication devices.

[0186] In an alternative embodiment, receiving second data information sent by at least one second communication device includes: receiving different target data blocks sent by each second communication device through different unreliable multicast paths characterized by different multicast trees.

[0187] wherein each target data block sent by each second communication device includes a plurality of data blocks obtained by splitting the second data information to be sent by the second communication device; each node on the multicast tree corresponds to a switch and a network interface of a corresponding communication device, and each multicast tree is used to characterize an unreliable multicast path between corresponding communication devices.

[0188] In an alternative embodiment, based on the received second data information, performing data processing corresponding to the second phase of the to-be-processed task may include:

[0189] If all the second data information sent by the second communication devices is collected within the expected waiting time, performing data processing corresponding to the second phase of the to-be-processed task based on all the second data information;

[0190] If all the second data information sent by the second communication devices is not collected within the expected waiting time, performing a complement processing on the missing second data information, and performing data processing corresponding to the second phase of the to-be-processed task based on the complemented second data information.

[0191] In an alternative embodiment, the above method may further include:

[0192] If all the second data information sent by the second communication devices is collected within the expected waiting time, performing an update processing on the expected waiting time based on the time taken to collect all the second data information.

[0193] The data communication method provided in this embodiment corresponds to the first communication device provided in the corresponding embodiment above. For related similarities, reference may be made to the description of the first communication device in the corresponding embodiment above, which will not be elaborated here.

[0194] Corresponding to the above second communication device, an embodiment of the present application further provides a data communication method applied to a second communication device. Refer to Figure 10 , and this data communication method includes the following processing steps:

[0195] Step 1001: In the first stage, send beacon information to at least one first communication device via reliable unicast, and send first data information to at least one first communication device via reliable unicast, so that each of the first communication devices can determine an expected waiting time; the expected waiting time represents the waiting time required to collect all the currently synchronized data information after the beacon information arrives.

[0196] Step 1002: In the second stage, send beacon information to at least one first communication device via reliable unicast, and send second data information to at least one first communication device via unreliable multicast, so that the at least one communication device can receive each of the synchronized second data information based on the expected waiting time.

[0197] Wherein, the first data information and the second data information include relevant data in the model training or inference process.

[0198] In an alternative embodiment, sending beacon information to at least one first communication device includes:

[0199] If the second communication device is a beacon source device, send beacon information from the second communication device to at least one first communication device.

[0200] Wherein, the beacon source device is a communication device selected as the beacon source.

[0201] In an alternative embodiment, sending second data information to at least one first communication device includes: sending each target data block to each first communication device via different unreliable multicast paths represented by corresponding different multicast trees.

[0202] Wherein, each of the target data blocks includes multiple data blocks obtained by splitting the second data information to be sent by the second communication device; each node on the multicast tree corresponds to a network interface of a switch and a corresponding communication device, and each multicast tree is used to represent an unreliable multicast path between corresponding communication devices.

[0203] The data communication method provided in this embodiment corresponds to the second communication device provided in the corresponding embodiment above. For relevant similarities, reference can be made to the description of the second communication device in the corresponding embodiment above, which will not be elaborated here.

[0204] Corresponding to the above switch, an embodiment of the present application further provides a data communication method applied to a switch. Refer to Figure 11 , and this data communication method includes the following processing steps:

[0205] Step 1101: Receive a target data block sent by a second communication device; the target data block includes corresponding data blocks obtained by dividing second data information to be sent by the second communication device, and a multicast address;

[0206] Step 1102: Determine the target multicast tree corresponding to the multicast address;

[0207] Step 1103: Transmit the target data block to the corresponding first communication device through the unreliable multicast path represented by the target multicast tree.

[0208] Among them, each node on the multicast tree corresponds to a network interface of a switch and a corresponding communication device, respectively, and each multicast tree is used to represent an unreliable multicast path between the corresponding communication devices; a communication device pair formed by the corresponding two communication devices corresponds to multiple multicast trees, and the multiple multicast trees corresponding to the communication device pair formed by the two communication devices are used to represent multiple unreliable multicast paths between the two communication devices, and the multiple unreliable multicast paths are respectively used to transmit different target data blocks.

[0209] Optionally, during the communication initialization phase of the distributed cluster, the switch is used to perceive the network topology of the distributed cluster and complete the multicast tree construction based on the network topology information to support UD multicast communication based on multi-tree interweaving in the distributed cluster, thereby avoiding as much as possible the continuous data loss or delay caused by a specific link represented by a single tree.

[0210] The data communication method provided in this embodiment corresponds to the switch provided in the corresponding embodiment above. For relevant similarities, please refer to the description of the switch in the corresponding embodiment above, which will not be repeated here.

[0211] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referenced to each other.

[0212] For the convenience of description, the above systems or devices are described as being divided into various modules or units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0213] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a creative contribution, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0214] Finally, it should also be noted that in this text, relational terms such as first, second, third, and fourth are used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0215] The above are only the preferred embodiments of this application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A first communication device, comprising: A first transceiver; And A first processor coupled to the first transceiver, wherein the first processor is configured to: In a first phase, receive beacon information sent by at least one second communication device via reliable unicast, and receive first data information sent by at least one second communication device via reliable unicast; Determine an expected waiting time according to the reception time of the beacon information and the reception time of each of the first data information; the expected waiting time represents the waiting time required to collect all the data information currently synchronized by each of the second communication devices after the beacon information arrives; In a second phase, receive beacon information sent by at least one second communication device via reliable unicast; Based on the expected waiting time, receive second data information sent by at least one second communication device via unreliable multicast; Wherein, the first data information and the second data information include relevant data in the model training or inference process.

2. The first communication device according to claim 1, wherein the first processor is further configured to: Based on each received first data information, perform data processing corresponding to the to-be-processed task in the first phase; Based on each received second data information, perform data processing corresponding to the to-be-processed task in the second phase; Among them, The to-be-processed task is a model training task or a model inference task, and the first communication device and each of the second communication devices are used for distributed processing of the to-be-processed task.

3. The first communication device according to claim 1, when the first processor receives beacon information sent by at least one second communication device, it is further configured to: Receive beacon information sent by a beacon source device; Among them, The beacon source device is a device selected as the beacon source among each of the second communication devices.

4. The first communication device according to claim 1, when the first processor receives second data information sent by at least one second communication device, it is further configured to: Receive different target data blocks sent by each second communication device via different unreliable multicast paths represented by different multicast trees; Among them, Each target data block sent by each second communication device includes a plurality of data blocks obtained by splitting the second data information to be sent by the second communication device; each node on the multicast tree corresponds to a network interface of a switch and a corresponding communication device, and each multicast tree is used to represent an unreliable multicast path between corresponding communication devices.

5. The first communication device according to claim 2, when the first processor performs data processing corresponding to the to-be-processed task in the second phase based on each received second data information, it is further configured to: If all the second data information sent by each of the second communication devices is collected within the expected waiting time, perform data processing corresponding to the to-be-processed task in the second phase based on all the second data information; If all the second data information sent by each of the second communication devices is not collected within the expected waiting time, the missing second data information is complemented, and the data processing corresponding to the to-be-processed task in the second stage is performed based on the complemented second data information of each device.

6. The first communication device according to claim 5, wherein the first processor is further configured to: If all the second data information sent by each of the second communication devices is collected within the expected waiting time, the expected waiting time is updated based on the time taken to collect all the second data information.

7. A second communication device, comprising: A second transceiver; And A second processor coupled to the second transceiver, wherein the second processor is configured to: In a first stage, send beacon information to at least one first communication device by reliable unicast, and send first data information to at least one first communication device by reliable unicast, so that each of the first communication devices determines an expected waiting time; The expected waiting time represents the waiting time required to collect all the data information currently synchronized after the beacon information arrives; In a second stage, send beacon information to at least one first communication device by reliable unicast, and send second data information to at least one first communication device by unreliable multicast, so that the at least one communication device receives each of the second data information synchronized based on the expected waiting time; Wherein, the first data information and the second data information include relevant data in the model training or inference process.

8. The second communication device according to claim 7, wherein when the second processor sends beacon information to at least one first communication device, it is further configured to: If the second communication device is a beacon source device, the second processor sends beacon information to at least one first communication device; Among them, The beacon source device is a communication device selected as the beacon source.

9. The second communication device according to claim 7, wherein when the second processor sends second data information to at least one first communication device, it is further configured to: Send each target data block to each of the first communication devices through different unreliable multicast paths represented by corresponding different multicast trees; Among them, Each of the target data blocks includes a plurality of data blocks obtained by splitting the second data information to be sent by the second communication device; each node on the multicast tree corresponds to a network interface of a switch and a corresponding communication device, and each multicast tree is used to represent an unreliable multicast path between corresponding communication devices.

10. A switch, comprising: A third transceiver; [[ID= ​ ​ ​ ​ Among them, each node on the multicast tree corresponds to the network interfaces of a switch and corresponding communication devices respectively. Each multicast tree is used to represent an unreliable multicast path between corresponding communication devices. A communication device pair formed by corresponding two communication devices corresponds to multiple multicast trees. The multiple multicast trees corresponding to the communication device pair formed by two communication devices are used to represent multiple unreliable multicast paths between the two communication devices. The multiple unreliable multicast paths are respectively used to transmit different target data blocks.