Distributed Transform training-oriented gradient bounded loss tolerance transmission optimization method

By adopting the UDP protocol and the application-layer packet loss detection mechanism in the distributed Transformer model training, the congestion and tail delay problems of TCP protocol in high bandwidth and low latency environments are solved, and a more efficient and stable training process is achieved.

CN119996533AActive Publication Date: 2025-05-13NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510449151.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-13
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The traditional distributed Transformer model based on TCP has congestion and tail delay problems in a network environment with high bandwidth and low latency, which affects the training efficiency.

Method used

The UDP protocol is used for gradient transmission, and the packet loss detection and retransmission mechanism is implemented at the application layer. The lost packets are managed through custom header information and RS retransmission tables, and the transmission rate is dynamically adjusted to alleviate link congestion.

Benefits of technology

It effectively avoids congestion and tail delay problems of TCP protocol in high bandwidth and low latency environments, and improves the efficiency and stability of distributed Transformer model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996533A_ABST
    Figure CN119996533A_ABST
Patent Text Reader

Abstract

The invention discloses a gradient bounded loss tolerance transmission optimization method for distributed Transform training. The gradient bounded loss tolerance transmission optimization method comprises the following steps of: firstly, carrying out gradient bounded loss tolerance; according to the method, packet loss rate detection, a retransmission mechanism, congestion control and reliable transmission are realized in an application layer by utilizing the low delay characteristic of a UDP protocol and combining the tolerance of Transform large model training to gradient bounded loss, and a user-defined header is added in a UDP data part. Specifically, the gradient tensor is blocked and numbered, and is transmitted through UDP, then the packet loss rate is calculated, and if the packet loss rate exceeds a threshold value, a retransmission mechanism is triggered; then aggregating received data blocks to update model parameters; and finally, performing congestion control according to the ACK receiving rate, and dynamically adjusting the sending rate. According to the method, congestion and tail delay can be effectively reduced, and the efficiency of distributed Transform large model training is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to communication technology and relates to distributed machine learning and deep learning, and in particular to a gradient bounded loss tolerant transmission optimization method for distributed Transformer large model training. Background Art

[0002] The Transformer model has achieved remarkable results in natural language processing, computer vision and other fields, but its training process requires a lot of computing resources and communication bandwidth. Distributed training is a common method to accelerate the training of the Transformer model, but it faces problems of network congestion and tail delay. Traditional distributed training usually uses the TCP protocol for communication. Although TCP's reliable transmission mechanism ensures data integrity, its congestion control algorithm may lead to insufficient bandwidth utilization in a high-bandwidth, low-latency network environment, and the retransmission mechanism will aggravate the tail delay problem.

[0003] The UDP protocol is a connectionless transport layer protocol with low latency and overhead. However, UDP itself does not provide reliability guarantees and requires the application layer to implement packet loss detection and retransmission. The gradient of Transformer model training is usually bounded, and the training process is tolerant to a certain degree of gradient loss. This feature makes it possible to use UDP for training. Summary of the invention

[0004] The purpose of the present invention is to provide a gradient bounded loss tolerant transmission optimization method for distributed Transformer training to solve the congestion and tail delay problems existing in the traditional TCP protocol in distributed training.

[0005] In order to achieve the above-mentioned purpose of the invention, the technical solution provided by the present invention is as follows:

[0006] A gradient bounded loss tolerant transmission optimization method for distributed Transformer training includes the following steps:

[0007] S1. Divide the gradient tensor into multiple data blocks and number them, use the UDP protocol to transmit the data blocks to the parameter server, and add custom header information to the data part of the UDP data packet;

[0008] S2. Establish a packet loss tolerance transmission module. The transport layer is based on the UDP protocol for transmission, and the application layer is based on the ACK mechanism. When the parameter server receives the data packet, it sends an ACK message to the working node. The ACK carries the sequence number of the corresponding data packet. After sending the data packet, the working node sets a timer to wait for the ACK reply. If the working node does not receive a reply within the time set by the timer, the data packet pointer is inserted into the RS retransmission table to prepare for the subsequent retransmission strategy;

[0009] S3. Establish convergence detection and gradient retransmission module: The parameter server determines the lost data packets by receiving the gradients from the working nodes. The parameter server calculates the packet loss rate and records the information when the packet loss rate exceeds the packet loss tolerance threshold.

[0010] When all working nodes have completed the gradient transmission, the parameter server will record all nodes that exceed the packet loss tolerance threshold, and use the custom header information in step S1 to indicate to the working nodes that the type of data packet sent is a request retransmission data packet;

[0011] When a worker node receives a retransmission request message, the worker node retrieves its RS retransmission table and checks whether any packets are recorded as lost. If lost packets are found, the worker retransmits these packets to the parameter server through the packet loss tolerance transmission module;

[0012] S4. Establish a dynamic sending rate control module: while transmitting gradient data packets, the working node calculates the rate at which ACK packets are received, evaluates the current link congestion by comparing its own data packet sending rate and ACK packet receiving rate, and adjusts the sending rate based on the current link congestion until the difference between the two exceeds a set threshold. The threshold is used to evaluate whether the working node is stable.

[0013] Furthermore, the custom header includes:

[0014] The shard flag is used to indicate whether the current data block is the last shard of the gradient tensor;

[0015] Message type, used to distinguish data packets, ACK messages, and retransmission request messages.

[0016] Furthermore, the custom header structure includes:

[0017] Flag, used to distinguish different types of UDP data packets. A Flag of 0 indicates a data packet, a Flag of 1 indicates an ACK message, and a Flag of 2 indicates a request for retransmission.

[0018] Fragment flag is used to indicate whether the current data block is the last fragment of the gradient tensor, so that the server can determine whether the transmission has been completed. When Fragment flag=1, it means that the work node has completed the transmission, and its corresponding Sequence number indicates the total number of gradient fragments transmitted by this work node;

[0019] Worker id is used to distinguish which worker node the data packet received by the parameter server comes from. At this time, the Flag flag should be 0. If the Flag flag is not 0, the Worker id is invalid;

[0020] Sequence number, used for data packet numbering;

[0021] The Acknowledgment number is valid only when Flag=1 and is used by the parameter server to notify the working node that the data block has been received.

[0022] Furthermore, step S1 is to slice the gradient tensor before transmission, determine the unique number of each slice, and initialize the custom header information of the corresponding data packet according to the information including the gradient slice number and data type.

[0023] In the above scheme, in the packet loss tolerant transmission module established in step S2, each working node establishes an RS retransmission table, which includes two operations, PUT and GET. The working node will insert the pointer of the data packet that has not received ACK confirmation into the RS retransmission table through the PUT operation. At the same time, during the retransmission phase, the working node will use the GRT operation to read the pointers of all data packets in the table.

[0024] In the above scheme, step S3 establishes a convergence detection and gradient retransmission module, which specifically includes the following steps:

[0025] S301, the parameter server configures the packet loss threshold and calculates the packet loss rate of the working node after the gradient transmission is completed: the parameter server will initialize the packet loss threshold before each round of gradient transmission. During the gradient transmission process of each working node, the parameter server will send an ACK message to the working node after receiving the data; after receiving the last data packet of all working nodes, the parameter server calculates the packet loss rate of each working node;

[0026] S302, the parameter server sends a retransmission request to the working nodes that exceed the threshold, and the working nodes that receive the request retransmit the data packets in the RS retransmission table: after the parameter server calculates the packet loss rate of each working node, it sends a retransmission request message to the working nodes whose packet loss rate is greater than the packet loss threshold; for this retransmission request message, the Flag of the custom header is set to 2, which serves as a sign for the working node to determine whether it is a retransmission request. After receiving the retransmission request message, the working node retrieves its corresponding RS retransmission table and uses the packet loss tolerance transmission module established in step S2 to retransmit these data packets to the parameter server;

[0027] This step includes synchronously updating the RS retransmission table to remove previous data and recording the packet loss in the current transmission to avoid packet loss during the retransmission process.

[0028] In the above scheme, step S4 establishes a dynamic transmission rate control module, and uses the inherent relationship between the ACK packet receiving rate and the data packet sending rate to perform rate regulation to alleviate link congestion. The specific process includes:

[0029] Calculate the sending rate at time point t and ACK reception rate :

[0030]

[0031]

[0032] represents the sending rate at time point t, which represents the sending situation within a specific time window. This means that in the time interval [ ], where t represents the current time point, Indicates the length of the time window; represents the ACK reception rate at time point t, Indicates that in the time interval [ ]The number of ACKs received within ];

[0033] When the sending rate is calculated and receiving rate Then, the current link status is estimated based on the difference between the two, and dynamic rate control is performed. First, the rate difference needs to be calculated according to the following formula:

[0034] in The sending rate and receiving rate The ratio of ,in ,when The current link state is considered to be unobstructed, and the speed is increased by the following formula:

[0035]

[0036] in is the proportional factor (0<η<1), which is used to control the adjustment amplitude. It is a detection increment, which is used to tentatively increase the bandwidth when the link is unobstructed;

[0037] when When the link is congested, it means that the sending rate needs to be reduced. The reduction formula is as follows:

[0038]

[0039] because is the receiving rate of ACK packets, which must be higher than The value of is small, so at this time A negative value is equivalent to a speed reduction;

[0040] Set rate boundary conditions to prevent abnormal transmission of working nodes under extreme conditions, thereby affecting system operation. The constraints are as follows:

[0041]

[0042] in is the minimum sending rate, and in order to ensure that the sending rate of the working node will not drop to 0, Greater than 0, then This is the maximum sending rate of the working node, which needs to be configured according to the actual link conditions.

[0043] Furthermore, in the dynamic sending rate control module, The threshold is set to the interval ,when Between , then the working node i is considered to be in a stable state and there is no need to increase or decrease the speed.

[0044] Compared with the prior art, the method of the present invention has the following significant effects:

[0045] (1) Compared with traditional TCP-based distributed training, the present invention adopts the UDP protocol for gradient transmission, avoiding the additional delay caused by TCP connection establishment, congestion control and confirmation mechanism.

[0046] (2) The present invention can be easily integrated into existing mainstream distributed training frameworks (such as PyTorch DDP and TensorFlow Distributed) without the need for large-scale modifications to the framework.

[0047] (3) The present invention implements packet loss detection and retransmission mechanism at the application layer, which can be customized and optimized according to the specific model, data set and network environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 This is a flow chart of the transmission optimization method of the present invention.

[0049] Figure 2 It is a schematic diagram of the custom UDP data header format of the present invention.

[0050] Figure 3 This is a diagram of the distributed training structure of the present invention.

[0051] Figure 4 It is a schematic diagram of the transmission optimization method of the present invention. DETAILED DESCRIPTION

[0052] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments.

[0053] Combination Figure 1 The invention provides a gradient bounded loss tolerant transmission optimization method for distributed Transformer large model training, comprising the following steps:

[0054] Step S1: Before each worker node transmits gradient data to the parameter server, each worker node divides the calculated gradient tensor into multiple data blocks of equal size. At the same time, each data block is numbered sequentially from 1, and the data blocks are transmitted from the worker node to the parameter server using the UDP protocol. In order to implement custom control at the application layer, custom header information is added to the data part of each UDP data packet, such as Figure 2 shown.

[0055] Before the worker node sends a UDP data packet to the parameter server (PS), it sets the Flag to 0, indicating that the data being transmitted is a gradient. At the same time, the worker node determines whether the sent block is the last block. If not, the Fragmentflag is set to 0. If the sent block is the last, the Flag flag is set to 1, informing the parameter server that the gradient tensor of the node has been transmitted. The Worker id is set to the worker node's own number. At the same time, the Sequence number is set to the unique number of the block to facilitate the subsequent reorganization and packet loss judgment of the parameter server.

[0056] Step S2: Establishing a packet loss tolerant transmission module

[0057] Existing transmission control mechanisms usually rely on a "one-size-fits-all" reliability strategy. For example, TCP requires that all packets must be received, and even a small number of packet delays may cause transmission obstruction; UDP does not provide any reliability guarantee. However, due to the bounded loss tolerance of large model training, some packets can be discarded without affecting model accuracy, thereby alleviating the long tail problem. Therefore, the high-speed transmission capability of the UDP protocol can be utilized, and the ACK confirmation mechanism can be implemented at the application layer. Figure 3 The framework shown in the figure provides basic reliability guarantees, ensuring the accuracy of large model training while also speeding up the training.

[0058] Step S201: The working node establishes an RS retransmission table

[0059] The working node maintains an RS resend table to record the information of data packets that may be lost (data block number). If the data packet times out and no ACK is received, the memory address information of the data packet is added to the RS resend table. Considering that the RS resend table only involves insertion and deletion operations, and the time complexity of these two operations is required to be as low as possible, a single linked list is adopted. The single linked list structure supports dynamic resizing, and the time complexity of insertion and deletion operations (when the node position is known) is O(1). It is simple to implement and has high memory utilization. These features make it very suitable for RS resend tables, which are usually small in scale, require frequent insertion and deletion, and do not require extreme search performance. Compared with structures such as arrays and hash tables, single linked lists achieve a good balance between simplicity, efficiency, and memory overhead.

[0060] Step S202: During the transmission process, the working node inserts the lost data packet pointer into the RS retransmission table.

[0061] In order to efficiently transmit the gradient data from the working node to the parameter server, the present invention adopts UDP to send data packets, and uses the custom header configured in step S1 to implement bounded loss of data packets.

[0062] After successfully receiving the data packet, the parameter server will immediately send an ACK message to the corresponding worker node through the UD protocol. In the custom header of the ACK message, the Flag field is set to 1 to clearly indicate that this is an acknowledgment packet, and the Acknowledgment Number field is set to the Sequence Number of the received data packet, thereby informing the worker node that the data packet has been successfully received.

[0063] To ensure timely delivery of data packets, the working node starts a short timer (e.g., 10ms) after sending each data packet. If the corresponding ACK message is not received before the timer expires, the working node will record the information of the data packet in the RS retransmission table it maintains as a basis for subsequent possible retransmission.

[0064] It is worth noting that this method adopts a delayed retransmission strategy. Even if a data packet is not confirmed in time, the working node will not initiate retransmission immediately, but continue to send subsequent data packets. The advantage of this is that it can avoid frequent retransmissions caused by short-term network fluctuations, thereby reducing unnecessary network overhead. Only when the parameter server detects that the packet loss rate exceeds the preset threshold, it will actively send a retransmission request to the working node, triggering the working node to selectively retransmit the lost data packets recorded in the RS retransmission table.

[0065] Through this carefully designed "packet marking, ACK confirmation, and delayed retransmission" mechanism, this method fully utilizes the high-speed transmission characteristics of the UDP protocol while effectively controlling network overhead, ensuring the reliable transmission of gradient data, and improving the overall efficiency and stability of distributed Transformer large model training.

[0066] Step S3, establishing a convergence detection and gradient retransmission module:

[0067] Since UDP is used for data transmission, packet loss is inevitable. In order to avoid excessive packet loss that affects the training accuracy of large models, the parameter server needs to calculate the packet loss rate after receiving the gradient tensor from each working node. If the packet loss rate is too high, it needs to request the lost data packets from the corresponding working node.

[0068] Step S301: The parameter server configures the packet loss threshold and calculates the packet loss rate of the working node after the gradient transmission is completed:

[0069] Before each round of gradient transmission, the parameter server will initialize the packet loss threshold according to the manual settings. After that, the parameter server receives the gradient data blocks from each working node and detects whether there is data packet loss based on the data block number. After all working nodes complete a round of gradient transmission, the parameter server calculates the packet loss rate of each working node. The calculation formula is as follows:

[0070]

[0071] in Represents a working node The packet loss rate, Represents a working node The total number of packets expected to be sent, represents the total number of packets actually received by the parameter server from worker node i. The value should be the same as the Sequence number value in the last data packet of worker node i.

[0072] Step S302: The parameter server sends a retransmission request to the working nodes that exceed the threshold, and the working nodes that receive the request retransmit the data packets in the RS retransmission table:

[0073] Assume that the total gradient loss threshold for large model training is , we only need to set constraints to ensure that the packet loss rate of each working node i does not exceed , the accuracy of model training can be guaranteed, and the constraints are expressed as follows:

[0074]

[0075] When the packet loss rate of worker node i is greater than , the parameter server considers that the gradient transmission of the working node is unreliable and needs to trigger retransmission. The parameter server will send a retransmission request message (via UDP) to the working node i whose packet loss rate exceeds the limit. At this time, the parameter server will send a request retransmission data packet to the working node i. In the custom header of the request retransmission data packet, the Flag field is set to 2 to clearly indicate that this is a retransmission request from the parameter server.

[0076] Once the worker node receives the retransmission request message, it retrieves its RS retransmission table to find the data packets that were not successfully sent, and then the worker node will resend these data packets to the parameter server using the packet loss tolerance transmission module in step S2. In addition, data packet loss may still occur during the retransmission process. Therefore, the RS retransmission table must be updated synchronously to remove previous data and record the packet loss in the current transmission.

[0077] Step S4, establishing a dynamic sending rate control module:

[0078] The main purpose of this module is to allow the working node to dynamically adjust its packet sending rate according to the network conditions (mainly reflected by the ACK receiving rate) to avoid network congestion and make full use of the available bandwidth as much as possible. The core idea is:

[0079] (1) Continuously monitor the ACK receiving rate and its own sending rate.

[0080] (2) Compare these two rates and evaluate the degree of network congestion.

[0081] (3) Adaptively adjust the sending rate according to the congestion level.

[0082] The main calculation formula is as follows:

[0083]

[0084]

[0085] Indicates the sending rate at time point t. This is the average rate, which is used to indicate the sending situation within a specific time window. This means that in the time interval [ ] is the number of packets sent within t. Where t represents the current time point, Indicates the length of the time window. For example, if W = 1 second, then this interval refers to the time from ( -1) Seconds Seconds is the time within this 1 second. It represents the ACK receiving rate at time point t. This is also an average rate. Indicates that in the time interval [ ] The number of ACKs received within ].

[0086] When the sending rate is calculated and receiving rate After that, the current link status can be estimated based on the difference between the two, and dynamic rate control can be performed. First, the rate difference needs to be calculated according to the following formula:

[0087] in The sending rate and receiving rate At the same time, we need to set the difference interval ,in .when When the link is in a smooth state, the following formula can be used to speed up the link:

[0088]

[0089] in is the proportional factor (0<η<1), which controls the adjustment range. It is recommended to obtain a reasonable value. Too large a value will cause rate oscillation, while too small a value will cause the rate to converge too slowly. is the detection increment ( >0), used to tentatively increase the bandwidth when the link is unobstructed.

[0090] when If the link is congested, the sending rate needs to be reduced. The reduction formula is as follows:

[0091]

[0092] because is the receiving rate of ACK packets, which must be higher than The value of is small, so at this time A negative value is equivalent to a speed reduction. At the same time, we also need to add rate boundary conditions to prevent abnormal transmission of working nodes under extreme conditions, thereby affecting system operation. The specific constraints are as follows:

[0093]

[0094] in is the minimum sending rate, and in order to ensure that the sending rate of the working node will not drop to 0, Should be greater than 0, and secondly This is the maximum sending rate of the working node, which needs to be configured according to the actual link conditions.

[0095] At the same time, in order to make the sending rate of all working nodes converge to a stable state, in this module, The threshold is set to an interval ,when Between We can assume that the working node i is already in a stable state and there is no need to increase or decrease the speed.

[0096] This method can dynamically adjust the sending rate according to the network conditions to adapt to different network environments and load changes. Using window-based rate calculation and an adjustment formula with a proportional factor can make the rate change smoother.

Claims

1. A gradient bounded loss tolerant transmission optimization method for distributed Transformer training, characterized in that: The following steps are involved: S1. Divide the gradient tensor into multiple data blocks and number them, use the UDP protocol to transmit the data blocks to the parameter server, and add custom header information to the data part of the UDP data packet; S2. Establish a packet loss tolerance transmission module: The transport layer is based on the UDP protocol for transmission, and the application layer is based on the ACK mechanism. When the parameter server receives the data packet, it sends an ACK message to the working node. The ACK carries the sequence number of the corresponding data packet. After sending the data packet, the working node sets a timer to wait for the ACK reply. If the working node does not receive a reply within the time set by the timer, the data packet pointer is inserted into the RS retransmission table to prepare for the subsequent retransmission strategy; S3. Establish convergence detection and gradient retransmission module: The parameter server determines the lost data packets by receiving the gradients from the working nodes. The parameter server calculates the packet loss rate and records the information when the packet loss rate exceeds the packet loss tolerance threshold. When all working nodes have completed the gradient transmission, the parameter server will record all nodes that exceed the packet loss tolerance threshold, and will use the custom header information in step S1 to indicate to the working nodes that the type of data packet sent is a request retransmission data packet; when the working node receives the retransmission request message, the working node will retrieve its RS retransmission table and check whether there are any data packets recorded as lost. If lost data packets are found, the worker will retransmit these data packets to the parameter server through the packet loss tolerance transmission module; S4. Establish a dynamic sending rate control module: while transmitting gradient data packets, the working node calculates the rate at which ACK packets are received, evaluates the current link congestion by comparing its own data packet sending rate and ACK packet receiving rate, and adjusts the sending rate based on the current link congestion until the difference between the two exceeds a set threshold. The threshold is used to evaluate whether the working node is stable.

2. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 1, characterized in that: The custom header includes: The shard flag is used to indicate whether the current data block is the last shard of the gradient tensor; Message type, used to distinguish data packets, ACK messages, and retransmission request messages.

3. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 1 or 2, characterized in that: The custom header structure includes: Flag, used to distinguish different types of UDP data packets. A Flag of 0 indicates a data packet, a Flag of 1 indicates an ACK message, and a Flag of 2 indicates a request for retransmission. Fragment flag is used to indicate whether the current data block is the last fragment of the gradient tensor, so that the server can determine whether the transmission has been completed. When Fragment flag=1, it means that the work node has completed the transmission, and its corresponding Sequence number indicates the total number of gradient fragments transmitted by this work node; Worker id is used to distinguish which worker node the data packet received by the parameter server comes from. The Flag flag should be 0. If the Flag flag is not 0, the worker id is invalid. Sequence number, used for data packet numbering; The Acknowledgment number is valid only when Flag=1 and is used by the parameter server to notify the working node that the data block has been received.

4. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 1, characterized in that: Step S1 is to slice the gradient tensor before transmission, determine the unique number of each slice, and initialize the custom header information of the corresponding data packet according to the information including the gradient slice number and data type.

5. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 1, characterized in that: In the packet loss tolerant transmission module established in step S2, each working node establishes an RS retransmission table, which includes two operations, PUT and GET. The working node will insert the pointer of the data packet that has not received ACK confirmation into the RS retransmission table through the PUT operation. At the same time, during the retransmission phase, the working node will use the GRT operation to read the pointers of all data packets in the table.

6. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 1, characterized in that: Step S3 establishes a convergence detection and gradient retransmission module, which specifically includes the following steps: S301, the parameter server configures the packet loss threshold and calculates the packet loss rate of the working node after the gradient transmission is completed: the parameter server will initialize the packet loss threshold before each round of gradient transmission. During the gradient transmission process of each working node, the parameter server will send an ACK message to the working node after receiving the data; after receiving the last data packet of all working nodes, the parameter server calculates the packet loss rate of each working node; S302, the parameter server sends a retransmission request to the working nodes that exceed the threshold, and the working nodes that receive the request retransmit the data packets in the RS retransmission table: after the parameter server calculates the packet loss rate of each working node, it sends a retransmission request message to the working nodes whose packet loss rate is greater than the packet loss threshold; for this retransmission request message, the Flag of the custom header is set to 2, which serves as a sign for the working node to determine whether it is a retransmission request. After receiving the retransmission request message, the working node retrieves its corresponding RS retransmission table and uses the packet loss tolerance transmission module established in step S2 to retransmit these data packets to the parameter server; This step includes synchronously updating the RS retransmission table to remove previous data and recording the packet loss in the current transmission to avoid packet loss during the retransmission process.

7. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 1, characterized in that: Step S4 establishes a dynamic transmission rate control module, and uses the inherent relationship between the ACK packet receiving rate and the data packet sending rate to perform rate regulation to alleviate link congestion. The specific process includes: Calculate the sending rate at time point t and ACK reception rate : , , represents the sending rate at time point t, It means that in the time interval [ ] The number of packets sent within Indicates the length of the time window; represents the ACK reception rate at time point t, Indicates that in the time interval [ ]The number of ACKs received within ]; When the sending rate is calculated and receiving rate Then, the current link status is estimated based on the difference between the two, and dynamic rate control is performed. The rate difference is calculated according to the following formula: , in The sending rate and receiving rate The ratio of ,in ,when The current link state is considered to be unobstructed, and the speed is increased by the following formula: , in is the proportional factor, used to control the adjustment amplitude, It is a detection increment, which is used to tentatively increase the bandwidth when the link is unobstructed; when When the link is congested, it means that the sending rate needs to be reduced. The reduction formula is as follows: , because is the receiving rate of ACK packets, which must be higher than The value of is small, so at this time A negative value is equivalent to a speed reduction; Set rate boundary conditions to prevent abnormal transmission of working nodes under extreme conditions, thereby affecting system operation. The constraints are as follows: , in is the minimum sending rate, and in order to ensure that the sending rate of the working node will not drop to 0, Greater than 0, then This is the maximum sending rate of the working node, which needs to be configured according to the actual link conditions.

8. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 7, characterized in that: In the dynamic sending rate control module, The threshold is set to the interval ,when Between , then the working node i is considered to be in a stable state and there is no need to increase or decrease the speed.

Citation Information

Patent Citations

  • Packet loss detection system and packet loss detection method based on Quic protocol

    CN117792578A

  • Multi-path data transmission method for distributed machine learning

    CN118945112A

  • Network security situation element extraction method and system based on hybrid deep learning

    CN119030767A

  • Layered Gradient Accumulation and Modular Pipeline Parallelism for Improved Training of Machine Learning Models

    US20220383084A1