Gradient Bounded Loss Tolerant Transmission Optimization Method for Distributed Transformer Training

By using the UDP protocol and the application-layer packet loss detection mechanism in the training of distributed Transformer model, the network congestion and tail delay problems caused by TCP protocol are solved, and more efficient and stable gradient transmission is achieved.

CN119996533BActive Publication Date: 2025-06-13NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510449151.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-13
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The traditional distributed Transformer model based on TCP has problems of congestion and tail latency in a high-bandwidth and low-latency network environment.

Method used

The UDP protocol is used for gradient transmission, and the packet loss detection and retransmission mechanism is implemented at the application layer. The lost packets are managed through custom header information and RS retransmission tables, and the transmission rate is dynamically adjusted to alleviate link congestion.

Benefits of technology

It effectively avoids the problem of insufficient delay and bandwidth utilization caused by TCP protocol in distributed training, improves training efficiency and stability, and is suitable for high bandwidth and low latency network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996533B_ABST
    Figure CN119996533B_ABST
Patent Text Reader

Abstract

The present invention discloses a transmission optimization method for gradient-bounded loss tolerance in distributed Transformer training. This method utilizes the low-latency feature of the UDP protocol and combines the tolerance of the Transformer large model training to gradient-bounded loss to implement packet loss rate detection, retransmission mechanism, congestion control, and reliable transmission at the application layer, and adds a custom header to the UDP data part. Specifically, the gradient tensor is divided into blocks and numbered, transmitted through UDP, and then the packet loss rate is calculated. If it exceeds the threshold, the retransmission mechanism is triggered; then the received data blocks are aggregated to update the model parameters; finally, congestion control is performed according to the ACK reception rate, and the sending rate is dynamically adjusted. The present invention can effectively reduce congestion and tail latency and improve the efficiency of distributed Transformer large model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communication technologies, relates to distributed machine learning and deep learning, and in particular relates to a transmission optimization method with bounded gradient loss tolerance for distributed Transformer large model training. Background Art

[0002] The Transformer model has achieved remarkable results in the fields of natural language processing, computer vision, etc. However, its training process requires a large amount of computing resources and communication bandwidth. Distributed training is a common method to accelerate the training of the Transformer model, but it faces network congestion and tail latency problems. Traditional distributed training usually uses the TCP protocol for communication. Although the reliable transmission mechanism of TCP ensures data integrity, in a high-bandwidth, low-latency network environment, its congestion control algorithm may lead to insufficient bandwidth utilization, and the retransmission mechanism will exacerbate the tail latency problem.

[0003] The UDP protocol is a connectionless transport layer protocol with low latency and overhead. However, UDP itself does not provide reliability guarantees, and the application layer needs to implement packet loss detection and retransmission. The gradients in Transformer model training usually have boundedness, and the training process is tolerant of a certain degree of gradient loss. This characteristic makes it possible to use UDP for training. Summary of the Invention

[0004] The purpose of the present invention is to provide a transmission optimization method with bounded gradient loss tolerance for distributed Transformer training to solve the congestion and tail latency problems existing in traditional TCP protocols in distributed training.

[0005] In order to achieve the above invention purpose, the technical solution provided by the present invention is as follows:

[0006] A transmission optimization method with bounded gradient loss tolerance for distributed Transformer training includes the following steps:

[0007] S1. Divide the gradient tensor into multiple data blocks and number them, use the UDP protocol to transmit the data blocks to the parameter server, and add custom header information to the data part of the UDP packet;

[0008] S2. Establish a packet loss tolerance transmission module. The transport layer performs transmission based on the UDP protocol, and the application layer executes based on the ACK mechanism. When the parameter server receives a data packet, it sends an ACK message to the worker node. The ACK carries the sequence number of the corresponding data packet. The worker node sets a timer after sending the data packet to wait for the ACK reply. If the worker node does not receive a reply within the time set by the timer, it inserts the pointer of the data packet into the RS retransmission table to prepare for subsequent retransmission strategies;

[0009] S3. Establish a convergence detection and gradient retransmission module: The parameter server determines the lost data packets by receiving gradients from the worker nodes. The parameter server calculates the packet loss rate and records this information when the packet loss rate exceeds the packet loss tolerance threshold;

[0010] When all worker nodes complete gradient transmission, the parameter server will record all nodes that exceed the packet loss tolerance threshold. At the same time, it will use the custom header information in step S1 to indicate to the worker nodes that the type of the sent data packet is a request for retransmission packet;

[0011] When the worker node receives a retransmission request message, the worker node will retrieve its RS retransmission table and check whether there are data packets recorded as lost. If lost data packets are found, the worker transmits these data packets to the parameter server again through the packet loss tolerance transmission module;

[0012] S4. Establish a dynamic transmission rate control module: While transmitting gradient data packets, the worker node calculates the rate of receiving ACK packets. By comparing its own data packet transmission rate and the ACK packet reception rate, it evaluates the current link congestion situation and adjusts the transmission rate based on the current link congestion situation until the difference between the two exceeds the set threshold. The threshold is used to evaluate whether the worker node is stable.

[0013] Further, the custom header includes:

[0014] Fragmentation identifier, used to indicate whether the current data block is the last fragment of the gradient tensor;

[0015] Message type, used to distinguish data packets, ACK messages, and retransmission request messages.

[0016] Further, the custom header structure includes:

[0017] Flag, used to distinguish different types of UDP data packets. Flag being 0 indicates a data packet, Flag being 1 indicates an ACK message, and Flag being 2 indicates a request for retransmission message;

[0018] Fragment flag, which is used to indicate whether the current data block is the last shard of the gradient tensor, so that the server can determine whether the transmission has ended. When Fragment flag = 1, it means that this worker node has completed the transmission, and its corresponding Sequence number represents the total number of gradient shards transmitted by this worker node;

[0019] Woker id, which is used to distinguish from which worker node the data packet received by the parameter server comes. At this time, the Flag flag should be 0. If the Flag flag is not 0, then this Woker id is invalid;

[0020] Sequence number, which is used for data packet numbering;

[0021] Acknowledgment number is only valid when Flag = 1 and is used for the parameter server to notify the worker node that the data block has been received.

[0022] Furthermore, step S1 is to shard the gradient tensor before transmission, determine the unique number of each shard, and initialize the custom header information of the corresponding data packet according to the information including the gradient shard number and data type.

[0023] In the above solution, in the packet loss tolerance transmission module established in step S2, each worker node establishes an RS retransmission table, which includes two operations: PUT and GET. The worker node will insert the pointer of the data packet that has not received the ACK confirmation into the RS retransmission table through the PUT operation. At the same time, in the retransmission stage, the worker node will use the GRT operation to read the pointers of all data packets in the table.

[0024] In the above solution, step S3 establishes a convergence detection and gradient retransmission module, which specifically includes the following steps:

[0025] S301. The parameter server configures the packet loss threshold and calculates the packet loss rate of the worker node after the gradient transmission ends: Before each round of gradient transmission, the parameter server initializes the packet loss threshold. During the gradient transmission of each worker node, after receiving the data, the parameter server will send an ACK message to the worker node; after receiving the last data packet of all worker nodes, the parameter server calculates the packet loss rate of each worker node;

[0026] S302. The parameter server sends a retransmission request to the worker nodes exceeding the threshold, and the worker nodes that receive the request retransmit the data packets in the RS retransmission table: When the parameter server calculates the packet loss rate of each worker node, it sends a retransmission request message to the worker nodes whose packet loss rate is greater than the packet loss threshold; for this retransmission request message, the Flag in the custom header is set to 2 as the flag for the worker node to determine whether it is a retransmission request. After receiving the retransmission request message, the worker node retrieves its corresponding RS retransmission table and uses the packet loss tolerance transmission module established in step S2 to retransmit these data packets to the parameter server;

[0027] This step includes synchronously updating the RS retransmission table to remove the previous data and recording the packet loss situation in the current transmission to avoid packet loss during the retransmission process.

[0028] In the above solution, step S4 establishes a dynamic transmission rate control module, and the specific process of using the internal relationship between the ACK packet reception rate and the data packet transmission rate for rate regulation to relieve link congestion includes:

[0029] Calculate the transmission rate at time point t and the ACK reception rate :

[0030]

[0031]

[0032] represents the transmission rate at time point t, indicating the transmission situation within a specific time window, then represents the number of data packets sent within the time interval , where t represents the current time point, represents the length of the time window; represents the ACK reception rate at time point t, represents the number of ACKs received within the time interval ;

[0033] When the transmission rate and the reception rate are calculated, the current link state is estimated and judged based on the size difference between the two, and dynamic rate control is performed. First, the rate difference needs to be calculated according to the following formula:

[0034] where is the ratio of the transmission rate and the reception rate , and at the same time, the difference interval is set, where When occurs, it is considered that the current link state is unobstructed, and the speed is increased through the following formula:

[0035]

[0036] where is the scaling factor (0 < η < 1), which is used to control the adjustment amplitude, is the detection increment, which is used to tentatively increase the bandwidth when the link is unobstructed;

[0037] When occurs, it indicates that the link is congested, and the transmission rate needs to be reduced. The deceleration formula is as follows:

[0038]

[0039] Since is the acceptance rate of ACK packets, which must be smaller than the value of , so at this time is negative, which is equivalent to reducing the speed;

[0040] Set the rate boundary conditions to prevent abnormal transmission of the working node under extreme conditions, thereby affecting the system operation. The constraint conditions are as follows:

[0041]

[0042] where is the minimum transmission rate. At the same time, in order to ensure that the transmission rate of the working node will not be reduced to 0, is greater than 0. Secondly, is the maximum transmission rate of the working node, which needs to be configured according to the actual link conditions.

[0043] Furthermore, in the dynamic transmission rate control module, the threshold of is set to the interval When is between the interval , it is considered that the working node i is in a stable state at this time, and there is no need to increase or decrease the speed.

[0044] Compared with the prior art, the significant effect of the method of the present invention is that:

[0045] (1) Compared with the traditional TCP-based distributed training, the present invention uses the UDP protocol for gradient transmission, avoiding the additional delay caused by the TCP connection establishment, congestion control, and acknowledgment mechanism.

[0046] (2) The present invention can be conveniently integrated into existing mainstream distributed training frameworks (such as PyTorch DDP, TensorFlow Distributed) without large-scale modification of the frameworks.

[0047] (3) The present invention implements a packet loss detection and retransmission mechanism at the application layer, which can be customized and optimized according to specific models, datasets, and network environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a flowchart of the transmission optimization method of the present invention.

[0049] Figure 2 It is a schematic diagram of the custom UDP data header format of the present invention.

[0050] Figure 3 It is a distributed training structure diagram of the present invention.

[0051] Figure 4 It is a schematic diagram of the transmission optimization method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0053] In combination with Figure 1 , a transmission optimization method for gradient-bounded loss tolerance in distributed Transformer large model training of the present invention includes the following steps:

[0054] Step S1, before each worker node transmits gradient data to the parameter server, each worker node divides the calculated gradient tensor (Gradient Tensor) into multiple data blocks of equal size. At the same time, each data block is sequentially numbered from 1, and these data blocks are transmitted from the worker node to the parameter server using the UDP protocol. In order to implement custom control at the application layer, custom header information is added to the data part of each UDP packet, such as Figure 2 shown.

[0055] Before the worker node sends a UDP packet to the parameter server (PS), it sets the Flag to 0, indicating that the transmitted data is a gradient. At the same time, the worker node determines whether the sent block is the last block. If not, it sets the Fragmentflag to 0. If the sent block is the last one, it sets the Flagment flag to 1, informing the parameter server that the gradient tensor of this node has been transmitted. The Worker id is set to the number of the worker node itself. At the same time, the Sequence number is set to the unique number of this block, which is convenient for the parameter server to perform subsequent recombination and packet loss judgment.

[0056] Step S2, establish a packet loss tolerance transmission module

[0057] Existing transmission control mechanisms usually rely on a "one-size-fits-all" reliability strategy. For example, TCP requires that all data packets must be received, and even a small number of packet delays may cause the transmission to be blocked; while UDP does not provide any reliability guarantee. However, due to the bounded loss tolerance characteristics of large model training, some data packets can be discarded without affecting the model accuracy, thus alleviating the long-tail problem. Therefore, the high-speed transmission ability of the UDP protocol can be utilized, and at the same time, an ACK confirmation mechanism is implemented at the application layer, combined with the Figure 3 framework shown to provide basic reliability guarantees, while ensuring the accuracy of large model training and also accelerating the training speed.

[0058] Step S201, the worker node establishes an RS retransmission table

[0059] The worker node maintains an RS retransmission table (Resend Table) to record information about data packets that may be lost (chunk numbers). If a data packet does not receive an ACK within the timeout period, the memory address information of the data packet is added to the RS retransmission table. Considering that the RS retransmission table only involves insertion and deletion operations, and requires the time complexity of these two operations to be as low as possible, a singly linked list form is adopted. The singly linked list structure supports dynamic size adjustment, and the time complexity of insertion and deletion operations (when the node position is known) is O(1), which is simple to implement and has high memory utilization. These characteristics make it very suitable for the RS retransmission table, which is usually small in scale, requires frequent insertion and deletion, and has no extreme requirements for search performance. Compared with structures such as arrays and hash tables, the singly linked list achieves a good balance among simplicity, efficiency, and memory overhead.

[0060] Step S202, the worker node inserts the lost data packet pointer into the RS retransmission table during transmission

[0061] To efficiently transmit gradient data from the worker node to the parameter server, the present invention uses UDP to send data packets, and at the same time utilizes the custom header configured in Step S1 to achieve bounded loss of data packets.

[0062] After successfully receiving a data packet, the parameter server immediately sends an ACK message to the corresponding worker node through the UD protocol. In the custom header of the ACK message, the Flag field is set to 1 to clearly indicate that this is an acknowledgment packet, and the Acknowledgment Number field is set to the Sequence Number of the received data packet, thereby informing the worker node that the data packet has been successfully received.

[0063] To ensure the timely delivery of data packets, after sending each data packet, the worker node starts a short timer (e.g., 10 ms). If the corresponding ACK message is not received before the timer times out, the worker node records the information of the data packet in the RS retransmission table it maintains as the basis for possible subsequent retransmissions.

[0064] It should be noted that this method adopts a delayed retransmission strategy. Even if a data packet is not confirmed in time, the worker node will not immediately initiate a retransmission but continue to send subsequent data packets. The advantage of this is that it can avoid frequent retransmissions caused by short network fluctuations, thereby reducing unnecessary network overhead. Only when the parameter server detects that the packet loss rate exceeds the preset threshold will it actively send a retransmission request to the worker node, triggering the worker node to selectively retransmit the lost data packets recorded in the RS retransmission table.

[0065] Through this carefully designed "data packet marking, ACK confirmation, delayed retransmission" mechanism, while making full use of the high-speed transmission characteristics of the UDP protocol, this method effectively controls network overhead, ensures the reliable transmission of gradient data, and improves the overall efficiency and stability of distributed Transformer large model training.

[0066] Step S3, establish a convergence detection and gradient retransmission module:

[0067] Since UDP is used for data transmission, it is inevitable that data packets will be lost. Therefore, to avoid excessive packet loss affecting the training accuracy of the large model, after the parameter server receives the gradient tensors from each worker node, it needs to calculate the packet loss rate. If the packet loss rate is too high, it needs to request the lost data packets from the corresponding worker node.

[0068] Step S301, the parameter server configures the packet loss threshold and calculates the packet loss rate of the worker node after the gradient transmission ends:

[0069] Before each round of gradient transmission, the parameter server initializes the packet loss threshold according to manual settings. Then the parameter server receives the gradient data blocks from each worker node and detects whether there are lost data packets according to the data block numbers. After all worker nodes complete a round of gradient transmission, the parameter server calculates the packet loss rate of each worker node. The calculation formula is as follows:

[0070]

[0071] Where represents the packet loss rate of the worker node , represents the total number of data packets that the worker node is expected to send, Indicates the total number of data packets actually received by the parameter server from worker node i. It is obvious that The value of should be the same as the Sequence number value in the last data packet of worker node i.

[0072] Step S302, the parameter server sends a retransmission request to worker nodes exceeding the threshold, and the worker nodes that receive the request retransmit the data packets in the RS retransmission table:

[0073] Assume that the total gradient packet loss threshold for large model training is , we only need to set constraints to ensure that the packet loss rate of each worker node i does not exceed , then the accuracy of model training can be guaranteed. The constraints are expressed as follows:

[0074]

[0075] When the packet loss rate of worker node i is greater than , the parameter server considers that the gradient transmission of this worker node is unreliable and needs to trigger retransmission. The parameter server will send a retransmission request message (via UDP) to worker node i with an excessive packet loss rate. At this time, the parameter server will send a request to retransmit the data packet to worker node i. In the custom header of the request to retransmit the data packet, the Flag field is set to 2 to clearly indicate that this is a retransmission request from the parameter server.

[0076] Once the worker node receives the retransmission request message, it retrieves the unsuccessfully sent data packets from its RS retransmission table, and then the worker node will use the packet loss tolerance transmission module in Step S2 to retransmit these data packets to the parameter server. In addition, packet loss may still occur during the retransmission process. Therefore, the RS retransmission table must be updated synchronously to remove the previous data and record the packet loss situation during the current transmission.

[0077] Step S4, establish a dynamic transmission rate control module:

[0078] The main purpose of this module is to enable the worker node to dynamically adjust its data packet transmission rate according to the network condition (mainly reflected by the ACK reception rate) to avoid network congestion and make full use of the available bandwidth as much as possible. The core idea is:

[0079] (1) Continuously monitor the ACK reception rate and its own transmission rate.

[0080] (2) Compare these two rates and evaluate the degree of network congestion.

[0081] (3) Adaptively adjust the transmission rate according to the degree of congestion.

[0082] The main calculation formula is as follows:

[0083]

[0084]

[0085] Represents the transmission rate at time point t. This is an average rate and is used to represent the transmission situation within a specific time window. Represents the number of data packets sent within the time interval . Where t represents the current time point, represents the length of the time window. For example, if W = 1 second, then this interval refers to the time from ( -1) second to second within this 1-second period. Represents the ACK reception rate at time point t. This is also an average rate. Represents the number of ACKs received within the time interval .

[0086] When the transmission rate and the reception rate are calculated, the current link state can be estimated and judged based on the difference between the two, and dynamic rate control can be performed. First, the rate difference needs to be calculated according to the following formula:

[0087] Where is the ratio of the transmission rate and the reception rate . At the same time, we need to set the difference interval , where . When , it can be considered that the current link state is in a smooth state, and the speed is increased through the following formula:

[0088]

[0089] Where is the scaling factor (0 < η < 1), which controls the adjustment amplitude. It is recommended to take a reasonable value. Too large will cause rate oscillation, and too small will cause slow rate convergence. And is the probing increment ( > 0), which is used to tentatively increase the bandwidth when the link is smooth.

[0090] When , it indicates that the link is congested, and the transmission rate needs to be reduced. The rate reduction formula is as follows:

[0091]

[0092] Since the acceptance rate of ACK packets must be smaller than the value of , so at this time is negative, which is equivalent to reducing the speed. At the same time, we also need to add rate boundary conditions to prevent abnormal transmission of worker nodes under extreme conditions, thus affecting the system operation. The specific constraint conditions are as follows:

[0093]

[0094] Among them is the minimum transmission rate. At the same time, in order to prevent the transmission rate of the worker node from dropping to 0, should be greater than 0. Secondly is the maximum transmission rate of the worker node, which needs to be configured according to the actual link conditions.

[0095] At the same time, in order to make the transmission rates of all worker nodes finally converge to a stable state, in this module, the threshold of is set to an interval . When is between the interval , we can consider that the worker node i is in a stable state at this time and does not need to increase or decrease the speed.

[0096] Through this method, the transmission rate can be dynamically adjusted according to the network conditions to adapt to different network environments and load changes. Using window-based rate calculation and adjustment formulas with scale factors can make the rate change smoother.

Claims

1. A gradient bounded loss tolerant transmission optimization method for distributed Transformer training, characterized in that: The following steps are involved: S1. Divide the gradient tensor into multiple data blocks and number them, use the UDP protocol to transmit the data blocks to the parameter server, and add custom header information to the data part of the UDP data packet; S2. Establish a packet loss tolerance transmission module: The transport layer is based on the UDP protocol for transmission, and the application layer is based on the ACK mechanism. When the parameter server receives the data packet, it sends an ACK message to the working node. The ACK carries the sequence number of the corresponding data packet. After sending the data packet, the working node sets a timer to wait for the ACK reply. If the working node does not receive a reply within the time set by the timer, the data packet pointer is inserted into the RS retransmission table to prepare for the subsequent retransmission strategy; S3. Establish convergence detection and gradient retransmission module: The parameter server determines the lost data packets by receiving the gradients from the working nodes. The parameter server calculates the packet loss rate and records the information when the packet loss rate exceeds the packet loss tolerance threshold. When all working nodes have completed the gradient transmission, the parameter server will record all nodes that exceed the packet loss tolerance threshold, and will use the custom header information in step S1 to indicate to the working nodes that the type of data packet sent is a request retransmission data packet; when the working node receives the retransmission request message, the working node will retrieve its RS retransmission table and check whether there are any data packets recorded as lost. If lost data packets are found, the worker will retransmit these data packets to the parameter server through the packet loss tolerance transmission module; S4. Establish a dynamic sending rate control module: while transmitting gradient data packets, the working node calculates the rate at which ACK packets are received, evaluates the current link congestion by comparing its own data packet sending rate and ACK packet receiving rate, and adjusts the sending rate based on the current link congestion until the difference between the two exceeds a set threshold. The threshold is used to evaluate whether the working node is stable.

2. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 1, characterized in that: The custom header includes: The shard flag is used to indicate whether the current data block is the last shard of the gradient tensor; Message type, used to distinguish data packets, ACK messages, and retransmission request messages.

3. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 1 or 2, characterized in that: The custom header structure includes: Flag, used to distinguish different types of UDP data packets. A Flag of 0 indicates a data packet, a Flag of 1 indicates an ACK message, and a Flag of 2 indicates a request for retransmission. Fragment flag is used to indicate whether the current data block is the last fragment of the gradient tensor, so that the server can determine whether the transmission has been completed. When Fragment flag=1, it means that the work node has completed the transmission, and its corresponding Sequence number indicates the total number of gradient fragments transmitted by this work node; Worker id is used to distinguish which worker node the data packet received by the parameter server comes from. The Flag flag should be 0. If the Flag flag is not 0, the worker id is invalid. Sequence number, used for data packet numbering; The Acknowledgment number is valid only when Flag=1 and is used by the parameter server to notify the working node that the data block has been received.

4. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 1, characterized in that: Step S1 is to slice the gradient tensor before transmission, determine the unique number of each slice, and initialize the custom header information of the corresponding data packet according to the information including the gradient slice number and data type.

5. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 1, characterized in that: In the packet loss tolerant transmission module established in step S2, each working node establishes an RS retransmission table, which includes two operations, PUT and GET. The working node will insert the pointer of the data packet that has not received ACK confirmation into the RS retransmission table through the PUT operation. At the same time, during the retransmission phase, the working node will use the GRT operation to read the pointers of all data packets in the table.

6. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 1, characterized in that: Step S3 establishes a convergence detection and gradient retransmission module, which specifically includes the following steps: S301, the parameter server configures the packet loss threshold and calculates the packet loss rate of the working node after the gradient transmission is completed: the parameter server will initialize the packet loss threshold before each round of gradient transmission. During the gradient transmission process of each working node, the parameter server will send an ACK message to the working node after receiving the data; after receiving the last data packet of all working nodes, the parameter server calculates the packet loss rate of each working node; S302, the parameter server sends a retransmission request to the working nodes that exceed the threshold, and the working nodes that receive the request retransmit the data packets in the RS retransmission table: after the parameter server calculates the packet loss rate of each working node, it sends a retransmission request message to the working nodes whose packet loss rate is greater than the packet loss threshold; for this retransmission request message, the Flag of the custom header is set to 2, which serves as a sign for the working node to determine whether it is a retransmission request. After receiving the retransmission request message, the working node retrieves its corresponding RS retransmission table and uses the packet loss tolerance transmission module established in step S2 to retransmit these data packets to the parameter server; This step includes synchronously updating the RS retransmission table to remove previous data and recording the packet loss in the current transmission to avoid packet loss during the retransmission process.

7. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 1, characterized in that: Step S4 establishes a dynamic transmission rate control module, and uses the inherent relationship between the ACK packet receiving rate and the data packet sending rate to perform rate regulation to alleviate link congestion. The specific process includes: Calculate the sending rate at time point t and ACK reception rate : , , represents the sending rate at time point t, It means that in the time interval [ ] The number of packets sent within Indicates the length of the time window; represents the ACK reception rate at time point t, Indicates that in the time interval [ ]The number of ACKs received within ]; When the sending rate is calculated and receiving rate Then, the current link status is estimated based on the difference between the two, and dynamic rate control is performed. The rate difference is calculated according to the following formula: , in The sending rate and receiving rate The ratio of ,in ,when The current link state is considered to be unobstructed, and the speed is increased by the following formula: , in is the proportional factor, used to control the adjustment amplitude, It is a detection increment, which is used to tentatively increase the bandwidth when the link is unobstructed; when When the link is congested, it means that the sending rate needs to be reduced. The reduction formula is as follows: , because is the receiving rate of ACK packets, which must be higher than The value of is small, so at this time A negative value is equivalent to a speed reduction; Set rate boundary conditions to prevent abnormal transmission of working nodes under extreme conditions, thereby affecting system operation. The constraints are as follows: , in is the minimum sending rate, and in order to ensure that the sending rate of the working node will not drop to 0, Greater than 0, then This is the maximum sending rate of the working node, which needs to be configured according to the actual link conditions.

8. The gradient bounded loss tolerant transmission optimization method for distributed Transformer training according to claim 7, characterized in that: In the dynamic sending rate control module, The threshold is set to the interval ,when Between , then the working node i is considered to be in a stable state and there is no need to increase or decrease the speed.

Citation Information

Patent Citations

  • Network security situation element extraction method and system based on hybrid deep learning

    CN119030767A

  • Layered Gradient Accumulation and Modular Pipeline Parallelism for Improved Training of Machine Learning Models

    US20220383084A1