Gradient compression and communication optimization method for distributed training of large models
Patent Information
- Application Number
- CN202610614064.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-07
- Publication Date
- 2026-08-28
AI Technical Summary
[0005]有鉴于此,本发明实施例希望提供用于大模型分布式训练的梯度压缩与通信优化方法,以解决或缓解现有技术中存在的技术问题
[0007]本发明与现有技术相比,具有以下优点:本发明通过双维度实时状态感知与分块级差异化压缩 - 路由协同决策,实现了压缩策略的精细化动态适配;同时采用硬件级压缩决策与执行机制,使决策过程与反向传播计算重叠执行,无需 CPU 调度干预,大幅提升计算通信并行度,解决了现有技术依赖软件协议栈导致的计算资源闲置问题;配套闭环误差反馈补偿机制,有效抑制压缩带来的精度损失,保障模型收敛性能。
Smart Images

Figure CN122654067A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of distributed training technology for large models, and particularly to gradient compression and communication optimization methods for distributed training of large models. Background Technology
[0002] In distributed synchronous training of large-scale deep learning models, each computing node needs to go through All in each iteration step. Reduce and other ensemble communication operations synchronize gradient data. As the number of model parameters and node size increases dramatically, the amount of gradient communication in a single iteration can reach hundreds of GB. The proportion of communication time to end-to-end training time increases significantly, resulting in a large number of idle computing units and severely limiting the linear speedup of the cluster.
[0003] To alleviate the aforementioned communication bottlenecks, existing technologies generally employ gradient compression techniques, such as Top compression. k-sparseness and non-uniform quantization reduce communication latency by decreasing the amount of gradient transmission data; there are also related technologies (publication number: CN120186057A) that propose a distributed training scheme that can dynamically adjust the compression ratio according to network bandwidth and transmission latency to improve the environmental adaptability of compression strategies.
[0004] However, these methods face a common challenge in practical deployment: their compression strategies are typically globally statically configured or coarse-grained globally dynamically adjusted. They fail to perceive the dynamic evolution of gradient sparsity and amplitude distribution characteristics during training, cannot respond to real-time transmission quality differences caused by instantaneous congestion or hardware state fluctuations in RDMA links within the cluster, and do not achieve coordinated optimization of gradient compression decisions and communication routing scheduling. Furthermore, the compression decision-making and execution processes largely rely on general software protocol stacks, resulting in insufficient overlap with the hardware execution process of backpropagation computation. These factors collectively lead to the difficulty of existing methods achieving a real-time optimal balance between model accuracy constraints and maximizing communication throughput in heterogeneous network environments, and computational... The problem of communication serialization has not been effectively solved, becoming a key bottleneck restricting further improvement in the efficiency of distributed training of large models. Summary of the Invention
[0005] In view of this, embodiments of the present invention aim to provide a gradient compression and communication optimization method for distributed training of large models, so as to solve or alleviate the technical problems existing in the prior art.
[0006] The technical solution of this invention is implemented as follows: a gradient compression and communication optimization method for distributed training of large models, applied to a distributed synchronous training cluster of large models containing multiple training nodes equipped with GPUs and RDMA network interface controllers, characterized by including the following steps: S1. Each training node directly reads the GPU hardware register and NIC hardware register through a resident monitoring process at a preset short cycle in microseconds, collects the statistical characteristics of the gradient tensor to be sent and the real-time status parameters of the RDMA link from this node to other training nodes, and generates a fixed-length communication feature vector accordingly. S2. The sending node divides the gradient tensor to be sent in the GPU memory into multiple gradient blocks according to a preset size that matches the granularity of the NIC's scatter-collect transmission, and assigns an independent compression strategy slot to each gradient block that can be dynamically written by hardware instructions. S3. Based on the communication feature vector, the sending node, after a preset synchronization waiting window, uses a deterministic state machine with no probability model prediction or iterative optimization to match compression strategies and communication routes for each gradient block, and converts the decision results into hardware compression instructions and writes them into the corresponding compression strategy slots. Specifically, for gradient blocks corresponding to RDMA links determined to be in a congested state, the state machine decides to use a preset highest compression ratio combination compression instruction and triggers the NIC hardware to switch the transmission path of that gradient block to a backup routing table entry directly maintained by the NIC hardware. For gradient blocks determined to have good link quality, the system decides to match compression instructions with the corresponding compression ratio based on the gradient magnitude distribution characteristics of that gradient block. The hardware compression instructions are executed in a GPU streaming multiprocessor or a dedicated data stream accelerator, without CPU scheduling, and the decision-making process overlaps with the backpropagation calculation. S4. The receiving node calls the hardware decompression unit to recover the gradient data and accumulate it into the receiving buffer according to the compression strategy slot instruction carried in the header of the received compressed data packet; and before the sending end performs the next gradient compression transmission, it subtracts the accumulated compression error value stored in the local residual buffer from the original gradient tensor to be compressed, and updates the error value in the residual buffer after the compression operation is completed; each training node synchronizes residual control messages through a dedicated RDMA virtual channel to compensate for the approximate error introduced by compression.
[0007] Compared with existing technologies, this invention has the following advantages: It achieves refined dynamic adaptation of compression strategies through dual-dimensional real-time state awareness and block-level differentiated compression-routing collaborative decision-making; simultaneously, it employs a hardware-level compression decision-making and execution mechanism, allowing the decision-making process and backpropagation computation to overlap, eliminating the need for CPU scheduling intervention and significantly improving computational and communication parallelism, thus solving the problem of idle computing resources caused by existing technologies relying on software protocol stacks; and it is equipped with a closed-loop error feedback compensation mechanism to effectively suppress the accuracy loss caused by compression and ensure model convergence performance.
[0008] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the invention will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a system structure block diagram of the present invention; Figure 3 This is a timing diagram of the present invention; Figure 4 This is the state transition diagram of the deterministic state machine of the present invention. Detailed Implementation
[0011] The present invention will be further described below through several exemplary embodiments. Those skilled in the art can make various modifications or substitutions to the following embodiments without departing from the technical principles and core concepts of the present invention, and the modified solutions will still fall within the protection scope of the present invention. The accompanying drawings and descriptions of the embodiments in this specification are for illustrative purposes only and are not intended to limit the protection scope of the present invention.
[0012] This invention provides a gradient compression and communication optimization method for distributed training of large models. It is applied to a distributed synchronous training cluster for large models, which includes multiple training nodes equipped with graphics processing units (GPUs) and network interface controllers (NICs) supporting remote direct memory access (RDMA). The method involves each training node collecting corresponding data from GPU and NIC hardware registers via a resident monitoring process to generate a communication feature vector. The sending node divides the gradient tensor to be transmitted into multiple gradient blocks and assigns an independent compression strategy slot to each gradient block. Based on the communication feature vector, the sending node uses a deterministic state machine to match compression strategies and communication routes for each gradient block, converting the decision results into hardware compression instructions and writing them into the corresponding compression strategy slots, thus achieving coordinated optimization of hardware-level compression and communication routing. The receiving node calls a hardware decompression unit to recover the gradient data according to the received compression strategy instructions, and compensates for the approximation error introduced by compression by synchronizing residual control messages through an error feedback compensation mechanism and a dedicated RDMA virtual channel.
[0013] In some embodiments of the present invention, the method is implemented through the following steps: S1. Each training node directly reads parameter data from the GPU hardware register and NIC hardware register through a resident monitoring process at a preset short cycle of microseconds, collects the statistical characteristics of the gradient tensor to be sent, as well as the real-time status parameters of the RDMA link from this node to other training nodes, and generates a fixed-length communication feature vector based on the collected data.
[0014] To adapt to the deployment requirements of training clusters with different hardware architectures, this step can be completed using the following two alternative implementation methods: In one implementation, each training node is configured with an independent hardware module. This module polls a specific memory region in the GPU and NIC that maps the values of the hardware registers at a fixed frequency, preprocesses the read data, calculates the basic statistics of the gradient values, and packages the processed data to generate a communication feature vector.
[0015] In another implementation, each training node deploys a lightweight, resident monitoring process that is granted direct user-mode access to the GPU and NIC hardware registers, allowing it to directly read the raw data from the corresponding hardware registers in microsecond-level short cycles. The read raw data is processed by a preset algorithm to calculate the histogram, skewness, kurtosis, and other statistical features of the gradient values, and combined with the real-time status parameters of the RDMA link to finally generate a fixed-length communication feature vector.
[0016] S2. The sending node divides the gradient tensor to be sent in the GPU memory into multiple gradient blocks according to a preset size that matches the granularity of the NIC's scatter-collect transmission, and assigns an independent compression strategy slot to each gradient block that can be dynamically written by hardware instructions.
[0017] In one implementation, the sending node utilizes the parallel computing capabilities of the GPU to logically divide the gradient tensor into blocks in the GPU memory, defining the block boundaries through an index or pointer array. At the same time, a dedicated storage area is reserved in the GPU memory, which is divided into multiple independent storage units. Each storage unit corresponds to a gradient block and is used to store the compression strategy of that block. The storage unit can be directly accessed and modified by GPU-specific instructions.
[0018] In another implementation, the sending node physically divides the gradient tensor to be sent in the GPU memory into a preset size that matches the granularity of the NIC's scatter-collect transmission, so that each gradient block occupies a contiguous storage space in the memory that is aligned with the address boundary of the NIC hardware transmission unit; at the same time, it allocates an independent compression strategy slot for each gradient block that can be dynamically written by hardware instructions. The slot is set in the control register mapping area of the GPU or NIC, or a dedicated memory area that can be directly accessed by hardware.
[0019] The compression strategy slots described in this embodiment are fully compatible with the hardware architecture of commercial GPUs and RDMANICs, requiring no customized modifications to the core hardware design. Those skilled in the art can implement it directly based on general commercial hardware. Specific details are as follows: 1. Storage location of compression strategy slots and hardware address mapping The compression strategy slot adopts a dual-copy architecture of "GPU main memory + NIC register mirroring" to ensure that both GPU and NIC hardware can access it with zero latency; The main slot is stored in a dedicated locked page memory area of the GPU memory. This area is allocated during the training initialization phase and is set as a PCIe peer-to-peer memory that can be directly accessed by the GPU streaming multiprocessor and NIC hardware, without the need for CPU relay. Each slot occupies a fixed 32 bytes of contiguous memory space, which is perfectly aligned with the cache line boundaries of the GPU memory and the DMA transfer address boundaries of the NIC, ensuring the highest efficiency of hardware access. The mirror copy of the slot is stored in the user-accessible control register space of the NIC hardware and is bound one by one to the scatter-gather transfer descriptor of the corresponding gradient block, which can be directly read and parsed by the NIC hardware.
[0020] 2. Hardware binding mechanism between slots and gradient blocks Each gradient block corresponds to a unique compression strategy slot, and the two are hardware-level bound through a fixed address offset relationship: the starting address of the gradient block in the GPU memory, plus a preset fixed offset, is the starting address of the corresponding slot. This offset is configured during the training initialization phase and remains fixed during the runtime phase. In the distributed-collection transmission list of the NIC hardware, the transmission descriptor corresponding to each gradient block contains the mirror register address of the corresponding slot, realizing the one-to-one binding of the transmission descriptor, gradient block, and compression strategy slot, ensuring that the NIC hardware can directly read the compression strategy of the corresponding block before transmission. Once the binding relationship is configured during the training initialization phase, it remains fixed throughout the entire process and does not require dynamic modification, thus avoiding the introduction of additional overhead.
[0021] Hardware compression command write timing and execution triggering mechanism: The writing timing of the compression strategy slot strictly matches the execution flow of the deterministic state machine: when the state machine issues the state under the S3 instruction, it writes the encoded hardware compression instruction directly into the slot of the corresponding gradient block according to the address offset relationship, and at the same time updates the mirror copy in the NIC register. The hardware compression instructions stored in the slot adopt a fixed-length encoding format, which corresponds one-to-one with the compression instruction set in claim 4 of the original solution. It includes three core fields: compression type, sparsity level, and quantization level, which can be directly parsed by both GPU and NIC hardware. The execution triggering mechanism adopts hardware-level gating: the GPU compression execution unit will only perform compression operation on the corresponding gradient block when the instruction valid bit in the slot is set to 1; after compression is completed, the hardware automatically sets the transfer ready bit in the slot to 1, triggering the NIC hardware to start DMA transfer, without the need for CPU scheduling intervention.
[0022] Synchronization mechanism between slots and GPU / NIC pipelines: The slot is equipped with a hardware-level read-write mutex lock. When the state machine writes a command, the read permission of the slot is locked; after the write is completed, the read permission is unlocked to ensure that the command read by the GPU / NIC is complete and valid, and there is no read-write conflict. The lifecycle of the slot is completely synchronized with the transmission lifecycle of the corresponding gradient block: after the transmission of the gradient block in the next iteration is completed, the hardware automatically resets all fields in the slot and waits for the instruction to be written in the next iteration, without the need for the CPU to clean up. The slots in the multi-gradient block are independent of each other, with no address space overlap. Slot reading, writing, and instruction execution in different blocks can be completely parallelized without resource conflicts, fully adapting to the parallel processing pipelines of GPUs and NICs.
[0023] S3. Based on communication feature vectors, the sending node, after a preset synchronization waiting window, uses a deterministic state machine with no probability model prediction or iterative optimization to match compression strategies and communication routes for each gradient block, and converts the decision results into hardware compression instructions written to the corresponding compression strategy slots. Specifically, for gradient blocks corresponding to RDMA links that are determined to be in a congested state, the state machine decides to use the preset highest compression ratio combination compression instructions, and triggers the NIC hardware to switch the transmission path of the gradient block to the backup routing table entry directly maintained by the NIC hardware; for gradient blocks that are determined to have good link quality, the system decides to match compression instructions with the corresponding compression ratio based on the gradient magnitude distribution characteristics of the gradient block; the hardware compression instructions are executed in GPU streaming multiprocessors or dedicated data stream accelerators, without CPU scheduling, and the decision-making process overlaps with the backpropagation calculation.
[0024] In one implementation, the transmitting node integrates a programmable gate array (FPGA) or application-specific integrated circuit (ASIC) module. This module receives the communication feature vector as input and, within a preset fixed time window, determines the compression strategy and communication route for each gradient block through an internally fixed lookup table or finite state machine logic. The decision result is encoded into a dedicated control signal, which directly drives the routing selection logic of the hardware compression unit and the NIC.
[0025] In another implementation, the sending node makes a decision based on the communication feature vector and after a preset synchronization waiting window timed by the hardware clock, through a deterministic state machine. The decision result is converted into a hardware compression instruction and directly written into the compression strategy slot of the corresponding gradient block. The hardware compression instruction is executed in a GPU streaming multiprocessor or a dedicated data stream accelerator. The entire decision-making process does not go through CPU scheduling and overlaps with the backpropagation calculation process.
[0026] The core state definition and state transition logic of a deterministic state machine: A deterministic state machine is a finite state machine embedded in the GPU hardware instruction set or FPGA logic. It involves no probabilistic prediction or iterative optimization. All state transitions are triggered by explicit hardware acquisition parameters. It defines 5 core working states, and the state transition logic is as follows: S0: State Acquisition State. After the state machine is powered on / training iteration initialization, it enters this state and continuously receives the communication feature vector data collected in step S1. After each gradient block of feature data is received, the block is marked as "decision ready". The state transition to state S1 is triggered only when the synchronization waiting window ends.
[0027] S1: Global Decision Ready State. The state machine completes the collection of communication feature vectors for all gradient blocks to be sent within the window period, and completes the global summary of the link state. There is no additional computational overhead. After entering the state, the transition to the S2 state is triggered immediately.
[0028] S2: The core execution state of the strategy matching decision state machine. Based on a preset fixed mapping rule, it matches the corresponding compression strategy and communication routing decision for each gradient block. All matching logic is a lookup-style deterministic operation without iterative calculation. After a single block decision is completed, the decision result is immediately encoded into hardware compression instructions and routing selection instructions. After all block decisions are completed, the state transitions to the S3 state.
[0029] S3: Instruction Issuance State The state machine writes the encoded hardware compression instruction into the compression strategy slot of the corresponding gradient block, and writes the routing selection instruction into the path selection register of the NIC hardware; after all instructions are written, the state transitions to the S4 state.
[0030] S4: The execution feedback state machine continuously monitors the execution status of the compression command and the link transmission status. After all gradient blocks are transmitted in the current iteration, it immediately triggers a state transition back to the S0 state and enters the acquisition process of the next iteration.
[0031] Deterministic mapping rules from communication feature vectors to compression strategies: The mapping rules are entirely based on the two types of parameters collected in step S1, without any dynamic fitting or probability prediction. All gears correspond one-to-one with the hardware compression instruction set in claim 4 of the original solution. The specific rules are as follows: Link state priority determination: Based on the congestion determination formula preset in the original scheme, the link state is divided into 3 levels. The link state priority is higher than the gradient distribution feature, as follows: Congestion status: MAC retransmission count > retransmission count threshold, corresponding to the highest compression ratio combination compression command in the original scheme; Mild congestion state: MAC retransmission count ≤ retransmission count threshold, and instantaneous congestion window size < preset window lower limit, corresponding to medium compression ratio strategy; Link in good condition: MAC retransmission count ≤ retransmission count threshold, and instantaneous congestion window size ≥ preset window lower limit value, and corresponding compression strategy is matched based on gradient distribution characteristics.
[0032] The gradient feature-compression strategy mapping rule under good link conditions is entirely based on the gradient statistical features collected in step S1, and corresponds completely to the original compression instruction set: If the proportion of zero values in the gradient block is ≥80%, match the "sparseness instruction that retains the first 10% of the magnitude"; If the gradient block zero value ratio is ≥50% and <80%, match the "4-level nonlinear quantization instruction"; If the proportion of zero values in the gradient block is less than 50%, and the absolute value of gradient skewness is greater than 2 and kurtosis is greater than 5 (the gradient distribution is highly concentrated), match the "combined compression instruction of retaining the first 25% of the amplitude and nonlinear quantization of 8 levels"; If the proportion of zero values in the gradient block is less than 50%, and the gradient magnitude is evenly distributed with an average magnitude greater than the preset magnitude threshold, then the "no compression instruction" or "16-level nonlinear quantization low compression ratio instruction" is matched.
[0033] Deterministic mapping rules from communication feature vectors to routing decisions Routing decisions are based solely on link congestion levels, with no CPU intervention throughout the process. The specific rules are as follows: When the link is determined to be in a congested state, the state machine immediately triggers the NIC hardware to switch to the pre-configured backup route path; When a link is determined to be in a state of mild congestion, the state machine maintains the current main route path, only adjusts the compression strategy of the corresponding gradient block, and does not trigger route switching; When the link is determined to be in good condition, the state machine maintains the current main route path and does not trigger a route switch.
[0034] The computation and communication overlap execution mechanism described in this embodiment is fully compatible with the backpropagation computation process and GPU CUDA scheduling architecture of mainstream deep learning frameworks. It does not require modification of the underlying core logic of the framework and can be directly implemented by those skilled in the art. The specific details are as follows: 1. Core mechanism for dual-stream scheduling based on GPU multi-stream This solution employs an asynchronous stream scheduling mechanism supported by GPU hardware, creating two independent CUDA streams for each GPU device during the training initialization phase: Main computation flow: Dedicated to performing backpropagation gradient calculation tasks, corresponding to the backpropagation calculation process of the original scheme; Communication optimization flow: Dedicated to performing gradient statistical feature acquisition, compression strategy decision-making, gradient compression, and data transmission tasks, corresponding to the original solution's S1-S4 full process; Two CUDA streams can be executed in parallel on the GPU hardware, sharing GPU computing resources but without execution order dependencies. Precise synchronization is achieved through the CUDA event mechanism, without CPU scheduling intervention, realizing hardware-level parallel overlap of the entire process of backpropagation computation and communication optimization.
[0035] 2. Timing of Statistical Feature Collection During Gradient Generation Layer by Layer To address the inherent characteristics of "layer-by-layer computation during backpropagation and layer-by-layer gradient generation" in large model training, this scheme employs a layer-by-layer data collection timing design, fully achieving overlap with backpropagation computation, as detailed below: The network layers of the large model perform gradient calculations sequentially from the output layer to the input layer in the order of backpropagation. After the gradient calculation of each layer is completed, the main computation flow immediately triggers the corresponding CUDA event to notify the communication optimization flow. Upon receiving an event notification, the communication optimization flow immediately initiates the block processing and statistical feature acquisition of the gradient tensor of that layer, without waiting for the entire backpropagation calculation to complete; The acquisition process is executed in parallel within the communication optimization flow. At this time, the main computation flow is performing the backpropagation gradient calculation of the next layer network. The two are completely parallel with no time overlap or conflict, ensuring that the statistical feature acquisition of all gradient blocks is completed before the entire backpropagation calculation is finished, which is in complete agreement with the requirements of step S1 of the original scheme.
[0036] 3. Synchronization logic of backpropagation computation and compression decision-making, and data transmission. The entire process achieves precise synchronization through three hardware-level CUDA events, with no CPU scheduling latency. The specific logic is as follows: Layer gradient ready event: Triggered after the gradient calculation of a certain layer in the main computation flow is completed, notifying the communication optimization flow to start the statistical feature acquisition of that layer; Full feature ready event: Triggered after the gradient statistical features of all layers have been collected, notifying the deterministic state machine to start the global compression strategy and routing decision. The triggering time of this event is no later than the time when the entire backpropagation calculation is completed. Compression Ready Event: Triggered when the compression instruction for a certain gradient block is written to the slot and the compression is completed, it notifies the NIC hardware to start the DMA data transfer of that block without waiting for all blocks to be compressed, thus achieving pipeline overlap between compression and transmission.
[0037] 4. Resource conflict avoidance mechanism Computational resource isolation: The main computing stream only uses the GPU's streaming multiprocessor cores to perform backpropagation calculations, while the communication optimization stream only uses the GPU's dedicated functional units (such as tensor cores and data transport units) to perform statistical calculations and compression operations. The two have no conflict in hardware resource usage and will not cause a performance degradation of the main computing stream. Memory access isolation: The main storage area of the gradient tensor is set to a copy-on-write mechanism. While the main computation stream writes gradient data, the communication optimization stream can read the gradient blocks that have been written, without read-write conflicts. Bus bandwidth isolation: The PCIe bus bandwidth usage for gradient statistical feature acquisition and compression instruction writing is controlled within a preset threshold, so as not to affect the GPU memory access bandwidth for backpropagation calculation and ensure the priority of the main computing task.
[0038] S4. The receiving node calls the hardware decompression unit to recover the gradient data and accumulate it into the receiving buffer according to the compression strategy slot instruction carried in the header of the received compressed data packet. Before the sending end performs the next gradient compression transmission, it subtracts the accumulated compression error value stored in the local residual buffer from the original gradient tensor to be compressed, and updates the error value in the residual buffer after completing this compression operation. The training nodes synchronize residual control messages through a dedicated RDMA virtual channel to compensate for the approximate error introduced by compression.
[0039] In one implementation, the receiving node is equipped with a hardware decompression module. This module can identify the compression identifier in the data packet header and execute a preset decompression algorithm. The decompressed data is directly written to the receiving buffer. Before each compression transmission, the sending end calculates the difference between the original gradient and the accumulated compression error value in the local residual buffer using the GPU kernel function, and sends the calculation result to the compression unit. The error value in the residual buffer is updated synchronously after the compression operation is completed. The training nodes synchronize residual control messages through independent logical channels.
[0040] In another implementation, the receiving node calls the hardware decompression unit to recover the gradient data based on the compression strategy slot instruction carried in the header of the received compressed data packet, and accumulates it to the receiving buffer through direct memory access (DMA). Before the sending end performs the next gradient compression transmission, it performs error feedback compensation on the original gradient tensor to be compressed, that is, subtracts the cumulative compression error value stored in the local residual buffer. After the compression operation is completed, the error value of the residual buffer is updated synchronously. The training nodes synchronize residual control messages through a dedicated RDMA virtual channel to compensate for the approximation error introduced by compression and ensure the model training accuracy.
[0041] To achieve efficient gradient compression and communication optimization, the accuracy of the communication feature vector representation and the real-time generation directly determine the effectiveness of subsequent compression strategies and routing decisions. To further improve the execution performance of this step, this embodiment adopts the following scheme for the parameter acquisition and feature generation process in step S1: First, to accurately characterize the distribution characteristics of gradient data and provide data support for subsequent differentiated compression strategy decisions, the statistical features of the gradient tensor include the histogram of absolute gradient values, skewness, kurtosis, and the proportion of zero values, calculated in real time in GPU memory; Second, to perceive the end-to-end link transmission quality in real time and provide a basis for dynamic adjustment of routing strategies, the real-time status parameters of the RDMA link include the instantaneous congestion window size and the retransmission count of the media access control layer; Finally, to avoid introducing additional end-to-end latency during acquisition and ensure the overall training throughput, the acquisition operation of the monitoring process and the backpropagation calculation process are executed in parallel, and all statistical features are acquired before the backpropagation is completed.
[0042] Specifically, the statistical characteristics of the gradient tensor described above can comprehensively and meticulously characterize the data distribution properties of the gradient tensor: the histogram of the absolute value distribution of gradient values can present the overall distribution of gradient values, reflecting the sparsity and concentration of the gradient; skewness is used to measure the asymmetry of the distribution, indicating the positive and negative distribution tendency of the gradient values; kurtosis is used to measure the steepness and flatness of the distribution, reflecting the concentration of gradient values near the mean; the proportion of zero values directly characterizes the sparsity of the gradient and is a core indicator for adapting sparsity compression strategies. All of the above statistical characteristics are calculated in real time in GPU memory, directly utilizing the parallel processing capabilities of the GPU, avoiding data transfer overhead between the GPU and CPU, and improving computational efficiency; specifically, CUDA kernel functions can be used to perform scanning statistics immediately after the gradient tensor is generated, and the statistical results are stored in GPU registers or shared memory for direct reading by the monitoring process.
[0043] Among the real-time status parameters of an RDMA link, the instantaneous congestion window size is a core parameter for flow control in the underlying RDMA protocol, directly reflecting the degree of network congestion and the receiving end's processing capacity. The Media Access Control (MAC) retransmission count directly represents the number of times a data packet is retransmitted at the MAC layer, serving as a direct basis for determining link quality degradation or network congestion. All of these parameters are obtained directly from the NIC hardware registers by the monitoring process, ensuring the timeliness and accuracy of data acquisition and avoiding delays and overhead introduced by the operating system or driver layer. Specifically, direct access to hardware counters and status bits can be achieved through dedicated register addresses or API interfaces provided by the NIC vendor.
[0044] Furthermore, the data acquisition operation of the monitoring process is executed in parallel with the backpropagation calculation process, and all statistical features are acquired before backpropagation is completed. This parallel execution mechanism avoids the data acquisition operation from blocking or delaying the backpropagation calculation, ensuring high training throughput; at the same time, it ensures that all feature information required for decision-making has been acquired before the gradient tensor enters the compression and transmission stage, ensuring that compression and communication optimization decisions are based on the latest valid data. This parallel mechanism can be implemented through asynchronous task scheduling: while the backpropagation calculation generates gradients, the monitoring process completes the gradient statistical feature calculation in parallel through an independent GPU kernel function or GPU asynchronous streaming mechanism without interfering with the main backpropagation calculation flow; the reading of NIC hardware registers is executed through an independent thread or interrupt mechanism, ensuring complete parallelism with GPU calculation.
[0045] Based on the above implementation method, in order to further optimize the transmission efficiency of gradient block and avoid the CPU overhead and transmission delay introduced by data reassembly, this embodiment further explains the gradient block method in step S2: the gradient block is physically divided in the GPU memory according to a predefined fixed block mask, the storage address of each gradient block is aligned with the address boundary of the scatter-collection transmission unit of the NIC hardware, and the compressed gradient block can be directly sent by the NIC hardware through DMA without the CPU participating in data reassembly.
[0046] Specifically, physical partitioning refers to the contiguous storage of gradient blocks in GPU memory according to a predefined hardware-friendly layout. A predefined fixed partitioning mask serves as a pre-set partitioning rule template, used to limit the partition size and memory layout of the gradient tensor. In practice, during the training initialization phase, memory address offsets and partition size parameters are pre-calculated based on the gradient tensor structure of the large model and the granularity of the NIC hardware's spread-collection transmission, forming a partitioning mask. After backpropagation calculations are completed, the gradient data is directly written to the pre-planned GPU memory area, completing the physical partitioning and avoiding dynamic memory allocation and data reassembly at runtime.
[0047] Aligning the storage address of each gradient block with the address boundary of the NIC hardware's scatter-collect transfer unit means that the starting address of each gradient block in the GPU memory is a multiple of the address that the NIC hardware can efficiently process for DMA transfers, including but not limited to multiples of cache line size, page size, or dedicated DMA transfer unit size. When allocating GPU memory, alignment requirements can be specified using memory allocation functions of APIs such as CUDA or OpenCL to ensure that the allocated memory addresses meet the NIC's hardware alignment specifications.
[0048] Based on the above physical segmentation and address alignment design, the compressed gradient blocks can be sent directly by the NIC hardware via DMA without the need for CPU to participate in data reassembly, which greatly reduces the CPU load and avoids the cross-bus memory copying overhead between GPU memory and host memory, thus improving transmission efficiency.
[0049] Based on the above implementation, in order to ensure the hardware execution efficiency and parallel processing capability of the compression strategy and maximize the gradient processing throughput, this embodiment adopts the following scheme for the hardware compression instruction set corresponding to the compression strategy slot: In order to cover the compression requirements under different link states and gradient distribution scenarios and achieve a dynamic trade-off between compression ratio and training accuracy, the hardware compression instruction set that can be written to the compression strategy slot includes no compression instruction, multi-level amplitude sorting sparsification instruction, multi-level nonlinear quantization instruction, and combined compression instruction of first sparsification and then quantization; Based on the above multi-level instruction set design, in order to achieve block-level fine compression control and high throughput processing, the compression strategy slots corresponding to each gradient block are independent of each other, and the compression operations of different blocks can be executed in parallel in the GPU streaming multiprocessor.
[0050] Specifically, the no-compression instruction instructs the hardware not to perform compression operations on gradient blocks, directly transmitting the original gradient data. This is suitable for gradient blocks with extremely high accuracy requirements and where information loss is unacceptable, or for scenarios with sufficient network bandwidth and excellent link quality, avoiding the introduction of additional compression / decompression latency. The multi-level amplitude sorting sparsity instruction is used to retain the elements with the largest amplitudes in the gradient block, setting the remaining elements with smaller amplitudes to zero to reduce the amount of data transmitted. "Multi-level" refers to the ability to select different sparsity retention ratios according to actual needs, including but not limited to 1%, 5%, 10%, and 25% of the total number of gradient elements, to achieve a dynamic trade-off between compression ratio and accuracy loss, especially suitable for scenarios with sparse gradient distribution and a large number of near-zero values. Multi-level nonlinear quantization instructions are used to map floating-point values in gradient blocks to finite discrete integer values, reducing the storage bits of a single gradient element. "Nonlinear" means the mapping relationship can be optimized according to the gradient value distribution characteristics; for example, higher quantization precision can be configured for small-amplitude gradients, while precision can be appropriately reduced for large-amplitude gradients. "Multi-level" refers to support for multiple quantization precisions, including but not limited to 8-bit, 4-bit, and 2-bit quantization, to adapt to different precision requirements and compression needs. The combined sparsification-then-quantization compression instruction performs amplitude-sorted sparsification on the gradient blocks first, and then performs nonlinear quantization on the remaining non-zero elements after sparsification. This achieves a higher compression ratio, minimizing the amount of data transmitted while ensuring controllable precision.
[0051] Furthermore, the compression strategy slots for each gradient block are independent of each other. Different compression instructions can be configured for each block based on its own gradient magnitude distribution characteristics and the real-time network status of the corresponding link, enabling fine-grained control of the compression strategy. This avoids using a globally uniform static compression configuration for all gradient blocks, allowing for a more efficient balance between compression efficiency and model training accuracy. Simultaneously, compression operations for different blocks can be executed in parallel on GPU streaming multiprocessors. After the compression instructions are written to the slots, GPU streaming multiprocessors or dedicated data stream accelerators can simultaneously execute corresponding compression operations on multiple gradient blocks, fully utilizing the parallel computing resources of the GPU, significantly reducing compression processing time, avoiding the additional overhead introduced by CPU scheduling, and improving the overall gradient processing throughput. Based on the above implementation methods, in order to achieve dynamic and accurate adaptation of the number of sparsity retention elements and maximize compression efficiency while ensuring model convergence, this embodiment further explains the calculation method of the number of retention elements for multi-level amplitude sorting sparsity instructions: the number of gradient elements retained by multi-level amplitude sorting sparsity instructions. It is determined by the following formula pre-defined in the GPU hardware:
[0052] In the formula: This represents the total number of gradient elements contained in the current gradient block. The pre-defined sparsity retention ratio values are provided for the hardware. ; To monitor the proportion of zero values in the current gradient block collected from the GPU hardware registers, ; The reference zero value scaling constant is preset for the hardware. ; The non-negative adjustment coefficient is preset by the hardware. ; This indicates the rounding up operation.
[0053] Specifically, the number of gradient elements k is the number of gradient elements retained in the current gradient block after magnitude sorting sparsity operation. Its value directly determines the compression ratio and the degree of information loss, and is the core control parameter for gradient compression. The calculation process of the above formula is completed directly within the GPU hardware without CPU scheduling intervention, which can ensure high computational efficiency and low latency.
[0054] The technical meaning of each parameter in the formula is as follows: N is the total number of gradient elements in the current gradient block, which is the basic parameter for sparsification calculation; The hardware has a pre-set sparsity retention ratio level, which corresponds to the preset base sparsity ratio. The base level can be selected according to the overall compression requirements or model characteristics. To monitor the proportion of zero values in the current gradient block collected from the GPU hardware registers, and to characterize the sparsity of the blocks in real time; A reference zero-value scaling constant pre-set for hardware, used as a benchmark to assess the degree of deviation of the current block sparsity; A pre-set non-negative adjustment coefficient for hardware, used to adjust the zero-value ratio. Number of elements to retain The influence weights are used to achieve fine-grained control of the sparsity strategy; This is a floor function used to ensure the accuracy of the calculation result. It is a positive integer, which meets the requirements of actual engineering implementation.
[0055] In distributed large-scale model training, the timing of gradient compression and communication routing decisions directly affects the overall system performance. To enable the synchronization waiting window to adapt to the dynamic changes of RDMA network links and balance the integrity of decision information with waiting delay, this embodiment adopts the following scheme for the synchronization waiting window in step S3: the preset synchronization waiting window duration is timed by a hardware clock, and its value is dynamically determined based on the preset percentile of the historical communication delay of this node; within the synchronization waiting window, the deterministic state machine continuously collects the ready gradient blocks and their corresponding link state information, and makes a global unified decision based on all collected information at the end of the window.
[0056] The synchronization wait window is timed by a hardware clock. Compared to the operating system or CPU software timer, the hardware clock has higher timing accuracy and lower operating overhead, and can provide timing accuracy at the microsecond to nanosecond level, ensuring precise control of the synchronization wait window and avoiding the uncertainty introduced by software scheduling delay.
[0057] Furthermore, to balance the integrity of decision information with waiting delays, the duration of the synchronization waiting window is not a fixed value, but is dynamically adjusted based on the historical communication delay data between this node and other training nodes. The specific implementation logic is as follows: First, the system continuously monitors and records the RDMA communication delay between this node and the target receiving node. The measurement can be completed by measuring the round-trip time or one-way transmission delay of messages in the queue through RDMA. Secondly, the window duration is determined based on the preset percentile of historical communication delays (including but not limited to the 90th percentile and 95th percentile), so that the window duration can cover most historical communication delay scenarios, ensuring the completeness of decision information collection while avoiding the problem of excessively long windows caused by extreme delay values. Based on this, the dynamic adjustment mechanism enables the system to adaptively respond to changes in network status, extending the window when the network is congested to ensure complete information collection, and shortening the window when the network is idle to reduce waiting delay.
[0058] During the synchronous waiting window, the deterministic state machine continuously collects relevant information on all ready-to-be-compressed gradient blocks, as well as the real-time RDMA link status information of the target receiving node corresponding to each block. The gradient block information includes statistical characteristics collected in step S1, such as block size, gradient magnitude distribution, and the proportion of zero values. The link status information includes the instantaneous congestion window size read from the NIC hardware registers and the media access control layer retransmission count. At the end of the window, the deterministic state machine performs a globally unified decision based on all the information collected during the window period. It comprehensively considers the characteristics of all ready-to-be-transmitted gradient blocks, the real-time status of the target link, and their mutual influences to match the optimal compression strategy and communication route for each gradient block. This avoids global performance degradation caused by independent block decisions, achieving efficient utilization of communication resources and optimization of overall communication latency.
[0059] Based on the above implementation methods, in order to achieve accurate, real-time, and efficient determination of RDMA link congestion status and ensure the timeliness and effectiveness of compression and routing decisions, this embodiment adopts the following scheme for determining RDMA link congestion status: The criterion for determining whether an RDMA link is in a congested state is the following inequality pre-defined in the NIC hardware:
[0060] In the formula, The retransmission count value is directly read from the NIC hardware registers by the monitoring process. This value directly represents the number of packet retransmissions at the NIC's NIC level and is a direct hardware indicator for measuring link reliability and congestion. The monitoring process reads the NIC hardware registers directly at a preset short period of microseconds, enabling low-latency acquisition of real-time retransmission count values and ensuring the timeliness and accuracy of link status awareness.
[0061] The retransmission count threshold is pre-configured in the NIC hardware register. It is a configurable hardware parameter used to set the upper limit of the acceptable number of retransmissions for the RDMA link. This threshold is configured in the NIC hardware register and can support direct hardware comparison and judgment without CPU intervention, thus improving judgment efficiency.
[0062] When the above inequality holds, the deterministic state machine determines that the corresponding RDMA link is in a congested or unreliable state. This determination mechanism, based on hardware register reads and hardware threshold comparisons, features determinism, high efficiency, and low latency. After determining the congestion state, the deterministic state machine can immediately trigger the corresponding response strategy, matching the highest compression ratio combination compression instruction to the corresponding gradient block, and instructing the NIC hardware to switch the transmission path to the backup routing table entry, thereby achieving rapid avoidance of congested links.
[0063] To further improve the adaptability of compression strategies under different link states, fully utilize communication resources, and optimize overall training performance, this embodiment makes the following refined designs for compression strategies and routing management mechanisms under different scenarios: The highest compression ratio combined compression instruction specifically involves: first, performing a sparsification operation that retains the top 5% of gradient elements with the largest amplitude; then, performing a non-linear quantization operation that maps the retained non-zero gradient elements to two levels. The sparsification operation of retaining the top 5% of gradient elements with the largest absolute values in the gradient block, along with their position information, while setting the remaining 95% of gradient elements to zero, minimizes the amount of data transmitted while preserving the key gradient information that contributes most to model updates. The non-linear quantization operation that maps the retained non-zero gradient elements after sparsification to two discrete levels (such as +1 and -1) based on their signs, significantly reducing the storage bits for a single gradient element and achieving an extremely high compression ratio.
[0064] For gradient blocks with good link quality, the compression strategy is adaptively adjusted according to the gradient characteristics of the blocks: if the proportion of zero values in a block exceeds a dynamically set threshold, a quantization command with a medium compression ratio is matched, and more than 2 levels of quantization, such as 4-level or 8-level, can be used to achieve a fine representation of non-zero gradient values while ensuring the compression ratio; if the gradient amplitude of a block is evenly distributed and the amplitude is greater than a preset threshold, a low compression ratio command or no compression command is matched, and low compression ratio quantization, such as 8-level or 16-level, can be used, or the original gradient data can be transmitted directly to preserve the integrity and accuracy of gradient information to the maximum extent.
[0065] Furthermore, the backup routing table entries are directly maintained and updated by the NIC hardware based on real-time link status. This routing switching mechanism is compatible with standard RDMA protocols such as InfiniBand and RoCEv2. Routing switching only modifies the physical transmission path of packets, without changing the source / destination queue pair (QP), GID address, or packet sequence number, maintaining the RDMA reliable connection (RC) state unchanged, and not triggering the protocol stack reconnection process. The routing switching operation requires no CPU intervention, and the transmission path after switching is forwarded through intermediate nodes, with the end-to-end logical path remaining unchanged. Specifically, the control logic built into the NIC hardware continuously monitors the real-time status of each RDMA link and autonomously and dynamically maintains and updates the backup routing table entries based on the link status. When the deterministic state machine determines link congestion and triggers routing switching, the switching action is entirely completed autonomously by the NIC hardware without CPU scheduling, avoiding the latency introduced by CPU context switching and software processing, and achieving low-latency and fast switching of congested links. After routing switching, only the physical transmission path changes, and the end-to-end logical path from the sender to the receiver remains unchanged. Upper-layer applications and protocol stacks do not need to be aware of the underlying path changes, ensuring the logical integrity and order of data flow.
[0066] The routing switching mechanism described in this embodiment is fully compatible with the IBTA standard RDMA protocol specification and adapts to commercial InfiniBand / RoCE network architectures without requiring modification to the underlying RDMA protocol stack. Specific implementation details are as follows: Routing table entry generation and hardware maintenance mechanism: The backup routing table entries are pre-configured by the subnet manager (SM) of the RDMA network during the training cluster initialization phase. The SM pre-configures one primary routing path and two to four equal-cost multipath (ECMP) backup routing paths for each end-to-end RDMA queue pair (QP). All paths are legal and available paths calculated by the SM based on the network topology and conform to the RDMA protocol specification. The pre-configured primary and backup routing table entries are all written into the dedicated routing register space of the NIC hardware. This space only supports logical reading and status marking by the NIC hardware. The CPU can only configure it during the initialization phase and has no right to modify it during the running phase. This fully conforms to the original invention concept of "direct maintenance of NIC hardware without CPU intervention". The NIC hardware's built-in link monitoring logic polls the MAC layer retransmission count and link error rate parameters of all pre-configured paths at the same microsecond interval as step S1, marking the availability status of each path in real time and completing the validity maintenance of backup routing table entries without CPU involvement.
[0067] Compatibility with RDMA protocol stack: Routing only changes the physical forwarding path of RDMA packets. The end-to-end logical address of the packets (source / destination GID, queue pair number, and packet sequence number) remains completely unchanged. The upper-layer RDMA protocol stack and collection communication library do not need to be aware of the underlying path changes, and are fully compatible with the standard RDMA reliable connection (RC) service type. The routing switching command only applies to the packet output port selection logic of the NIC hardware. It does not modify the core header fields of the RDMA packet, does not trigger the state change of the RDMA queue pair, and will not cause connection interruption or reconnection. It is fully compatible with existing mainstream distributed training communication libraries such as NCCL and MPI.
[0068] Packet order preservation and lossless transmission mechanisms during path switching: Before the route switch is triggered, the NIC hardware waits for all packets currently sent to the sending queue to complete transmission and receive an acknowledgment response, ensuring that all packets before the switch arrive at the receiving end in order. The message sending queue is equipped with a hardware-level gating switch. During the route switching process, the sending queue pauses the sending of new messages and resumes sending immediately after the switch is completed, ensuring that messages before and after the switch will not be out of order. The entire handover process is completed in microseconds, without triggering the retransmission mechanism of the RDMA protocol, resulting in no data packet loss and achieving lossless path handover.
[0069] Routing switching trigger and fallback mechanism: When the deterministic state machine determines that the primary routing link is congested, it immediately writes a switching instruction to the path selection register of the NIC hardware. The NIC hardware selects the optimal available path to complete the switching based on the pre-marked available backup routing table entries, without the need for CPU intervention. After the path switch is completed, the NIC hardware continuously monitors the link status of the original main route. When the MAC retransmission count of the original main route is lower than the threshold for three consecutive collection cycles, it automatically switches back to the main route path. The switching back process also follows the above-mentioned order preservation and lossless mechanism.
[0070] Gradient compression is a lossy operation. If the approximation errors it introduces are not effectively controlled and compensated, they can lead to a decrease in model convergence speed and even affect the final accuracy of the model. The cumulative effect of errors is particularly significant in high compression ratio scenarios. To address these issues and ensure the convergence and accuracy stability of model training, this embodiment proposes the following error feedback compensation mechanism. The error feedback compensation operation performed by the sending end before each compression transmission satisfies the following pre-defined relationship in the GPU hardware:
[0071]
[0072] The two formulas above together constitute the closed-loop calculation logic for error feedback compensation in this scheme. The technical meanings of each parameter in the formulas are defined in the order of calculation execution as follows: First, the basic iteration identifier parameters: This represents the number of training iterations. Input parameters for the first step of error compensation calculation: For the first The original gradient vector of the current gradient block in the next iteration; The core correction parameters for the first step of error compensation calculation: The first one stored in the residual buffer The cumulative error vector of the next iteration; Output parameters of the first step error compensation calculation: This is the gradient vector fed into the compression unit after error compensation; Input parameters for the second step of error calculation: This is the approximate gradient vector recovered after compression and decompression; The output parameters from the second step of error calculation are used for error compensation in the next iteration: This is the error vector generated during this compression.
[0073] Specifically, the core logic of the error feedback compensation operation is as follows: before each gradient compression transmission, the accumulated error generated by the previous compression operation is subtracted from the current original gradient, and error compensation is completed during the current compression process. This mechanism ensures that even if a single compression is a lossy operation, in the long run, all gradient information can be transmitted and aggregated, maintaining training accuracy. The error feedback compensation operation is executed entirely at the GPU hardware level, without CPU intervention. The relevant computational logic is embedded in the GPU computing unit or a dedicated accelerator, enabling efficient execution and parallel overlapping with the backpropagation computation process without introducing additional latency.
[0074] The technical meanings of each parameter in the formula are as follows: The number of training iterations is used to distinguish the gradient and error vector at different iteration steps; For the first The original gradient vector of the current gradient block in the next iteration, the unprocessed complete gradient block data obtained by backpropagation, serves as the input parameter for the error compensation operation. The first one stored in the residual buffer The cumulative error vector of the next iteration stores the gradient error information that was not transmitted in the compression operation of the previous iteration, and is the core parameter for error compensation. The gradient vector is fed into the compression unit after error compensation, and the original gradient is the result of subtracting the accumulated error. This can realize the compensation of compression error in the previous iteration. Let be the approximate gradient vector recovered after compression and decompression. Gradient data recovered after compression and decompression; The error vector generated in this compression is the difference between the compensated transmitted gradient and the recovered approximate gradient. This value will be stored in the residual buffer for error compensation in the next iteration.
[0075] Based on the above implementation methods, to further improve system performance in high-speed, high-concurrency distributed training environments, ensure gradient data writing efficiency, isolation between control flow and data flow transmission, and fine-grained management of compression errors, as a feasible specific implementation method, the receiving end processing flow and residual management mechanism adopt the following scheme: After the hardware decompression unit of the receiving end parses the data packet header, it writes the decompressed gradient data into the receiving buffer through direct memory access; the residual control message and the regular gradient communication message occupy different RDMA virtual channel identifiers to avoid blocking of control flow and data flow transmission; the residual buffer maintains the cumulative compression error in gradient blocks, and the error value of each block is updated independently.
[0076] Specifically, the hardware decompression unit is a dedicated hardware module integrated inside the NIC or GPU, used to perform the decompression operation of gradient data. After decompression, the hardware decompression unit can directly write the recovered gradient data to the pre-allocated receive buffer in the GPU video memory with low latency through the DMA mechanism, avoiding unnecessary data copying between the GPU and the CPU, as well as data reorganization and scheduling operations of the CPU, thus greatly improving data processing throughput and real-time performance.
[0077] Meanwhile, by assigning different RDMA virtual channel identifiers to residual control messages and regular gradient communication messages, physical isolation between control flow and data flow can be achieved. Even if regular gradient communication messages experience transmission delays due to network congestion or high load, residual control messages can still be transmitted in real time through independent virtual channels, avoiding the blocking of control information by data flow and ensuring the real-time performance and accuracy of the error compensation mechanism.
[0078] Furthermore, the residual buffer maintains the cumulative compression error in units of gradient blocks. The error value of each block is calculated, stored, and updated independently without affecting each other. This fine-grained error management method can perform precise error compensation for gradient blocks using different compression strategies. When the sending end updates the residual buffer, it only needs to modify the error value of the corresponding block, which improves the flexibility and parallelism of error management and avoids the system complexity and performance bottleneck caused by global error management.
[0079] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in the present invention, and these should all be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A gradient compression and communication optimization method for distributed training of large models, applied to a distributed synchronous training cluster of large models containing multiple training nodes equipped with GPUs and RDMA network interface controllers, characterized in that... Includes the following steps: S1. Each training node directly reads the GPU hardware register and NIC hardware register through a resident monitoring process at a preset short cycle in microseconds, collects the statistical characteristics of the gradient tensor to be sent and the real-time status parameters of the RDMA link from this node to other training nodes, and generates a fixed-length communication feature vector accordingly. S2. The sending node divides the gradient tensor to be sent in the GPU memory into multiple gradient blocks according to a preset size that matches the granularity of the NIC's scatter-collect transmission, and assigns an independent compression strategy slot to each gradient block that can be dynamically written by hardware instructions. S3. Based on communication feature vectors, the sending node, after a preset synchronization waiting window, uses a deterministic state machine with no probability model prediction or iterative optimization to match compression strategies and communication routes for each gradient block, and converts the decision results into hardware compression instructions and writes them into the corresponding compression strategy slots. Specifically, for gradient blocks corresponding to RDMA links that are determined to be in a congested state, the state machine decides to use a preset highest compression ratio combination compression instruction and triggers the NIC hardware to switch the transmission path of the gradient block to a backup routing table entry directly maintained by the NIC hardware. For gradient blocks that are determined to have good link quality, the state machine decides to match compression instructions with the corresponding compression ratio based on the gradient magnitude distribution characteristics of the gradient block. The hardware compression instructions are executed in a GPU streaming multiprocessor or a dedicated data stream accelerator, without CPU scheduling, and the decision-making process overlaps with the backpropagation calculation. S4. The receiving node calls the hardware decompression unit to recover the gradient data and accumulate it into the receiving buffer according to the compression strategy slot instruction carried in the header of the received compressed data packet; and before the sending end performs the next gradient compression transmission, it subtracts the accumulated compression error value stored in the local residual buffer from the original gradient tensor to be compressed, and updates the error value in the residual buffer after completing this compression operation. Each training node synchronizes residual control messages through a dedicated RDMA virtual channel to compensate for the approximation error introduced by compression.
2. The gradient compression and communication optimization method for distributed training of large models according to claim 1, characterized in that, In step S1, the statistical features of the gradient tensor include the histogram of the absolute value distribution of gradient values, skewness, kurtosis, and the proportion of zero values calculated in real time in the GPU memory; the real-time status parameters of the RDMA link include the instantaneous congestion window size and the retransmission count of the media access control layer; the acquisition operation of the monitoring process and the backpropagation calculation process are executed in parallel, and all statistical features are acquired before the backpropagation is completed.
3. The gradient compression and communication optimization method for distributed training of large models according to claim 1, characterized in that, In step S2, the gradient blocks are physically divided in the GPU memory according to a predefined fixed block mask. The storage address of each gradient block is aligned with the address boundary of the scatter-collection transmission unit of the NIC hardware. The compressed gradient blocks can be sent directly by the NIC hardware via DMA without the CPU participating in data reassembly.
4. The gradient compression and communication optimization method for distributed training of large models according to claim 1, characterized in that, In step S3, the set of hardware compression instructions that can be written to the compression strategy slot includes: no compression instructions, multi-level amplitude sorting sparsification instructions, multi-level nonlinear quantization instructions, and combined compression instructions that first sparsify and then quantize; the compression strategy slots corresponding to each gradient block are independent of each other, and the compression operations of different blocks can be executed in parallel in the GPU streaming multiprocessor.
5. The gradient compression and communication optimization method for distributed training of large models according to claim 4, characterized in that, The number of gradient elements retained by the multi-level amplitude sorting sparsity instruction It is determined by the following formula pre-defined in the GPU hardware: ; In the formula: This represents the total number of gradient elements contained in the current gradient block. The pre-defined sparsity retention ratio values are provided for the hardware. ; The proportion of zero values in the current gradient block collected by the monitoring process from the GPU hardware registers. ; The reference zero value scaling constant is preset for the hardware. ; The non-negative adjustment coefficient is preset by the hardware. ; This indicates the rounding up operation.
6. The gradient compression and communication optimization method for distributed training of large models according to claim 1, characterized in that, In step S3, the time length of the preset synchronization waiting window is timed by the hardware clock, and its value is dynamically determined based on the preset percentile of the historical communication delay of this node. Within the synchronization waiting window, the deterministic state machine continuously collects the arrival of each gradient block and its corresponding link state information, and makes a global unified decision based on the collected full information at the end of the window.
7. The gradient compression and communication optimization method for distributed training of large models according to claim 1, characterized in that, In step S3, the criterion for determining whether the RDMA link is in a congested state is the following inequality pre-defined in the NIC hardware: ; In the formula, The media access control layer retransmission count value is read directly from the NIC hardware register by the monitoring process. The retransmission count threshold is pre-configured in the NIC hardware register; when this inequality is true, the deterministic state machine determines that the corresponding RDMA link is in a congested or unreliable state.
8. The gradient compression and communication optimization method for distributed training of large models according to claim 1, characterized in that, In step S3, the highest compression ratio combined compression instruction is: first, perform a sparsification operation that retains the first 5% of the largest amplitude value, and then perform a non-linear quantization operation that maps to 2 levels; for gradient blocks with good link quality, if the proportion of zero values in the block exceeds the dynamically set threshold, then match a quantization instruction with a medium compression ratio. If the block gradient magnitude distribution is uniform and the magnitude is greater than the preset threshold, then a low compression ratio instruction or no compression instruction is matched; the backup routing table entries are directly maintained and updated by the NIC hardware based on the real-time link status, the route switching operation does not require CPU intervention, the transmission path after switching is forwarded through intermediate nodes, and the end-to-end logical path remains unchanged.
9. The gradient compression and communication optimization method for distributed training of large models according to claim 1, characterized in that, In step S4, the error feedback compensation operation performed by the sending end before each compressed transmission satisfies the following pre-defined relationship in the GPU hardware: ; In the formula, To the number of training iterations, For the first The original gradient vector of the current gradient block in the next iteration The first one stored in the residual buffer The cumulative error vector of the next iteration This is the gradient vector fed into the compression unit after error compensation. This is the approximate gradient vector recovered after compression and decompression. This is the error vector generated during this compression.
10. The gradient compression and communication optimization method for distributed training of large models according to claim 1, characterized in that, In step S4, after the hardware decompression unit of the receiving end parses the data packet header, it writes the decompressed gradient data into the receiving buffer through direct memory access. The residual control message and the regular gradient communication message occupy different RDMA virtual channel identifiers to avoid blocking of control flow and data flow transmission. The residual buffer maintains the cumulative compression error in gradient blocks, and the error value of each block is updated independently.
Citation Information
Patent Citations
Distributed training method, device, system, equipment and medium
CN120186057A