Distributed AllReduce gradient aggregation method based on homomorphic compression
By using HG-Sketch data structure and index sharing technology in the decentralized AllReduce architecture, direct aggregation of compressed gradients and improved aggregation efficiency within the network are solved, and the problems of poor algorithm adaptability and error accumulation in the architecture are significantly improved.
Patent Information
- Application Number
- CN202510180817.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
AI Technical Summary
In the decentralized AllReduce architecture, the existing homomorphic gradient compression framework is difficult to adapt, resulting in poor algorithm adaptability and error accumulation problems.
Using HG-Sketch data structure and index sharing technology, the compression gradient is realized through multi-layer index tables, and the deployment strategy of programmable switches is optimized through integer linear planning models to enhance the capability of in-network aggregation.
It effectively reduces communication overhead, improves gradient aggregation efficiency, is suitable for training large-scale deep learning models, and significantly improves gradient aggregation speed and throughput.
Smart Images

Figure CN120124716A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer networks and distributed deep learning, and particularly relates to the research on realizing efficient gradient aggregation in the decentralized architecture AllReduce. Specifically, it relates to a distributed AllReduce gradient aggregation method based on homomorphic compression. Background Art
[0002] Gradient aggregation is widely used in distributed deep learning for tasks such as model training synchronization, communication optimization, and computing resource scheduling. These tasks require efficient data structures and algorithms to adapt to communication latency and hardware resource limitations (such as memory and network bandwidth).
[0003] In distributed deep learning, to support the efficient training of large-scale deep neural networks, gradient synchronization is the core and foundation. The purpose of gradient synchronization is to coordinate model updates among multiple computing nodes to ensure the consistency of the global model. However, with the continuous growth of model scale and data complexity, the high communication overhead caused by gradient synchronization has become the main bottleneck in distributed training, seriously affecting the training efficiency. Existing distributed training mainly adopts two architectures to achieve gradient synchronization: the Parameter Server (PS) architecture and the AllReduce architecture. In the PS architecture, all computing nodes perform gradient aggregation through a centralized server. Although it is easy to scale, its centralized design is prone to causing communication bottlenecks, especially in large-scale models and high-frequency synchronization scenarios. In contrast, the AllReduce architecture adopts a decentralized mechanism and uses collective communication to directly exchange and aggregate gradients among nodes, overcoming the communication bottleneck problem of the PS architecture and becoming the mainstream architecture of modern distributed training with lower latency and higher synchronization efficiency.
[0004] With the growth of the number of model parameters, the communication overhead of gradient synchronization has further increased. For this reason, various gradient compression techniques have been proposed to reduce the amount of data exchange. Among them, homomorphic gradient compression is an emerging solution. By directly performing gradient aggregation in the compressed state, it can not only reduce the communication overhead but also avoid the additional computational overhead of the decompression process. However, existing homomorphic gradient compression frameworks are mainly designed for the PS architecture, which relies on a central node to perform a single aggregation of the global gradient, thereby effectively controlling the range of quantization errors. However, in the decentralized AllReduce architecture, due to the lack of a central node for unified coordination, gradient aggregation needs to be carried out in a multi-round iterative manner among multiple nodes, which brings problems of algorithm adaptability and error accumulation. Summary of the Invention
[0005] Aiming at the defects and deficiencies of the above existing technologies, the purpose of the present invention is to provide a distributed AllReduce gradient aggregation method based on homomorphic compression, which is used to support efficient gradient synchronization in distributed deep learning. By designing the HG-Sketch data structure and optimizing the deployment strategy of programmable switches, it can support efficient gradient aggregation in distributed learning frameworks during large-scale model training, providing a solid technical foundation for reducing communication overhead and improving training speed. It is applicable to efficiently implementing gradient synchronization in distributed deep learning frameworks, significantly reducing communication overhead and enhancing gradient aggregation efficiency in environments with limited network bandwidth and hardware resources.
[0006] With the continuous growth of the scale of models and datasets, the high communication overhead of gradient exchange has become the main bottleneck in distributed training. Although existing homomorphic compression frameworks can effectively reduce communication overhead, they rely on centralized architectures and are difficult to adapt to the mainstream decentralized AllReduce architecture. Therefore, the present invention uses HG-Sketch technology to directly aggregate compressed gradients in the network through a multi-layer index table, thereby eliminating additional computational overhead. In addition, by adopting index sharing technology, the memory usage of programmable switches is significantly optimized, and the deployment strategy of switches is optimized through an integer linear programming (ILP) model to further enhance the in-network aggregation ability. This method provides a high-throughput and low-communication-overhead gradient synchronization solution for supporting distributed deep learning, providing strong support for accelerating the training efficiency of large-scale deep learning models.
[0007] The present invention effectively solves the application problem of homomorphic compression methods in the distributed AllReduce architecture by constructing the HG-Sketch (multi-layer index table) structure and combining index sharing technology and programmable switch deployment optimization technology. First, the HG-Sketch structure adopts the design of a multi-layer index table and does not need to rely on a central node to aggregate the quantization range. Nodes can independently complete the compression and aggregation operations of gradients with Sketch, thus adapting to the decentralized AllReduce architecture. Second, the characteristics of HG-Sketch effectively suppress the gradual accumulation of errors during the iterative process of gradient aggregation through decentralized storage and calculation. To further optimize resource usage, the index sharing technology enables multiple gradient layers to share a unified index table, reducing the memory occupancy of the multi-layer index table on programmable switches. At the same time, the programmable switch deployment optimization technology realizes low latency and high throughput of gradient aggregation by determining the optimal aggregation position of switches and making full use of network resources.
[0008] The technical solutions specifically adopted by the present invention to solve its technical problems are as follows:
[0009] A Distributed AllReduce Gradient Aggregation Method Based on Homomorphic Compression: By means of the HG-Sketch data structure and an optimized deployment strategy of programmable switches, it supports gradient aggregation in large-scale model training for a distributed learning framework; the HG-Sketch data structure realizes in-network direct aggregation of compressed gradients through a multi-layer index table to eliminate additional computational overhead; the optimized deployment strategy of programmable switches optimizes the memory usage of programmable switches by adopting index sharing technology and optimizes the deployment strategy of switches through an integer linear programming model to enhance the in-network aggregation ability.
[0010] Furthermore, each computing node generates local gradient data, and after data compression and format conversion, it is mapped to the HG-Sketch data structure; in the HG-Sketch stage, the data is compressed and stored through a multi-layer index table and a Count-Sketch table, realizing decentralized storage and quantization aggregation of gradients; in the index sharing stage, multiple gradient layers share the same index table, and the memory occupancy of the multi-layer index table on the programmable switch is reduced through index sharing technology; in the deployment optimization stage, an optimal strategy for the switch deployment location is determined through an optimization model based on integer linear programming to maximize the communication efficiency of in-network gradient aggregation.
[0011] Furthermore, the HG-Sketch data structure is constructed by multiple groups of independent hash functions to build a multi-layer index table and a Count-Sketch table; each computing node maps the gradient value into the multi-layer index table through data compression, and each layer of the index table stores independent gradient elements. The index table is used to record the position where the gradient is mapped to the Count-Sketch table to support direct addition operations of gradients in the compressed state; each gradient element is hash-mapped through multiple groups of independent hash functions, and the mapping results are stored in the index table; the index positions calculated by each hash function are recorded and weighted and accumulated in the Count-Sketch table.
[0012] Furthermore, the index sharing technology specifically means that multiple gradient layers share the same index table, and the data mapping results of all gradient layers are stored in the same index table. The capacity of the index table is determined by defining it as the number of indexes required by the largest gradient layer.
[0013] Furthermore, the optimization of the switch deployment strategy through the integer linear programming model is specifically as follows:
[0014] Model the network topology structure, and describe the connection relationship between nodes and switches in matrix form;
[0015] Set constraint conditions, including the bandwidth requirements, distance limitations between nodes and switches, and the load capacity of switches;
[0016] Calculate the objective function value using an optimization solver and generate an optimal deployment plan subject to all constraints; the objective of the objective function is to minimize the total communication cost of the system, which is achieved by optimizing the connection relationship between nodes and switches.
[0017] Further, the constraints are specifically as follows: The variable y j represents whether the j-th switch is enabled, where y j =1 indicates that the switch is enabled, and y j =0 indicates that the switch is not enabled; at least one switch is enabled, that is: The variable x ij represents whether node i is connected to switch j, where x ij =1 indicates that node i is already connected to switch j, and x ij =0 indicates not connected; each node must and can only be connected to one switch, that is: The variable b ij represents the bandwidth between node i and switch j; the bandwidth between the connected node and the switch is not less than the minimum bandwidth requirement b min , that is: b ij ·x ij ≥b min , The variable d ij represents the physical distance between node i and switch j; the distance between the connected node and the switch shall not exceed the maximum communication distance d max , that is: d ij ·x ij ≤d max , In addition, the variable C j represents the capacity of switch j, which represents the maximum number of node connections that the switch can support simultaneously; it is required that the number of nodes connected to a certain switch j does not exceed its capacity limit, that is:
[0018] Further, the variable L ij represents the communication cost between node i and switch j; the variable x ij whether node i is connected to switch j, where x ij =1 indicates connection, and x ij =0 indicates not connected; the mathematical form of the objective function is: Among them, represents the sum over all n nodes, represents the sum over all m switches; by minimizing the sum of L ij ·x ij .
[0019] Furthermore, the deployment optimization of the programmable switch determines the optimal switch locations and connection relationships through an ILP model. The optimization objective is to minimize the total communication latency, and the constraints include:
[0020] 1) Each computing node needs to be connected to at least one switch;
[0021] 2) The bandwidth between the node and the switch meets the minimum requirement;
[0022] 3) The distance between the node and the switch does not exceed the maximum limit;
[0023] 4) The load capacity of the switch does not exceed the preset threshold.
[0024] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of a distributed AllReduce gradient aggregation method based on homomorphic compression as described above.
[0025] A non-transitory computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a distributed AllReduce gradient aggregation method based on homomorphic compression as described above.
[0026] Compared with the prior art, the present invention and its preferred solutions use the HG-Sketch technology to achieve homomorphic compression and aggregation of gradients, support direct aggregation in the compressed state through a multi-layer index table structure, and optimize the memory usage of the programmable switch by combining the index sharing technology, thereby effectively reducing the communication overhead and improving the gradient aggregation efficiency. By constructing a two-dimensional counting sketch and a hash index table, and calculating the index positions of each gradient element through multiple groups of independent hash functions and inserting them into the counting sketch. The hash index multiplexing of multiple gradient layers is realized by sharing the index table. The shared index table records the hash mapping information of the maximum gradient layer to avoid generating an index table independently for each gradient layer, thereby reducing the hardware storage resource requirements; the deployment strategy of the programmable switch is optimized through integer linear programming (ILP) to minimize the communication latency of gradient aggregation while meeting the constraints of bandwidth, distance, and load capacity; by gradually aggregating the compressed gradients inside the switch and recovering them through the index table, high-precision reconstruction of the data is ensured.
[0027] Among them, the HG-Sketch technology uses a counting sketch to store compressed gradient values and generates position indexes through multiple groups of hash functions, effectively suppressing error propagation during the aggregation process. The index sharing technology stores the hash mapping results of all gradient layers in a unified index table and dynamically adjusts the capacity of the table to adapt to gradient data of different scales. Through the gradual aggregation process of compressed gradients, the addition operation of gradients is directly performed in the compressed state, and the global synchronization of data can be completed without decompression. The aggregated gradients are gradually peeled and restored through the position information in the index table, ensuring that the finally restored gradient data has high accuracy and consistency.
[0028] The present invention eliminates the dependence on the central node by designing the HG-Sketch structure, supports decentralized gradient aggregation, and ensures errors during the aggregation process; through the index sharing technology, it significantly reduces the occupancy of the multi-layer index table in the programmable switch memory; through the optimized switch deployment strategy, it maximizes the communication efficiency of gradient aggregation and reduces the communication overhead in distributed training; the experimental results show that the present invention improves the gradient aggregation speed by 3.8 times and the throughput by 4.2 times, and is suitable for large-scale deep learning distributed training application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0030] Figure 1 It is a schematic diagram and flowchart of the overall structure of the embodiment of the present invention.
[0031] Figure 2 It is an example diagram of the implementation method of gradient data compression in the embodiment of the present invention.
[0032] Figure 3 It is an example diagram of the implementation method of index sharing in the embodiment of the present invention.
[0033] Figure 4 It is a construction and implementation flowchart of the overall solution of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0035] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0036] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0037] As Figure 4 shown, in the embodiment of the present invention, the provided scheme construction process includes the following steps:
[0038] 1) Construct the HG-Sketch structure: Use multiple groups of independent hash functions and multi-layer index tables to construct the HG-Sketch structure. Each index table records the position mapping of gradient data through distributed storage, thereby supporting gradient aggregation in the compressed state, eliminating the dependence on the central node to aggregate the quantization range and suppressing the accumulation problem of quantization errors, enabling the method to adapt to the decentralized AllReduce architecture.
[0039] 2) Adopt index sharing technology: Share a unified index table through multiple gradient layers to reduce the occupancy of multi-layer index tables in the programmable switch memory and optimize the usage efficiency of hardware resources.
[0040] 3) Optimize the deployment of programmable switches: Based on the integer linear programming (ILP) model, model the deployment location and node connection relationship of programmable switches to ensure maximizing the communication efficiency of gradient aggregation while meeting bandwidth and load constraints.
[0041] In an embodiment of the present invention, first, by constructing the HG-Sketch structure, an independent hash position is assigned to each gradient element, and the gradient addition operation is directly performed in the compressed state. The HG-Sketch represents the index and update of the gradient through the following formula: H(i) = h 1 (i)+…+h d (i), where H(i) is the total index value of the gradient, and h d (i) is the index value generated by the d-th hash function, which ensures the distributed storage of each gradient and suppresses the gradual accumulation of quantization errors in multiple rounds of aggregation.
[0042] In an embodiment of the present invention, to further reduce the consumption of storage resources, the present invention adopts index sharing technology, and the hash mapping results of all gradient layers are stored in a unified index table. By dynamically adjusting the table capacity to adapt to the gradient data of different gradient layers, the memory overhead applied on the programmable switch is reduced.
[0043] In an embodiment of the present invention, to optimize the communication efficiency of gradient aggregation, the present invention determines the optimal deployment strategy for the aggregation location of programmable switches through an integer linear programming model (ILP). The objective function of the ILP model is to minimize the total communication delay, while satisfying the following constraints:
[0044] 1) Each computing node needs to be connected to at least one switch;
[0045] 2) The bandwidth between the node and the switch meets the minimum requirement, and the bandwidth between the node and the switch meets the minimum requirement to ensure that the data transmission rate is not lower than the requirement of gradient synchronization;
[0046] 3) The load of the switch does not exceed the hardware preset threshold to ensure that the processing capacity of each switch can support the number of connected nodes;
[0047] 4) The distance between the node and the switch does not exceed the maximum limit to reduce latency and improve the reliability of data transmission.
[0048] In an embodiment of the present invention, the in-network aggregation process of compressed gradients is as follows: Each node first performs initial compression on local gradient data to generate a Count-Sketch table. In the compressed state, gradient aggregation is gradually performed between nodes. After the aggregation is completed, the aggregation result is stripped through an index table, and the high-precision original gradient value is restored according to the hash mapping relationship to ensure the unbiasedness and consistency of the final result. During the restoration process, by verifying the data at each index position, the risk of error accumulation is further reduced.
[0049] The present invention provides a distributed AllReduce gradient aggregation method based on homomorphic compression. By designing and introducing the HG-Sketch structure, index sharing technology, and an optimized programmable switch deployment strategy, the problems existing in traditional homomorphic compression methods in a decentralized AllReduce architecture, such as dependence on a central node, quantization error accumulation, large occupancy of programmable switch storage resources, and low communication efficiency, are solved. The following combines Figures 1 to 3 to further illustrate the specific implementation process of the present invention.
[0050] (1) Overall overview
[0051] Please refer to Figure 1 , the overall design structure of the method of the present invention is divided into three main stages: HG-Sketch construction, index sharing, and DHC deployment optimization. The figure shows the complete process of gradient aggregation from data compression to optimized deployment.
[0052] First, each computing node generates local gradient data, converts it into a format adapted to the method of the present invention through data compression, and then maps it to the HG-Sketch structure (the first stage). In the HG-Sketch stage, the data is compressed and stored through multiple levels of index tables and Count-Sketch tables, achieving decentralized storage and quantization aggregation of gradients. Next, in the index sharing stage (the second stage), multiple gradient layers share the same index table (H-Table), and the memory occupancy of the multiple levels of index tables on the programmable switch is reduced through the shared index technology. Finally, in the DHC deployment optimization stage (the third stage), an optimization model based on integer linear programming (ILP) is used to determine the optimal strategy for the switch deployment location, thereby maximizing the communication efficiency of gradient aggregation within the network and further improving the overall performance of the system.
[0053] Figure 1 It clearly shows the phased design of the present invention. The three stages complement each other from gradient compression, storage optimization to the improvement of aggregation communication efficiency, providing a complete solution for gradient synchronization in large-scale decentralized distributed training scenarios.
[0054] (2) Construction of the HG-Sketch structure
[0055] In the present invention, constructing the HG-Sketch structure is the first stage. Multiple independent hash functions are used to construct multiple levels of index tables and Count-Sketch tables. In data compression, each computing node maps the gradient value to the multiple levels of index tables, and the index table is responsible for recording the position where the gradient is mapped to the Count-Sketch, as Figure 2 shown. Through the HG-Sketch structure, it is possible to directly complete the addition operation of gradients in the compressed state, and this process does not require the participation of a central node, adapting to the decentralized AllReduce architecture.
[0056] The specific implementation is as Figure 2 shown. Each gradient element is hashed through multiple independent hash functions, and the mapping results are stored in the H-Table index table. The index positions calculated by each hash function are recorded and weighted and accumulated in the Count-Sketch table. Through this design, HG-Sketch can directly complete the aggregation addition operation of compressed gradients in the compressed state, thus avoiding the need to rely on a central node for global quantization range summarization in traditional methods. In addition, HG-Sketch stores gradient information in multiple levels of index tables, significantly reducing the storage complexity and operation overhead. Each level of index table stores independent gradient elements, supporting distributed multi-round gradient aggregation, and effectively suppressing the problem of gradual accumulation of errors in the iterative process.
[0057] (3) Index sharing technology
[0058] To further optimize the use of storage resources, the present invention proposes an index sharing technology for reducing the memory occupation of the multi-layer index table in the programmable switch. Please refer to Figure 3 , multiple gradient layers share the same H-Table, and the capacity of the H-Table is determined by defining the number of indexes required by the largest gradient layer. The index sharing technology not only effectively saves memory overhead but also improves the parallel efficiency of data processing.
[0059] Specifically, in the index sharing technology, the data mapping results of all gradient layers are stored in the same H-Table. Through the hash algorithm, the data in the H-Table is evenly distributed, avoiding the problem of uneven data distribution in the index table. In addition, the index sharing technology realizes unified capacity management by setting the capacity of the largest gradient layer (i.e., Figure 3 ), enabling the H-Table to meet the needs of the most complex gradient layer while avoiding waste of redundant resources.
[0060] (4) Deployment optimization
[0061] After the gradient data is compressed and index optimized, the communication efficiency within the network becomes the key bottleneck. The present invention further improves the communication efficiency of gradient aggregation through the switch deployment optimization technology based on integer linear programming (ILP). The specific process includes:
[0062] 1) Model the network topology structure, and describe the connection relationship between nodes and switches in matrix form to comprehensively depict the communication paths of the network;
[0063] 2) Set constraint conditions, including the bandwidth requirements, distance limitations between nodes and switches, and the load capacity of switches; The following constraint conditions are aimed at optimizing the switch deployment strategy to ensure that the communication requirements in the distributed AllReduce architecture are met while achieving efficient use of resources. Specifically: The variable y j represents whether the j-th switch is enabled, where y j =1 indicates that the switch is enabled, and y j =0 indicates that the switch is not enabled. To ensure the normal operation of the system, the constraint conditions require at least one switch to be enabled, that is: The variable x ij represents whether node i is connected to switch j, where x ij =1 indicates that node i is connected to switch j, and x ij =0 indicates not connected. To ensure the uniqueness of node connections, it is required that each node must and can only be connected to one switch, that is: In terms of communication performance, the variable b ijDenotes the bandwidth between node i and switch j. To ensure the communication rate for gradient synchronization, it is required that the bandwidth between the connected node and the switch is not less than the minimum bandwidth requirement b min , that is: b ij ·x ij ≥b min , The variable d ij Denotes the physical distance between node i and switch j. To ensure that the latency is within an acceptable range, the distance between the connected node and the switch shall not exceed the maximum communication distance d max , that is: d ij ·x ij ≤d max , In addition, the variable C j Denotes the capacity of switch j, that is, the maximum number of node connections that this switch can support simultaneously. To prevent the switch from being overloaded, it is required that the number of nodes connected to a certain switch j does not exceed its capacity limit, that is:
[0064] 3) Use an optimization solver to calculate the objective function value and generate an optimal deployment plan under the premise of satisfying all constraints. The purpose of the objective function is to minimize the total communication cost of the system, which is achieved by optimizing the connection relationship between nodes and switches. The variable L ij Denotes the communication cost between node i and switch j, which may be related to distance, bandwidth, or other transmission cost factors; the variable x ij Whether node i is connected to switch j, where x ij = 1 indicates connection, x ij = 0 indicates non-connection. The mathematical form of the objective function is: Among them, Denotes the summation over all n nodes, Denotes the summation over all m switches. By minimizing the sum of L ij ·x ij , this objective function optimizes the connection layout between nodes and switches under the premise of satisfying the system constraints, thereby reducing the communication cost in the distributed system and improving the utilization efficiency of resources.
[0065] Through the above optimization and deployment strategy, the aggregation location of the programmable switch is accurately determined, the communication delay is effectively reduced, and the gradient aggregation efficiency is significantly improved. This strategy ensures that the gradient aggregation operation is completed with high efficiency and low latency, and is applicable to large-scale distributed training scenarios. Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.
[0066] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may be any combination of one or more computer-readable media. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.
[0067] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those with ordinary skills in the field to which the present invention pertains. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to indicate relative position relationships, and when the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0068] As described above, these are only the preferred embodiments of the present invention, and the present invention is not limited to other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
[0069] This patent is not limited to the above best implementation mode. Anyone inspired by this patent can obtain various other forms of a distributed AllReduce gradient aggregation method based on homomorphic compression. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the coverage scope of this patent.
Claims
1. A distributed AllReduce gradient aggregation method based on homomorphic compression, characterized by: The HG-Sketch data structure and the deployment strategy of the optimized programmable switch are used to support the gradient aggregation of the distributed learning framework in large-scale model training; the HG-Sketch data structure realizes the direct aggregation of compressed gradients in the network through a multi-layer index table to eliminate the additional computing overhead; the deployment strategy of the optimized programmable switch optimizes the memory usage of the programmable switch by adopting the index sharing technology, and optimizes the deployment strategy of the switch by an integer linear programming model to enhance the aggregation capability within the network.
2. The distributed AllReduce gradient aggregation method based on homomorphic compression according to claim 1, characterized in that: Each computing node generates local gradient data and maps it to the HG-Sketch data structure after data compression and format conversion. In the HG-Sketch stage, data is compressed and stored through multi-layer index tables and Count-Sketch tables, realizing decentralized storage and quantitative aggregation of gradients. In the index sharing stage, multiple gradient layers share the same index table, and the shared index technology is used to reduce the memory usage of the multi-layer index table on the programmable switch; In the deployment optimization phase, the optimal strategy for switch deployment locations is determined through an optimization model based on integer linear programming to maximize the communication efficiency of gradient aggregation within the network.
3. The distributed AllReduce gradient aggregation method based on homomorphic compression according to claim 1, characterized in that: The HG-Sketch data structure is obtained by constructing a multi-layer index table and a Count-Sketch table through multiple sets of independent hash functions; each computing node maps the gradient value to the multi-layer index table through data compression, and each layer of the index table stores independent gradient elements. The index table is used to record the position where the gradient is mapped to the Count-Sketch table to support the direct completion of the gradient addition operation in a compressed state; each gradient element is hash mapped through multiple sets of independent hash functions, and the mapping results are stored in the index table; the index position calculated by each hash function is recorded in the Count-Sketch table and weightedly accumulated.
4. The distributed AllReduce gradient aggregation method based on homomorphic compression according to claim 1, characterized in that: The index sharing technology is specifically that multiple gradient layers share the same index table, the data mapping results of all gradient layers are stored in the same index table, and the capacity of the index table is determined by the number of indexes required by the largest gradient layer.
5. The distributed AllReduce gradient aggregation method based on homomorphic compression according to claim 1, characterized in that: The deployment strategy of optimizing the switch through the integer linear programming model is specifically as follows: Model the network topology and describe the connection relationship between nodes and switches in matrix form; Set constraints, including bandwidth requirements between nodes and switches, distance limits, and load capacity of switches; The optimization solver is used to calculate the objective function value and generate the optimal deployment plan under the premise of satisfying all constraints; the purpose of the objective function is to minimize the total communication cost of the system, which is achieved by optimizing the connection relationship between nodes and switches.
6. The distributed AllReduce gradient aggregation method based on homomorphic compression according to claim 5, characterized in that: The constraint condition is specifically: using variable y j Indicates whether the jth switch is enabled, where y j =1 means the switch is enabled, y j =0 means the switch is not enabled; Enable at least one switch, namely: variable x ij Indicates whether node i is connected to switch j, where x ij =1 indicates that node i is connected to switch j, x ij =0 means not connected; each node must and can only be connected to one switch, that is: Using variable b ij represents the bandwidth between node i and switch j; the bandwidth between the connected node and the switch is not less than the minimum bandwidth requirement b min ,Right now: Variable d ij Represents the physical distance between node i and switch j; the distance between the connected node and the switch must not exceed the maximum communication distance d max ,Right now: In addition, the variable C j Indicates the capacity of switch j, which represents the maximum number of node connections that the switch can support simultaneously; it requires that the number of nodes connected to a switch j must not exceed its capacity limit, that is:
7. The distributed AllReduce gradient aggregation method based on homomorphic compression according to claim 6, characterized in that: With variable L ij represents the communication cost between node i and switch j; variable x ij Whether node i is connected to switch j, where x ij =1 means connection, x ij =0 means no connection; the mathematical form of the objective function is: in, It means to sum all n nodes. represents the sum of all m switches; by minimizing L ij ·x ij The sum of .
8. The distributed AllReduce gradient aggregation method based on homomorphic compression according to claim 1, characterized in that: The deployment optimization of programmable switches determines the optimal switch location and connection relationship through the ILP model. The optimization goal is to minimize the total communication delay. The constraints include: 1) Each computing node needs to be connected to at least one switch; 2) The bandwidth between nodes and switches meets the minimum requirements; 3) The distance between the node and the switch does not exceed the maximum limit; 4) The load capacity of the switch does not exceed the preset threshold.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of a distributed AllReduce gradient aggregation method based on homomorphic compression as described in any one of claims 1 to 8 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a distributed AllReduce gradient aggregation method based on homomorphic compression as described in any one of claims 1 to 8 are implemented.