Gradient polymerization method and device
By adopting a two-layer memory structure and an optimized memory management strategy in distributed learning training, the problem of inefficient cache efficiency during gradient aggregation is solved, and efficient gradient data processing and system stability are achieved.
Patent Information
- Application Number
- CN202510525769.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-18
AI Technical Summary
During distributed learning training, gradient aggregation requires frequent access to a large amount of gradient data, which is limited by the inherent access rate of storage resources, resulting in inexpensive cache efficiency.
Using a two-layer memory structure, the first memory is used to store the gradient data to be aggregated and the gradient data to be retransmitted, and the second memory is used to store all received gradient data, and optimizes the data storage and access policies through the memory management module to reduce frequent data read and write operations.
Improve cache utilization efficiency, reduce memory access conflicts, improve the computing efficiency of gradient aggregation and the fault tolerance of the system.
Smart Images

Figure CN120336208A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of electronic devices, and particularly to a gradient aggregation method and apparatus. Background Art
[0002] During the process of distributed learning training, calculating gradient aggregation requires frequent access to a large amount of gradient data. Limited by the inherent access rate of storage resources, frequent data read and write operations are difficult to execute efficiently, resulting in low cache efficiency. Summary of the Invention
[0003] In view of this, the present disclosure provides a gradient aggregation method and apparatus.
[0004] According to a first aspect of the present disclosure, there is provided a gradient aggregation method, including:
[0005] Receiving an aggregation request;
[0006] Parsing the gradient data to be aggregated in the aggregation request and storing it in a second memory;
[0007] Obtaining the gradient data to be aggregated;
[0008] Performing gradient aggregation on the gradient data to be aggregated in a first memory to obtain an aggregation result;
[0009] Returning the aggregation result;
[0010] Wherein, the first memory includes a first memory space and a second memory space, the first memory space is used to store the gradient data to be aggregated, and the second memory space is used to store the gradient data to be retransmitted.
[0011] According to an embodiment of the present disclosure, the obtaining the gradient data to be aggregated includes:
[0012] Responding to a gradient aggregation instruction, obtaining, in the first memory, the target gradient data to be aggregated indicated by the gradient aggregation instruction, where the target gradient data is used for gradient aggregation in a target round;
[0013] In the case where the target gradient data is not hit in the first memory, obtaining the target gradient data in the second memory and storing the target gradient data in the first memory.
[0014] According to an embodiment of the present disclosure, the storing the target gradient data in the first memory includes:
[0015] Determining the gradient data to be aggregated and the gradient data to be retransmitted in the target gradient data;
[0016] Store the gradient data to be aggregated in the target gradient data in the first memory space;
[0017] Store the gradient data to be retransmitted in the target gradient data in the second memory space.
[0018] According to an embodiment of the present disclosure, the method further includes:
[0019] When the storage space of the first memory is insufficient, obtain the gradient data with the lowest access frequency in the first memory;
[0020] Store the gradient data with the lowest access frequency in the second memory.
[0021] According to an embodiment of the present disclosure, the first memory further includes a doubly linked list, the doubly linked list includes a plurality of nodes, and the nodes correspond to the storage addresses of the first memory space;
[0022] The storing the gradient data with the lowest access frequency in the second memory includes:
[0023] Store the gradient data corresponding to the tail node in the doubly linked list in the second memory.
[0024] According to an embodiment of the present disclosure, after storing the gradient data to be aggregated in the target gradient data in the first memory space, the method further includes:
[0025] Determine a target node corresponding to the target gradient data in the doubly linked list;
[0026] Move the target node to the head of the doubly linked list.
[0027] According to an embodiment of the present disclosure, the method further includes:
[0028] In response to the end of the gradient aggregation instruction, clear the gradient data to be retransmitted in the second memory space.
[0029] According to an embodiment of the present disclosure, the obtaining the target gradient data to be aggregated indicated by the gradient aggregation instruction in the first memory includes:
[0030] Obtain the target gradient data to be aggregated indicated by the gradient aggregation instruction in the first memory space;
[0031] When the target gradient data is not found in the first memory, obtain the target gradient data in the second memory space.
[0032] According to an embodiment of the present disclosure, the performing gradient aggregation on the gradient data to be aggregated in the first memory to obtain an aggregation result includes:
[0033] Perform gradient aggregation on the target gradient data in the first memory space to obtain an aggregation result.
[0034] The second aspect of the present disclosure provides another gradient aggregation device, including:
[0035] A first memory and a second memory;
[0036] A network interface for receiving an aggregation request;
[0037] A parser for parsing the gradient data to be aggregated in the aggregation request and storing it in the second memory;
[0038] A memory management module for obtaining the gradient data to be aggregated, and performing gradient aggregation on the gradient data to be aggregated in the first memory to obtain an aggregation result;
[0039] The network interface is further configured to return the aggregation result;
[0040] Wherein, the first memory includes a first memory space and a second memory space, the first memory space is used to store the gradient data to be aggregated, and the second memory space is used to store the gradient data to be retransmitted.
[0041] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0042] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0043] Figure 1 Schematically shows a structural diagram of a gradient aggregation system provided by an embodiment of the present disclosure;
[0044] Figure 2 Schematically shows a structural diagram of a gradient aggregation device provided by an embodiment of the present disclosure;
[0045] Figure 3 Schematically shows a flowchart of a gradient aggregation method provided by an embodiment of the present disclosure;
[0046] Figure 4 Schematically shows a data transfer diagram during a gradient aggregation process provided by an embodiment of the present disclosure. Detailed Embodiments
[0047] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.
[0048] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0049] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0050] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0051] In the embodiments of the present disclosure, in terms of the collection, update, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the involved data (for example, including but not limited to user personal information), they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and to safeguard the security of user personal information, network security, and national security.
[0052] The embodiments of the present disclosure provide a gradient aggregation method and device. Before introducing the technical solutions provided by the embodiments of the present disclosure, the related technologies involved in the present disclosure will be described first.
[0053] Currently, during the distributed learning training process, gradient aggregation is a core operation, and its execution efficiency directly affects the overall training performance. With the continuous expansion of the scale of deep learning models, the amount of gradient data to be processed in distributed training has increased explosively. The gradient aggregation process needs to frequently access a large amount of gradient data, which is usually distributed across multiple computing nodes and needs to be transmitted through the network and summarized for calculation.
[0054] Related technologies have proposed in-network computing technologies based on programmable switches, migrating the gradient aggregation operation from computing nodes to the network level. However, programmable switches have defects such as extremely limited storage and computing resources, inability to support floating-point operations, and difficulty in being compatible with high-speed network protocol stacks, which limit their application in high-performance distributed training.
[0055] With the development of hardware technology, field-programmable gate array (FPGA) accelerators have been introduced into in-network computing systems, providing more powerful computing capabilities and more flexible storage management solutions for gradient aggregation. However, the storage space of the static random-access memory (SRAM) of FPGAs is limited and difficult to accommodate the massive gradient data generated by large-scale training models. Although the storage capacity can be extended through external memories, the access latency of external memories is significantly higher than that of SRAM, resulting in inefficient execution of frequent data read and write operations during the gradient aggregation process and low cache efficiency.
[0056] Embodiments of the present disclosure provide a gradient aggregation method and apparatus. The method includes: receiving an aggregation request; parsing the gradient data to be aggregated in the aggregation request and storing it in a second memory; obtaining the gradient data to be aggregated; performing gradient aggregation on the gradient data to be aggregated in a first memory to obtain an aggregation result; and returning the aggregation result. The first memory includes a first memory space and a second memory space, where the first memory space is used to store the gradient data to be aggregated, and the second memory space is used to store the gradient data to be retransmitted.
[0057] By implementing the gradient aggregation method and apparatus of the embodiments of the present disclosure, the first memory includes a first memory space for storing the gradient data to be aggregated and a second memory space for storing the gradient data to be retransmitted. The storage structure of the first memory can efficiently manage the gradient data to be aggregated participating in the aggregation calculation and the data that needs to be retransmitted at the same time, avoiding the impact of frequent data read and write operations in traditional methods on the aggregation efficiency. By storing gradient data with different uses in dedicated memory areas, memory access conflicts are reduced, thereby improving the cache utilization efficiency.
[0058] The following introduces a gradient aggregation system 100 provided by the embodiments of the present disclosure.
[0059] Figure 1 Shows a schematic structural diagram of the gradient aggregation system 100.
[0060] As Figure 1 shown, the gradient aggregation system 100 includes a gradient aggregation device 101, a training device 102, and a controller 103.
[0061] During the distributed learning training process, the training device 102 is responsible for generating and sending gradient data to the gradient aggregation device 101; the gradient aggregation device 101 is responsible for receiving requests, processing the gradient data, and returning the aggregation result; the controller 103 is responsible for configuring, managing, and monitoring the status of the gradient aggregation device 101 and the training device 102 to ensure the normal operation of the system.
[0062] Among them, the gradient aggregation device 101 includes a switch 111 and a processor 112. The switch 111 is responsible for identifying and forwarding network data packets, and forwarding the network packets containing gradient data to the processor 112; the processor 112 receives the data forwarded by the switch 111, performs gradient aggregation calculations, and returns the aggregation result to the training device through the switch 111. The two work together to achieve efficient reception, processing, and result transmission of gradient data.
[0063] In some embodiments, the switch 111 can be a general Ethernet switch, a programmable switch, a router, or a network controller, etc., which are network devices capable of implementing the network data packet forwarding function.
[0064] In some embodiments, the processor 112 can be a Field-Programmable Gate Array (FPGA) accelerator, or it can also be a Graphics Processing Unit (GPU), an Application-Specific Integrated Circuit (ASIC), a Smart Network Interface Card (SmartNIC), or a Data Processing Unit (DPU), etc., which are hardware accelerators capable of performing gradient aggregation operations.
[0065] Among them, the processor 111 may integrate memories with relatively high access rates, such as Static Random-Access Memory (SRAM) and On-Chip Cache. The processor 111 may also be externally connected to memories with relatively large storage capacities, such as High Bandwidth Memory (HBM) and Dynamic Random-Access Memory (DRAM).
[0066] It should be noted that the processor 112 may be integrated into the switch 111 or connected to the outside of the switch 111 through the interface 113. The embodiments of the present disclosure do not limit this.
[0067] In some embodiments, the training device 102 may be a Graphics Processing Unit (GPU) server, or may also be a computing device capable of performing machine learning training tasks, such as a Central Processing Unit (CPU) server, a Tensor Processing Unit (TPU) server, or a Data Processing Unit (DPU) server.
[0068] In some embodiments, the training device 102 may include a training framework module 121, a collective communication library module 122, an in-network computing client library module 123, a GPU card 124, and a Remote Direct Memory Access (RDMA) network card 125.
[0069] The training framework module 121 is responsible for executing the training algorithm of the machine learning model and generating gradient data that needs to be aggregated. The training framework module 121 transfers the gradient data to the collective communication library module 122 for distributed communication processing.
[0070] The collective communication library module 122 receives the gradient data transmitted by the training framework module 121 and calls the interface provided by the in-network computing client library module 123 to encapsulate and process the gradient data.
[0071] The in-network computing client library module 123 is responsible for converting the gradient data into a network packet format adapted to the gradient aggregation device 101, sending it to the gradient aggregation device 101 through the Remote Direct Memory Access (RDMA) network card 125, and receiving the aggregation result and returning it to the collective communication library module 122.
[0072] The GPU card 124 provides computing resources for the training framework module 121 and executes the computing tasks of model training.
[0073] The Remote Direct Memory Access network card 125 is responsible for efficiently transmitting gradient data and receiving aggregation results, and conducts direct network communication with the gradient aggregation device 101.
[0074] In the interaction with the gradient aggregation device 101, the training device 102 sends gradient data to the gradient aggregation device 101 through the in-network computing client library module 123 and the Remote Direct Memory Access network card 125, and receives the returned aggregation results.
[0075] In the interaction with the controller 103, the training device 102 receives the training task configuration and management instructions issued by the controller 103, and reports the training status and resource usage to the controller 103 for unified scheduling and management.
[0076] In some embodiments, the controller 103 may include a node management module 131, a topology management module 132, an aggregation tree management module 133, a resource management module 134, a switch configuration module 135, a processor configuration module 136, a log configuration module 137, and a status monitoring module 138.
[0077] The resource management module 134, as the central module of the controller 103, is responsible for coordinating the operation of each module and managing the allocation and scheduling of system resources. The topology management module 132 detects the network topology structure and provides the information to the aggregation tree management module 133, which generates an optimized gradient aggregation path based on this. The node management module 131 is responsible for identifying and managing the training device 102, collecting its resource status and reporting it to the resource management module 134.
[0078] The switch configuration module 135 and the processor configuration module 136 are respectively responsible for configuring the switch and the processor in the gradient aggregation device 101 to implement optimized data forwarding and processing logic. The log configuration module 137 collects the system operation logs to provide a basis for debugging and optimization. The status monitoring module 138 monitors the running status of the entire system in real time and notifies the relevant modules to handle it when an abnormality is found.
[0079] In the interaction with the gradient aggregation device 101, the controller 103 configures the forwarding rules of the switch through the switch configuration module 135, configures the aggregation logic and memory management policy of the processor through the processor configuration module 136, and receives the running status information of the gradient aggregation device 101 through the status monitoring module 138.
[0080] In the interaction with the training device 102, the controller 103 issues the training task configuration to the training device 102 through the node management module 131, schedules the computing resources of the training device 102 through the resource management module 134, and receives the running status information reported by the training device 102 through the status monitoring module 138.
[0081] The following introduces a gradient aggregation device 101 provided by an embodiment of the present disclosure.
[0082] Figure 2 A schematic structural diagram of the gradient aggregation device 101 is shown.
[0083] As Figure 2 shown, the gradient aggregation device 101 includes a network interface 201, a parser 202, a first memory 203, a second memory 204, a memory management module 205, and a packet encapsulator (not shown in the figure).
[0084] Among them, the network interface 201 is responsible for communicating with the external network, receiving the aggregation requests sent by each computing node in the distributed training system, and returning the aggregation results through the same interface after processing. After receiving the aggregation request, the network interface 201 transfers the data to the parser 202 for processing.
[0085] In some embodiments, the network interface 201 may be a high-speed network interface such as QSFP-DD, QSFP28, or SFP+, capable of handling high-bandwidth data transmission requirements.
[0086] The parser 202 is responsible for parsing the received aggregation request, extracting the gradient data to be aggregated therein, and initially storing the gradient data to be aggregated in the second memory 204. After the parsing is completed, the parser 202 notifies the memory management module 205 to start the data processing process.
[0087] The memory management module 205 is the core control unit of the entire gradient aggregation device 101, responsible for coordinating the data exchange between the first memory 203 and the second memory 204. It obtains the gradient data from the second memory 204 according to the aggregation requirements and loads it into the first memory 203 for efficient processing. When the space in the first memory 203 is insufficient, the memory management module 205 is also responsible for implementing a cache replacement strategy to migrate the data with low-frequency access back to the second memory 204.
[0088] The first memory 203 is a high-speed cache area, which is divided into a first memory space 207 and a second memory space 208, storing the gradient data to be aggregated and the gradient data to be retransmitted respectively. The gradient aggregation calculation is directly performed in the first memory space 207 of the first memory 203 to improve the cache speed.
[0089] The second memory 204 is a large-capacity storage area, storing all the received gradient data to make up for the insufficient capacity of the first memory 203. When the gradient data that the memory management module 205 needs to process is not in the first memory 203, it is read from the second memory 204.
[0090] The data packet encapsulator is responsible for repackaging the aggregation result generated by the memory management module 205 into the network data packet format, and then handing it over to the network interface 201 to send back to each computing node of the distributed training system.
[0091] The following introduces a gradient aggregation method provided in the embodiments of the present disclosure.
[0092] Figure 3 A flowchart of the gradient aggregation method is shown.
[0093] As Figure 3 shown, the gradient aggregation method may further include operation S301 to operation S305.
[0094] Operation S301, receiving an aggregation request;
[0095] Operation S302, parsing the gradient data to be aggregated in the aggregation request and storing it in the second memory;
[0096] Operation S303, obtaining the gradient data to be aggregated;
[0097] Operation S304, performing gradient aggregation on the gradient data to be aggregated in the first memory to obtain an aggregation result, where the first memory includes a first memory space and a second memory space, the first memory space is used to store the gradient data to be aggregated, and the second memory space is used to store the gradient data to be retransmitted;
[0098] Operation S305, returning the aggregation result.
[0099] In operation S301, the gradient aggregation device 101 receives an aggregation request from the training device 102 through the network interface 201. Among them, the aggregation request refers to a network data packet sent by the training device 102 in the distributed training system and containing the gradient data to be aggregated. In the embodiments of the present disclosure, it can be understood that after each round of training iteration, the training device 102 encapsulates the locally calculated gradient data into a network data packet in a specific format and sends it to the gradient aggregation device 101 as a request message. The aggregation request is used to instruct the gradient aggregation device to perform an aggregation operation on the gradient data generated by multiple training devices 102 so that all training devices 102 can obtain consistent model update information.
[0100] In some embodiments, the aggregation request may include: a data packet header, control information, and gradient data.
[0101] Among them, the gradient data contains the gradient values that actually need to be aggregated and is the core content of the aggregation request. These data are usually a series of numerical values representing the partial derivatives of the parameters of the deep learning model with respect to the loss function.
[0102] Among them, the data packet header is the basic part of network communication and mainly contains the following fields:
[0103] Source device identifier, which is used to identify the training device that sends the aggregation request. In a distributed training system, each training device 102 has a unique ID to distinguish different computing nodes. The gradient aggregation device 101 can track the source of gradient data through this identifier to ensure correct aggregation of data from different nodes;
[0104] Gradient aggregation device identifier, which is used to identify the gradient aggregation device that receives the request. In a complex network topology, there may be multiple gradient aggregation devices, and the gradient aggregation device identifier ensures that the data packet is correctly routed to the specified processing unit;
[0105] Data packet type identifier, which is used to indicate that it is an aggregation request data packet;
[0106] Data packet length, which is used to indicate the byte length of the data packet;
[0107] Sequence number, which is used to identify the position of the data packet in the transmission sequence to help the receiver detect lost packets and duplicate packets;
[0108] Checksum, which is used to verify whether an error occurs during the transmission of the data packet;
[0109] Specifically, when the gradient aggregation device 101 receives a data packet, it first parses the data packet header, verifies the legality and integrity of the data packet, and determines the subsequent processing flow according to the data packet type identifier. If it is found that the data packet is damaged or incomplete, it may request retransmission or discard the data packet.
[0110] Among them, the control information part contains instructions and parameters related to the gradient aggregation operation, mainly including:
[0111] Aggregation operation type: which is used to indicate what kind of aggregation operation to perform on the gradient data;
[0112] Epoch identifier, which is used to identify the epoch of the current training iteration;
[0113] Group identifier: including Worker ID and Packet ID, which is used to uniquely identify a gradient data group.
[0114] Worker ID: which is used to identify the training device that generates the gradient data.
[0115] Packet ID: which is used to identify different data packets sent by the same training device. Since the amount of gradient data may be very large, a training device may need to send multiple data packets to transmit the complete gradient data.
[0116] A data format identifier, used to indicate the numerical format of gradient data, such as 32-bit floating-point numbers. Different data formats require different processing logics.
[0117] In operation S302, the gradient data to be aggregated refers to the model parameter gradient values generated by each training device in a distributed training system after one round of training iteration and that need to be summarized and calculated. The gradient data to be aggregated is used to achieve consistent updates of model parameters in the distributed training system, ensuring that the models on each training node can be optimized based on global information, thereby improving training efficiency and model convergence.
[0118] Among them, the second memory 204 refers to the storage component in the gradient aggregation device 101 for storing large-capacity data. In some embodiments, the second memory 204 can be a high-bandwidth memory (HBM) externally connected to the processor 112 or a dynamic random-access memory (DRAM) and other memories with relatively large storage capacities. It is used to store all the gradient data received from the training device 102, serving as the main repository for data, making up for the limited capacity of the first memory 203, and at the same time providing reliable persistent storage for the gradient data to ensure that a complete dataset can be quickly accessed when the system needs to reload or retransmit data.
[0119] Specifically, the parser 202 receives the aggregation request data packet from the network interface 201 and performs parsing processing. The parser 202 first extracts information such as the source device identifier, data packet type identifier, and sequence number from the data packet header to verify the legality and integrity of the data packet. Subsequently, the parser 202 extracts parameters such as the round identifier, group identifier (including WorkerID and Packet ID), and data format identifier from the control information section. Finally, the parser 202 separates the gradient data section from the data packet, determines the parsing method of the data according to the data format identifier, and parses the gradient data into a format that can be processed internally by the system.
[0120] After parsing, the memory management module 205 stores the gradient data parsed by the parser 202 together with the relevant control information into the second memory 204. The second memory 204 has a relatively large storage capacity and can accommodate the complete gradient data generated by all training devices in a distributed training system in one iteration, avoiding data loss due to insufficient storage space. Although the first memory 203 has a faster access speed, its capacity is limited and it cannot store all the gradient data at the same time. By first storing all the gradient data in the second memory 204, the system can establish a complete data backup, providing a basis for subsequent possible data retransmission and recovery operations and improving the fault tolerance of the system.
[0121] In operation S303, the memory management module 205 obtains the gradient data to be aggregated from the second memory 204 and loads it into the first memory space 207 of the first memory 203 to prepare for subsequent gradient aggregation calculations.
[0122] In operation S304, the first memory space 207 refers to the storage area divided in the first memory 203 for storing the gradient data to be aggregated. The second memory space 208 refers to the storage area divided in the first memory 203 for storing the gradient data to be retransmitted. There is a logically separated but physically unified relationship between the first memory space 207 and the second memory space 208. They share the physical resources of the first memory 203 but play different functional roles, and are uniformly scheduled and allocated by the memory management module 205. Data can be exchanged between the two spaces to improve the overall efficiency of the system.
[0123] Among them, the gradient data to be retransmitted refers to the gradient data that needs to be resent to the gradient aggregation device due to network fluctuations, packet loss, or checksum failure in a distributed training system. In the embodiments of the present disclosure, it can be understood as the gradient data copy temporarily retained by the system to ensure data integrity during the gradient aggregation process. These data may need to be requested or processed again due to communication anomalies in the current aggregation round. It is used to ensure that the aggregation process can be quickly restored in case of communication errors, and to prevent the entire distributed training process from being interrupted or restarted due to the loss of a single packet, thereby improving the fault tolerance and operating efficiency of the system.
[0124] Specifically, the memory management module 205 first ensures that the required gradient data has been loaded into the first memory space 207 of the first memory 203, and then performs gradient aggregation calculations. The gradient aggregation process involves arithmetically aggregating the gradient values at the same positions sent by multiple training devices 102 to generate a global gradient update value. The above process is directly carried out in the first memory space 207, making full use of the high access speed of the first memory 203, and significantly improving the aggregation calculation efficiency.
[0125] Furthermore, in an actual distributed training environment, network fluctuations and data transmission errors occur frequently, and it is required to be able to quickly handle data retransmission requirements. The first memory 203 is divided into two memory spaces with different functions, so that the gradient data to be aggregated and the gradient data to be retransmitted are stored in the first memory space 207 and the second memory space 208 respectively, which can avoid the interference between the access patterns of the two types of data and improve the cache utilization efficiency. The gradient data in the first memory space 207 frequently participates in calculation operations, and the access pattern has spatial and temporal locality; while the gradient data in the second memory space 208 is used as a backup and is only accessed when a retransmission request occurs, and the access pattern is relatively scattered and irregular. This memory division method enables the system to adopt an optimized memory management strategy for different types of data, reduce cache conflicts and jitters, and achieve more efficient use of memory resources.
[0126] Operation S305, after completing the gradient aggregation calculation, the memory management module 205 passes the aggregation result to the packet encapsulator 206 for processing. The packet encapsulator 206 is responsible for repackaging the aggregation result into a packet format suitable for network transmission to ensure that the result can be correctly received and parsed by the training device 102. During the encapsulation process, the packet encapsulator 206 constructs a response packet according to the control information in the original aggregation request, including setting the correct source device identifier (the identifier of the gradient aggregation device 101 at this time), the target device identifier (set to the broadcast address or the address of a specific training device 102), the packet type identifier (marked as the aggregation result type), the round identifier, and the necessary check information. After encapsulation, the packet is passed to the network interface 201, which is responsible for sending the aggregation result back to all training devices 102 participating in the current training round through the network.
[0127] Based on the above embodiments, the following introduces a data transfer during the gradient aggregation provided by the embodiments of the present disclosure.
[0128] Figure 4 The schematic diagram of data transfer during the gradient aggregation is shown.
[0129] As Figure 4 shown, the gradient aggregation method provided by the embodiments of the present disclosure adopts a data transmission and processing strategy in batches, realizing efficient gradient data aggregation. In an actual distributed training environment, the number of model parameters is huge, and the amount of gradient data generated is huge. If all data is transmitted at once, it will cause network congestion and overload of the processing unit. Therefore, the embodiments of the present disclosure adopt a segmented transmission and processing method to improve the system throughput and response speed while ensuring data integrity.
[0130] Specifically, after each training device 102 (Worker) completes a round of training iteration, it divides and packs the locally computed gradient data into n Request data packets, and each data packet contains k gradient values in the FP32 (32-bit floating point) format. The above data structure design takes into account the processing capabilities and memory characteristics of the processor 112. The FP32 format ensures the accuracy of the gradient data, and the choice of the k value balances the data packet size and processing efficiency.
[0131] Each training device 102 adopts a batch sending strategy. Each time, it selects s data packets to form a batch and sends them to the switch 111. The switch 111 forwards these data packets to the processor 112 according to the pre-configured forwarding rules.
[0132] Subsequently, the processor 112 receives the data packets from all training devices 102. At this time, the processor 112 first stores the received gradient data in the second memory 204, and loads the data that needs to be processed immediately into the first memory space 207 of the first memory 203 through the memory management module 205. The processor 112 maintains a complete reception status table to record the reception status of each data packet sent by each training device 102 to ensure data integrity.
[0133] After the processor 112 confirms that it has received the gradient data from all training devices 102 at the same location, it immediately starts to perform the aggregation operation. During the aggregation process, the processor 112 performs an addition operation on the gradient values at each location, adds the k gradient values at each location respectively, and forms new k aggregated gradient values. This operation is efficiently executed in the first memory space 207 of the first memory 203, making full use of the access characteristics of the cache.
[0134] After completing the aggregation calculation, the processor 112 reorganizes the aggregation result into s Response data packets, and each Response data packet also contains k aggregated gradient values in the FP32 format. Then the processor 112 sends these Response data packets to the switch 111, and the switch 111 sends the aggregation result to all training devices 102 participating in the training at one time through the pre-configured broadcast mechanism. This broadcast method significantly reduces the number of network transmissions and the overall latency compared with the traditional point-to-point transmission. After all training devices 102 receive the aggregation result, they can synchronously update their respective model parameters and immediately prepare for the aggregation operation of the next batch of s data packets.
[0135] By adopting the present embodiment, when a batch of data is subjected to aggregation calculation on the processor 112, the next batch of data can be transmitted on the network simultaneously, improving the resource utilization rate and processing throughput of the overall system. Secondly, the batch processing mechanism reduces the amount of data processed at a single time, enabling the processor 112 to efficiently complete the calculation within the limited space of the first memory 203, avoiding frequent data exchange and performance degradation caused by insufficient memory.
[0136] Thirdly, the synchronous batch processing mode enables all training devices 102 to maintain a consistent training progress, reducing the waiting time between nodes and improving the overall training efficiency. Finally, considering the instability of the actual network environment, even if there is a problem with the transmission of a certain batch of data, it only affects the processing of that batch of data and does not cause the entire training process to be interrupted, improving the fault tolerance of the system.
[0137] In the actual application scenario of large-scale distributed training, the amount of gradient data generated in a single round of training may reach several gigabytes, far exceeding the capacity limit of the cache. Secondly, the time when different training devices send gradient data is uncertain, resulting in inconsistent arrival timings of gradient data and increasing the complexity of data management. When processing a large amount of gradient data, frequent access to storage devices will cause bandwidth congestion and increased energy consumption.
[0138] In view of the above problems, on the basis of the above embodiments, the present disclosure embodiments disclose a gradient data acquisition mechanism based on a multi-level storage architecture.
[0139] Specifically, for operation S303, it may specifically include:
[0140] Operation S401, in response to a gradient aggregation instruction, obtain target gradient data to be aggregated indicated by the gradient aggregation instruction in the first memory, where the target gradient data is used for gradient aggregation of a target round;
[0141] Operation S402, in the case where the target gradient data is not found in the first memory, obtain the target gradient data in the second memory and store the target gradient data in the first memory.
[0142] Among them, the gradient aggregation instruction refers to an internal control signal generated by the control module 202 according to the requirements of gradient aggregation. This signal contains specific command codes and parameter information, and is used to instruct the memory management module 205 to perform gradient data aggregation operations for a specific round and group. In the present disclosure embodiments, the gradient aggregation instruction carries key information required for gradient aggregation, such as a target round identifier (indicating which round of training the current aggregation is for the generated gradient), a data grouping identifier (indicating which batch of s data packets need to be processed), a list of identifiers of the training devices 102 participating in the aggregation (indicating which training node data needs to participate in the aggregation), and the type of aggregation operation.
[0143] Similarly, the target gradient data refers to a specific gradient data set to be processed by the current gradient aggregation instruction. In the disclosed embodiment, the target gradient data can be understood as a data set consisting of gradient values generated by multiple training devices 102 at specific grouping positions in the target round.
[0144] Exemplarily, when the gradient aggregation instruction indicates that the data of the fifth group in the tenth round of training needs to be processed, the target gradient data includes the gradient values of the fifth group generated by all the training devices 102 participating in the training in the tenth round of training.
[0145] Specifically, the memory management module 205 first parses the gradient aggregation instruction received from the control module 202 and extracts the key information contained therein. Subsequently, the memory management module 205 searches for the target gradient data indicated by the gradient aggregation instruction in the first memory 203. Since the first memory 203 has a high access rate, obtaining data therein can greatly reduce data access delay and significantly improve the speed of aggregation calculation.
[0146] During this process, the memory management module 205 maintains a memory mapping table to record the gradient data information currently cached in the first memory 203, including round identifier, group identifier, and data validity mark, etc., so as to quickly locate the target gradient data through this information.
[0147] There are usually two situations in which the target gradient data is not hit in the first memory 203: one is that the gradient data of the current round is requested for processing for the first time; the other is that the previously cached target gradient data has been replaced due to the limited capacity of the first memory 203.
[0148] When the memory management module 205 confirms that the target data is not hit in the first memory 203, it will immediately initiate a data request to the second memory 204. The memory management module 205 uses the same target round identifier and group identifier to locate the target gradient data in the second memory 204. Once found, it will be immediately transferred to the first memory 203, and the memory mapping table will be updated to record the newly cached data information.
[0149] By adopting the above embodiment, the target gradient data is preferentially searched in the high-speed first memory 203, and the time locality feature of gradient data access is utilized to reduce data access delay. When the first memory 203 does not hit, the system obtains data from the second memory 204 with a larger capacity and caches it, overcoming the capacity limitation of the first memory 203, so that the system can process large-scale gradient data.
[0150] In the practical application of a distributed training system, due to the instability of the network environment, situations such as packet transmission failures, delays, or out-of-order arrivals often occur. This requires the system to have an effective data retransmission mechanism to ensure the continuity of the training process. Secondly, during the gradient aggregation process, some data is frequently involved in computational operations, while another part of the data is only used as a backup to handle possible communication anomalies.
[0151] To achieve a balance between supporting high-performance computing through limited cache resources and ensuring system fault tolerance, based on the above embodiments, as an alternative embodiment, in operation S402, storing the target gradient data in the first memory may specifically further include the following operations:
[0152] Operation S501, determining the gradient data to be aggregated and the gradient data to be retransmitted in the target gradient data;
[0153] Operation S502, storing the gradient data to be aggregated in the target gradient data in the first memory space;
[0154] Operation S503, storing the gradient data to be retransmitted in the target gradient data in the second memory space.
[0155] Specifically, in a distributed training system, gradient data is usually transmitted in groups, and each round of training of each training device 102 generates multiple data packets. The memory management module 205 maintains a reception status table to record the reception status of each data packet. When a data packet is successfully received and its integrity is verified, the gradient data in the data packet is determined to be the gradient data to be aggregated; at the same time, the memory management module 205 will, according to a preset fault tolerance strategy, mark the gradient data under specific conditions (such as data of critical nodes or data at critical stages during the aggregation process) as the gradient data to be retransmitted for quick recovery in case of network anomalies.
[0156] After the data characteristic analysis is completed, the memory management module 205 stores the gradient data to be aggregated in the target gradient data in the first memory space 207. The first memory space 207 adopts an optimized memory layout and access mechanism, which is designed specifically for frequent read and write operations and aggregation calculations, and can provide the best access performance. At the same time, the gradient data to be retransmitted is stored in the second memory space 208. The second memory space 208 adopts a different memory management strategy, giving priority to data integrity and fast retrieval capabilities rather than access speed, because this data is usually only accessed when a retransmission request occurs.
[0157] By adopting the embodiments of the present disclosure, the differential storage management for different characteristic data reduces memory access conflicts and avoids the cache thrashing problem caused by the mixed storage of frequently accessed computing data and occasionally accessed backup data. Secondly, the gradient data to be retransmitted is independently stored in the second memory space 208, ensuring that the gradient data to be retransmitted can be quickly retrieved and used when needed, and also avoiding its premature elimination during the conventional cache replacement process.
[0158] In the actual distributed training process, due to factors such as network transmission delay, processing speed differences of computing nodes, and system scheduling, the arrival time of gradient data sent by different training devices 102 at the gradient aggregation device 101 is uncertain. This uncertainty may cause some gradient data to be marked as to be retransmitted and stored in the second memory space 208. After the subsequent network environment returns to normal, these data can participate in the aggregation calculation without retransmission. To efficiently utilize the cached data and avoid unnecessary re-acquisition operations, the embodiments of the present disclosure disclose a two-level cache query mechanism.
[0159] Based on the above embodiments, as an optional embodiment, in operation S401, obtaining the target gradient data to be aggregated indicated by the gradient aggregation instruction in the first memory may specifically further include the following operations:
[0160] Obtain the target gradient data to be aggregated indicated by the gradient aggregation instruction in the first memory space; in the case where the target gradient data is not found in the first memory, obtain the target gradient data in the second memory space.
[0161] Specifically, when the memory management module 205 receives the gradient aggregation instruction, the memory management module 205 first accesses the first memory space 207. The memory management module 205 queries the memory mapping table and retrieves the storage location of the target gradient data in the first memory space 207 with the round identifier and group identifier as keys. If the query is successful, that is, the target gradient data is found in the first memory space 207, the memory management module 205 directly obtains the data and prepares for the aggregation operation without further query, thereby maximizing the use of the high-speed access characteristics of the first memory space 207.
[0162] However, in the dynamic environment of distributed training, due to reasons such as network fluctuations and node delays, some gradient data previously marked for retransmission may actually have been completely received and stored in the second memory space 208. In the case where the target gradient data is not found in the first memory space 207, the memory management module 205 will continue to search for the target data in the second memory space 208 instead of directly accessing the slower second memory 204. The memory management module 205 queries the round identifier and group identifier of the same round and retrieves the target gradient data in the data index of the second memory space 208. If the target data is found in the second memory space 208, the memory management module 205 will move or copy it to the first memory space 207, update the memory mapping table, and then continue with the subsequent aggregation operation. Only when the target gradient data cannot be found in the entire first memory 203 will the memory management module 205 access the second memory 204 to obtain the data.
[0163] By adopting the embodiments of the present disclosure, the gradient aggregation device 101 can provide a strong fault tolerance while maintaining a high cache efficiency. This mechanism reduces the data access latency during the gradient aggregation process and reduces the access frequency to the second memory 204, thereby improving the overall system throughput and energy efficiency ratio. Especially in complex deployment scenarios with frequent network environment fluctuations, this design can better adapt to the uncertainty of data transmission and ensure the stability and efficiency of the distributed training process.
[0164] Based on the above embodiments, as an alternative embodiment, gradient aggregation of the gradient data to be aggregated in the first memory to obtain an aggregation result may specifically include gradient aggregation of the target gradient data in the first memory space to obtain an aggregation result.
[0165] Specifically, during the aggregation process, the memory management module 205 first confirms that all relevant gradient data of the target round has been successfully loaded into the first memory space 207, and then directly executes the aggregation algorithm at the memory locations where these data are located. For example, when performing additive aggregation, the memory management module 205 will sequentially access the gradient values at the same positions sent by each training device 102 stored in the first memory space 207, accumulate them to obtain the aggregation result, and directly write the result back to the same memory location, overwriting the original data, thus avoiding additional memory allocation and data copy operations.
[0166] Since the first memory space 207 is specifically used to store the gradient data to be aggregated, its memory layout has been optimized for the aggregation operation. For example, it may use a continuous memory area to store relevant gradient values, or use a specific data structure to organize the gradient values at the same positions of different training devices 102, enabling the aggregation algorithm to execute in the optimal memory access mode and further improving the computing efficiency.
[0167] In a distributed training system, as the training rounds progress, a large amount of historical data will gradually accumulate in memory resources, especially cache resources, leading to a continuous increase in memory pressure. Particularly in the first memory 203, due to its limited capacity, if the processed data is not cleared in a timely manner, the available space will be reduced, thereby affecting the loading and processing efficiency of new data. In addition, during the communication process of distributed training, to cope with network instability, the memory management module 205 will retain the gradient data to be retransmitted in the second memory space 208 as a backup. If these data are not cleared in a timely manner after the aggregation of the target round is completed, they will not only occupy valuable cache resources but may also be confused with the data of the new round, increasing the complexity of memory management.
[0168] In view of the above problems, on the basis of the above embodiments, as an optional embodiment, in the gradient aggregation method, it can also be executed that: in response to the end of the gradient aggregation instruction, the gradient data to be retransmitted in the second memory space is cleared.
[0169] Specifically, when the memory management module 205 receives the signal indicating the end of the gradient aggregation instruction, it first identifies the target round and data grouping information corresponding to the instruction, and then looks up the position indexes of all the gradient data to be retransmitted in the second memory space 208 related to this round and grouping in the memory mapping table. Subsequently, the memory management module 205 clears the corresponding gradient data in the second memory space 208 in an orderly manner according to these indexes, and updates the memory mapping table, marking these memory positions as available states to reserve space for the data processing of the next round. In the above manner, the memory management module 205 can release the memory resources that are no longer needed in a timely manner while ensuring data security.
[0170] It should be noted that the embodiments of the present disclosure only clear the gradient data to be retransmitted in the second memory space 208, and do not involve the data in the first memory space 207. This is because the data in the first memory space 207 usually participates in the actual aggregation calculation, and its life cycle and replacement strategy are automatically managed by the cache mechanism. However, the data to be retransmitted in the second memory space 208 is a backup specifically reserved for fault tolerance purposes, and its usage pattern is more special, and it needs to be actively cleared when it is confirmed that it is no longer needed to improve the memory utilization efficiency.
[0171] By implementing the embodiments of the present disclosure, efficient recycling of the second memory space 208 can be achieved, avoiding waste of memory resources. By timely clearing the gradient data to be retransmitted that is no longer needed, valuable cache space is released, enabling the second memory space 208 to accommodate more data in new rounds and improving the cache hit rate. In addition, since there is no dependency relationship between gradient data in different rounds, timely clearing of the data in completed rounds does not affect the correctness of training, and can also reduce interference between data in different rounds, reducing the risk of processing errors caused by data confusion.
[0172] The following introduces a cache mechanism provided by the embodiments of the present disclosure during the gradient aggregation process.
[0173] During the actual operation of the distributed training system, as gradient data is continuously generated and transmitted, the first memory 203 is easily reached the upper limit of the storage capacity. Especially when dealing with large-scale models, the total amount of gradient data generated in a single round of training iteration usually far exceeds the capacity of the first memory 203. At this time, the newly received gradient data cannot be directly loaded into the first memory 203 for processing. If the existing data is simply cleared using the first-in-first-out or random replacement strategy, it may cause hot data that is frequently accessed to be prematurely eliminated, thereby triggering the cache thrashing problem and seriously affecting the system performance. Therefore, the embodiments of the present disclosure disclose a cache replacement strategy based on access frequency to optimize the space utilization efficiency of the first memory 203.
[0174] Based on the above embodiments, as an optional embodiment, the above gradient aggregation method may further include the following operations:
[0175] Operation S601, when the storage space of the first memory is insufficient, obtain the gradient data with the lowest access frequency in the first memory;
[0176] Operation S602, store the gradient data with the lowest access frequency in the second memory.
[0177] Specifically, when the memory management module 205 detects that the available space of the first memory 203 is not enough to accommodate new gradient data, a cache replacement operation will be triggered. The memory management module 205 first scans the access records of all gradient data in the first memory 203, and determines the gradient data with the lowest access frequency by analyzing the historical access frequency and the most recent access time of each group of data. These low-frequency access data are usually gradient data in earlier historical rounds or data that has been used up in the current aggregation process, and the probability of their being accessed again is relatively low.
[0178] The memory management module 205 uses a counter to record the access times of each group of gradient data and maintains a timestamp to record the most recent access time. Based on the access times and access time, the heat value of each group of data is calculated. The gradient data with the lowest heat value is identified as the gradient data with the lowest access frequency and becomes the preferred target for cache replacement.
[0179] The memory management module 205 migrates the gradient data with the lowest access frequency from the first memory 203 to the second memory 204 to free up space in the first memory 203. During the migration process, the memory management module 205 first ensures data integrity by copying the target data completely to the designated location in the second memory 204 and verifying the copy result. Then, it updates the memory mapping table to record the new storage location of the data to ensure that subsequent operations can correctly locate the data. Finally, the memory management module 205 frees up the corresponding space in the first memory 203 so that it can be used to store new gradient data.
[0180] It should be noted that this cache replacement strategy mainly acts on the data in the first memory space 207 and does not directly affect the data to be retransmitted in the second memory space 208. This is because the data to be retransmitted usually does not participate in the regular cache replacement process, and its life cycle is determined by the retransmission requirements and the aggregation completion status, and it is cleared uniformly at the end of the aggregation instruction.
[0181] By implementing the embodiments of the present disclosure, the first memory 203 can retain the hot data with a higher access frequency and migrate the cold data to the second memory 204, forming a hierarchical cache architecture of "hot data stored hot, cold data stored cold". This architecture makes full use of the locality characteristics of gradient data access, significantly improves the cache hit rate, and reduces the access frequency to the second memory 204. Since the access speed of the first memory 203 is much faster than that of the second memory 204, a high cache hit rate directly translates into lower data access latency and higher processing throughput.
[0182] In addition, the cache replacement strategy based on access frequency can also adapt to the data access patterns in different training stages. In the initial stage of training, the model parameters change greatly, and the gradient data in different rounds are significantly different. At this time, the cache strategy will replace the historical data faster; while in the later stage of training, the model gradually converges, the gradient changes are relatively stable, and the gradient data at similar positions are frequently accessed. The cache strategy will tend to retain these hot data to further improve the system efficiency.
[0183] In the related art, it is usually necessary to traverse the entire cache space to determine the data with the lowest access frequency. The above operations will cause performance loss when the capacity of the first memory 203 is large, affecting the real-time response ability of gradient aggregation. To solve this problem, the embodiments of the present disclosure introduce a Least Recently Used (LRU) cache management mechanism based on a doubly linked list, which realizes efficient and low-overhead cache replacement operations.
[0184] Based on the above embodiments, as an alternative embodiment, the first memory further includes a doubly linked list, the doubly linked list includes a plurality of nodes, and the nodes correspond to the storage addresses of the first memory space;
[0185] In operation S602, storing the gradient data with the lowest access frequency in the second memory may specifically include the following operations:
[0186] Operation S603, storing the gradient data corresponding to the tail node in the doubly linked list in the second memory.
[0187] Specifically, in addition to the storage space for storing actual gradient data, the first memory 203 also maintains a doubly linked list structure for tracking and recording the access order of each gradient data in the first memory space 207. This doubly linked list is composed of multiple nodes, each node corresponds to a storage address in the first memory space 207, and stores specific gradient data. The doubly linked list is organized in the order of data access time. The node of the most recently accessed data is located at the head of the list, and the node of the least recently accessed data is located at the tail of the list, forming an ordered structure reflecting the access popularity of the data.
[0188] When the storage space of the first memory 203 is insufficient and cache replacement is required, the memory management module 205 does not need to scan the entire storage space to calculate and compare the access frequencies of each data, but directly accesses the tail node of the doubly linked list. The tail node of the doubly linked list naturally corresponds to the gradient data in the first memory space 207 that has not been accessed for the longest time. These data are marked as having the lowest access frequency under the LRU algorithm and are most suitable to be replaced out of the cache. The memory management module 205 obtains the storage address pointed to by the tail node, directly reads the gradient data stored at this address, and transfers it completely to the appropriate position in the second memory 204.
[0189] Furthermore, in terms of computational complexity, the operation of determining the data with the lowest access frequency is reduced from an O(n) traversal complexity to an O(1) direct access, greatly reducing the time overhead of cache replacement decisions. The maintenance operations of the doubly linked list (such as node movement, insertion, and deletion) also have a complexity of only the O(1) level. No matter how large the cache size is, it can maintain an efficient update speed. This enables low overhead and high responsiveness of cache management while processing a large amount of gradient data.
[0190] During the data migration process, the memory management module 205 first obtains the corresponding storage address from the tail node of the doubly linked list, and then reads the complete information of the gradient data from this address, including the round identifier, the group identifier, and the actual gradient value. The memory management module 205 allocates appropriate storage space in the second memory 204 and writes the complete gradient data into this space. After the writing is completed, the memory management module 205 updates the memory mapping table to record the new storage location of the gradient data so that subsequent operations can locate it correctly. Finally, the memory management module 205 removes the tail node from the doubly linked list and releases the corresponding storage area in the first memory space 207, making it available for storing new gradient data.
[0191] By adopting the embodiments of the present disclosure, while ensuring the cache replacement quality, the management overhead is minimized. The LRU algorithm of the doubly linked list can accurately capture and retain the recently and frequently used data, eliminate the data that has not been accessed for a long time, and achieve the optimal utilization of the cache space.
[0192] During the gradient aggregation process of the distributed training system, the access pattern of the gradient data usually presents significant temporal locality characteristics. That is, the probability that the gradient data that has been recently accessed will be accessed again in a short period of time is relatively high. Especially in the same round of iterative training, the gradient data at some key positions may be frequently read and updated. In order to make full use of this access characteristic and improve the cache efficiency, the embodiments of the present disclosure further improve the cache update strategy on the basis of the LRU cache mechanism, realizing the dynamic adjustment and optimization of the cache content.
[0193] On the basis of the above embodiments, as an optional embodiment, the above gradient aggregation method may further include the following operations:
[0194] Operation S604, determining the target node corresponding to the target gradient data in the doubly linked list;
[0195] Operation S605, moving the target node to the head of the doubly linked list.
[0196] Specifically, when the memory management module 205 successfully obtains the target gradient data, regardless of whether the data is directly obtained from the first memory space 207 or loaded from the second memory space 208 or the second memory 204, it is necessary to update the state of the data in the cache to reflect its latest access situation. The memory management module 205 first searches for the target node corresponding to the target gradient data in the doubly linked list. This search process is usually implemented through a hash table, mapping the unique identifier of the gradient data to the specific node in the doubly linked list to achieve fast positioning with low complexity.
[0197] After determining the target node corresponding to the target gradient data, the memory management module 205 immediately performs a linked list update operation, moving the target node from its current position to the head position of the doubly linked list. This operation first disconnects the target node from the linked list, adjusts the connection relationships of its previous and next nodes, and then inserts the target node into the head of the linked list to become the new first node. The entire process only involves simple pointer operations and has a low computational complexity. Even in high-frequency data access scenarios, it can maintain extremely low operation overheads.
[0198] By adopting the embodiments of the present disclosure, it is ensured that the most frequently used data is always retained in the high-speed first memory 203, while the less frequently used data is evicted to the second memory 204. This not only improves the cache hit rate, reduces the number of accesses to the high-latency second memory 204, but also reduces the overall energy consumption of data transmission, enabling the system to more efficiently execute the gradient aggregation task and providing strong performance support for distributed training.
[0199] The following introduces a gradient aggregation device provided by the embodiments of the present disclosure.
[0200] As Figure 2 shown, the gradient aggregation device 101 includes:
[0201] A first memory 203 and a second memory 204;
[0202] A network interface 201 for receiving an aggregation request;
[0203] A parser 202 for parsing the gradient data to be aggregated in the aggregation request and storing it in the second memory 204;
[0204] A memory management module 205 for obtaining the gradient data to be aggregated and performing gradient aggregation on the gradient data to be aggregated in the first memory 203 to obtain an aggregation result;
[0205] The network interface 201 is further configured to return the aggregation result;
[0206] Among them, the first memory 203 includes a first memory space 207 and a second memory space 208. The first memory space 207 is used to store the gradient data to be aggregated, and the second memory space 208 is used to store the gradient data to be retransmitted.
[0207] Based on the above embodiments, as an alternative embodiment, the memory management module 205 is further configured to respond to a gradient aggregation instruction, obtain target gradient data to be aggregated indicated by the gradient aggregation instruction in the first memory 203, where the target gradient data is used for gradient aggregation in a target round; in the case that the target gradient data is not found in the first memory 203, obtain the target gradient data in the second memory 204, and store the target gradient data in the first memory 203.
[0208] Based on the above embodiments, as an alternative embodiment, the memory management module 205 is further configured to determine gradient data to be aggregated and gradient data to be retransmitted in the target gradient data; store the gradient data to be aggregated in the target gradient data in the first memory space 207; store the gradient data to be retransmitted in the target gradient data in the second memory space 208.
[0209] Based on the above embodiments, as an alternative embodiment, the memory management module 205 is further configured to, in the case that the storage space of the first memory 203 is insufficient, obtain the gradient data with the lowest access frequency in the first memory 203; store the gradient data with the lowest access frequency in the second memory 204.
[0210] Based on the above embodiments, as an alternative embodiment, the first memory 203 further includes a doubly linked list, the doubly linked list includes a plurality of nodes, and the nodes correspond to the storage addresses of the first memory space 207; the memory management module 205 is further configured to store the gradient data corresponding to the tail node in the doubly linked list in the second memory 204.
[0211] Based on the above embodiments, as an alternative embodiment, the memory management module 205 is further configured to determine a target node corresponding to the target gradient data in the doubly linked list; move the target node to the head of the doubly linked list.
[0212] Based on the above embodiments, as an alternative embodiment, the memory management module 205 is further configured to respond to the end of the gradient aggregation instruction and clear the gradient data to be retransmitted in the second memory space 208.
[0213] Based on the above embodiments, as an alternative embodiment, the memory management module 205 is further configured to obtain target gradient data to be aggregated indicated by the gradient aggregation instruction in the first memory space 207;
[0214] In the case that the target gradient data is not found in the first memory, obtain the target gradient data in the second memory space 208.
[0215] Based on the above embodiments, as an alternative embodiment, the memory management module 205 is further configured to perform gradient aggregation on the target gradient data in the first memory space 207 to obtain an aggregation result.
[0216] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features recited in various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0217] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A gradient aggregation method, comprising: Receiving an aggregation request; Parsing the gradient data to be aggregated in the aggregation request and storing it in a second memory; Obtaining the gradient data to be aggregated; Performing gradient aggregation on the gradient data to be aggregated in a first memory to obtain an aggregation result; Returning the aggregation result; Wherein, the first memory includes a first memory space and a second memory space, the first memory space is used to store the gradient data to be aggregated, and the second memory space is used to store the gradient data to be retransmitted.
2. The method according to claim 1, wherein the obtaining the gradient data to be aggregated includes: Responding to a gradient aggregation instruction, obtaining, in the first memory, the target gradient data to be aggregated indicated by the gradient aggregation instruction, where the target gradient data is used for gradient aggregation in a target round; In the case where the target gradient data is not found in the first memory, obtaining the target gradient data in the second memory and storing the target gradient data in the first memory.
3. The method according to claim 2, wherein the storing the target gradient data in the first memory includes: Determining the gradient data to be aggregated and the gradient data to be retransmitted in the target gradient data; Storing the gradient data to be aggregated in the target gradient data in the first memory space; Storing the gradient data to be retransmitted in the target gradient data in the second memory space.
4. The method according to claim 2, further comprising: In the case where the storage space of the first memory is insufficient, obtaining the gradient data with the lowest access frequency in the first memory; Storing the gradient data with the lowest access frequency in the second memory.
5. The method according to claim 4, wherein the first memory further includes a doubly linked list, the doubly linked list includes a plurality of nodes, and the nodes correspond to the storage addresses of the first memory space; The storing the gradient data with the lowest access frequency in the second memory includes: Storing the gradient data corresponding to the tail node in the doubly linked list in the second memory.
6. The method according to claim 5, after storing the gradient data to be aggregated in the target gradient data in the first memory space, further comprising: Determining a target node corresponding to the target gradient data in the doubly linked list; Moving the target node to the head of the doubly linked list.
7. The method according to claim 3, further comprising: Responding to the end of the gradient aggregation instruction, clearing the gradient data to be retransmitted in the second memory space.
8. The method according to claim 2, wherein the obtaining, in the first memory, the target gradient data to be aggregated indicated by the gradient aggregation instruction includes: Obtaining, in the first memory space, the target gradient data to be aggregated indicated by the gradient aggregation instruction; In the case where the target gradient data is not found in the first memory, obtaining the target gradient data in the second memory space.
9. The method according to claim 3, wherein aggregating the gradient data to be aggregated in the first memory to obtain an aggregation result includes: Aggregating the target gradient data in the first memory space to obtain an aggregation result.
10. A gradient aggregation device, comprising: A first memory and a second memory; A network interface for receiving an aggregation request; A parser for parsing the gradient data to be aggregated in the aggregation request and storing same in the second memory; A memory management module for obtaining the gradient data to be aggregated and aggregating the gradient data to be aggregated in the first memory to obtain an aggregation result; The network interface is further configured to return the aggregation result; wherein the first memory includes a first memory space and a second memory space, the first memory space is used for storing the gradient data to be aggregated, and the second memory space is used for storing the gradient data to be retransmitted.