Aggregation transmission method suitable for selecting on-network aggregation
Through adaptive memory management and congestion control mechanisms that are aware of on-network aggregation, the existing on-network aggregation transmission mechanism solves the problems of network congestion and hash conflicts when facing a large number of training tasks, and realizes efficient memory resource allocation and network transmission, improving the overall performance of distributed training.
Patent Information
- Application Number
- CN202510107492.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
The existing network aggregation transmission mechanism is prone to network congestion and hash conflicts when facing a large number of training tasks, resulting in very low task throughput.
Adaptive memory management module is used to allocate dedicated memory areas for online aggregation and allow non-on-network aggregation companies to use idle aggregation for aggregation. At the same time, an on-network aggregation-aware congestion control mechanism is introduced, giving priority to reducing the transmission rate of non-on-network aggregation to avoid network congestion.
It effectively avoids the competition for memory resources between jobs, reduces the probability of memory conflicts, improves the efficiency of network aggregation, and maintains optimal performance in a multi-task environment, reducing the impact of network congestion on training tasks.
Smart Images

Figure CN120050330A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer network systems, and in particular, to an aggregation transmission method suitable for in-network aggregation selection. Background Art
[0002] Distributed training techniques have been widely adopted because they make it possible to train complex models using large-scale datasets. The most commonly used distributed training technique is data parallel training. In data parallel training, the model is replicated across all computing nodes, and the data is divided into multiple subsets, with each computing node responsible for processing a different subset. During training, each computing node needs to perform gradient aggregation after forward and backward propagation to synchronize the model parameters; based on the location of gradient aggregation, the parameter server architecture and the ring all-reduce architecture are two of the most typical data parallel architectures. In the parameter server architecture, the computing nodes send the gradients to a centralized parameter server, which aggregates the gradients and broadcasts the aggregated results to all computing nodes. In the ring all-reduce architecture, each computing node receives gradients from adjacent nodes, performs gradient aggregation locally, and then sends the results to the next adjacent node. However, both of these architectures generate a large amount of network traffic, leading to network congestion. In-network aggregation offloads gradient aggregation from the server to a programmable switch, thereby accelerating the gradient aggregation process. Specifically, each computing node synchronously sends gradient packets to the switch, which organizes an array of aggregators. The switch locates the aggregator index based on the metadata carried on the packets (such as the packet sequence number (PSN) and job identifier (JobID)) and performs the aggregation operation. When the aggregation is complete, the switch sends the data to downstream devices (such as the parameter server or all computing nodes). Therefore, in-network aggregation reduces the occupancy of network bandwidth, expands the training scale of deep learning models, and eliminates network bottlenecks. For example, SwitchML has demonstrated a 5.5-fold increase in the training speed for a single deep neural network job when in-network aggregation is enabled.
[0003] Existing in-network aggregation transmission mechanisms usually adopt customized transmission mechanisms. One method is the self-clocked scheme used by SwitchML. Initially, SwitchML sets the sending window to S data packets (for example, the total number of bytes is the size of the bandwidth-delay product (BDP)), and sends the S data packets to the switch. When the switch receives and aggregates these data packets, it copies the aggregation result into the aggregated packet and adds an ACK flag to it. This means that the data packets returned by the switch not only carry data but also act as ACK packets. When the i-th ACK packet is received, SwitchML immediately sends the (S + i)-th data packet, even if there is queuing in the network. However, this method does not respond to network congestion, so it is only applicable to a limited number of training tasks. When the number of tasks increases, severe network congestion will lead to packet loss, which is an important issue that cannot be ignored; ATP and subsequent research designed a customized in-network aggregation transmission mechanism for multi-tenant learning tasks. Specifically, each computing node calculates a hash value for each gradient data packet, which represents the aggregator index on the switch. Due to possible hash collisions of the aggregator index, gradient data packets may be aggregated on the switch (for example, in the case of no collision) or fallback to the parameter server for aggregation (for example, in the case of a collision). At the same time, ATP enables the ECN marking function in the switch to reflect the network congestion status to the computing nodes. ATP merges the ECN markings of the gradient data packets from the computing nodes and broadcasts the ECN signal to all computing nodes. In this way, all computing nodes can receive the same congestion signal to ensure that they remain synchronized (i.e., the same window size and adjust the transmission rate simultaneously). When a computing node detects congestion through the ECN marking on the parameter ACK or three consecutive out-of-order ACKs, it halves the window size. Otherwise, the computing node increases the window size by one maximum transmission unit (MTU) within each round-trip time (RTT).
[0004] However, it is found that the existing best-effort aggregation will cause serious hash collisions, resulting in the need to fallback to the parameter server for aggregation. Therefore, under the dual influence of memory contention and network congestion, the throughput of the best-effort in-network aggregation task is very low. Thus, the present invention proposes an aggregation transmission method suitable for selective in-network aggregation to solve the above problems. Summary of the Invention
[0005] Aiming at the defects existing in the above-mentioned prior art, the purpose of the present invention is to provide an aggregation transmission method suitable for selective in-network aggregation.
[0006] To achieve the above purpose, the present invention provides an aggregation transmission method suitable for selective in-network aggregation, including the following steps:
[0007] Step 1: At the worker node side, when a new job arrives, the worker node by default sends the gradient data packet to the parameter server (PS) for aggregation and continuously requests memory resources from the switch during this process;
[0008] Step 2: The adaptive memory management module at the switch side allocates memory for the in-network aggregation job according to the request of the memory request module and allows non-in-network aggregation jobs to use the idle aggregator left by the in-network aggregation job for aggregation;
[0009] Step 3: When the data packet arrives at the PS, the PS performs in-server aggregation on non-in-network aggregation jobs and performs aggregation degree feedback on in-network aggregation jobs;
[0010] Step 4: All jobs use the in-network aggregation-aware congestion control mechanism to manage the sending rate of data packets, preferentially reducing the sending rate of non-in-network aggregation jobs to avoid the decline of the memory aggregation efficiency of in-network aggregation jobs.
[0011] Furthermore, the memory request module includes the following: A new training job is defaulted to a non-in-network aggregation job, and the gradient data packet performs opportunistic in-network aggregation on the switch; The non-in-network aggregation job periodically sends REQUEST data packets to request switch memory resources; The aggregation degree feedback includes the following: The PS determines the INA degree by parsing the bitmap of the received gradient data packet; Each data packet contains an N-bit bitmap indicating whether the gradient from the i-th worker node has been aggregated in the data packet; The PS performs a bitwise "OR" operation on the bitmaps in the received M data packets to calculate the INA degree d as follows;
[0012]
[0013] The PS carries the calculated INA degree d in the aggregated data packet and broadcasts it to all worker nodes; For non-in-network aggregation jobs, the INA degree d is set to 0.
[0014] Furthermore, the congestion control includes the following: The switch uses ECN to mark network congestion; If the queue length of the switch port exceeds the preset threshold, the gradient data packet sent by the worker node will be marked with the CE code point; The PS receives the marked data packet, aggregates the CE marks, and updates the link congestion level factor a i ;
[0015] a i =(1 - g)×a i-1 +g×F;
[0016] where F is the proportion of data packets marked with ECN in the previous window, and 0 < g < 1 is the weight factor;
[0017] If a iIf it is 0, then in each RTT period, the sliding window size cwnd is increased by 1 MTU; otherwise, the window size cwnd is decreased as follows:
[0018]
[0019] For an INA job with an INA degree d greater than 0, it follows the DCTCP window reduction strategy; for a non-INA job with an INA degree d of 0, the window size is always halved. The adaptive memory allocation includes the following steps: The switch organizes the memory into addressable aggregators and maintains the number c of jobs currently allocated to the memory. If c is less than the memory requirement q of the job, the switch responds to a new REQUEST packet, allocates a dedicated area for the job, and adds an entry in the Address table. The packets in the sliding window of an INA job are statically mapped to the switch memory, using the position of the packet in the sliding window as the offset of its aggregation index. The packets of a non-INA job are hashed to any aggregator index, and the switch checks whether the transmit window size of the corresponding area is less than this offset. If so, the packet is allowed to be aggregated in this aggregator.
[0020] Furthermore, the method also includes implementing FlexINA and its baseline method based on the OMNET++ simulator to evaluate the overall performance improvement. It is applicable to a multi-task environment and can maintain the best performance under different numbers of tasks. Specifically, it includes the following steps: Evaluating the performance of the method in a multi-task environment by restarting multiple identical jobs; increasing the total number of jobs from 1 to 8 (i.e., 1, 2, 4, 8 jobs) while keeping the total memory capacity unchanged; each job consists of 8 worker nodes randomly distributed across multiple racks; using the average iteration time as the evaluation metric for FlexINA and the baseline method.
[0021] Furthermore, the in-network aggregation-aware congestion control mechanism is as follows. When the switch detects network congestion, it preferentially reduces the transmission rate of non-in-network aggregation jobs. Through the ECN marking mechanism, the switch feeds back the network congestion status to the worker nodes, and the worker nodes adjust the transmit window size according to the received ECN marking to reduce network congestion. The adaptive memory management module is as follows. The switch maintains a memory allocation table to record the memory allocation status of each job. When the switch receives a memory request, it decides whether to allocate memory according to the memory allocation table and the current memory usage. When allocating memory, the switch preferentially allocates it to in-network aggregation jobs and records the allocated memory area. When the switch detects free memory, it allows non-in-network aggregation jobs to utilize the free memory for aggregation.
[0022] Further, the in-network aggregation includes the following steps: The switch assigns a fixed memory area to each in-network aggregation operation. The switch maps data packets to specific aggregators in the assigned memory area according to the positions of the data packets in the sliding window. When the switch receives a data packet, it aggregates the data packet into the corresponding aggregator according to the mapping relationship. The opportunistic memory preemption includes the following steps: When the switch detects a data packet of a non-in-network aggregation operation, it calculates the hash value and offset of the data packet. The switch checks whether the size of the send window in the corresponding area is less than the offset. If so, it allows the data packet to be aggregated in the aggregator. When the switch allows a data packet to be aggregated, it updates the memory allocation table and records the memory usage of the non-in-network aggregation operation.
[0023] Further, the method further includes using a 100Gbps leaf-spine topology, including two leaf switches and one spine switch. All switches are enabled with ECN marking, and the ECN threshold is set to 100 data packets. The leaf switches allow in-network aggregation and initialize 340 aggregators to meet the peak throughput requirements of a single task. Multiple cross-rack training tasks are run, each task including multiple worker nodes and one parameter server. By adjusting the maximum window size and hash space size, the performance under different configurations is evaluated.
[0024] Further, the method further includes statically mapping the data packets in the sliding window of the in-network aggregation operation to the switch memory, using the position of the data packet in the sliding window as the offset of its aggregation index. When the switch receives a data packet, it obtains the base address of the memory area by querying the allocation table and combines it with the offset carried in the data packet to obtain the physical address. After the aggregation is completed, the switch sends the aggregation result to downstream devices, such as the parameter server or all computing nodes. In a multi-task environment, through the task scheduling mechanism, the in-network aggregation operation is preferentially scheduled to reduce memory conflicts and network congestion. The task scheduling mechanism dynamically adjusts the execution order of tasks according to the current network state and memory usage. Through task scheduling, it is ensured that the in-network aggregation operation can fully utilize the allocated aggregator resources and improve the aggregation efficiency.
[0025] Further, the method further includes evaluating the performance stability under different network load conditions. By simulating three different network load scenarios of high, medium, and low, multiple training tasks are respectively run, and the changes in key performance indicators such as the average iteration time, packet loss rate, and aggregation efficiency of the method under different loads are recorded and analyzed. Specifically, in the high-load scenario, multiple large-scale training tasks are started simultaneously to make the network traffic exceed 80% of the switch bandwidth; in the medium-load scenario, an appropriate number of training tasks are started to keep the network traffic between 50% and 80% of the switch bandwidth; in the low-load scenario, only a small number of training tasks are started to make the network traffic lower than 50% of the switch bandwidth. The performance differences between the method and the prior art method under different load scenarios are compared and analyzed to ensure that the method can maintain good performance under various network load conditions, effectively reduce the impact of network congestion on training tasks, and improve the overall efficiency of distributed training.
[0026] Further, based on the performance evaluation results, the optimization objectives are clarified, such as reducing the average iteration time, and the key factors affecting the objectives are identified, such as the memory allocation amount and the congestion control window size. A mathematical model is established to describe the relationship between the factors and the objectives, the existing data is used to determine the parameters in the model, and new data is used to verify whether the model is accurate. Finally, the system parameters are adjusted according to the results of the model to establish a performance optimization model for guiding the parameter configuration and system optimization in actual deployment. Through mathematical modeling and data analysis, the optimal memory allocation strategy, congestion control parameters, and key parameters of the task scheduling strategy are determined. For example, the functional relationship between the memory allocation amount and the scale of the training task is fitted according to the experimental data, and the dynamic adjustment formula between the congestion control window size and the network load is obtained. Using this performance optimization model, users can quickly and accurately configure the system parameters according to the requirements of the actual application scenario, give full play to the advantages of the method, achieve the efficient execution of distributed training tasks, and further improve the overall performance and resource utilization rate of the system.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] 1. Through the adaptive memory management module, the present invention allocates a dedicated memory area for in-network aggregation operations and allows non-in-network aggregation operations to use the idle aggregators left by in-network aggregation operations for aggregation, effectively avoiding the competition for memory resources between operations, reducing the occurrence probability of memory conflicts, and thus improving the efficiency of in-network aggregation;
[0029] 2. For in-network aggregation operations, the present invention adopts the in-network static aggregation method, which statically maps the data packets in the sliding window to the switch memory, enabling the data packets to be accurately and efficiently aggregated into the specified aggregator, further improving the efficiency and accuracy of the aggregation operation. For non-in-network aggregation operations, the opportunistic memory preemption mechanism is adopted to reasonably utilize the idle memory for aggregation, making full use of network resources and improving the overall aggregation efficiency.
[0030] 3. The present invention introduces an in-network aggregation-aware congestion control mechanism. When network congestion occurs, it preferentially reduces the sending rate of non-in-network aggregation operations to ensure that the memory aggregation efficiency of in-network aggregation operations is not affected, thereby ensuring that in-network aggregation operations can be carried out efficiently and stably, and improving the overall performance of distributed training tasks.
[0031] 4. The present invention combines ECN marking and link congestion level factors, and flexibly adjusts the size of the sliding window according to different job types (INA jobs and non-INA jobs) and their aggregation degree feedback, accurately controlling the sending rate of data packets, effectively alleviating network congestion, reducing data packet loss, and improving the reliability of network transmission.
[0032] 5. The switch of the present invention organizes the memory into addressable aggregators, and dynamically allocates memory for jobs according to the memory requirements of jobs and the current memory usage situation, realizing the reasonable allocation and full utilization of memory resources, avoiding waste of memory resources, and improving the resource utilization rate of the system.
[0033] 6. The present invention is applicable to multi-task environments and can maintain the best performance under different numbers of tasks. Through the task scheduling mechanism, in-network aggregation operations are preferentially scheduled, reducing memory conflicts and network congestion, ensuring that in-network aggregation operations can make full use of the allocated aggregator resources, improving the aggregation efficiency, enabling multiple tasks to run efficiently and collaboratively, and enhancing the overall performance and resource utilization rate of the system.
[0034] 7. The present invention evaluates the performance stability of the method by simulating three different network load scenarios of high, medium, and low. It can maintain good performance under various network load conditions, effectively reducing the impact of network congestion on training tasks, improving the overall efficiency of distributed training, and enhancing the stability and reliability of the system.
[0035] 8. Based on the performance evaluation results, the present invention establishes a performance optimization model, comprehensively considers various factors, and determines the optimal memory allocation strategy, congestion control parameters, and key parameters of the task scheduling strategy. This provides a scientific basis for parameter configuration and system optimization in actual deployment, enabling users to quickly and accurately configure system parameters according to the requirements of the actual application scenario, giving full play to the advantages of this method, achieving the efficient execution of distributed training tasks, further improving the overall performance and resource utilization rate of the system, and ensuring that the system can operate stably and efficiently under different conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the solutions in the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings described below are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0037] Figure 1 is a flowchart provided by the present invention;
[0038] Figure 2 is a schematic diagram of the FlexINA architecture provided by the present invention;
[0039] Figure 3 is a schematic diagram of INA degree feedback provided by the present invention;
[0040] Figure 4 is a schematic diagram of sliding window static mapping provided by the present invention;
[0041] Figure 5 is a schematic diagram of adaptive memory allocation provided by the present invention;
[0042] Figure 6 is a schematic diagram of the overall performance results provided by the present invention;
[0043] Figure 7 is a schematic diagram of the performance results under different numbers of tasks provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The following will elaborate on the preferred embodiments of the present invention in conjunction with the drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0045] The terms "comprising" and "having" and any variations thereof in the description, claims and the above drawings of the present invention are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the description, claims or the above drawings of the present invention are used to distinguish different objects, rather than to describe a specific order.
[0046] Please refer to Figures 1-7 , an aggregation transmission method suitable for in-network aggregation selection, comprising the following steps:
[0047] Step 1: At the worker node side, when a new job arrives, the worker node defaults to sending gradient packets to the parameter server (PS) for aggregation, and continuously requests memory resources from the switch during this process;
[0048] Step 2: The adaptive memory management module at the switch side allocates memory for in-network aggregation jobs according to the requests of the memory request module, and allows non-in-network aggregation jobs to use the idle aggregators left by in-network aggregation jobs for aggregation;
[0049] Step 3: When the packets arrive at the PS, the PS performs in-server aggregation on non-in-network aggregation jobs and performs aggregation degree feedback on in-network aggregation jobs;
[0050] Step 4: All jobs use an in-network aggregation-aware congestion control mechanism to manage the packet sending rate, preferentially reducing the sending rate of non-in-network aggregation jobs to avoid a decrease in the memory aggregation efficiency of in-network aggregation jobs.
[0051] As an improvement to the above technical solution, the memory request module includes the following: A new training job is defaulted to a non-in-network aggregation job, and gradient packets are opportunistically aggregated within the network on the switch; Non-in-network aggregation jobs periodically send REQUEST packets to request switch memory resources; If the memory request is successful, the job is converted to an in-network aggregation job; The aggregation degree feedback includes the following: The PS determines the INA degree by parsing the bitmap of the received gradient packets; Each packet contains an N-bit bitmap indicating whether the gradient from the i-th worker node has been aggregated in the packet; The PS performs a bitwise "OR" operation on the bitmaps in the M received packets to calculate the INA degree d; The PS carries the calculated INA degree d in the aggregation packet and broadcasts it to all worker nodes; For non-in-network aggregation jobs, the INA degree d is set to 0.
[0052] As an improvement to the above technical solution, the congestion control includes the following: The switch uses ECN to mark network congestion; If the queue length of the switch port exceeds a preset threshold, the gradient packets sent by the worker node will be marked with the CE code point; The PS receives the marked packets, aggregates the CE marks, and updates the link congestion level factor a i ; If ai If it is 0, then in each RTT period, the sliding window size cwnd increases by 1 MTU; for INA jobs with an INA degree d greater than 0, the DCTCP window reduction strategy is followed; for non-INA jobs with an INA degree d of 0, the window size is always halved; Adaptive memory allocation includes the following steps: The switch organizes the memory into addressable aggregators and maintains the number c of jobs currently allocated to the memory; if c is less than the memory requirement q of the job, the switch responds to a new REQUEST packet, allocates a dedicated area for the job, and adds an entry in the Address table. Adaptive memory allocation: When in-network aggregation jobs have to reduce the sending rate, adaptive memory allocation enables non-in-network aggregation jobs to utilize idle aggregators without affecting the performance of in-network aggregation jobs; The packets in the sliding window of INA jobs are statically mapped to the switch memory, using the position of the packet in the sliding window as the offset of its aggregation index; The packets of non-INA jobs are hashed to any aggregator index, and the switch checks whether the sending window size of the corresponding area is less than this offset. If so, the packet is allowed to be aggregated in this aggregator. The concept of INA degree is introduced, which represents the degree of aggregation in the network. Based on this, FlexINA tends to preferentially reduce the sending rate of jobs with a lower INA degree (such as non-INA jobs) to prevent serious interference to INA jobs.
[0053] As an improvement to the above technical solution, the following steps are also included: FlexINA and its benchmark method are implemented based on the OMNET++ simulator to evaluate the overall performance improvement. It is applicable to a multi-task environment and can maintain the best performance under different numbers of tasks. Specifically, it includes the following steps: Evaluate the performance of the method in a multi-task environment by restarting multiple identical jobs; Increase the total number of jobs from 1 to 8 (i.e., 1, 2, 4, 8 jobs) while keeping the total memory capacity unchanged; Each job consists of 8 worker nodes randomly distributed across multiple racks; Use the average iteration time as the evaluation metric for FlexINA and the benchmark method.
[0054] As an improvement to the above technical solution, the congestion control mechanism for in-network aggregation awareness includes the following: when the switch detects network congestion, it preferentially reduces the transmission rate of non-in-network aggregation operations. Through the ECN marking mechanism, the switch feeds back the network congestion status to the working nodes, and the working nodes adjust the transmission window size according to the received ECN markings to reduce network congestion. Congestion control for in-network aggregation awareness: To protect in-network aggregation operations, FlexINA introduces a congestion control mechanism for in-network aggregation awareness. Its goal is to preferentially reduce the transmission rate of non-in-network aggregation operations during congestion to avoid a decline in the memory aggregation efficiency of in-network aggregation operations. The adaptive memory management module includes the following: The switch maintains a memory allocation table to record the memory allocation status of each job. When the switch receives a memory request, it decides whether to allocate memory based on the memory allocation table and the current memory usage. When the switch allocates memory, it preferentially allocates it to in-network aggregation operations and records the allocated memory area. When the switch detects free memory, it allows non-in-network aggregation operations to utilize the free memory for aggregation.
[0055] As an improvement to the above technical solution, static in-network aggregation includes the following steps: The switch allocates a fixed memory area for each in-network aggregation operation. The switch maps the data packets to specific aggregators in the allocated memory area according to the positions of the data packets in the sliding window. When the switch receives a data packet, it aggregates the data packet into the corresponding aggregator according to the mapping relationship. Opportunistic memory preemption includes the following steps: When the switch detects a data packet of a non-in-network aggregation operation, it calculates the hash value and offset of the data packet. The switch checks whether the transmission window size of the corresponding area is less than the offset. If so, it allows the data packet to be aggregated in the aggregator. When the switch allows a data packet to be aggregated, it updates the memory allocation table to record the memory usage of the non-in-network aggregation operation.
[0056] As an improvement to the above technical solution, the method further includes the following steps: using a 100Gbps leaf-spine topology structure, including two leaf switches and one spine switch. All switches are enabled with ECN markings, and the ECN threshold is set to 100 data packets. The leaf switches allow in-network aggregation and initialize 340 aggregators to meet the peak throughput requirements of a single task. Run multiple cross-rack training tasks, each task containing multiple working nodes and one parameter server. Evaluate the performance under different configurations by adjusting the maximum window size and hash space size.
[0057] As an improvement to the above technical solution, the method further includes the following steps: statically map the data packets in the sliding window of the in-network aggregation operation to the switch memory, use the position of the data packet in the sliding window as the offset of its aggregation index. When the switch receives a data packet, it obtains the base address of the memory area by querying the allocation table, and combines it with the offset carried in the data packet to obtain the physical address. After the switch completes the aggregation, it sends the aggregation result to downstream devices, such as a parameter server or all computing nodes. In a multi-task environment, through the task scheduling mechanism, the in-network aggregation operation is preferentially scheduled to reduce memory conflicts and network congestion. The task scheduling mechanism dynamically adjusts the execution order of tasks according to the current network status and memory usage. Through task scheduling, it is ensured that the in-network aggregation operation can make full use of the allocated aggregator resources and improve the aggregation efficiency.
[0058] As an improvement to the above technical solution, the method further includes the following steps: further evaluate the performance stability under different network load conditions. By simulating three different network load scenarios of high, medium, and low, run multiple training tasks respectively, record and analyze the changes in key performance indicators such as the average iteration time, packet loss rate, and aggregation efficiency of the method under different loads. Specifically, in the high-load scenario, start multiple large-scale training tasks simultaneously to make the network traffic exceed 80% of the switch bandwidth; in the medium-load scenario, start an appropriate number of training tasks to keep the network traffic between 50% and 80% of the switch bandwidth; in the low-load scenario, only start a small number of training tasks to make the network traffic less than 50% of the switch bandwidth. Compare and analyze the performance differences between the method and the existing technology method under different load scenarios to ensure that the method can maintain good performance under various network load conditions, effectively reduce the impact of network congestion on training tasks, and improve the overall efficiency of distributed training.
[0059] As an improvement to the above technical solution, based on the performance evaluation results, clarify the optimization objectives, such as reducing the average iteration time, identify the key factors affecting the objectives, such as the memory allocation amount and the congestion control window size, establish a mathematical model to describe the relationship between the factors and the objectives, use the existing data to determine the parameters in the model, verify whether the model is accurate with new data, and finally adjust the system parameters according to the results of the model to establish a performance optimization model for guiding the parameter configuration and system optimization in actual deployment. Through mathematical modeling and data analysis, determine the optimal memory allocation strategy, congestion control parameters, and key parameters of the task scheduling strategy. For example, fit the functional relationship between the memory allocation amount and the training task scale based on experimental data, as well as the dynamic adjustment formula between the congestion control window size and the network load. Using this performance optimization model, users can quickly and accurately configure system parameters according to the requirements of the actual application scenario, give full play to the advantages of the method, achieve the efficient execution of distributed training tasks, and further improve the overall performance and resource utilization rate of the system.
[0060] Experimental setup: Adopt a 100Gbps leaf-spine topology structure, including two leaf switches and one spine switch. All switches are enabled with ECN marking, and the ECN threshold is set to 100 packets. The leaf switches allow in-network aggregation and initialize 340 aggregators to meet the demand for peak throughput of a single task. Run two cross-rack AlexNet training tasks (Task A and Task B), each task containing 8 worker nodes and 1 parameter server. The worker nodes are evenly distributed in two racks, with 4 worker nodes in each rack. Under this setup, these two tasks compete for memory resources on the leaf switches and bandwidth resources on the core network at the same time. Test the performance of ATP and A TP, and focus on the aggregated throughput of each task. Set the hash space size to the same number as the peak throughput aggregator, and set the maximum window size to the bandwidth-delay product of the 100Gbps network (i.e., 340 packets or 104KB).
[0061] Experimental Observation: The experimental results show that the average throughput of ATP and A TP is almost the same but very low (i.e., 20 Gbps, compared with the ideal 50 Gbps), resulting in a completion time of about 300 ms for both tasks. It is believed that best-effort aggregation triggers multiple hash conflicts, leading to network congestion. Specifically, each ATP task initializes a hash space of size 340 on the end host, which can cover all aggregators. Then, it determines the aggregator index of the data packet based on the following formula: Index = Hash(JobID, PSN % CWND_max), where PSN represents the data packet sequence number, JobID represents the task identifier, and CWND_max is the maximum window size (in bytes, equal to the bandwidth-delay product). Since two tasks may map to the same hash value, this causes hash conflicts. These conflicts cause the data packets to queue in the switch buffer, triggering ECN marking, and ultimately causing the task to reduce the sending rate to alleviate network congestion. To verify this, CWND_max is reduced to 170 data packets (i.e., the aggregators are evenly distributed among the two tasks). This setting makes the hash values of the data packets map to only half of the hash space, thus significantly reducing the probability of hash conflicts. The results show that, ideally, each task can achieve a throughput close to 50 Gbps and the task completion time is about 125 ms. ATP decouples memory and bandwidth contention through two sliding windows (CWND and ACW). For each window, ATP adopts a congestion adjustment mechanism similar to DCTCP to handle memory conflicts and network congestion. However, these two types of resources are still interrelated. For example, increasing the ACW window can aggregate more traffic in the network and significantly reduce the network queue length.
[0062] The above experimental results show that due to hash conflicts, best-effort aggregation cannot fully utilize the performance of in-network aggregation. To avoid memory contention, a static allocation scheme can be adopted, such as SwitchML, NetReduce, and Panama, where the switch memory is evenly allocated to all tasks.
[0063] Improve performance by jointly optimizing congestion control and memory management. Figure 1Shows the architecture of FlexINA, including worker nodes, switches, and a parameter server (PS). At the worker node side, the memory request module is used to request memory from the switch. When a new job arrives, the worker node by default sends gradient packets to the PS for aggregation. During this process, the worker node continuously requests switch memory. Once the memory request is successful, the non-in-network aggregation job will be permanently converted to an in-network aggregation job. All jobs use an in-network aggregation-aware congestion control mechanism to manage the packet sending rate. At the switch side, the adaptive memory management module allows in-network aggregation jobs to perform aggregation in the allocated aggregators. For non-in-network aggregation jobs, this module allows them to utilize the idle aggregators left by in-network aggregation jobs for aggregation. When the packets arrive at the PS, the PS performs in-server aggregation for non-in-network aggregation jobs and provides aggregation degree feedback for in-network aggregation jobs. FlexINA fundamentally improves in-network aggregation performance through the following two simple but efficient design principles;
[0064] FlexINA adopts selective INA, which allows jobs to perform in-network aggregation under memory constraints. Suppose a new training job consists of N worker nodes, and these N worker nodes are distributed across one or more racks. By default, this new job is considered a non-INA job, and gradient packets will be opportunistically aggregated in the network on the switch. During this process, non-INA jobs will periodically send REQUEST packets to request switch memory resources. The REQUEST packets contain necessary data, such as a job identifier (globally unique), for memory allocation. Instead of using a centralized controller to implement selective INA, FlexINA uses a distributed and lightweight memory allocation mechanism to implement it on programmable switches. If the memory request is successful, FlexINA considers the job as an INA job, and subsequent gradient packets will be aggregated on the switch;
[0065] FlexINA adopts a distributed memory allocation strategy, which may lead to partial success in memory allocation. For example, in Figure 2 Switch 1 allocates memory for the job, while Switch 2 fails to allocate memory due to memory overflow. In this case, FlexINA still considers the job as an INA job. As long as at least one rack switch allocates memory, the gradient data sent by the worker nodes will eventually be sent to the PS. As a result, the PS can determine the INA degree by parsing the M received gradient packets. The PS only calculates the INA degree of INA jobs, and non-INA jobs directly return 0. Specifically, each packet contains an N-bit bitmap. If the i-th bit is set to 1, it means that the gradient from the i-th worker node has been aggregated in the packet. In Figure 1In it, the PS will receive three data packets P1, P2, and P3. Since Switch 1 performs in-network aggregation, the bitmap of Packet P1 is 1100. The switch will perform a bitwise "OR" operation on the bitmaps in the M received data packets. Once all bits of all bitmaps are set to 1, the PS calculates the INA degree d as follows:
[0066]
[0067] Specifically, if all switches have allocated memory for Job A, then d = 1. After that, the PS will carry the INA degree d in the aggregated data packet and broadcast it to all worker nodes. For non-INA jobs, although their data packets may be opportunistically aggregated in the network, d is still set to 0;
[0068] FlexINA still uses ECN to mark network congestion. If the queue length of a switch port exceeds a preset threshold, the gradient data packets sent by the worker nodes will be marked with the CE code point. The PS receives the marked data packets, aggregates the CE marks, and broadcasts them to all worker nodes. Specifically, FlexINA updates the link congestion level factor αi for the i-th RTT as follows:
[0069] a i = (1 - g) × a i-1 + g × F;
[0070] where F is the proportion of data packets marked as ECN in the previous window, and 0 < g < 1 is the weight factor. If α = 0, then in each RTT period, the sliding window size cwnd increases by 1 MTU, otherwise, the window size cwnd decreases as follows:
[0071]
[0072] According to the definition of d, the d of INA jobs is always greater than 0, while the d of non-INA jobs is 0. Therefore, INA jobs with d = 1 follow the DCTCP window reduction strategy, while non-INA jobs always halve the window size. In this way, FlexINA can relieve congestion by actively reducing the rate of non-INA jobs;
[0073] The switch organizes the memory into addressable aggregators (i.e., registers). For each job, its memory requirement is fixed at q, which is the amount of memory required for a job to achieve its maximum throughput. Therefore, for a total of A aggregators, the number of jobs it can support is p. The switch maintains the number of jobs c currently allocated to the memory and decides whether to respond to a new REQUEST packet by comparing the values of c and q. If c < q, the switch mirrors the REQUEST packet to the control plane. The control plane parses the packet and allocates a dedicated area for the job. Then, the control plane adds an entry {Key:JobId; Value:Base_Address,Region_id} to the Address table in the data plane. Subsequently, the gradient packet can query the Address table to obtain the Base_Address of the allocated memory region. If there is a table hit, the switch performs static in-network aggregation for INA jobs; otherwise, it performs opportunistic memory preemption for non-INA jobs.
[0074] Traditional best-effort aggregation methods calculate a hash value for each packet, ranging from 0 to A - 1. This method makes it difficult for the switch to quickly identify wasted aggregators. Therefore, FlexINA proposes a static aggregation scheme. As Figure 3 shown, FlexINA statically maps the packets in the sliding window to the switch memory. Specifically, FlexINA uses the position of the packet in the sliding window as the offset of its aggregation index. For example, for the sliding window cwnd at the current RTT, the first packet in the sliding window will be mapped to the first aggregator in the allocated area, and so on. In fact, existing work also adopts static mapping. However, different from these works, FlexINA does not map based on the packet sequence number, but based on the sliding window itself. In this way, FlexINA ensures that INA jobs preferentially use the aggregators at the head of the memory region. In other words, if the sending window of an INA job is decreasing, this job will not use the aggregators at the tail of the region. When the switch receives a packet from an INA job, it obtains the Base_Address by querying the allocation_table table and combines it with the offset carried in the packet to get the physical address.
[0075] During communication, the rate drops. When there is bandwidth competition among INA jobs, they must reduce their send windows to alleviate network congestion. Due to the static allocation of INA jobs, the idle memory in the switch is distributed at the tail of each region. Zero aggregation during the calculation process. When an INA job enters the calculation stage, the switch memory will be completely idle. This can be regarded as a special case of Scenario 1, where cwnd is equal to 0. Therefore, by maintaining the current window size cwnd of INA jobs in each memory region, it is determined whether non-INA jobs can be aggregated. Specifically, FlexINA allows packets of non-INA jobs to be hashed to any aggregator index, such as Figure 4 shown. FlexINA calculates a region_id and an offset for the packets of non-INA jobs. When a packet arrives at the switch, the switch checks whether the send window size of the corresponding region is less than the offset. If so, it means that the memory at the specified offset is idle, and the packet is allowed to be aggregated in that aggregator.
[0076] III. Experimental Results
[0077] FlexINA and its baseline method were implemented based on the OMNET++ simulator. The simulation environment used an adjustable number of racks with a leaf-spine topology. In this simulator, the servers of each job (i.e., worker nodes and parameter server PS) can be placed in the same rack or distributed across multiple racks. The bandwidth of all links is 100 Gbps, the peak throughput aggregator (PTA) is 340 aggregators for a single job, and for the switch, the ECN marking threshold is set to 100 packets;
[0078] The overall performance improvement of FlexINA was evaluated. The topology shown in the figure was used, and two identical models were placed. Figure 5Shows the throughput, latency, and aggregation efficiency results of four models. Compared with the baseline method, FlexINA can increase the aggregation throughput by up to 2.2 times and 2.8 times. ATP and A2TP adopt the best-effort INA method, which introduces memory contention and thus reduces the memory aggregation efficiency. In contrast, FlexINA avoids memory contention through selective scheduling. Based on this, FlexINA also adopts INA-aware congestion control to reduce aggregator waste and, through adaptive memory allocation, enables non-INA jobs to fill the available memory. The figure shows that FlexINA can reduce the average iteration time by up to 50.3% and 59.5% compared with the two baseline methods. This is because FlexINA adopts an unfair bandwidth and memory allocation strategy that prioritizes the completion of INA jobs. The fair competition model used by ATP and A2TP causes the two jobs to compete with each other during the communication phase, extending the communication time. In addition, the aggregate count in the switch and the total aggregation count of the parameter server are recorded every 10 ms. FlexINA achieves the highest aggregate count on the switch and the lowest aggregation count on the PS. This shows that FlexINA not only improves the in-network aggregation efficiency but also helps avoid bottlenecks on the PS. It is observed that A2TP is slightly higher than ATP in terms of switch and PS aggregate counts because, to avoid memory conflicts, ATP may send more packets to the PS for aggregation;
[0079] The performance of FlexINA in a multi-task environment is evaluated by restarting multiple identical jobs. Specifically, the total number of jobs is increased from 1 to 8 (i.e., 1, 2, 4, 8 jobs), while keeping the total memory capacity unchanged. Each job consists of 8 worker nodes randomly distributed across multiple racks. The average iteration time is used as the evaluation metric for FlexINA and the baseline methods. Figure 6 Shows the average iteration time of each model under varying numbers of jobs. Overall, FlexINA achieves the best performance in all cases. Compared with ATP and A2TP, when the number of jobs is 8, FlexINA reduces the latency by 64.8% and 44.8% respectively. As the number of jobs increases, the average iteration time of ATP grows exponentially. This is because the increase in the number of jobs leads to more severe memory conflicts, which in turn exacerbates network congestion. The TCP-like congestion control adopted by ATP fails to alleviate memory conflicts in this case. When congestion occurs, due to the drastic fluctuation of the congestion window, bandwidth is wasted. ATP avoids severe memory conflicts through decoupling. However, as before, ATP sends more packets to the PS. The network bottleneck on the PS side then leads to a decrease in the sending rate of ATP. For the ResNet50 model with low communication requirements and dominant computational time, the iteration time hardly changes as the number of jobs increases.
[0080] Working principle and usage method of the present invention:
[0081] Working principle
[0082] (I) System architecture
[0083] The system architecture mainly includes working nodes, switches, and a parameter server (PS). The working nodes are responsible for sending gradient data packets, the switch performs the aggregation operation of the data packets, and the parameter server performs the final aggregation processing of the data packets and feeds back the aggregation degree information.
[0084] (II) Core mechanism
[0085] 1. Memory request and allocation
[0086] Memory request module: When a new training job arrives, it is defaulted to a non-in-network aggregation job, and its gradient data packets are opportunistically aggregated within the network on the switch. The non-in-network aggregation job periodically sends REQUEST packets to request switch memory resources. If the request is successful, the job is converted to an in-network aggregation job.
[0087] Adaptive memory management module: This module on the switch side allocates memory for in-network aggregation jobs according to the memory requests, and allows non-in-network aggregation jobs to use the idle aggregators left by in-network aggregation jobs for aggregation. The switch organizes the memory into addressable aggregators and maintains the number c of jobs currently allocated to the memory. If c is less than the memory requirement q of the job, the switch responds to the new REQUEST packet, allocates a dedicated area for the job, and adds an entry in the Address table.
[0088] 2. Aggregation degree feedback
[0089] PS processing: When the data packet arrives at the PS, the PS performs in-server aggregation on non-in-network aggregation jobs and performs aggregation degree feedback on in-network aggregation jobs. The PS determines the INA degree by parsing the bitmap of the received gradient data packet. Each data packet contains an N-bit bitmap indicating whether the gradient from the i-th working node has been aggregated in the data packet. The PS performs a bitwise "OR" operation on the bitmaps in the M received data packets to calculate the INA degree d, and carries the calculated INA degree d in the aggregated data packet and broadcasts it to all working nodes. For non-in-network aggregation jobs, the INA degree d is set to 0.
[0090] 3. Congestion control
[0091] ECN Marking and Window Adjustment: The switch uses ECN to mark network congestion. If the queue length of the switch port exceeds the preset threshold, the gradient packets sent by the working node will be marked with the CE code point. The PS receives the marked packets, aggregates the CE marks, and updates the link congestion level factor ai. If ai = 0, the sliding window size cwnd increases by 1 MTU every RTT period; otherwise, the window size cwnd decreases according to a specific formula. For INA jobs with an INA degree d greater than 0, the DCTCP window reduction strategy is followed; for non-INA jobs with an INA degree d of 0, the window size is always halved.
[0092] 4. Static Intra-network Aggregation and Opportunistic Memory Preemption
[0093] Static Intra-network Aggregation: The switch assigns a fixed memory area to each in-network aggregation job and maps the data packets to specific aggregators in the allocated memory area according to their positions in the sliding window. When the switch receives a data packet, it aggregates the data packet into the corresponding aggregator according to the mapping relationship.
[0094] Opportunistic Memory Preemption: When the switch detects a data packet of a non-in-network aggregation job, it calculates the hash value and offset of the data packet, checks whether the send window size of the corresponding area is less than the offset. If so, it allows the data packet to be aggregated in this aggregator and updates the memory allocation table to record the memory usage of the non-in-network aggregation job.
[0095] (III) Task Scheduling and Performance Optimization
[0096] Task Scheduling Mechanism: In a multi-task environment, the in-network aggregation jobs are preferentially scheduled through the task scheduling mechanism to reduce memory conflicts and network congestion. The task scheduling mechanism dynamically adjusts the execution order of tasks according to the current network state and memory usage, ensuring that the in-network aggregation jobs can make full use of the allocated aggregator resources and improve the aggregation efficiency.
[0097] Performance Optimization Model: Based on the performance evaluation results of the method, a performance optimization model is established, comprehensively considering factors such as network topology, switch performance, the number of computing nodes, and the scale of training tasks. Through mathematical modeling and data analysis, the optimal memory allocation strategy, congestion control parameters, and key parameters of the task scheduling strategy are determined to guide the parameter configuration and system optimization in actual deployment.
[0098] II. Usage Method
[0099] (I) Environment Setup
[0100] Network topology: Build a 100 Gbps leaf-spine topology, including two leaf switches and one spine switch. All switches are enabled with ECN marking, and the ECN threshold is set to 100 packets. The leaf switches allow in-network aggregation and initialize 340 aggregators to meet the peak throughput requirements for a single task.
[0101] Task configuration: Run multiple cross-rack training tasks, each task containing multiple worker nodes and one parameter server. The worker nodes are evenly distributed across multiple racks.
[0102] (2) Method implementation steps
[0103] When a new training job arrives, the worker nodes by default send gradient data packets to the parameter server (PS) for aggregation, and continuously request memory resources from the switch during this process. The out-of-network aggregation jobs periodically send REQUEST packets to request switch memory resources. The switch decides whether to allocate memory based on the memory allocation table and the current memory usage. If the memory request is successful, the job is converted into an in-network aggregation job. The switch allocates a fixed memory area for each in-network aggregation job and maps the data packets to specific aggregators in the allocated memory area according to the position of the data packets in the sliding window. When the switch receives a data packet, it aggregates the data packet into the corresponding aggregator according to the mapping relationship. When the switch detects a data packet of an out-of-network aggregation job, it calculates the hash value and offset of the data packet, checks whether the size of the sending window in the corresponding area is less than the offset. If so, it allows the data packet to be aggregated in this aggregator, updates the memory allocation table, and records the memory usage of the out-of-network aggregation job. When the data packet arrives at the PS, the PS performs in-server aggregation on the out-of-network aggregation job and performs aggregation degree feedback on the in-network aggregation job. The PS determines the INA degree by parsing the bitmap of the received gradient data packet, carries the calculated INA degree d in the aggregated data packet, and broadcasts it to all worker nodes. The switch uses ECN to mark network congestion. If the queue length of the switch port exceeds the preset threshold, the gradient data packets sent by the worker nodes will be marked with the CE code point. The PS receives the marked data packets, aggregates the CE marks, and updates the link congestion level factor ai. According to the value of ai, it adjusts the sliding window size cwnd to manage the sending rate of data packets. This method and its benchmark method are implemented through the OMNET++ simulator to evaluate the overall performance improvement. In a multi-task environment, the total number of jobs is increased from 1 to 8 (i.e., 1, 2, 4, 8 jobs), while keeping the total memory capacity unchanged. Using the average iteration time as the evaluation metric, based on the performance evaluation results, a performance optimization model is established to determine the key parameters of the optimal memory allocation strategy, congestion control parameters, and task scheduling strategy. According to the requirements of the actual application scenario, the system parameters are configured to achieve the efficient execution of distributed training tasks and improve the overall performance and resource utilization rate of the system.
[0104] The above is only used to illustrate the technical solution of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; any equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present invention, directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present invention.
Claims
1. A method for selecting an aggregate transmission method for in-network aggregation, characterized in that: It includes the following steps: Step 1: At the worker node side, when a new job arrives, the worker node defaults to sending gradient data packets to the parameter server (PS) for aggregation and continuously requests memory resources from the switch during this process; Step 2: The adaptive memory management module at the switch side allocates memory for in-network aggregation jobs according to the requests of the memory request module and allows non-in-network aggregation jobs to use the idle aggregators left by in-network aggregation jobs for aggregation; Step 3: When the data packet arrives at the PS, the PS performs in-server aggregation on non-in-network aggregation jobs and performs aggregation degree feedback on in-network aggregation jobs; Step 4: All jobs use an in-network aggregation-aware congestion control mechanism to manage the sending rate of data packets, preferentially reducing the sending rate of non-in-network aggregation jobs to avoid a decrease in the memory aggregation efficiency of in-network aggregation jobs.
2. The method for selecting an aggregation transmission method for in-network aggregation according to claim 1, characterized in that: The memory request module includes the following: A new training job is defaulted to a non-in-network aggregation job, and gradient data packets perform opportunistic in-network aggregation on the switch; Non-in-network aggregation jobs periodically send REQUEST data packets to request switch memory resources; The aggregation degree feedback includes the following: The PS determines the INA degree by parsing the bitmap of the received gradient data packet; Each data packet contains an N-bit bitmap indicating whether the gradient from the i-th worker node has been aggregated in the data packet; The PS performs a bitwise "OR" operation on the bitmaps in the M received data packets to calculate the INA degree d as follows; The PS carries the calculated INA degree d in the aggregation data packet and broadcasts it to all worker nodes; For non-in-network aggregation jobs, the INA degree d is set to 0.
3. The method for selecting an aggregation transmission method for in-network aggregation according to claim 2, characterized in that: The congestion control includes the following: the switch uses ECN to mark network congestion; if the queue length of the switch port exceeds a preset threshold, the gradient data packet sent by the working node will be marked as a CE code point; the PS receives the marked data packet, aggregates the CE mark, and updates the link congestion level factor a i ; a i =(1-g)×a i-1 +g×F; Where F is the proportion of data packets marked as ECN in the previous window, and 0 < g < 1 is the weight factor; If a i =0, the sliding window size cwnd increases by 1 MTU in each RTT period; otherwise, the window size cwnd decreases as follows: The INA degree d of an INA job is greater than 0, and it follows the DCTCP window reduction strategy; The INA degree d of a non-INA job is 0, and the window size is always halved; Adaptive memory allocation includes the following steps: The switch organizes the memory into addressable aggregators and maintains the number c of jobs currently allocated to the memory; The data packets in the sliding window of an INA job are statically mapped to the switch memory, using the position of the data packet in the sliding window as the offset of its aggregation index; The data packets of non-INA jobs are hashed to any aggregator index, and the switch checks whether the sending window size of the corresponding area is less than this offset.
4. The method for selecting an aggregation transmission method for in-network aggregation according to claim 3, characterized in that: This method further includes implementing FlexINA and its baseline method based on the OMNET++ simulator to evaluate the overall performance improvement. It is applicable to a multi-task environment and can maintain the best performance under different numbers of tasks. Specifically, it includes the following steps: Evaluating the performance of the method in a multi-task environment by restarting multiple identical jobs; Increasing the total number of jobs from 1 to 8 (i.e., 1, 2, 4, 8 jobs) while keeping the total memory capacity unchanged; Each job consists of 8 worker nodes randomly distributed on multiple racks; Using the average iteration time as the evaluation metric for FlexINA and the baseline method.
5. The method for selecting an aggregation transmission method for in-network aggregation according to claim 4, characterized in that: The congestion control mechanism of network aggregation awareness includes the following: when the switch detects network congestion, it gives priority to reducing the sending rate of non-network aggregation jobs. Through the ECN marking mechanism, the switch feeds back the network congestion status to the working node. The working node adjusts the sending window size according to the received ECN mark to reduce network congestion. The adaptive memory management module includes the following: the switch maintains a memory allocation table to record the memory allocation status of each job. When the switch receives a memory request, it decides whether to allocate memory according to the memory allocation table and the current memory usage. When allocating memory, the switch gives priority to the network aggregation job and records the allocated memory area. When the switch detects that the memory is idle, it allows the non-network aggregation job to use the idle memory for aggregation.
6. The method for selecting an aggregation transmission method for in-network aggregation according to claim 5, characterized in that: Static in-network aggregation includes the following steps: the switch allocates a fixed memory area for each in-network aggregation job, the switch maps the data packet to a specific aggregator in the allocated memory area according to the position of the data packet in the sliding window, and when the switch receives the data packet, it aggregates the data packet to the corresponding aggregator according to the mapping relationship. Opportunistic memory preemption includes the following steps: when the switch detects a data packet of a non-in-network aggregation job, it calculates the hash value and offset of the data packet.
7. The method for selecting an aggregation transmission method for in-network aggregation according to claim 6, characterized in that: The method further includes using a 100Gbps leaf-spine topology, including two leaf switches and one spine switch, all switches have ECN marking enabled, and the ECN threshold is set to 100 packets, the leaf switches allow in-network aggregation, and initialize 340 aggregators to meet the peak throughput requirements of a single task, run multiple cross-rack training tasks, each task contains multiple working nodes and a parameter server, and evaluate the performance under different configurations by adjusting the maximum window size and hash space size.
8. The method for selecting an aggregation transmission method for in-network aggregation according to claim 7, characterized in that: The method further includes statically mapping the data packets in the sliding window of the network aggregation operation to the switch memory, using the position of the data packet in the sliding window as the offset of its aggregation index, and when the switch receives the data packet, it obtains the base address of the memory area by querying the allocation table, and combines it with the offset carried in the data packet to obtain the physical address. After the aggregation is completed, the switch sends the aggregation result to the downstream device. In a multi-task environment, the task scheduling mechanism is used to prioritize the scheduling of network aggregation operations to reduce memory conflicts and network congestion. The task scheduling mechanism dynamically adjusts the execution order of tasks according to the current network status and memory usage, and ensures that the network aggregation operations can fully utilize the allocated aggregator resources through task scheduling.
9. The method for selecting an aggregation transmission method for in-network aggregation according to claim 8, characterized in that: The method further includes evaluating the performance stability under different network load conditions, simulating three different network load scenarios of high, medium and low, running multiple training tasks respectively, recording and analyzing the changes in the average iteration time, packet loss rate and aggregation efficiency key performance indicators of the method under different loads, and comparing and analyzing the performance differences between the method and the prior art methods under different load scenarios.
10. The method for selecting an aggregation transmission method for in-network aggregation according to claim 9, characterized in that: Based on the performance evaluation results, existing data is used to determine the parameters in the model, and new data is used to verify whether the model is accurate. Finally, the system parameters are adjusted according to the results of the model to establish a performance optimization model. Through mathematical modeling and data analysis, the optimal memory allocation strategy, congestion control parameters, and key parameters of the task scheduling strategy are determined.
Citation Information
Cited By
Model training acceleration method and system based on in-network gradient data annular full reduction
CN121660028A