Method and system for accelerating hybrid expert model training on multiple GPUs through dynamic sample placement
By dynamically adjusting the position of the training samples, combining data locality and network locality, the All-to-All communication of the hybrid expert model is optimized, solving the problem of high communication costs and achieving more efficient model training.
Patent Information
- Application Number
- CN202510070284.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-13
AI Technical Summary
During the training of hybrid expert models, frequent All-to-All communication leads to high time costs and affects training efficiency.
By leveraging the data locality in expert routing and the network locality between training devices, the location of training samples is dynamically adjusted to reduce the traffic between nodes and optimize the All-to-All communication to intra-node communication.
Without affecting model convergence and adding additional overhead, All-to-All communication is significantly accelerated, and the training speed of hybrid expert models is improved, which is better than existing popular MoE training systems.
Smart Images

Figure CN120144273A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and particularly relates to a method and system for accelerating the training of a mixture of experts model on multiple GPUs through dynamic sample placement. Background Art
[0002] In recent years, due to the continuous increase in model scale, large language models (LLMs) have shown excellent performance in language understanding and generation. However, larger models are usually accompanied by higher computational costs, making further expansion difficult. To address this issue, the mixture of experts model (MoE) has been introduced, which can significantly expand the model scale without increasing the computational cost. Combining MoE with Transformer-based models can achieve excellent performance in various tasks, including natural language processing, computer vision, recommendation systems, and speech recognition.
[0003] The MoE model usually replaces the feed-forward network (FFN) layer with an MoE layer, which consists of a gating network and several small feed-forward networks representing different experts. In the MoE layer, each token is routed by the gating network to several selected experts, and the final output is obtained by weighted summation of the calculation results of the selected experts. Through this method, the number of experts can be increased to expand the model scale while keeping the computational complexity unchanged, thereby obtaining better performance.
[0004] The most important challenge in effectively training the MoE model lies in how to efficiently perform All-to-All communication. Because each All-to-All communication requires all tokens to participate, and considering forward and backward propagation, each MoE layer needs to perform four All-to-All communications per training iteration. Such frequent and heavy communication will incur a huge time cost. Therefore, accelerating All-to-All communication is crucial for improving training efficiency.
[0005] Existing distributed MoE training systems generally adopt an expert parallel training method, where expert parameters are allocated on different GPU devices and other parameters are replicated. In each MoE layer, each token will be routed by the gating network to the top-k different experts for processing, where k is a hyperparameter, usually a very small value, such as 1 or 2, which helps to reduce the computational complexity. After obtaining the gating routing, the MoE layer will send the tokens to the devices where the corresponding experts are located according to the routing. Then, the calculation results of the experts will be sent back to the original device where the tokens are located. Since the experts are distributed on different devices, all devices need to send and receive information from each other during this process, which is the All-to-All communication mentioned above.
[0006] Based on expert parallelism, the researchers further proposed the following methods to accelerate training:
[0007] 1) Dynamic expert placement: Some studies have found that during training, expert routing exhibits data locality, that is, training data often shows a preference for certain experts. Accordingly, the researchers proposed a method to dynamically adjust the positions of experts during training based on historical routing results to reduce communication volume. For example, popular experts can be replicated and placed on more devices, so that the communication volume related to them will be reduced.
[0008] 2) Modify the model definition: Due to the dynamic nature of expert routing, the number of tokens received by different experts is usually uneven. To this end, some researchers have achieved better load balancing by modifying the model structure. Some methods modify the routing mechanism to balance the load among experts; some methods notice the network locality between GPU devices, that is, the communication speed within a node is much faster than the communication speed between nodes. Accordingly, a routing topology loss is introduced to preferentially process routing tokens within the same node, thereby reducing inter-node communication; some methods compress the tokens to a smaller hidden layer dimension before performing All-to-All communication to reduce the total communication load.
[0009] When training the MoE model in a distributed manner, the method of dynamic expert placement will incur additional overhead for transmitting and storing expert parameters between GPU devices. In addition, these methods only heuristically adjust the positions of experts and theoretically cannot guarantee to what extent the communication speed can be accelerated; while the method of modifying the model definition will affect the convergence of the original model. When applying these methods, multiple experiments are usually required to adjust the hyperparameters to ensure model convergence, which affects the practicality of the methods. Summary of the Invention
[0010] The present invention addresses the above problems and provides a method and system for accelerating the training of a mixture-of-experts model on multiple GPUs through dynamic sample placement.
[0011] The technical solution adopted by the present invention is as follows:
[0012] A method for accelerating the training of a mixture-of-experts model on multiple GPUs through dynamic sample placement, comprising the following steps:
[0013] Utilize the data locality in expert routing and the network locality between training devices to dynamically adjust the positions of training samples during the training of the mixture-of-experts model according to the expert routing results;
[0014] Utilize the positions of the dynamically adjusted training samples to accelerate All-to-All communication and optimize the training speed of the mixture-of-experts model.
[0015] Further, model the cost of All-to-All communication, and formulate the dynamic adjustment of the positions of training samples as a combinatorial optimization problem, and find the optimal sample placement position that maximizes efficiency given the expert routing.
[0016] Further, the modeling of the cost of All-to-All communication includes:
[0017] For each All-to-All communication, its time cost is expressed by the following formula:
[0018] t = max(t intra , t inter )
[0019]
[0020] where t intra represents the time cost of the intra-node channel, and t inter represents the time cost of the inter-node channel; s intra represents the traffic of the intra-node channel, and s inter represents the traffic of the inter-node channel; v intra represents the bandwidth of the intra-node channel, and v inter represents the bandwidth of the inter-node channel; α intra represents the delay cost of the intra-node channel, and α inter represents the delay cost of the inter-node channel; β intra represents the bandwidth cost of the intra-node channel, and β inter represents the bandwidth cost of the inter-node channel;
[0021] S intra and S inter are calculated by the device numbers of the experts and samples:
[0022] S intra = {(i, e)|Node(Smpdev(i)) = Node(Expdev(e)) ∧ Smpdev(i) ≠ Expdev(e)}
[0023] S inter = {(i, e)|Node(Smpdev(i)) ≠ Node(Expdev(e))}
[0024] where Expdev(e) is the device number where the e-th expert is located, Smpdev(i) is the device number to which the i-th sample should be routed, and Node(j) is the node number of the j-th device.
[0025] Further, the combinatorial optimization problem is:
[0026]
[0027] Among them, t (l,gather) represents the time of the Gather operation at the l-th layer, and t (l+1,scatter) represents the time of the Scatter operation at the (l + 1)-th layer, represents the intra-node communication time of the Gather operation at the l-th layer, represents the inter-node communication time of the Gather operation at the l-th layer, represents the intra-node communication time of the Scatter operation at the (l + 1)-th layer, represents the inter-node communication time of the Scatter operation at the (l + 1)-th layer.
[0028] Furthermore, the combinatorial optimization problem is solved by adopting a two-stage solution strategy: in the first stage, t inter is optimized globally, and in the second stage, without affecting t inter , the t intra within each node is minimized.
[0029] Furthermore, the two-stage solution strategy includes the following steps:
[0030] Suppose there are N nodes, and each node consists of J / N devices. The optimization formula in the first stage is written as the following ILP problem:
[0031]
[0032] After obtaining the optimal solution in the first stage, in the second stage, considering separately rearranging the samples within each node, for the n-th node, let be the set of samples assigned to this node after solving the optimization formula in the first stage, and let {J} n = {j|j ∈ {J} ∧ Node(j) = n} be the set of experts on this node. To optimize the n-th node, solve the following ILP problem:
[0033]
[0034] The second stage consists of N ILP problems, and each problem corresponds to solving the sample placement scheme on each node.
[0035] Furthermore, the combinatorial optimization problem is solved by adopting a polynomial-time algorithm. The polynomial-time algorithm transforms the ILP problem into a weighted bipartite graph matching problem, and then performs polynomial-time solution based on the KM algorithm.
[0036] A system for accelerating the training of a mixture-of-experts model on multiple GPUs through dynamic sample placement, which includes:
[0037] A dynamic sample placement module, which is used to dynamically adjust the positions of training samples during the training of a mixture-of-experts model according to the expert routing results by leveraging the data locality in expert routing and the network locality among training devices.
[0038] A training module, which is used to accelerate the All-to-All communication and optimize the training speed of the mixture-of-experts model by using the positions of the dynamically adjusted training samples.
[0039] The beneficial effects of the present invention are as follows:
[0040] Existing work has not explored methods for optimization from the data perspective. However, the present invention can start from the data perspective without affecting convergence and introducing additional overhead, combine the data locality in expert routing with the network locality among training devices to accelerate the All-to-All communication, and thus optimize the training speed of MoE. Description of the Drawings
[0041] Figure 1 is an example diagram of an expert exchange.
[0042] Figure 2 is an example of a weighted bipartite graph. Detailed Embodiments
[0043] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below through specific embodiments and the accompanying drawings.
[0044] The present invention notices two characteristics during the MoE training process: data locality and network locality. Data locality means that training data often shows a preference for certain experts (the experts in the present invention are also called expert networks, which are neural networks, such as feed-forward networks FFN, etc.), and the corresponding distribution of this preference is usually unbalanced; network locality is an inherent characteristic of modern clusters for deep learning training. Specifically, there are various communication channels in modern clusters. For example, devices within a node usually communicate through PCIe or NVLINK, while devices between nodes use Ethernet or InfiniBand. Intra-node communication is usually faster than inter-node communication.
[0045] Accordingly, the present invention proposes a new training system CommMoE, which can combine data locality in expert routing with network locality among training devices. CommMoE dynamically adjusts the positions of data samples during training according to the expert routing results, enabling more tokens to be transmitted through high-speed channels rather than low-speed channels. However, there are several difficulties in implementing dynamic sample placement: firstly, how to adjust the placement positions to maximize efficiency is a complex and unexplored problem; secondly, since adjustments need to be made to each layer in each iteration, it is crucial to design an efficient algorithm to instantaneously derive the placement positions. To solve these problems, the present invention models the cost of All-to-All communication and formulates the dynamic sample placement problem as a combinatorial optimization problem. Subsequently, it is divided into two stages to simplify the solution, and a corresponding polynomial-time algorithm is designed to ensure that the sample placement positions can be efficiently solved.
[0046] The main contributions of the present invention are as follows:
[0047] (1) Firstly, the present invention firstly proposes a method to accelerate All-to-All communication through dynamic sample placement by utilizing data locality and network locality.
[0048] (2) Secondly, the present invention formulates the dynamic sample placement problem as a combinatorial optimization problem, aiming to find the optimal sample placement positions that maximize efficiency given the expert routing.
[0049] (3) Thirdly, the present invention decomposes the problem into two stages and develops a polynomial-time solution to efficiently derive the sample placement positions during training.
[0050] (4) Finally, the present invention is 1.67 times superior to the current popular MoE training system in terms of training efficiency.
[0051] First, the way to model All-to-All communication in MoE training will be described below, and how to model the optimization problem will be elaborated. Then, how to solve this optimization problem will be explained. Finally, the necessary implementation details will also be introduced. For clarity, Table 1 will list the common symbols used in the present invention.
[0052] Table 1 Common symbols used in the present invention
[0053] L The number of tokens per sample. H The hidden dimension size per token. E The number of experts in the MoE layer. K The number of experts to which each token is routed. I The number of samples per iteration (also known as the global batch size). J The number of devices (i.e., GPUs). N The number of nodes (machines). <![CDATA[1 [·] > Indicator function. {n} The set of natural numbers less than n, i.e., {0, 1, …, n - 1}.
[0054] (1) Problem Modeling
[0055] (1.1) Communication Modeling: First, discuss the optimization goal of CommMoE - All-to-All communication. Use the α-β model to model All-to-All communication, where α represents the latency cost and β represents the bandwidth cost. Specifically, divide communication into three categories: intra-device communication, intra-node communication, and inter-node communication, and each type of communication uses a different channel. Since intra-device communication is usually implemented through memory copying, it is much faster than the other two types of communication, so it is not considered in the modeling of this invention. Therefore, the communication time is determined by the maximum time required to transmit data across intra-node and inter-node channels. The bandwidths of the intra-node channel and the inter-node channel are represented by v intra and v inter respectively. Therefore, for each All-to-All communication, its time cost can be expressed by the following formula, where s. represents the traffic of the corresponding channel.
[0056]
[0057] The bandwidth (v.) and latency (α.) can be obtained by performing a performance analysis of the hardware environment before training, while the traffic (s.) needs to be dynamically determined according to the routing results within the MoE layer. The calculation method of the traffic is as follows: Let route ∈ N I×L×K be the token routing result of the gating network, representing that each token will be sent to K experts. Then the number of tokens that the i-th sample needs to send to the e-th expert can be denoted as:
[0058]
[0059] where route i,l,k represents the number of the k-th selected expert to which the l-th token of the i-th sample is sent by the gating network.
[0060] Next, num ∈ N I×E can be used to model the traffic between different channels: Let Expdev(e) be the device number where the e-th expert is located, Smpdev(i) be the device number to which the i-th sample should be routed, and Node(j) be the node number of the j-th device. Considering the traffic as the number of tokens to be transmitted, we can get:
[0061]
[0062] where S intra and S inter can be calculated through the device numbers of the experts and samples:
[0063]
[0064] (1.2) Rationality of dynamic sample placement: Based on the above modeling, it can be found that the time cost of All-to-All communication is highly related to the placement of experts and samples. Considering that network locality is prevalent in distributed clusters, i.e., v intra >v inter , it is possible to reduce the communication volume between nodes by adjusting the sample placement, even if the intra-node communication volume increases slightly.
[0065] Take Figure 1 as an example to explain the advantages of dynamic sample placement in more detail, where I = 4, L = 4, E = 4, K = 1, and both experts and samples are placed in sequence, i.e., Expdev = [0, 1, 2, 3], Smpdev = [0, 1, 2, 3]. Figure 1 In (a), it shows that a MoE layer will experience two All-to-All operations during forward propagation, denoted by All-to-All scatter (fully-to-fully distribution) and All-to-All gather (fully-to-fully aggregation) respectively. Figure 1 In (b), it shows the communication volume without changing the sample placement. According to Formula 3 and Formula 4, if the present invention only considers the transmission volume of node 0, then S inter ={(0, 2), (0, 3), (1, 2), (1, 3)}, s inter = 5. However, when the sample position is optimized (as shown in Figure 1 (c)), i.e., Smpdev = [3, 1, 2, 0], the corresponding inter-node communication volume becomes S inter ={(3, 2), (3, 3), (1, 2), (1, 3)}, i.e., s inter = 2. In addition, the sample position adjustment can be combined with the All-to-All Gather operation, that is, instead of restoring the tokens to their original positions, the tokens are directly placed at the new positions according to the changed sample positions. This way directly optimizes the current communication operation without introducing any additional communication.
[0066] (1.3) Problem modeling: After determining the sample placement adjustment, it can be seen that changing Smpdev will affect two All-to-All operations: the Gather operation of the current MoE layer and the Scatter operation of the next MoE layer. Therefore, the optimization of the present invention targets these two operations. For the l-th layer, the optimization problem can be written in the following form:
[0067]
[0068] where, t (l,gathrt) represents the time of the Gather operation of the l-th layer, t (l+1,scatter)Denotes the time of the Scatter operation at the (l + 1)-th layer, Denotes the intra-node communication time of the Gather operation at the l-th layer, Denotes the inter-node communication time of the Gather operation at the l-th layer, Denotes the intra-node communication time of the Scatter operation at the (l + 1)-th layer, Denotes the inter-node communication time of the Scatter operation at the (l + 1)-th layer.
[0069] Since a change in sample placement affects two All-to-All operations, both communication times are included in the optimization objective. Additionally, to ensure computational and memory balance between devices, each device should retain the same number of samples before and after sample placement adjustment. This forms the basis of the optimization constraints of the present invention.
[0070] (2) Problem Solving
[0071] Equation (5) is a complex combinatorial optimization problem for which an optimal solution cannot be obtained in polynomial time. As the cluster size increases, it may take a considerable amount of time to find even an approximate solution. Since this problem needs to be solved before each Gather operation at each MoE layer, directly solving it would result in unaffordable additional time costs. To solve this problem, the present invention designs an effective method to obtain an approximate solution. Specifically, first, the optimization problem is decomposed into two stages, and a polynomial-time algorithm is developed to implement the solution, the details of which are as follows:
[0072] Two-stage solution: Although Equation (1) takes the maximum of two communication costs, in practice, due to the large difference in bandwidth between inter-node and intra-node connections, the most time-consuming is usually the inter-node communication time. Therefore, the present invention proposes a two-stage solution strategy: The first stage optimizes t inter globally, while the second stage minimizes t inter within each node without affecting t intra .
[0073] Formally, assuming there are N nodes, each node consisting of J / N devices, the optimization formula for the first stage can be written as the following integer linear programming (ILP) problem:
[0074]
[0075] Since the first stage focuses on inter-node communication, the constraint of cross-device balance in Equation (5) is transformed into cross-node balance. After obtaining the optimal solution of the first stage, the second stage considers separately rearranging the samples within each node. For the n-th node, let The sample set assigned to this node after solving formula 6. Let {J} n ={j|j∈{J}∧Node(j)=n} be the set of experts on this node ({J} n is determined by the device location, rather than obtained through formula 6). Next, to optimize the nth node, the following ILP problem should be solved:
[0076]
[0077] The second stage consists of N ILP problems, each corresponding to the sample placement scheme on each node. They are independent of each other and can be solved concurrently.
[0078] Polynomial solution method: By decomposing the original combinatorial optimization problem, N + 1 ILP problems are obtained. Since these problems need to be solved for each layer in each training iteration, the efficiency of solving the problems is crucial. However, since the ILP problem is NP-Hard, the time cost of brute-force solution exceeds the time cost of the Scatter operation and expert calculation, so this is unrealistic.
[0079] Considering that each sample must be assigned to a device, the present invention regards the ILP problem as an assignment problem, transforms them into a weighted bipartite graph matching problem, and then develops a polynomial-time solver based on the widely used Kuhn-Munkres (KM) algorithm. First, introduce how to transform the ILP problem into a weighted bipartite matching problem: Let c i,n and c′ i,j respectively represent the inter-node and intra-node communication volumes generated by placing the i-th sample on the j-th device of the nth node, and they can be calculated by the following formulas:
[0080]
[0081] where S represents the set of device numbers not located at the nth node, and S′ represents the set of device numbers located at the nth node and not numbered j.
[0082] To make the expression clearer, let p i,n represent whether the i-th sample is placed on the nth node, and p′ i,j represent whether the i-th sample is placed on the j-th device. Then the optimization objective can be expressed as:
[0083]
[0084] After modifying the corresponding constraints, the ILP problem is transformed into a (0,1)-ILP problem. For example, the following gives the optimization problem in the first stage after transformation:
[0085]
[0086] The optimization problem in the second stage is similar to that in the first stage, which will not be elaborated here. This (0,1)-ILP problem can be modeled as a weighted bipartite matching problem: Consider a bipartite graph with two sets of graph nodes P and Q. Set P represents all training samples, and |P| = I. Set Q represents all training nodes (machines), where each training node can process B := I / N training samples. To model this, each graph node in Q is replicated B times, resulting in |Q| = I. There are weighted edges between each pair of graph nodes from P and Q. Let P i represent the i-th training sample, and Q n represent the -th training node. The weight of the edge between P i and Q n is denoted as This transformation simplifies the problem of finding the minimum-weight perfect matching in this bipartite graph, and the Kuhn-Munkres (KM) algorithm can be used to efficiently solve it in polynomial time to obtain the optimal solution. Figure 2 illustrates an example of constructing the bipartite graph in the Figure 1 first stage. The graph nodes on the left represent set P, and the graph nodes on the right represent set Q. Each pair of graph nodes is connected by a weighted edge, represented by a dashed line. The red edges represent the final matching scheme, where the total weight of all matching edges is minimized.
[0087] (3) Implementation of the CommMoE System
[0088] CommMoE is implemented on top of PyTorch, and some related custom operations (e.g., the calculation of num, c, c′ and the KM algorithm) are implemented in C++ and CUDA. The complete workflow of CommMoE includes the following steps:
[0089] (3.1) For each sub-module in the model, perform the following operations:
[0090] (3.2) If the sub-module is a MoE layer, perform the following steps (3.3) - (3.9):
[0091] (3.3) Obtain the routing information from the gating network and calculate the parameter num through formula (2).
[0092] (3.4) Call the Solve(num) function in a background thread to perform the solution process.
[0093] (3.5) Obtain the input data through All-to-All Scatter.
[0094] (3.6) Perform expert calculations to obtain output data.
[0095] (3.7) Add the input and output data to obtain a new output.
[0096] (3.8) Obtain the best sample placement result from the background thread.
[0097] (3.9) Perform an All-to-All Gather operation according to the best sample placement.
[0098] (3.10) If the sub-module is not an MoE layer, perform the following step (3.11):
[0099] (3.11) Normally execute the training process of the sub-module.
[0100] Among them, for the Solve(num) function in step (3.4), it includes the following steps:
[0101] (3.12) Obtain parameters c and c′ according to formula (8) and construct a bipartite graph.
[0102] (3.13) Obtain the optimal solutions p and p′ through the Kuhn-Munkres algorithm.
[0103] (3.14) Determine the best allocation of samples according to the optimal solutions p and p′.
[0104] In addition to problem solving, CommMoE is also optimized in the following ways:
[0105] Expert residual inlining: In the classical MoE model, the residual connection is independent of the MoE layer. However, in CommMoE, the position of the training data changes after the All-to-All gather operation, while the samples in the residual connection remain in the original position. To ensure the correctness of the model, the present invention inlines the residual connection into the expert calculation, as shown in step (3.7) above. This optimization ensures the consistency of the model accuracy before and after applying the algorithm.
[0106] Solver offloading: The KM algorithm is difficult to parallelize and is not suitable for highly parallel accelerators such as GPUs. Therefore, the present invention offloads the solving process to the CPU for execution. As shown in step (3.4) above, after obtaining the routing result of the current layer, each device calculates and transfers num to the CPU memory. The predicted routing of the next layer required by the optimization algorithm can be obtained by putting the current layer input into the gating network of the next layer. The solver only needs to solve the new sample positions before the All-to-All Gather operation. In this way, the solving process can overlap with the All-to-All Scatter operation and the expert calculation without introducing any additional overhead.
[0107] (4) Technical effects
[0108] The present invention proposes CommMoE to optimize the main bottleneck in training MoE models, i.e., the All-to-All communication. By leveraging data and network locality, the method of the present invention dynamically adjusts the placement of training samples during training, transforming the inter-node communication into intra-node communication to improve the efficiency of All-to-All communication. The method proposed by the present invention can achieve higher training throughput compared with existing popular training systems.
[0109] Experimental settings: In the experiment, CommMoE was compared with existing methods based on dynamic expert placement, FasterMoE and SmartMoE, as well as FastMoE which did not adjust the positions of experts or samples. All experiments were conducted on a cluster consisting of 4 nodes, with each node equipped with 8 NVIDIA A800-SXM4-40GB GPUs. The GPUs within each node were connected via NVLink with a bandwidth of 400GB / s, while the nodes were interconnected via InfiniBand with a bandwidth of 100GB / s. The GPT model architecture was selected, and all FFN layers in each model were replaced with MoE layers. The number of experts was set to E = 2×J, where J is the number of GPUs in the corresponding experiment, and the number of experts selected for each token was fixed at K = 2.
[0110] Experimental comparison results: Under different cluster scales from 16 cards to 32 cards, CommMoE is 1.67 times faster than FastMoE, 1.37 times faster than FasterMoE, and 1.33 times faster than SmartMoE. Specifically, FasterMoE has achieved significant optimization by overlapping expert calculations and supporting dynamic expert placement. However, as the model hidden dimension increases, the cost of communicating expert weights rises, making it difficult to maintain the same acceleration level. This results in the performance gap between FasterMoE and CommMoE. On the other hand, SmartMoE mainly focuses on balancing the computational load and does not emphasize the optimization of communication efficiency. When communication becomes the main bottleneck, the benefits of load balancing become less obvious. Therefore, by dynamically adjusting the sample placement, CommMoE always outperforms state-of-the-art systems.
[0111] Application scenarios of the present invention:
[0112] The present invention can be used in fields such as natural language processing, computer vision, recommendation systems, and speech recognition. For example, in natural language processing, the method of the present invention can be used to train an autoregressive language MoE model, such as SwitchTransformer, so as to achieve the effect of accelerating the training efficiency. For computer vision, the method of the present invention can be used to train an autoregressive vision MoE model, such as V-MoE, so as to achieve the effect of accelerating the training efficiency.
[0113] Another embodiment of the present invention provides a system for accelerating the training of a mixture-of-experts model on multiple GPUs through dynamic sample placement, which includes:
[0114] A dynamic sample placement module, configured to utilize the data locality in the expert routing and the network locality between training devices to dynamically adjust the positions of training samples according to the expert routing results during the training process of the mixture-of-experts model;
[0115] A training module, configured to accelerate the All-to-All communication and optimize the training speed of the mixture-of-experts model by using the positions of the dynamically adjusted training samples.
[0116] It should be understood that the methods and systems disclosed in the above embodiments provided by the present invention can be implemented in other ways. For example, the above module division can have other division methods in actual implementation, and multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0117] Each module in the present invention can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium, including several instructions for causing a computer device to execute part or all of the steps of the method described in the present invention. For example, an embodiment of the present invention provides a computer device (such as a computer or a computer cluster, a server or a server cluster, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present invention. For example, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disc, etc.), the computer-readable storage medium stores a computer program, and when the computer program is executed by the computer, each step of the method of the present invention is implemented.
[0118] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention is subject to the scope defined by the claims.
Claims
1. A method for accelerating the training of mixed expert models on multiple GPUs through dynamic sample placement, characterized in that: The following steps are involved: By utilizing the data locality in expert routing and the network locality between training devices, the positions of training samples are dynamically adjusted according to the expert routing results during the training process of the hybrid expert model. The dynamically adjusted positions of training samples are used to accelerate All-to-All communication and optimize the training speed of the mixture of experts model.
2. The method according to claim 1, characterized in that The cost of All-to-All communication is modeled, and the dynamic adjustment of the position of training samples is formulated as a combinatorial optimization problem to find the best sample placement position that maximizes efficiency when expert routing is known.
3. The method according to claim 2, characterized in that The cost of All-to-All communication is modeled, including: For each All-to-All communication, the time cost is expressed by the following formula: t=max(t intra ,t inter ) Among them, t intra represents the time cost of the channel within the node, t inter represents the time cost of the channel between nodes; t intra represents the communication volume of the channel within the node, s inter represents the communication volume of the channel between nodes; v intra represents the bandwidth of the channel within the node, v inter represents the bandwidth of the channel between nodes; α intra represents the delay cost of the channel within the node, α inter represents the delay cost of the channel between nodes; β intra represents the bandwidth cost of the channel within the node, β inter represents the bandwidth cost of the channel between nodes; S intra and S inter Calculated from the equipment numbers of the expert and the sample: S intra ={(i,e)|Node(Smpdev(i))=Node(Expdev(e))∧Smpdev(i)≠Expdev(e)} S inter ={(i,e)|Node(Smpdev(i))≠Node(Expdev(e))} Among them, Expdev(e) is the device number where the e-th expert is located, Smpdev(i) is the device number to which the i-th sample should be routed, and Node(j) is the node number of the j-th device.
4. The method according to claim 3, characterized in that The combinatorial optimization problem is: Among them, t (l,gather) represents the time of the Gather operation at layer l, t (l+1,scatter) Indicates the time of the Scatter operation at the l+1th layer, represents the intra-node communication time of the Gather operation at layer l, Indicates the inter-node communication time of the Gather operation at layer l, represents the intra-node communication time of the Scatter operation at the l+1th layer, Indicates the inter-node communication time of the Scatter operation at the l+1th layer.
5. The method according to claim 4, characterized in that The combinatorial optimization problem is solved using a two-stage solution strategy: the first stage is to optimize t inter , the second stage does not affect t inter Under the premise of minimizing t in each node intra .
6. The method according to claim 5, characterized in that The two-stage solution strategy includes the following steps: Assuming there are N nodes, each node consists of J / N devices, the optimization formula of the first stage is written as the following ILP problem: After obtaining the optimal solution in the first stage, the second stage considers rearranging the samples in each node separately. For the nth node, let To solve the optimization formula in the first stage, we assign {J} to the node. n ={j|j∈{J}∧Node(j)=n} is the expert set on the node. To optimize the nth node, solve the following ILP problem: The second stage consists of N ILP problems, each of which corresponds to solving a sample placement plan on each node.
7. The method according to claim 6, characterized in that The combinatorial optimization problem is solved using a polynomial time algorithm, which transforms the ILP problem into a weighted bipartite graph matching problem and then performs a polynomial time solution based on the KM algorithm.
8. A system for accelerating the training of mixed expert models on multiple GPUs through dynamic sample placement, characterized in that: include: A dynamic sample placement module is used to dynamically adjust the location of training samples according to expert routing results during the training process of the hybrid expert model by utilizing data locality in expert routing and network locality between training devices; The training module is used to accelerate the All-to-All communication and optimize the training speed of the hybrid expert model by utilizing the dynamically adjusted positions of the training samples.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.