A Dynamic and Efficient Routing Method and Device for Hybrid Expert Large Model
By dynamically adjusting the number of experts in the hybrid expert big model, the uneven resource allocation problem caused by fixed Top-K experts is solved, and more efficient resource utilization and performance improvement is achieved.
Patent Information
- Application Number
- CN202411348957.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-09-26
AI Technical Summary
The existing hybrid expert model uses a fixed number of Top-K experts to deal with it, which lacks adaptability to different complexity tokens, resulting in uneven resource allocation and low efficiency.
A dynamic and efficient routing method is introduced, and the number of experts in each word is dynamically adjusted according to the complexity and importance of the input data through the allocator module, and a policy gradient algorithm is used to train the allocator module to optimize resource allocation.
It improves the utilization efficiency of computing resources, improves the performance and adaptability of the model, reduces unnecessary computing overhead, and optimizes the overall efficiency of the model.
Smart Images

Figure CN119514638B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a dynamic and efficient routing method and device for a mixture-of-experts large model. Background Art
[0002] While the number of parameters of large language models (LLMs) has increased significantly, the computational costs of their training and inference have also increased significantly. The mixture-of-experts (MoE) architecture balances model scale and computational costs by introducing expert networks and routing strategies. However, most existing MoE large models use a fixed number of top-K experts for processing, lacking adaptability to different complexity tokens, resulting in uneven resource allocation and low efficiency. Summary of the Invention
[0003] The present invention provides a dynamic and efficient routing method and device for a mixture-of-experts large model, aiming to solve the defects in the prior art that a fixed number of top-K experts are used for processing, lacking adaptability to different complexity tokens, resulting in uneven resource allocation and low efficiency, and to optimize the model performance and efficiency. The technical solutions proposed by the present invention are as follows:
[0004] In a first aspect, the present invention provides a dynamic and efficient routing method for a mixture-of-experts large model, including:
[0005] Obtain input data and a pre-trained mixture-of-experts large model, where the mixture-of-experts large model includes an allocator module and a router;
[0006] Divide the input data into multiple tokens, and use the allocator module to determine the optimal number of experts for each token;
[0007] Use the router to select an expert combination for each token according to the optimal number of experts, and route each token to the corresponding expert combination.
[0008] Optionally, before dividing the input data into multiple tokens and using the allocator module to determine the optimal number of experts for each token, the method further includes:
[0009] Train the allocator module using a policy gradient algorithm.
[0010] Optionally, the training of the allocator module using a policy gradient algorithm includes:
[0011] For each training sample, decompose the training sample into multiple tokens;
[0012] Use the allocator module to generate a probability distribution of the number of experts for each token;
[0013] Sample according to the probability distribution to determine the number of activated experts;
[0014] Use the activated experts to perform inference on the tokens and calculate the corresponding performance metrics;
[0015] Input the number of activated experts and the performance metrics into a pre-defined reward function to calculate the reward value for each training sample;
[0016] Based on the reward value, use the policy gradient algorithm to update the policy parameters of the allocator module to maximize the expected return.
[0017] Optionally, the allocator module includes an input layer, at least one hidden layer, and an output layer;
[0018] The step of dividing the input data into multiple tokens and using the allocator module to determine the optimal number of experts for each token includes:
[0019] Divide the input data into multiple tokens and generate a representation vector for each token;
[0020] The input layer receives the representation vector of each token;
[0021] At least one hidden layer extracts the token features of each representation vector;
[0022] The output layer outputs the probability distribution of the number of experts corresponding to each token according to the token features of each representation vector, and samples based on the probability distribution of the number of experts to select the optimal number of experts for each token.
[0023] Optionally, the expert combination includes at least one expert model; the method further includes:
[0024] Process the tokens based on each expert model in the expert combination and generate the corresponding output results;
[0025] Aggregate the output results of different expert models to obtain the model output.
[0026] Optionally, before training the allocator module using the policy gradient algorithm, the method further includes:
[0027] Replace each MoE layer in the mixture-of-experts large model with a dynamic K-routing layer including the allocator module.
[0028] In a second aspect, the present invention also provides a dynamic and efficient routing device for a mixture-of-experts large model, and the device includes the following modules:
[0029] An acquisition module for acquiring input data and a pre-trained mixture-of-experts large model, where the mixture-of-experts large model includes an allocator module and a router;
[0030] An allocation module for dividing the input data into multiple tokens and determining the optimal number of experts for each token using the allocator module;
[0031] A routing module for using the router to select an expert combination for each token according to the optimal number of experts and route each token to the corresponding expert combination.
[0032] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the dynamic and efficient routing method for a mixture-of-experts large model as described in the first aspect above.
[0033] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the dynamic and efficient routing method for a mixture-of-experts large model as described in the first aspect above.
[0034] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the dynamic and efficient routing method for a mixture-of-experts large model as described in the first aspect above.
[0035] Based on the above technical solutions, the beneficial effects of the present invention compared with the prior art are:
[0036] The dynamic and efficient routing method and device for a mixture-of-experts large model provided by the present invention can adaptively adjust the number of activated experts according to the difficulty and importance of each token by dividing the input data into multiple tokens and using the allocator module to determine the optimal number of experts for each token. The optimal number of experts for each token is dynamically determined by a lightweight allocator module. This mechanism not only improves the utilization efficiency of computing resources but also further enhances the performance of the model through refined resource allocation.
[0037] Other features and advantages of the present invention will be described in the subsequent specification, and some will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification and the drawings.
[0038] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0039] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0040] Figure 1 It is a schematic flowchart of the dynamic and efficient routing method for the hybrid expert large model provided by the present invention.
[0041] Figure 2a It is an architecture diagram of the traditional Top-K routing mechanism.
[0042] Figure 2b It is an architecture diagram of the dynamic and efficient routing mechanism for the hybrid expert large model provided by the present invention.
[0043] Figure 3 It is a schematic diagram for the comparative analysis of the accuracy rate and the number of activated experts provided by the present invention.
[0044] Figure 4 It is a schematic diagram for the comparison of the activation situations at different hierarchical depths and task difficulties provided by the present invention.
[0045] Figure 5 It is a schematic diagram for the comparison of the activation situations before and after fine-tuning with different parts of speech provided by the present invention.
[0046] Figure 6 It is a schematic structural diagram of the dynamic and efficient routing device for the hybrid expert large model provided by the present invention.
[0047] Figure 7 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0048] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0049] Today, with the rapid development of large language models (LLMs), there is an unprecedented challenge: how to effectively control the computational cost while the scale of model parameters continues to expand. The traditional Mixture-of-Experts (MoE) architecture balances the scalability and computational efficiency of the model to a certain extent by introducing expert networks and routing strategies. However, these models generally adopt a fixed Top-K routing strategy, ignoring the differences in complexity among different tokens, resulting in suboptimal resource allocation efficiency.
[0050] To address this issue, the present invention proposes a novel adaptive routing learning mechanism - a dynamic routing strategy based on a low-parameter allocator. The core of this strategy lies in the introduction of a parameterized allocator that can dynamically adjust the number of activated experts according to the intrinsic difficulty and importance of each token. This mechanism not only improves the utilization efficiency of computational resources but also further enhances the performance of the model through refined resource allocation. The motivation of the present invention is to explore a more efficient routing mechanism to cope with the growing model scale and computational cost. Through efficient routing learning based on the allocator, not only can we promote the further improvement of the performance of LLMs, but also provide new ideas and methods for resource optimization in large-scale applications. With the continuous progress of artificial intelligence technology, this adaptive and efficient resource allocation strategy will have important theoretical and practical significance for promoting the wide application and in-depth research of LLMs.
[0051] The following combines Figures 1 - 6 to describe the dynamic and efficient routing method and device for a large model of mixed experts according to the present invention.
[0052] Before routing, each MoE layer in the original large model of mixed experts is first replaced with a dynamic K-routing layer including an allocator module and a router. The design of the dynamic K-routing layer aims to dynamically select the number of experts according to the complexity and importance of the input data. The allocator module is responsible for calculating the probability distribution of each input token being assigned to different experts. This is usually achieved through functions such as softmax to ensure the non-negativity and normalization of the probability distribution. The router routes the input tokens to the corresponding experts for processing according to the probability distribution calculated by the allocator module. The router needs to implement an efficient routing mechanism, such as a dynamic routing strategy, to minimize the number of activated experts while maintaining the model performance.
[0053] Therefore, the above-mentioned mixture-of-experts large model after replacement includes a dispatcher module and a router. The present invention introduces a lightweight and pluggable dispatcher module that works in cooperation with the original router to dynamically perform expert allocation for each token. The lightweight dispatcher module, as a pluggable component, works with the router to generate a probability distribution of the number of experts for each token.
[0054] After replacing the mixture-of-experts large model, in order to train the dispatcher module, the present invention adopts the Proximal Policy Optimization (PPO) algorithm. The goal is to reduce the number of activated experts while maintaining or improving the performance of the model in benchmark tests. During the training process, the parameters of other parts of the model are kept unchanged, and only the parameters of the dispatcher module are updated. The PPO algorithm, with its advantages in stability and efficiency, provides an optimization direction for the decisions of the dispatcher, enabling the model to minimize the number of activated experts while maintaining performance.
[0055] First, define the objective function. The objective function is designed to reduce the number of activated experts while maintaining or improving the performance of the model in benchmark tests. This is usually achieved by defining a loss function consisting of two parts: one part is the loss of model performance (such as cross-entropy loss), and the other part is the reward for reducing the number of experts (such as negative number of experts or entropy of expert activation probability).
[0056] Next, use the PPO algorithm to train the dispatcher module. The policy gradient algorithm updates the policy parameters by estimating the performance gradient of the policy, making the policy tend to select actions that can obtain higher rewards, that is, reducing the number of experts while maintaining the model performance. In each training iteration, first generate routing decisions for a batch of input data through the current policy, and calculate the corresponding losses and rewards. Then, use these losses and rewards to estimate the gradient of the policy and update the parameters of the dispatcher module. During the training process, except for the parameters of the dispatcher module, the parameters of other parts of the model, such as the router, remain unchanged. This helps ensure that the changes in model performance mainly come from the improvement of the dispatcher module. Regularly evaluate the performance of the model on benchmark tests during the training process to ensure that the performance does not decrease significantly. Adjust the objective function, the parameters of the policy gradient algorithm, or the structure of the dispatcher module according to the evaluation results to further optimize the model performance.
[0057] The dynamic K-routing layer of the present invention enables the model to adaptively adjust the number of experts according to the complexity and importance of different input data, enhancing the flexibility and adaptability of the model. By dynamically selecting the number of experts, unnecessary computational overhead is reduced, and the computational efficiency of the model is improved. The allocator module is trained by the PPO algorithm to ensure that while reducing the number of experts, the performance of the model in benchmark tests is maintained or improved. Therefore, replacing the MoE layer in the mixture-of-experts model with the dynamic K-routing layer and training the allocator module by the policy gradient algorithm is an effective optimization method that can maintain or improve the model performance while improving the computational efficiency.
[0058] The process of training the allocator module using the policy gradient algorithm specifically includes the following steps:
[0059] S101. For each training sample, decompose the training sample into multiple tokens.
[0060] For each training sample, use a text preprocessing tool such as a tokenizer to decompose the text of the training sample into a series of tokens. These tokens can be words, punctuation marks, numbers, etc. Each token is an independent unit in the text, facilitating subsequent processing.
[0061] S102. Use the allocator module to generate a probability distribution of the number of experts for each token.
[0062] The allocator module is a neural network whose input is the embedded representation of the token and whose output is a vector, with each element of the vector corresponding to a possible number of experts. First, the embedded representation of the token is transformed through one or more linear layers to generate an unnormalized score vector. The output layer of the allocator module uses the softmax function, which can convert the output of the linear layer into a probability distribution. Then, use the softmax function to convert this score vector into a probability distribution. The softmax function ensures that the sum of all probability values is 1, forming a valid probability distribution, and each probability value represents the possibility of selecting the corresponding number of experts.
[0063] S103. Sample according to the probability distribution generated in S102 to determine the activated number of experts.
[0064] Randomly select a number of experts according to the probability distribution as the activated number of experts for the current token. This process is random but guided by the probability distribution.
[0065] S104. Use the activated experts to perform inference on the token and calculate the corresponding performance metrics.
[0066] The above performance metrics can be a comprehensive consideration of various metrics such as accuracy, speed, resource utilization, etc. The expert reasoning process may involve complex calculations or model predictions, and the calculation of performance metrics depends on specific task requirements.
[0067] According to the number of activated experts, select the corresponding number of experts from the pre-trained expert library. Input the token into the selected combination of experts for complex calculations or model predictions. Calculate performance metrics such as accuracy, speed, resource utilization, etc. based on the reasoning results and the true labels.
[0068] S105. Input the number of activated experts and the performance metrics into a pre-defined reward function to calculate the reward value for each training sample.
[0069] First, define a reward function. The design of the reward function should be based on the task objective and the optimization objective. For example, if the goal is to maximize accuracy while minimizing the number of experts, the reward function may need to consider these two factors comprehensively. This function takes the number of activated experts and the performance metrics as inputs and outputs a scalar reward value. Then, input the number of activated experts and the performance metrics into the reward function to calculate the reward value for each training sample. In this invention, taking accuracy as the performance metric as an example, it shows how to construct the reward function:
[0070] The reward function includes two parts: accuracy reward and expert number penalty. Reward function = accuracy reward - expert number penalty
[0071] Regarding the accuracy reward: Accuracy is the proportion of correct predictions or processing. In a mixture of experts system, it can be defined as the ratio of the number of tokens correctly processed to the total number of tokens. For each training sample, calculate the accuracy of the training sample after being processed by the combination of experts. The calculation process of accuracy is as follows: Ensure that the training data has accurate labels or reference outputs for subsequent evaluation of the model performance. Use the mixture of experts model to predict the training data to obtain the prediction results for each token. Compare the prediction results with the true labels or reference outputs, and count the number of tokens correctly processed. Calculate the accuracy according to the number of tokens correctly processed and the total number of tokens. The accuracy can be directly used as the reward value, or it can be transformed in form such as linear transformation, exponential transformation, etc. to adjust the sensitivity of the reward.
[0072] Regarding the expert number penalty: To encourage the model to use as few experts as possible, a penalty term for the number of experts can be added to the reward function. Set a penalty value proportional to the number of experts, and subtract this value from the accuracy reward. The weight of the penalty term can be adjusted according to actual needs to balance the relationship between accuracy and the number of experts.
[0073] The above reward function R can be set as:
[0074]
[0075] Among them, Accuracy represents the accuracy of the model on the validation set, Average Experts Activated represents the average number of experts activated per token, and α and β are weight coefficients that balance the importance of the two. According to the experimental and performance evaluation results, the weights of the accuracy reward and the expert number penalty are continuously adjusted.
[0076] S106. Update the policy parameters of the allocator module using the PPO algorithm based on the reward value to maximize the expected return.
[0077] First, initialize the model and set the initial parameters of the allocator module and the expert model. Then, for each training sample, process it according to the steps of S101 to S104. Calculate the reward value of each training sample using a predefined reward function. Based on the reward value, update the policy parameters of the allocator module using the PPO algorithm. During the update process, the goal is to maximize the expected return, that is, the cumulative sum of the reward values. Regularly evaluate the performance of the model during the training process, including the accuracy and the usage of the number of experts. Adjust the parameters of the reward function or the configuration of the optimization algorithm according to the evaluation results. Repeat the training process and the performance evaluation steps until the model performance reaches a satisfactory level or reaches the preset number of training epochs. Through the above steps, a reward function that can maximize both the accuracy and minimize the number of experts can be designed, and the allocator module in the mixture-of-experts model can be optimized based on this reward function.
[0078] The PPO algorithm guides the update direction of the parameters by calculating the gradients of the policy parameters. Specifically, estimate the expectation of the reward value and consider the difference between the current policy and the old policy to limit the update amplitude. According to the calculated gradients, update the policy parameters of the allocator module along the gradient direction to maximize the expected return. Repeat the steps of S101 to S106 to continuously improve the performance of the allocator module through iterative optimization. Through the above steps, the allocator module can learn how to allocate the optimal number of experts for each token according to the characteristics of the input data and the task requirements, so as to achieve dynamic and efficient routing. During the training process, keep the parameters of the original mixture-of-experts large model unchanged and only train the newly introduced allocator module.
[0079] The present invention uses the PPO algorithm to optimize the decision-making of the allocator to maximize the overall performance while minimizing the number of activated experts. Since the PPO algorithm has strong flexibility and adaptability, through the iterative optimization of the PPO algorithm, the allocator module can learn a more reasonable expert allocation strategy, can automatically adjust the expert allocation strategy according to the needs of different tasks and scenarios, and improve the overall performance of the system. By precisely controlling the number of experts, the allocator module can avoid resource waste and improve resource utilization and computing efficiency.
[0080] Referring to Figure 1 As shown, the dynamic and efficient routing method for the hybrid expert large model provides a dynamic K-routing strategy for adaptively adjusting the number of activated experts according to the inherent difficulty and importance of each token. The method includes the following steps:
[0081] Step S201, obtain input data and a pre-trained hybrid expert large model.
[0082] The above input data is the original data processed by the model, which can be in various forms such as text, image, voice, etc. The hybrid expert large model consists of an allocator module and a router. The allocator module is responsible for determining how many experts are needed to process each token, while the router is responsible for routing these tokens to the corresponding expert combinations.
[0083] Step S202, divide the input data into multiple tokens and use the allocator module to determine the optimal number of experts for each token.
[0084] Decompose the input data into multiple tokens for independent processing. The allocator module determines an optimal number of experts for each token based on the characteristics of the input data and the task requirements. This number varies depending on the complexity of the token and the current state of the model.
[0085] Step S203, use the router to select an expert combination for each token according to the optimal number of experts and route each token to the corresponding expert combination.
[0086] The router selects the most suitable expert combination for each token from the pre-trained expert library according to the information provided by the allocator module. This selection process may be based on various factors such as the expertise of the experts, the current load, historical performance, etc. Once the expert combination is selected, the router routes each token to the corresponding expert combination for processing.
[0087] Through the allocator module, the present invention can flexibly allocate different numbers of experts to each token, thereby adapting to the complexity and processing requirements of different tokens. By dynamically selecting the number of experts, the allocator module can reduce unnecessary computational overhead and improve the computational efficiency of the overall model. This process can be easily extended to more complex model structures and larger datasets, providing the possibility to handle larger-scale and more complex NLP tasks. And through the optimization of the policy gradient algorithm (such as PPO), the allocator module can gradually learn a better expert allocation strategy, thereby further improving the performance and accuracy of the model. The method of the present invention can be applied to large models such as large language models (LLMs) and large vision models based on the MoE architecture.
[0088] Referring to Figure 2a As shown, in the traditional Top-K routing, each token always activates the same number of experts regardless of its complexity. This approach may lead to resource waste when dealing with simple tokens, and may affect the model performance due to insufficient resources when facing complex tokens. In contrast, referring to Figure 2b As shown, the routing strategy proposed by the present invention first dynamically determines the actual number of experts activated by each token through an allocator, and then routes each token through a router under the guidance of this number, thereby achieving more precise resource allocation. This ability to dynamically adjust enables the model to invest more computational resources in complex tokens, while reducing resource consumption for simple tokens, thus optimizing the efficiency and performance of the model as a whole.
[0089] In an optional embodiment, the expert combination includes at least one expert model; the method further includes:
[0090] S204. Process the token based on each expert model in the expert combination and generate corresponding output results.
[0091] In the foregoing step S203, each token has been routed to a specific expert combination. In this step, the tokens will be separately sent to each expert model in the combination according to the expert combination they are assigned to. After receiving the token, each expert model will independently process it. The processing process may include calculations, inferences, predictions, etc., depending on the type of the expert model and the task requirements. After the processing is completed, each expert model will generate a corresponding output result. These output results may be in the form of probability distributions, prediction labels, vector representations, etc., depending on the design of the output layer of the model.
[0092] S205. Aggregate the output results of different expert models to obtain the model output.
[0093] First, collect the output results of all expert models in the expert ensemble. Then, adopt an aggregation strategy to integrate these output results. The aggregation strategy can be simple averaging, weighted averaging, voting, attention mechanism, etc., depending on the task requirements and model design. Finally, the comprehensive result obtained according to the aggregation strategy will be used as the final output of the model. This output can be used for subsequent tasks such as classification, regression, generation, etc.
[0094] By combining multiple expert models, the mixture-of-experts large model can utilize the expertise and advantages of different experts, thereby enhancing the processing ability and performance of the overall model. Each expert model in the expert ensemble may be good at processing different types of data or features. Therefore, by aggregating their outputs, the model can better adapt to various input situations and improve the generalization ability. The mixture-of-experts large model can dynamically add or remove expert models according to needs to adapt to different tasks and datasets. This flexibility makes the model highly scalable. Since not all expert models need to process each token, the mixture-of-experts large model can utilize computing resources more effectively. Only when necessary will the relevant expert models be activated for processing. By optimizing the allocation of the expert ensemble and the aggregation strategy, the performance of the mixture-of-experts large model can be further improved. For example, using the attention mechanism to dynamically adjust the weights of different expert models can make the model pay more attention to the expert models that have an important impact on the final output.
[0095] In an optional embodiment, the above-mentioned allocator module includes an input layer, at least one hidden layer, and an output layer;
[0096] The step of dividing the input data into multiple tokens and using the allocator module to determine the optimal number of experts for each token as described in step S202 includes:
[0097] S301. Divide the input data into multiple tokens and generate a representation vector for each token.
[0098] First, the model preprocesses the input original text, including removing useless characters, standardizing the text format, etc. The preprocessed text will be split into smaller units, namely tokens. This process decomposes the text into words, subwords, or characters, etc. Then, a fixed-length representation vector (embedding vector) is generated for each token. The representation vector is the representation of the token in the continuous vector space, which can capture the semantic and context information of the token. This is usually achieved through a pre-trained embedding layer or lookup table, where each token is mapped to a unique vector.
[0099] The process of generating the representation vector is as follows:
[0100] Map the generated tokens to unique indices or identifiers in the vocabulary. The vocabulary is a list containing all possible tokens, and each token has a corresponding numerical representation. This step converts the tokens into a numerical format that the model can process. To capture the semantic relationships and context information between tokens, the model usually converts the numerical representations of tokens into representation vectors. These representation vectors are dense, high-dimensional numerical representations that contain rich semantic information, enabling the model to understand and generate coherent text.
[0101] In the MoE-based LLM model, after the representations of tokens are generated, they are fed into the MoE layer for further processing. In the MoE layer, each token is dynamically routed to different expert models for calculation based on its representation vector. The input layer of the dispatcher module receives the representation of the token (i.e., the above-mentioned representation vector) as input.
[0102] S302. The input layer receives the representation vectors of each token as input.
[0103] The input layer of the dispatcher module is responsible for receiving the representation vectors from step S301. These representation vectors, as the initial input data of the model, contain the semantic and context information of the tokens.
[0104] S303. At least one hidden layer extracts the token features of each representation vector.
[0105] After the input layer, one or more hidden layers perform non-linear transformations on the representation vectors to extract higher-level token features. These hidden layers usually contain activation functions such as ReLU, which can introduce non-linearity and enable the model to learn complex mapping relationships. Information is passed between hidden layers through forward propagation, and each layer is calculated based on the output of the previous layer, gradually extracting more abstract and useful features.
[0106] S304. The output layer outputs the probability distribution of the number of experts corresponding to each token based on the token features of each representation vector, and samples based on the probability distribution of the number of experts to select the optimal number of experts for each token.
[0107] The output layer receives the feature representations from the last hidden layer and applies the softmax function to convert these features into a probability distribution over the number of experts corresponding to each token. Each probability value represents the likelihood of selecting the corresponding number of experts. This probability distribution is used to guide how tokens are routed to different numbers of expert models for more efficient and accurate computations. Based on this probability distribution, a random number generator is used for sampling to select a specific number of experts as the optimal number for each token. This number will be used in subsequent steps to guide the router to route the tokens to the corresponding expert combinations.
[0108] Through the non-linear transformation of multiple hidden layers in the present invention, the allocator module can extract the deep features of tokens, which are crucial for subsequent decisions on the number of experts. Due to the adoption of the probability distribution approach, the allocator module can flexibly allocate different numbers of experts to each token, thus adapting to the complexity and processing requirements of different tokens. By dynamically selecting the number of experts, the allocator module can reduce unnecessary computational overhead and improve the computational efficiency of the overall model. This process can be easily extended to more complex model structures and larger datasets, providing the possibility for handling larger-scale and more complex tasks.
[0109] To evaluate the performance of the efficient routing learning strategy of the low-parameter allocator on different large language models (LLMs) and compare it with the conventional Top-K routing method. Three popular MoE-based LLMs were selected as the baseline models, namely Mixtral-8x7B, DeepSeek-MoE-16B, and Qwen1.5-MoE-A2.7B. These models are representative in architecture and cover different numbers of experts and routing strategies. In the experiment, the original structures of these models were maintained, and only the routing strategy was adjusted. Through extensive evaluations on three popular MoE-based LLMs, it was demonstrated that the proposed routing method significantly outperformed the traditional Top-K routing. It not only consistently reduced the number of activated experts by 30% to 40%, but also almost maintained the performance in multiple standard LLM benchmark tests and even achieved performance improvements in some cases.
[0110] Performance evaluation mainly focuses on two aspects: the accuracy of the model (Accuracy) and the number of activated experts (Activated Experts). Accuracy is a key metric for measuring the performance of the model on a specific task, while the number of activated experts is directly related to the computational efficiency of the model. The present invention integrates the proposed routing strategy into the above-mentioned baseline model and optimizes the allocator through the PPO algorithm. This strategy can dynamically allocate the most appropriate number of experts for each token, aiming to reduce the consumption of computing resources while maintaining or improving performance. Experimental results show that this routing strategy has achieved significant performance improvement on all baseline models. Specifically, referring to Figure 3 As shown, in terms of reducing the number of activated experts, the routing strategy of the present invention reduces the expert activation by an average of 30% to 40%, which directly translates into an increase in the inference speed. In terms of accuracy, the routing strategy proposed by the present invention almost maintains the original performance in most cases, and even achieves performance improvement in some tasks.
[0111] To deeply understand the impact of the dynamic routing strategy proposed by the present invention, the impacts of different task difficulties, different hierarchical depths, and different parts-of-speech on expert activation are further analyzed. Referring to Figure 4 As shown, it is the activation situation for different hierarchical depths and task difficulties. Referring to Figure 5 As shown, it is the activation situation for different parts-of-speech before and after fine-tuning. It can be found that when dealing with complex tasks, the middle layer, and content words, this routing tends to activate more experts, indicating that this strategy can perform effective resource allocation according to the complexity and importance of the input. Figure 4 In Figure 5 the Layer Index is the hierarchical index, Activated Experts is the number of activated experts, Frequency is the frequency,
[0112] In
[0113] the Expert Index is the expert index, and Activated Ratio is the activation ratio.
[0114] The dynamic and efficient routing device for a hybrid expert large model provided by the present invention, with reference to Figure 6 as shown, includes the following modules:
[0115] An acquisition module 401, configured to acquire input data and a pre-trained hybrid expert large model, wherein the hybrid expert large model includes a dispatcher module and a router;
[0116] A distribution module 402, configured to divide the input data into multiple tokens, and determine the optimal number of experts for each token by using the dispatcher module;
[0117] A routing module 403, configured to select an expert combination for each token according to the optimal number of experts by using the router, and route each token to the corresponding expert combination.
[0118] Figure 7 Illustrates a schematic diagram of the physical structure of an electronic device, as Figure 7 shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 550, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 550. The processor 510 may call logical instructions in the memory 530 to execute the dynamic and efficient routing method for the hybrid expert large model.
[0119] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0120] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the dynamic and efficient routing method for the hybrid expert large model provided by the above-mentioned various methods.
[0121] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the dynamic and efficient routing method for the hybrid expert large model provided by the above-mentioned various methods.
[0122] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0123] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A dynamic and efficient routing method for a large hybrid expert model, characterized in that Including: Obtain input data and a pre-trained mixture-of-experts large model, where the mixture-of-experts large model includes a dispatcher module and a router, and the input data is text; Divide the input data into multiple tokens, and use the dispatcher module to determine the optimal number of experts for each token; Use the router to select an expert combination for each token according to the optimal number of experts, and route each token to the corresponding expert combination; Before dividing the input data into multiple tokens and using the dispatcher module to determine the optimal number of experts for each token, the method further includes: Training the dispatcher module using a policy gradient algorithm; The training of the dispatcher module using the policy gradient algorithm includes: For each training sample, decompose the training sample into multiple tokens; Use the dispatcher module to generate a probability distribution of the number of experts for each token; Sample according to the probability distribution to determine the activated number of experts; Use the activated experts to perform inference on the tokens and calculate the corresponding performance metrics; Input the activated number of experts and the performance metrics into a predefined reward function to calculate the reward value of each training sample; Based on the reward value, use the policy gradient algorithm to update the policy parameters of the dispatcher module to maximize the expected return.
2. The dynamic and efficient routing method for a hybrid expert large model according to claim 1, characterized in that The dispatcher module includes an input layer, at least one hidden layer, and an output layer; The dividing the input data into multiple tokens and using the dispatcher module to determine the optimal number of experts for each token includes: Divide the input data into multiple tokens and generate a representation vector for each token; The input layer receives the representation vector of each token; At least one hidden layer extracts the token features of each representation vector; The output layer outputs the probability distribution of the number of experts corresponding to each token according to the token features of each representation vector, and samples based on the probability distribution of the number of experts to select the optimal number of experts for each token.
3. The dynamic and efficient routing method for the hybrid expert large model according to claim 1, wherein The expert combination includes at least one expert model; the method further includes: Process the tokens based on each expert model in the expert combination and generate corresponding output results; Aggregate the output results of different expert models to obtain the model output.
4. The dynamic and efficient routing method for a hybrid expert large model according to claim 1, wherein Before training the dispatcher module using the policy gradient algorithm, the method further includes: Replace each MoE layer in the mixture-of-experts large model with a dynamic K routing layer including a dispatcher module.
5. A dynamic and efficient routing device for a large model of hybrid experts, characterized in that, Including: An acquisition module for acquiring input data and a pre-trained mixture-of-experts large model, where the mixture-of-experts large model includes a dispatcher module and a router, and the input data is text; An allocation module for dividing the input data into multiple tokens and using the dispatcher module to determine the optimal number of experts for each token; before dividing the input data into multiple tokens and using the dispatcher module to determine the optimal number of experts for each token, it further includes: training the dispatcher module using a policy gradient algorithm; The training of the dispatcher module using the policy gradient algorithm includes: For each training sample, decompose the training sample into multiple tokens; use the allocator module to generate a probability distribution of the number of experts for each token; sample according to the probability distribution to determine the number of activated experts; use the activated experts to perform inference on the tokens and calculate the corresponding performance metrics; input the number of activated experts and the performance metrics into a predefined reward function to calculate the reward value for each training sample; update the policy parameters of the allocator module using the policy gradient algorithm based on the reward value to maximize the expected return; A routing module, configured to use a router to select an expert combination for each token according to the optimal number of experts and route each token to the corresponding expert combination.
6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the dynamic and efficient routing method for a hybrid expert large model according to any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the dynamic and efficient routing method for a hybrid expert large model according to any one of claims 1 to 4.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the dynamic and efficient routing method for a hybrid expert large model according to any one of claims 1 to 4.