Hybrid expert model optimization method and system based on dynamic token routing
By introducing a grouping attention mechanism of dynamic token routing and mixed weight sharing in the Transformer architecture, the problem of excessive computation and memory costs when processing long sequences is solved, more efficient computing and memory usage is achieved, and the stability of the model is improved.
Patent Information
- Application Number
- CN202510294499.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-17
AI Technical Summary
The existing Transformer architecture is too expensive to process long sequences due to its self-attention mechanism.
The hybrid expert model optimization method based on dynamic token routing is adopted, and the computing resources are dynamically allocated through dynamic scoring functions and expert routing mechanisms, and the grouping attention mechanism of mixed weight sharing is used for parallel processing. At the same time, the auxiliary loss function is used to optimize routing decisions during the training process.
It significantly reduces the computational complexity and memory overhead, improves the computational efficiency and memory usage efficiency, and ensures routing consistency in the training and inference stages, improving the stability and reliability of the model.
Smart Images

Figure CN120163187A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic information technology, and particularly to an optimization method and system for a hybrid expert model based on dynamic token routing. Background Art
[0002] The Transformer architecture has achieved remarkable results in many fields due to its efficient self-attention mechanism, including natural language processing, computer vision, and reinforcement learning. However, the computational and memory overhead of its self-attention mechanism grows quadratically with the increase in the length of the input sequence. Especially when dealing with long sequences, it significantly increases the computational cost. This problem limits the application of the Transformer model when dealing with large-scale data and long sequences.
[0003] To overcome these deficiencies, the present application proposes an optimization method and system for a hybrid expert model based on dynamic token routing to improve computational efficiency and memory usage efficiency. Summary of the Invention
[0004] The purpose of the present application is to provide an optimization method and system for a hybrid expert model based on dynamic token routing, aiming to solve the problem of excessive computational and memory costs caused by the self-attention mechanism of the existing Transformer architecture when dealing with long sequences.
[0005] To achieve the above purpose, the present application provides the following technical solutions:
[0006] In the first aspect, the present application provides an optimization method for a hybrid expert model based on dynamic token routing. The steps include:
[0007] Evaluating the basic units in the Transformer model through a dynamic scoring function to generate a scoring result for each basic unit;
[0008] According to the scoring result and the expert routing mechanism, allocating the basic unit to one of several predefined experts;
[0009] Adopting a grouped attention mechanism with hybrid weight sharing to map the basic units assigned by the experts to different groups of attention heads for parallel processing; wherein, the grouped attention mechanism with hybrid weight sharing is a combination of weight sharing and hybrid strategy in the attention mechanism;
[0010] During the training process of the hybrid expert model, optimizing the routing decision based on the auxiliary loss function to keep the routing in the inference stage consistent with that in the training stage.
[0011] In the second aspect, the present application provides an optimization system for a hybrid expert model based on dynamic token routing, specifically including:
[0012] Scoring module: used to evaluate the basic units in the Transformer model through a dynamic scoring function and generate the scoring results for each basic unit;
[0013] Routing and allocation module: used to allocate the basic units to one of several predefined experts according to the scoring results and the expert routing mechanism;
[0014] Parallel expert processing module: used to map the basic units assigned by experts to different groups of attention heads by adopting a grouped attention mechanism with hybrid weight sharing for parallel processing; wherein, the grouped attention mechanism with hybrid weight sharing is a combination of weight sharing and hybrid strategy in the attention mechanism;
[0015] Optimization module: used to optimize the routing decision based on the auxiliary loss function during the training process of the mixture-of-experts model, so that the routing in the inference stage is consistent with that in the training stage.
[0016] In a third aspect, the present application provides a computer device, which includes a processor and a memory coupled to the processor. Wherein, the memory stores program instructions for implementing an optimization method of a mixture-of-experts model based on dynamic token routing; the processor is used to execute the program instructions stored in the memory to implement an optimization of a mixture-of-experts model based on dynamic token routing.
[0017] In a fourth aspect, the present application provides a storage medium storing program instructions executable by a processor, and the program instructions are used to execute an optimization method of a mixture-of-experts model based on dynamic token routing.
[0018] The present application provides an optimization method and system of a mixture-of-experts model based on dynamic token routing, having the following beneficial effects:
[0019] (1) Through the dynamic scoring function and the expert routing mechanism, the computing resources are dynamically allocated according to the importance of the basic units, avoiding the waste of computing and memory caused by all basic units being uniformly processed in the traditional Transformer architecture. Especially when dealing with long sequences, the computing complexity and memory overhead are significantly reduced.
[0020] (2) In the traditional Transformer architecture, each attention head projects the key and value of the input token independently, resulting in a large number of model parameters. The grouped attention (GQA) mechanism with hybrid weight sharing proposed in the present application significantly reduces the amount of computation by sharing and grouping the projection weights. The basic units processed by each expert are mapped to different groups of attention heads, effectively reducing the computing and memory overhead of each expert.
[0021] (3) An auxiliary loss function is adopted to ensure the consistency of routing decisions of the mixture-of-experts model in the training stage and the inference stage, solve the problem of performance degradation caused by the inconsistency between training and inference in traditional methods, and improve the stability and reliability of the model.
[0022] (4) The technical solution proposed in this application can be extended to fields such as computer vision and reinforcement learning in addition to being applied to natural language processing tasks. By adjusting the expert capacity or the allocation strategy of the KV cache, the performance of specific tasks is optimized to meet the requirements of different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic flow chart of an optimization method for a mixture-of-experts model based on dynamic token routing according to Embodiment 1 of this application;
[0024] Figure 2 It is a schematic structural diagram of an optimization system for a mixture-of-experts model based on dynamic token routing according to Embodiment 2 of this application;
[0025] Figure 3 It is a schematic structural diagram of a computer device according to Embodiment 3 of this application;
[0026] Figure 4 It is a schematic structural diagram of a storage medium according to Embodiment 4 of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0028] The following analyzes the solutions in the prior art in combination with related technologies.
[0029] In the prior art, it mainly includes sparse attention mechanisms (such as Big Bird, Longformer, etc.), which reduce the computational complexity through fixed sparse patterns. In addition, Key-Value (KV) cache optimization methods (such as PyramidKV, DynamicKV, etc.) improve the memory usage efficiency by dynamically adjusting the cache. However, the prior art adopts static or overly simple strategies in the allocation of KV cache resources, ignoring the dynamic changes in the importance between tokens. Although the model based on MoE (Mixture-of-Experts) improves the computational efficiency, due to the uneven utilization rate of experts, it leads to waste of computational resources and lacks sufficient token-level adaptability. At the same time, there are also problems of inconsistency between the training and inference stages in the KV cache eviction strategy and the expert selection routing strategy. Especially in autoregressive tasks, the complete sequence information is used during training, while only the past context information can be accessed during inference.
[0030] The present application proposes an optimization method and system for a hybrid expert model based on dynamic token routing. By introducing a dynamic calculation and memory resource allocation mechanism based on token importance, the above problems are overcome, especially in improving the computational efficiency and memory usage efficiency while ensuring the contribution of low-priority tokens.
[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0032] Embodiment 1
[0033] Please refer to Figure 1 , which is a schematic flowchart of an optimization method for a hybrid expert model based on dynamic token routing according to Embodiment 1 of the present application; the steps include:
[0034] S1: Evaluate the basic units in the Transformer model through a dynamic scoring function to generate a scoring result for each basic unit.
[0035] In this embodiment, the dynamic importance of each basic unit is evaluated through a dynamic scoring function, and computing resources are allocated according to its importance, thereby improving the computational efficiency and reducing unnecessary computational overhead. Among them, the dynamic scoring function generates a scoring result for the basic unit through a trained linear layer, and uses a sigmoid activation function to limit the scoring result within the range of 0 to 1.
[0036] The expression of the dynamic scoring function is:
[0037] scores(x) = sigmoid(xW + b),
[0038] where x is the input basic unit; W and b are the weights and biases trained by the hybrid expert model respectively.
[0039] S2: According to the scoring result and the expert routing mechanism, allocate the basic unit to one of several predefined experts.
[0040] In this embodiment, the expert routing mechanism dynamically allocates basic units according to the computing power and resource limitations of each expert. To ensure that each basic unit can only be allocated to one expert, based on the scoring result of the basic unit, for each expert, the top k basic units are selected using a sparse masking function and the selected basic units are allocated to the corresponding expert. The expression of the sparse masking function mask is:
[0041] mask = indicator(top-k(scores(x))),
[0042] where the top-k operation is used to select the top k basic units; the indicator function is used to generate a sparse mask for the selected basic units to identify the allocation result.
[0043] Specifically, assume there are E = 3 experts, and the processing capacity ratios of each expert are [0.5, 0.3, 0.2]. According to the scoring results of the basic units, the top k high-scoring basic units are selected for each expert x. For example: Expert 1 processes the first 50% of the basic units (k1 = 0.5N), where N is the sequence length; Expert 2 processes the first 30% of the basic units (k1 = 0.3N); Expert 3 processes the first 20% of the basic units (k1 = 0.2N). The allocation result of each basic unit is identified through the sparse mask function mask, and the load of each expert is ensured to be balanced.
[0044] S3: Adopt a grouped attention mechanism with hybrid weight sharing to map the basic units assigned by experts into different groups of attention heads for parallel processing; wherein, the grouped attention mechanism with hybrid weight sharing is a combination of weight sharing and hybrid strategy in the attention mechanism.
[0045] In the traditional Transformer architecture, multiple attention heads process each token, and each head has independent query, key, and value parameters, resulting in a large number of model parameters. In this embodiment, in order to reduce the computational and memory consumption while ensuring the model efficiency, a grouped attention (GQA) mechanism with hybrid weight sharing is introduced.
[0046] First, it is necessary to project the key and value of the basic unit. The traditional method performs independent key-value projection for each attention head, while in this application, the projection weights are shared and grouped, thereby reducing the computational amount. Specifically, project the key and value of the basic unit, group multiple attention heads by experts, and each group shares the projection weights of the key and value. Assume the total number of attention heads H = 8, and perform head grouping for each expert: Expert 1 is assigned 4 heads (H / 2 1 = 4); Expert 2 is assigned 2 heads (H / 2 2 = 2); Expert 3 is assigned 1 head (H / 2 3 = 1); The attention heads within each group share the projection weights of the key and value, but the query weights are independent.
[0047] For each expert, a weighted projection mechanism is adopted to reduce the number of key-value pairs. Specifically, perform an aggregation calculation on the attention heads within each group to obtain an aggregated projection result. The formula for the aggregation calculation is expressed as:
[0048]
[0049] Among them, G e is the head group of expert e; represents the projection result of the h-th attention head, where i is; is the result obtained by grouped query attention, where j ∈ {k, v} is used to distinguish the results of keys and values; h ∈ {1,..., H} is the attention head index. By aggregating the projection results, the computational and content overhead of each expert is reduced.
[0050] S4: During the training process of the mixture-of-experts model, optimize the routing decision based on the auxiliary loss function to make the routing in the inference stage consistent with that in the training stage.
[0051] In traditional autoregressive language models, the complete input sequence is used for expert routing decisions during training, while only the partially generated sequence is used for inference during inference. To overcome this inconsistency, an auxiliary loss function is introduced in this embodiment to ensure the consistency between the training stage and the inference stage. The expression of the auxiliary loss function Laux is:
[0052] Laux(x) = cross_entropy(scores(x), argmax(T)),
[0053] where scores is the dynamic scoring function; x is the basic unit of the input; T is the expert mapping of each basic unit; argmax(T) extracts the expert index assigned to each basic unit; cross_entropy is the cross-entropy loss function.
[0054] In addition, in terms of the KV cache, the memory utilization efficiency is further improved by dynamically allocating the cache size. The key-value cache size of experts is dynamically adjusted through the expert routing mechanism, and different key-value cache sizes are allocated to each expert. This mechanism for dynamically adjusting the cache can ensure that the memory overhead is effectively controlled when processing long sequences.
[0055] It should be noted that the method proposed in this application can be extended to fields such as computer vision and reinforcement learning in addition to natural language processing. For specific application scenarios, adjust the expert capacity or the allocation strategy of the key-value cache as needed to optimize the performance of specific tasks.
[0056] In summary, the purpose of Embodiment 1 is to solve the problem of excessive computational and memory costs faced by the existing Transformer architecture when processing long sequences. By introducing a dynamic basic unit selection and expert routing mechanism, computational resources are dynamically allocated according to the importance of the basic units, improving computational efficiency. At the same time, a grouped attention mechanism with hybrid weight sharing is adopted to reduce computational and memory overheads. In addition, an auxiliary loss function is introduced to ensure that during the training process of the mixture-of-experts model, the routing decisions in the inference stage are consistent with those in the training stage, solving the problems caused by inconsistent training and inference.
[0057] Embodiment 2
[0058] Please refer to Figure 2 , which is a schematic structural diagram of an optimization system for a mixture-of-experts model based on dynamic token routing according to Embodiment 2 of the present application; the specific content includes:
[0059] Scoring module: used to evaluate the basic units in the Transformer model through a dynamic scoring function and generate the scoring results of each basic unit;
[0060] Routing allocation module: used to allocate the basic units to one of several predefined experts according to the scoring results and the expert routing mechanism;
[0061] Parallel expert processing module: used to map the basic units assigned by the experts to different groups of attention heads by adopting a grouped attention mechanism with hybrid weight sharing for parallel processing; wherein, the grouped attention mechanism with hybrid weight sharing is a combination of weight sharing and hybrid strategies in the attention mechanism;
[0062] Optimization module: used to optimize the routing decision based on the auxiliary loss function during the training process of the mixture-of-experts model to make the routing in the inference stage consistent with that in the training stage.
[0063] In this embodiment, the parallel expert processing module further includes: a weight sharing sub-module for sharing weights between different groups of attention heads to reduce model parameters; a hybrid strategy sub-module for combining the outputs of different groups of attention heads to generate the final attention result. And the system further includes a data preprocessing module for preprocessing the input data to meet the input requirements of the Transformer model; a post-processing module for processing the model output to generate the final application result.
[0064] Embodiment 3
[0065] Please refer to Figure 3 , which is a schematic structural diagram of a computer device according to Embodiment 3 of the present application. The computer device 50 includes a processor 51 and a memory 52 coupled to the processor 51.
[0066] The memory 52 stores program instructions for implementing the above-mentioned method for optimizing a hybrid expert model based on dynamic token routing.
[0067] The processor 51 is configured to execute the program instructions stored in the memory 52 to implement an optimization of a hybrid expert model based on dynamic token routing.
[0068] Among them, the processor 51 can also be referred to as a CPU (Central Processing Unit).
[0069] The processor 51 may be an integrated circuit chip with signal processing capabilities. The processor 51 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0070] Embodiment 4
[0071] Please refer to Figure 4 , which is a schematic structural diagram of the storage medium according to Embodiment 4 of the present application. The storage medium according to the embodiment of the present application stores a program file 61 capable of implementing all the above methods. Among them, the program file 61 can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods according to various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes, or devices such as a computer, a server, a mobile phone, or a tablet.
[0072] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, apparatus, article or method comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, apparatus, article or method comprising the element.
[0073] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
[0074] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present application. The scope of the present application is defined by the appended claims and their equivalents.
[0075] Certainly, the present invention can also have other various implementation manners. Based on this implementation manner, other implementation manners obtained by those of ordinary skill in the art without any creative work belong to the scope protected by the present invention.
Claims
1. A hybrid expert model optimization method based on dynamic token routing, characterized in that: include: The basic units in the Transformer model are evaluated through a dynamic scoring function to generate scoring results for each basic unit; Allocating the basic unit to one of several predefined experts according to the scoring result and the expert routing mechanism; A hybrid weight-sharing group attention mechanism is adopted to map the basic units assigned by experts to different attention head groups for parallel processing; wherein the hybrid weight-sharing group attention mechanism is a combination of weight-sharing and hybrid strategies in the attention mechanism; During the training process of the hybrid expert model, the routing decision is optimized based on the auxiliary loss function to keep the routing in the inference phase consistent with that in the training phase.
2. A hybrid expert model optimization method based on dynamic token routing according to claim 1, characterized in that: The step of evaluating the basic units in the Transformer model by using a dynamic scoring function to generate a scoring result for each basic unit specifically includes: The dynamic scoring function generates a scoring result of a basic unit through a trained linear layer, and uses a sigmoid activation function to limit the scoring result to a range of 0 to 1; The expression of the dynamic scoring function is: scores(x)=sigmoid(xW+b), Among them, x is the basic unit of input; W and b are the weight and bias of hybrid expert model training respectively.
3. A hybrid expert model optimization method based on dynamic token routing according to claim 1, characterized in that: The step of allocating the basic unit to one of a plurality of predefined experts according to the scoring result and the expert routing mechanism specifically includes: The expert routing mechanism dynamically allocates basic units according to the computing power and resource constraints of each expert; Based on the scoring results of the basic units, a sparse mask function is used to select the first k basic units for each expert, and the selected basic units are assigned to the corresponding experts; The expression of the sparse mask function mask is: mask=indicator(top-k(scores(x))), Among them, the top-k operation is used to select the first k basic units; the indicator function is used to generate a sparse mask for the selected basic units to identify the allocation results.
4. The hybrid expert model optimization method based on dynamic token routing according to claim 1, characterized in that: The hybrid weight-sharing group attention mechanism maps the basic units assigned by the experts to different attention head groups for parallel processing; wherein the hybrid weight-sharing group attention mechanism is a step of combining weight sharing and hybrid strategies in the attention mechanism, specifically including: Projecting the keys and values of the basic unit, grouping multiple attention heads by experts, and each group sharing the projection weights of the keys and values; A weighted projection mechanism is used to aggregate the attention heads in each group to obtain the aggregated projection results.
5. A hybrid expert model optimization method based on dynamic token routing according to claim 4, characterized in that: The formula for the aggregation calculation is expressed as: Among them, G e Group the heads of expert e; represents the projection result of the h-th attention head; is the result of grouped query attention, j∈{k,v} is used to distinguish the key and value results; h∈{1,...,H} is the attention head index.
6. A hybrid expert model optimization method based on dynamic token routing according to claim 1, characterized in that: The expression of the auxiliary loss function Laux is: Laux(x)=cross_entropy(scores(x),argmax(T)), Among them, scores is a dynamic scoring function; x is the basic unit of input; T is the expert mapping of each basic unit; argmax(T) extracts the expert index assigned to each basic unit; cross_entropy is the cross entropy loss function.
7. The hybrid expert model optimization method based on dynamic token routing according to claim 1, characterized in that: The method further includes: dynamically adjusting the key-value cache size of the experts through the expert routing mechanism, and allocating a different key-value cache size to each expert.
8. A hybrid expert model optimization system based on dynamic token routing, characterized in that: include: Scoring module: used to evaluate the basic units in the Transformer model through a dynamic scoring function and generate scoring results for each basic unit; Routing allocation module: used for allocating the basic unit to one of several predefined experts according to the scoring result and the expert routing mechanism; Parallel expert processing module: used to adopt a hybrid weight-sharing group attention mechanism to map the basic units assigned by experts to different attention head groups for parallel processing; wherein the hybrid weight-sharing group attention mechanism is a combination of weight sharing and hybrid strategies in the attention mechanism; Optimization module: used to optimize routing decisions based on the auxiliary loss function during the training process of the hybrid expert model, so that the routing in the inference phase is consistent with that in the training phase.
9. A computer device, characterized in that: The computer device includes a processor and a memory coupled to the processor, wherein the memory stores program instructions for implementing a hybrid expert model optimization method based on dynamic token routing as described in any one of claims 1-7; and the processor is used to execute the program instructions stored in the memory to implement a hybrid expert model optimization method based on dynamic token routing.
10. A storage medium, characterized in that: The invention stores program instructions executable by a processor, wherein the program instructions are used to execute a hybrid expert model optimization method based on dynamic token routing as described in any one of claims 1 to 7.
Citation Information
Cited By
Large language model training method and reasoning method
CN120875045A
Multi-task electromagnetic model based on hybrid expert network
CN121981190A