Expert pruning implementation, device and equipment for super-large-scale hybrid expert model based on few examples
By evaluating the influence of expert outputs and the degree of perturbation of word representations, redundant experts in the MoE model are pruned, solving the problems of high video memory consumption and waste of computing resources, and achieving efficient MoE model deployment and inference.
Patent Information
- Application Number
- CN202510832055.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-19
AI Technical Summary
During the deployment process, ultra-large-scale Mixed of Experts (MoE) models consume high graphics memory and require large computing resources. They also suffer from problems such as expert redundancy and inefficient utilization, making it difficult to optimize deployment costs and inference efficiency.
By evaluating the influence of expert outputs and the degree of disturbance of word representation, an expert global importance score is generated, and unnecessary experts are pruned. New equipment and methods, including an influence calculation module, a word perturbation calculation module and an expert determination module, are used to achieve expert pruning.
It significantly reduces memory requirements, improves the efficiency and practicality of model pruning, and ensures the stability and computational efficiency of the model after pruning.
Smart Images

Figure CN120671756A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of hybrid expert models, and in particular to an implementation, device and equipment for expert pruning of a super-large-scale hybrid expert model based on a small number of examples. Background Art
[0002] In the past two years, with the continued expansion of deep learning models, the Mixture of Experts (MWE) model has become a mainstream framework for large models due to its strong scalability. The MoE model uses a sparse activation strategy, activating only a subset of experts (sub-models) during each inference. This significantly reduces computational costs while ensuring efficient expression. The MoE model has been widely used in fields such as natural language processing and computer vision, demonstrating strong performance potential.
[0003] However, as model size surges, the memory consumption of MoE models also increases exponentially. For example, the DeepSeek-R1 model, with its 671 billion parameters, requires 1,500 GB of video memory at BF16 precision. Even with the more memory-efficient FP8 precision, the memory requirement still reaches 750 GB, equivalent to a high-performance GPU cluster of 4×8 A800s or 2×8 H800s. This significant memory overhead becomes a major bottleneck when deploying ultra-large MoE models.
[0004] Traditional MoE models typically rely on significant computing resources to support their large network structures, making them difficult to deploy in real-world applications, especially in resource-constrained environments. Furthermore, due to the large number of experts, the subsets of experts activated for different domain tasks vary significantly, leading to a large amount of unnecessary computation and memory waste, further increasing deployment costs.
[0005] In this context, exploring and proposing efficient and lightweight deployment strategies is particularly important. Reducing expert redundancy in MoE models, improving computational efficiency, and ensuring inference performance have become the focus of researchers and engineers. Summary of the Invention
[0006] The present application provides a method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples, which is characterized by comprising: According to the routing unit in the hybrid expert model, the expert output influence is generated through the expert's gating score and the expert output strength; According to the expert unit in the hybrid expert model, the hidden state of the word before and after the module is compared to generate the word representation disturbance degree; According to the influence of expert output and the perturbation of word unit representation, all word units are summed up to generate the expert global importance score for expert pruning.
[0007] Optionally, the method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples is characterized by: The expert's gated score is the probability distribution of all experts generated by converting the original hidden state of the current word into a normalized logical value through a fully connected layer; The expert output strength is obtained by receiving the original hidden state of all experts and calculating the output vector modulus through the L2 norm.
[0008] Optionally, comparing the hidden states of word units before and after passing through the module based on the expert units in the hybrid expert model to generate the word unit representation disturbance degree includes: The hidden states before and after the module are the original hidden state and the aggregated hidden state; The original hidden state is the hidden state from the initial embedding; The aggregated latent state is the latent state after being processed by the expert; The word-unit representation disturbance degree is calculated by calculating the difference between the cosine similarity and 1 to quantify the degree of change of the word-unit representation by the expert.
[0009] Optionally, the method of summing all word-units based on the influence of expert outputs and the perturbation of word-unit representations to generate an expert global importance score for expert pruning includes: Multiply the influence of all experts' outputs by the perturbation of word units to obtain the global importance score of all experts; Sort all experts according to their global importance score; A set number of experts with the highest ranking are retained, and the rest are all pruned.
[0010] Optionally, the method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples is characterized by: Expert output influence The calculation formula is: ; in, Score the gating for the experts, Output strength for experts, Output influence for experts, is the number of layers, Expert number, Number the word.
[0011] Optionally, the method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples is characterized by: Word unit representation disturbance The calculation formula is: ; in, is the original hidden state, is the hidden state after aggregation, is the word unit indicating the degree of disturbance, is the number of layers, is the word number, is the cosine similarity.
[0012] Optionally, the method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples is characterized by: Expert global importance score The calculation formula is: ; in, Score the gating for the experts, is the word unit indicating the degree of disturbance, Score the global importance of the experts, is the number of layers, is the word number, Expert number, is the length of the input sequence.
[0013] The present application also provides a device for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples, characterized in that the device includes: The influence calculation module is used to generate the expert output influence based on the expert's gating score and the expert output strength; The word perturbation calculation module is used to compare the hidden state of the word before and after passing through the module to generate the perturbation degree of the word representation; The expert determination module is used to sum all word-units and generate an expert global importance score for expert pruning.
[0014] Optionally, the expert determination module includes: Scoring unit, used to calculate the global importance of all experts one by one; Ranking unit, used to rank the global importance of all experts; The pruning unit is used to prune the experts except the retained number according to the ranking.
[0015] The present application also provides an electronic device, characterized in that it is used to implement any of the above-mentioned methods for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples, comprising: A processor for performing all computationally intensive tasks to implement a very large-scale expert pruning method for a mixture of experts model based on a small number of examples; Memory is used to store processor executable instructions and static storage data.
[0016] The beneficial effects of this application are: this application obtains the influence of expert outputs and the degree of disturbance of word unit representation to evaluate the local importance of experts and quantify the global contribution of experts, and performs comprehensive scoring to retain the experts required for the task to prune redundant experts, which significantly reduces memory requirements, and expert selection can be completed with only one forward propagation, thereby improving the efficiency and practicality of model pruning. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required in the embodiments or the description of the prior art. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0018] Figure 1 A flowchart of an embodiment of a method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples disclosed in this application is shown; Figure 2 An example diagram of the architecture of an embodiment of a method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples disclosed in this application is shown; Figure 3 The structural block diagram of a device for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples disclosed in this application is shown. DETAILED DESCRIPTION
[0019] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0020] The terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0021] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0022] In addition, numerous specific details are provided in the detailed description below to better illustrate the present application. Those skilled in the art will appreciate that the present application can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present application.
[0023] This paper proposes an expert pruning method for large-scale mixture of experts (MoE) models, which is used for memory optimization and accelerated inference in large-scale MoE models. Based on a small number of domain examples, this method proposes a phenomenon called Few-Shot Expert Localization, which can effectively identify the subset of experts most relevant to the current task under extremely small sample conditions. Therefore, an output-based expert importance assessment algorithm and an expert-granular word-unit contribution estimation algorithm are designed to comprehensively evaluate the role of experts in the inference process and to measure each expert's contribution to word-unit prediction in a fine-grained manner. Experimental results show that when only half of the experts are retained, the proposed method not only achieves performance close to that of the original model, but also significantly improves inference throughput.
[0024] Example 1 like Figure 1 FIG. 1 is a flowchart of a method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples according to an embodiment of the present application, which specifically includes the following contents: S100, based on the routing unit in the hybrid expert model, generates the expert output influence through the expert's gating score and the expert output strength.
[0025] Specifically, in the Mixture of Experts (MoE) model, expert output strength is used to quantify each expert's direct contribution to the current word. The routing unit receives the word's original hidden state and calculates a gating score to represent the probability distribution of the word assigned to each expert. The selected experts then independently calculate the original hidden state to generate an expert output vector. Finally, the L2 norm of the expert output is calculated to generate the expert output strength, which quantifies the output strength. The expert output strength is multiplied by the gating score to obtain the expert output influence, reflecting the expert's true influence on the word.
[0026] S200, based on the expert unit in the hybrid expert model, compare the hidden states of the word before and after passing through the module to generate the word representation disturbance degree.
[0027] Specifically, the word-unit representation perturbation measure measures the extent to which an expert has modified the word-unit's hidden state. Its core is to compare the difference between the hidden states before and after the expert's manipulation. Weighted aggregation is used to generate the aggregated hidden state. The cosine similarity between the pre- and post-aggregation hidden states is then calculated. To intuitively quantify the perturbation, the word-unit representation perturbation is calculated as the difference between the similarity and 1. A larger value indicates a more significant modification of the word-unit by the expert, and thus a higher level of importance for the expert.
[0028] S300, based on the expert output influence and the word unit representation disturbance, all word units are summed up to generate an expert global importance score for expert pruning.
[0029] Specifically, the expert global importance score integrates data from all word-units to provide a basis for expert selection. First, for each expert, the product of the word-unit representation perturbation and the expert's output influence is calculated to obtain the expert's contribution value for that word-unit. Next, the contribution values of all word-units are summed to obtain the expert's global importance score, which reflects the expert's importance within the entire input sequence. Finally, the experts are sorted based on their scores, retaining a set number of the top-ranked experts and removing the remaining experts.
[0030] In practical applications, not every expert in a MoE model is equally effective in every task; instead, it exhibits significant domain specialization. For some tasks, the model often requires only a small number of domain examples relevant to the task to reliably activate a sparse and highly relevant subset of experts. This phenomenon suggests that despite the large size of the MoE model, many experts are underutilized during inference. Therefore, identifying and retaining the most important experts for the task at hand is crucial for improving the inference efficiency of the MoE model. This application analyzes the expert distribution of large-scale mixed-expert models using DeepSeek-R1 as a representative model. It finds that current MoE models suffer from several issues when handling large-scale tasks. First, significant memory overhead: As the model size increases, the memory requirements of MoE models increase dramatically, requiring extremely high hardware support for large-scale MoE deployments. Second, expert redundancy and inefficient utilization: Although MoE models contain a large number of experts, not every expert is effectively activated for every task during inference, resulting in significant waste of computational and memory resources. Third, there is a lack of efficient, lightweight deployment strategies: Due to the complexity and scale of MoE models, existing deployment strategies often require extensive computing resources to ensure model performance, making it difficult to optimize deployment costs and inference efficiency. Experiments have found that the overlap between experts with the highest gating scores in different fields is very low, while the overlap between experts in different datasets within the same field is high. This suggests that the MoE model can stably activate a sparse and highly correlated subset of experts. While pruning high-scoring experts in a particular field does negatively impact performance in that field, it does not affect model performance in other fields. Taking these issues into account, this method proposes efficient pruning and optimization of experts in large MoE models to reduce the memory overhead of large-scale MoE models.
[0031] In summary, the dynamic evaluation-based Mixed-of-Experts (MoE) model pruning method proposed in this application achieves an optimal balance between model efficiency and performance by systematically analyzing the influence of expert outputs and the perturbation of word lemmas, offering multiple significant advantages. First, by combining a dual evaluation mechanism of gating scores and expert output strength, this method effectively avoids the "falsely active experts" problem caused by traditional methods that rely solely on gating probabilities, making expert selection more accurate and reliable. Experimental data demonstrates that this evaluation approach improves the accuracy of expert selection, particularly in complex reasoning tasks. Second, the innovative design of word lemma perturbation accurately identifies key experts who contribute substantially to the model output while filtering out redundant modules that, while activated, have little impact. This fine-grained evaluation mechanism minimizes model performance loss even when pruning a large number of experts, while significantly reducing graphics memory usage. Importantly, the entire evaluation process is based entirely on real-time data analysis via forward propagation, requiring no additional training or complex parameter adjustments. This ensures both the model's dynamic adaptability and computational efficiency. This method is not only applicable to the general MoE architecture, but also can automatically adjust the expert selection strategy according to different task characteristics, providing a practical solution for the efficient deployment of large-scale pre-trained models.
[0032] As an implementation method, in step S100, according to the routing unit in the hybrid expert model, the expert output influence is generated by the expert's gating score and the expert output strength, including: The gating score of the expert is the probability distribution of all experts generated by converting the original hidden state of the current word into a normalized logical value through a fully connected layer.
[0033] Specifically, such as Figure 2 As shown, the gating score The router receives the original hidden state of the current word through the input hidden state. , after the linear change layer and bias term processing, a logical value is generated, and then normalized by the Softmax function to a probability distribution, which is the gated score The gated score It reflects the probability weight assigned to each expert by the word unit, and the sum of all weights in the same layer should be 1, and the word unit will give priority to activating the gate score. Higher experts.
[0034] Among them, Input Hidden is the initial input representation of the current layer. When it is the first layer, its source is word embedding (Word Embedding) and positional encoding (Positional Encoding). When it is an intermediate layer, its source is the output hidden state (Output Hidden) of the previous layer.
[0035] The expert output strength is obtained by receiving the original hidden state of all experts and calculating the output vector modulus through the L2 norm.
[0036] Specifically, each expert Independent calculation of the hidden state to generate the output vector , the output vector of each expert The generated expert output strength is calculated by the L2 norm formula . According to the output strength of experts , to distinguish experts as active or inert experts.
[0037] The expert output influence is obtained by calculating the expert's gating score and the expert output strength to obtain the expert with the greatest influence on the current word.
[0038] Specifically, the process is as follows Figure 2 As shown in the Output-Aware Expert Importance section, based on the gated score of the currently selected expert and expert output strength , the expert output influence is calculated by the expert output influence calculation formula , and outputs the Expert Score. Its purpose is to assess each expert's influence on the current word, focusing primarily on the expert's gating score and output strength. The expert with the largest product of gating value and expert output modulus has a greater influence on the current word. Therefore, calculating this product estimates the importance of the expert.
[0039] Among them, suppose the expert's gating score is , the expert output strength is , the expert output influence is , the number of layers is , expert number is , the word number is , the calculation formula for expert output influence is: .
[0040] As an implementation method, in step S200, based on the expert unit in the hybrid expert model, the hidden states of the word before and after the word passes through the module are compared to generate the word representation disturbance degree, including: The hidden states before and after the module are the original hidden state and the aggregated hidden state; Specifically, such as Figure 2 As shown, the original hidden state , is the hidden state from Input Hidden, which may be the initial embedded hidden state or the output hidden state of the previous layer depending on the layer.
[0041] The aggregated hidden state , is the collaborative processing result of the selected experts on the input hidden state generated after weighted calculation by the expert units in the aggregated hidden state (Middle Hidden).
[0042] The word-unit representation disturbance degree is calculated by calculating the difference between the cosine similarity and 1 to quantify the degree of change of the word-unit representation by the expert.
[0043] Specifically, such as Figure 2 As shown in the Expert-Level Token Contribution section, by comparing the hidden states of an input token before and after it passes through an expert unit, we can measure the varying degrees of influence each expert has on the input token. Given the hidden state representations before and after a routing expert unit, we calculate the difference between the cosine similarity and 1 to measure these changes. A large change in similarity indicates a significant impact on the representation of the token; a small change in similarity indicates a small impact. By multiplying the two, we can identify the n experts with the highest scores for the current task and retain them. The remaining experts can be pruned to reduce memory usage. Experiments have demonstrated the feasibility of this approach in mathematics and programming. When applied to a specific vertical domain, this approach can use approximately 20 representative data points (domain example samples) from that domain to obtain statistical information on which experts should be pruned. These pruned experts are then recorded and the pruned models are output.
[0044] Among them, let the original hidden state be , the hidden state after aggregation is , the word unit represents the disturbance degree is , the number of layers is , the word number is , the cosine similarity is , the calculation formula of word unit representation disturbance degree is: .
[0045] As an implementation method, in step S300, all word-grams are summed based on the expert output influence and word-gram representation disturbance to generate an expert global importance score for expert tailoring, including: S301: Multiply the expert output influence and the word unit representation disturbance degree of all experts to obtain the global importance score of all experts.
[0046] Specifically, such as Figure 2 As shown in the Expert Score section, the influence of the expert output is calculated by each word unit. And the word unit represents the perturbation degree is The product of all word units is summed up to comprehensively evaluate the expert's performance on all word units.
[0047] Among them, suppose the expert's gating score is , the word unit represents the disturbance degree is , the expert's global importance score is , the number of layers is , the word number is , the expert number is i, and the length of the input sequence is , the expert global importance score calculation formula is: .
[0048] S302: Sort all experts according to their global importance scores.
[0049] Specifically, after obtaining the global importance scores of all experts through the expert global importance score calculation formula in step S301, they are sorted in descending order of scores.
[0050] S303: retain a set number of experts with the highest rankings, and remove all other experts.
[0051] Specifically, according to the expert ranking generated in step S302, Figure 2 As shown in the Expert Score section, the 128 experts with the highest scores were selected and retained, and the rest were trimmed.
[0052] Example 2 Based on the same principle as the above method, we also propose a method for implementing expert pruning of ultra-large-scale hybrid expert models based on a small number of examples. Figure 3 In an embodiment of the present disclosure, a device 100 for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples includes: An influence calculation module 110 is used to generate expert output influence based on the expert's gating score and expert output strength; A word perturbation calculation module 120 is used to compare the hidden states of a word before and after it passes through the module to generate a perturbation degree of the word representation; The expert determination module 130 is used to sum all word-units to generate an expert global importance score for expert pruning.
[0053] As an optional implementation scheme of the present application, optionally, the expert determination module 130 includes: Scoring unit 131, used to calculate the global importance of all experts one by one; A ranking unit 132 is used to rank the global importance of all experts; The pruning unit 133 is configured to prune the experts except for the retained number according to the ranking.
[0054] Obviously, those skilled in the art should understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned control methods. The modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Alternatively, they can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0055] Those skilled in the art will appreciate that all or part of the processes in the above-described embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When executed, the program can include the processes of the above-described control method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD). The storage medium can also include a combination of the above-mentioned types of memory.
[0056] Example 3 Furthermore, the present application proposes an electronic device, characterized in that it is used to implement any of the above-mentioned methods for expert pruning of a large-scale hybrid expert model based on a small number of examples, comprising: A processor for performing all computationally intensive tasks to implement a very large-scale expert pruning method for a mixture of experts model based on a small number of examples; Memory is used to store processor executable instructions and static storage data.
[0057] The electronic device of the embodiment of the present disclosure includes a processor and a memory for storing processor-executable instructions, wherein the processor is configured to implement any of the above-mentioned methods for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples when executing the executable instructions.
[0058] It should be noted that the number of processors can be one or more. Furthermore, the electronic device in the embodiments of the present disclosure may also include an input device and an output device. The processor, memory, input device, and output device may be connected via a bus or other means, which are not specifically limited herein.
[0059] The memory, as a computer-readable storage medium for implementing the method for expert pruning of a very large-scale hybrid expert model based on a small number of examples, can be used to store software programs, computer executable programs, and various modules, such as the program or module corresponding to the method for implementing the method for expert pruning of a very large-scale hybrid expert model based on a small number of examples in the embodiments of the present disclosure. The processor executes the software programs or modules stored in the memory to perform various functional applications and data processing in the electronic device.
[0060] The input device can be used to receive input numbers or signals. The signals can be key signals related to user settings and function control of the device / terminal / server. The output device can include a display device such as a display screen.
[0061] The embodiments of the present application have been described above. The above description is illustrative and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for implementing expert pruning in a large-scale hybrid expert model based on a small number of examples, characterized by: include: According to the routing unit in the hybrid expert model, the expert output influence is generated through the expert's gating score and the expert output strength; According to the expert unit in the hybrid expert model, the hidden state of the word before and after the module is compared to generate the word representation disturbance degree; According to the influence of expert output and the perturbation of word unit representation, all word units are summed up to generate the expert global importance score for expert pruning.
2. The method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples according to claim 1, characterized in that: The expert's gated score is the probability distribution of all experts generated by converting the original hidden state of the current word into a normalized logical value through a fully connected layer; The expert output strength is obtained by receiving the original hidden state of all experts and calculating the output vector modulus through the L2 norm.
3. The method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples according to claim 1, characterized in that: The method of comparing the hidden states of word units before and after passing through the module based on the expert units in the hybrid expert model to generate the word unit representation disturbance degree includes: The hidden states before and after the module are the original hidden state and the aggregated hidden state; The original hidden state is the hidden state from the initial embedding; The aggregated latent state is the latent state after being processed by the expert; The word-unit representation disturbance degree is calculated by calculating the difference between the cosine similarity and 1 to quantify the degree of change of the word-unit representation by the expert.
4. The method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples according to claim 1, characterized in that: The method of summing up all word-units based on the influence of expert output and the perturbation of word-unit representation to generate an expert global importance score for expert pruning includes: Multiply the influence of all experts' outputs by the perturbation of word units to obtain the global importance score of all experts; Sort all experts according to their global importance score; A set number of experts with the highest ranking are retained, and the rest are all pruned.
5. The method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples according to claim 2, characterized in that: Expert output influence The calculation formula is: ; in, Score the gating for the experts, Output strength for experts, Output influence for experts, is the number of layers, Expert number, Number the word.
6. The method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples according to claim 3, characterized in that: Word unit representation disturbance The calculation formula is: ; in, is the original hidden state, is the hidden state after aggregation, is the word unit indicating the disturbance degree, is the number of layers, is the word number, is the cosine similarity.
7. The method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples according to claim 4, characterized in that: Expert global importance score The calculation formula is: ; in, Score the gating for the experts, is the word unit indicating the disturbance degree, Score the global importance of the experts, is the number of layers, is the word number, Expert number, is the length of the input sequence.
8. A device for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples, characterized in that: The device comprises: The influence calculation module is used to generate the expert output influence based on the expert's gating score and the expert output strength; The word perturbation calculation module is used to compare the hidden state of the word before and after passing through the module to generate the perturbation degree of the word representation; The expert determination module is used to sum all word-units and generate an expert global importance score for expert pruning.
9. The apparatus for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples according to claim 8, wherein the expert determination module comprises: Scoring unit, used to calculate the global importance of all experts one by one; Ranking unit, used to rank the global importance of all experts; The pruning unit is used to prune the experts except the retained number according to the ranking.
10. An electronic device, characterized in that: The method for implementing expert pruning of a large-scale hybrid expert model based on a small number of examples as described in any one of claims 1 to 7 comprises: A processor for performing all computationally intensive tasks to implement a very large-scale expert pruning method for a mixture of experts model based on a small number of examples; Memory is used to store processor executable instructions and static storage data.
Citation Information
Patent Citations
Method and system for improving structure of language model based on hybrid expert model
CN118194917A
Hybrid expert language model optimization method and device, equipment, medium and product
CN118673992A
Mixture-of-experts layer with dynamic gating
US20240169463A1