A method for pruning a sparse hybrid expert model, a terminal device, and a storage medium

CN122797653APending Publication Date: 2026-09-22SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610971455.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

现有模型裁剪方法为逐层裁剪,每层裁剪产生的精度误差会沿网络深度持续累积,最终导致稀疏混合专家模型整体精度大幅衰减

Benefits of technology

[0009]第四方面,本申请实施例提供了一种计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,所述计算机程序被处理器执行时实现上述第一方面中任一项所述的稀疏混合专家模型的裁剪方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797653A_ABST
    Figure CN122797653A_ABST
Patent Text Reader

Abstract

The application is suitable for the field of compression technology, and provides a pruning method of a sparse mixed expert model, a terminal device and a storage medium. The method comprises the following steps: grouping and dividing mixed expert model layers in an initial sparse mixed expert model to obtain a local block comprising a plurality of mixed expert model layers; then, by comparing output features of adjacent mixed expert model layers in the local block, a candidate layer with high redundancy is screened out; then, in the candidate layer, experts in the candidate layer are pruned according to expert importance scores of the experts in the candidate layer, so that a target sparse mixed expert model is obtained. In the application, the feature similarity between the mixed expert model layers is determined in units of local blocks, the mixed expert model layers with high repetition between layers are identified, the mixed expert model layers and experts with large contribution to the overall output of the local block are preferentially retained, the feature transformation logic within the local block is maintained to be complete, and the accuracy of the pruned sparse mixed expert model is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of model compression technology, and in particular relates to a pruning method, terminal device and storage medium for sparse hybrid expert models. Background Technology

[0002] Sparse hybrid expert models are a type of neural network architecture that is characterized by "high parameters and low computation". Each input only dynamically activates a small portion of the "expert" subnetwork, thereby significantly increasing the model capacity while keeping the computational cost per inference low.

[0003] Currently, sparse hybrid expert models have a massive number of total parameters, requiring all expert parameters to reside entirely in GPU memory. Even if only a small number of experts are sparsely activated during the inference phase, the sheer size of the overall parameter set necessitates the deployment of multiple high-performance servers, resulting in extremely high GPU memory consumption and data transmission bandwidth overhead. To address these issues, the industry commonly employs expert structured pruning schemes to remove redundant experts from the hybrid expert model layers, thereby accelerating inference and reducing resource costs by decreasing the total number of experts. Existing model pruning methods involve layer-by-layer pruning, where the accuracy error generated by each layer accumulates continuously along the network depth, ultimately leading to a significant decrease in the overall accuracy of the sparse hybrid expert model. Summary of the Invention

[0004] This application provides a pruning method, terminal device, and storage medium for sparse hybrid expert models, which can solve the following problems.

[0005] In a first aspect, embodiments of this application provide a pruning method for sparse hybrid expert models, including: The hybrid expert model layers in the initial sparse hybrid expert model are divided into groups to obtain multiple local blocks, wherein a local block contains multiple consecutive hybrid expert model layers. Based on the similarity of the output features between adjacent hybrid expert model layers in the local block, the layer importance of each hybrid expert model layer in the local block is determined. Based on the importance of the layers, candidate layers with high redundancy are selected from the hybrid expert model layers within the local block; Based on the expert importance scores of the experts in the candidate layer, the experts in the candidate layer are pruned to obtain the target sparse hybrid expert model.

[0006] In this application, the hybrid expert model layers in the initial sparse hybrid expert model are grouped and divided into local blocks containing multiple hybrid expert model layers. Then, by comparing the output features of adjacent hybrid expert model layers in the local blocks, candidate layers with high redundancy are selected. Next, within the candidate layers, experts are pruned based on their expert importance scores to obtain the target sparse hybrid expert model. This application determines the feature similarity between hybrid expert model layers on a local block basis, identifies hybrid expert model layers with high repetition and functional redundancy, and performs pruning based on the feature similarity of adjacent hybrid expert model layers. Priority is given to retaining hybrid expert model layers and experts that contribute significantly to the overall output of the local block, maintaining the integrity of the feature transformation logic within the local block. The semantics and feature expressive power of the sparse hybrid expert model are not significantly attenuated after pruning, ensuring the accuracy of the pruned sparse hybrid expert model.

[0007] Secondly, embodiments of this application provide a pruning device for a sparse hybrid expert model, including: The block partitioning module is used to group and partition the hybrid expert model layers in the initial sparse hybrid expert model to obtain multiple local blocks, wherein a local block contains multiple and consecutive hybrid expert model layers; An importance determination module is used to determine the layer importance of each hybrid expert model layer within the local block based on the similarity of the output features between adjacent hybrid expert model layers in the local block. The layer filtering module is used to filter candidate layers with high redundancy from the hybrid expert model layers within the local block according to the importance of the layers. The model pruning module is used to prune the experts in the candidate layer according to their expert importance scores to obtain the target sparse hybrid expert model.

[0008] Thirdly, embodiments of this application provide a terminal device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the pruning method of the sparse hybrid expert model described in any one of the first aspects above.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the pruning method for the sparse hybrid expert model described in any one of the first aspects.

[0010] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the pruning method of the sparse hybrid expert model described in any one of the first aspects.

[0011] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart illustrating a pruning method for a sparse hybrid expert model provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a method for determining the importance of a layer according to an embodiment of this application; Figure 3 This is a flowchart illustrating a method for determining expert importance scores according to an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a pruning device for a sparse hybrid expert model provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0014] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0015] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0016] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.

[0017] Sparse hybrid expert models are a neural network architecture characterized by "high parameters and low computation." This model possesses a large number of total parameters, but only dynamically activates a small subset of "expert" subnetworks (referred to as experts) when processing each input data (i.e., tokens or sample data). This significantly increases model capacity while maintaining low computational cost per inference. However, because all expert parameters in a sparse hybrid expert model must reside fully in GPU memory, even if only a small number of experts are sparsely activated during inference, the sheer size of the overall parameter set necessitates the use of multiple high-performance server clusters for inference deployment, resulting in extremely high GPU memory consumption and data transmission bandwidth overhead. Furthermore, long-term fixed Top-k routing can easily lead to expert load imbalance, with many experts remaining idle for extended periods, resulting in low hardware computing power utilization and limited model inference throughput. Top-k represents the selection of the top K experts.

[0018] To address the above issues, the general approach is to reduce the number of experts, thereby decreasing resource consumption. Currently, the mainstream expert reduction methods fall into two categories: The first type is a globally unified threshold pruning scheme: a single expert importance threshold is set for all hybrid expert model layers (hereinafter referred to as layers) in the sparse hybrid expert model, and experts with global scores below the threshold are uniformly deleted. This scheme is simple to implement, but has significant drawbacks: different hybrid expert model layers perform different functions, with shallow layers focusing on basic semantic extraction and deep layers responsible for high-order logic generation. The distribution of experts, activation intensity, and routing scores vary greatly among layers. The globally unified threshold cannot adapt to the functional differences between layers, and it is easy to cause cross-layer imbalance problems such as over-pruning of shallow layers leading to the loss of basic features and insufficient pruning of deep layers leading to the inability to remove redundancy. The accuracy of model generation after pruning is significantly reduced.

[0019] The second type is the layer-by-layer independent pruning scheme: the importance of experts is calculated layer by layer, an independent pruning threshold is set for each layer, and experts are deleted. After each layer is pruned, the process moves directly to the next layer, and accuracy loss is only checked within a single layer. This scheme can adapt to a single layer of expert distribution, but the layer-by-layer sequential pruning will cause errors to accumulate layer by layer: the feature bias introduced by the previous layer will be fully transmitted to all subsequent layers, and severe semantic distortion will occur after multiple layers of errors are superimposed.

[0020] To address the aforementioned issues, this application proposes a pruning method for sparse hybrid expert models. This method involves grouping the hybrid expert model layers in the initial sparse hybrid expert model into multiple local blocks, where each local block contains multiple consecutive hybrid expert model layers. Then, based on the redundancy of the hybrid expert model layers within each local block, the hybrid expert model layers requiring pruning are determined. Furthermore, within these pruning layers, experts are pruned according to their importance scores to obtain the target sparse hybrid expert model. This application can prune redundant hybrid expert model layers, retaining those that contribute significantly to accuracy, thus reducing the accuracy degradation of the pruned sparse hybrid expert model.

[0021] The following combination Figure 1 The pruning method of the sparse hybrid expert model in the embodiments of this application will be described in detail.

[0022] Figure 1 A schematic flowchart of the pruning method for the sparse hybrid expert model provided in this application is shown, with reference to... Figure 1 The method is described in detail below: S101, the hybrid expert model layers in the initial sparse hybrid expert model are divided into groups to obtain multiple local blocks, wherein a local block contains multiple and consecutive hybrid expert model layers.

[0023] In this embodiment, the sparse hybrid expert model contains multiple hybrid expert model layers (MoE layers). Each hybrid expert model layer contains multiple experts. When processing a token, only the top-K experts are activated through a gating network / router, where K is set as needed. Although the total number of parameters is huge, only a small portion of the parameters are used in a single calculation, thus improving computational efficiency.

[0024] Hybrid expert model layer (MoE layer) is a neural network architecture paradigm. Its core idea is to replace the traditional single-layer feedforward network with multiple experts and dynamically allocate input through a gating network. It is a broad category that theoretically includes both sparse activation and dense activation paths.

[0025] A gated network (i.e., a router) is a learnable routing mechanism that dynamically calculates expert scores based on input data, denoted as routing scores, and selects the Top-K most suitable experts for processing each token.

[0026] Experts consist of multiple parallel subnetworks (usually feedforward neural networks, FFN), each expert being responsible for processing a specific type of data pattern or task feature.

[0027] In this embodiment, the undone Sparse Mixture of Experts (SMoE) model is denoted as the initial Sparse Mixture of Experts model. The MoE layers in the initial SMoE model are divided into multiple local blocks according to network depth. The number of MoE layers in a local block can be set as needed. For example, in a general large language model, a local block may include 4 to 6 consecutive MoE layers; in very large models or scenarios where recalibration costs are high, a local block may also include 6 to 8 consecutive MoE layers.

[0028] For example, the initial sparse hybrid expert model includes MoE layers a, b, c, d, e, and f connected in sequence. If three MoE layers form a local block, then MoE layers a, b, and c form a local block; MoE layers d, e, and f form a local block.

[0029] The purpose of local block partitioning is to transform the full model pruning problem into several locally verifiable problems. Compared to the method of performing only one unidirectional pruning from the first layer to the last layer, the progressive processing within local blocks can absorb local errors within each local block and back off in time when the threshold is not met.

[0030] S102, based on the similarity of the output features between adjacent hybrid expert model layers in the local block, determine the layer importance of each hybrid expert model layer in the local block.

[0031] In this embodiment, inputting the test samples from the sample set into the initial sparse hybrid expert model can yield the output features of each hybrid expert model layer.

[0032] Adjacent hybrid expert model layers are hybrid expert model layers that are connected. For example, if MoE layer a, MoE layer b, and MoE layer c are connected in sequence and form a local block, then the hybrid expert model layer adjacent to MoE layer a is MoE layer b; and the hybrid expert model layers adjacent to MoE layer b are MoE layer a and MoE layer c.

[0033] In this embodiment, the similarity of output features is used to measure the degree of similarity between the output feature vectors of two hybrid expert model layers after processing the same input data (i.e., test samples). If the similarity of the output features of the two hybrid expert model layers is relatively high, it indicates that there is functional redundancy between the two hybrid expert model layers.

[0034] Layer importance is a metric for evaluating the contribution or redundancy of a layer in a hybrid expert model within the overall model. Layers with lower importance may contain more redundant information and are therefore more suitable for pruning. Higher layer importance indicates a more critical layer in the hybrid expert model.

[0035] Specifically, such as Figure 2 As shown, methods for determining the importance of layers may include: S1021, Input the test samples in the sample set into the initial sparse hybrid expert model to obtain the layer output feature vectors of each layer of the hybrid expert model.

[0036] In this embodiment, at least one test sample exists in the sample set, and the type of the test sample is determined according to the type of the sparse hybrid expert model. For example, if the sparse hybrid expert model is a text processing model, the test sample is text data; if the sparse hybrid expert model is an image and text processing model, the test sample is image and text data.

[0037] When a test sample 'a' is input into the initial sparse hybrid expert model, the corresponding layer output feature vectors of each hybrid expert model layer for that test sample 'a' are obtained. Similarly, when a test sample 'b' is input into the initial sparse hybrid expert model, the corresponding layer output feature vectors of each hybrid expert model layer for that test sample 'b' are obtained. For each test sample, a set of layer output feature vectors for each hybrid expert model layer is obtained; if there are M test samples, each hybrid expert model layer will generate M layer output feature vectors, where M ≥ 1.

[0038] If there are multiple test samples, each hybrid expert model layer will generate multiple layer output feature vectors. For a hybrid expert model layer, the multiple layer output feature vectors are weighted and summed or averaged to obtain the final layer output feature vector of the hybrid expert model layer.

[0039] S1022, based on the layer output feature vector of the i-th hybrid expert model layer in the local block and the layer output feature vector of the target hybrid expert model layer, determine the similarity between the i-th hybrid expert model layer and the target hybrid expert model layer, where 1≤i≤N, N is the total number of hybrid expert model layers in the local block, and the target hybrid expert model layer is the hybrid expert model layer in the local block that is adjacent to the i-th hybrid expert model layer.

[0040] In this embodiment, the similarity can be determined using calculation methods such as cosine similarity or Euclidean distance.

[0041] Specifically, based on the layer output feature vector of the i-th hybrid expert model layer and the layer output feature vector of the (i-1)-th hybrid expert model layer in the local block, a first similarity between the i-th hybrid expert model layer and the (i-1)-th hybrid expert model layer is determined; based on the layer output feature vector of the i-th hybrid expert model layer and the layer output feature vector of the (i+1)-th hybrid expert model layer in the local block, a second similarity between the i-th hybrid expert model layer and the (i+1)-th hybrid expert model layer is determined.

[0042] As an example, for a hybrid expert model layer l (i.e., the i-th hybrid expert model), calculate the cosine similarity between hybrid expert model layer l and the previous hybrid expert model layer l-1 (i.e., the (i-1)-th hybrid expert model), that is... , The cosine similarity (i.e., the first similarity) between the hybrid expert model layer l and the previous hybrid expert model layer l-1 is used to characterize the hybrid expert model layer l. This is the layer output feature vector of layer l in the hybrid expert model. Let l be the layer output feature vector of the previous hybrid expert model layer l-1. Calculate the cosine similarity between hybrid expert model layer l and the next hybrid expert model layer l+1 (i.e., the (i+1)th hybrid expert model layer), i.e. , The cosine similarity (i.e., the second similarity) between the hybrid expert model layer l and the next hybrid expert model layer l+1 is used to characterize the hybrid expert model layer l. This will be used to output the feature vector for the next layer l+1 of the hybrid expert model.

[0043] In practical use, if the hybrid expert model layer is the first hybrid expert model layer in a local block, then the adjacent hybrid expert model layer is the next hybrid expert model layer l+1. If the hybrid expert model layer is the last hybrid expert model layer in a local block, then the adjacent hybrid expert model layer is the previous hybrid expert model layer l-1.

[0044] S1023, Based on the similarity, determine the layer importance of the i-th hybrid expert model layer.

[0045] In one approach, the maximum value between the first similarity and the second similarity is determined to obtain the redundancy of the i-th hybrid expert model layer; based on the redundancy, the layer importance of the i-th hybrid expert model layer is determined.

[0046] Specifically, the redundancy is subtracted from 1 to obtain the layer importance of the i-th hybrid expert model layer.

[0047] For example, if the first similarity is 0.9 and the second similarity is 0.7, then the redundancy is 0.9, and 1-0.9=0.1, so the layer importance is 0.1.

[0048] In another approach, the first similarity and the second similarity are weighted and summed or mean is calculated to obtain the equilibrium similarity. The layer importance corresponding to different equilibrium similarities is pre-stored, and the layer importance corresponding to the equilibrium similarity is searched to obtain the layer importance of the i-th hybrid expert model layer.

[0049] S103, based on the importance of the layers, candidate layers are selected from the hybrid expert model layers within the local block, wherein the candidate layers are the hybrid expert model layers that are redundant in the local block.

[0050] In this embodiment, the layer importance of each hybrid expert model layer in the local block is sorted, and the hybrid expert model layers with a layer importance greater than a preset threshold are extracted. The extracted hybrid expert model layers are then determined as candidate layers.

[0051] The preset threshold can be a fixed value set in advance as needed; or, based on the position of the local block in the initial sparse mixture expert model, the preset threshold corresponding to the local block can be found, and different local blocks can correspond to different preset thresholds.

[0052] S104, based on the expert importance scores of the experts in the candidate layer, the experts in the candidate layer are pruned to obtain the target sparse hybrid expert model.

[0053] In this embodiment, the expert importance score represents the contribution of each expert in the sparse hybrid expert model to the model. The higher the expert importance score, the more important the expert is, and the more carefully the expert should be pruned. The lower the expert importance score, the less important the expert is, and the expert can be pruned.

[0054] The experts in the candidate layer are sorted according to their importance scores. Experts with an importance score greater than a preset value are extracted to obtain the target experts. The target experts in the candidate layer are then pruned.

[0055] In this embodiment, after pruning the candidate layer in a local block, the next local block can be traversed and pruned until all local blocks are traversed, thus obtaining the target sparse hybrid expert model.

[0056] The target sparse hybrid expert model is a pruned version of the initial sparse hybrid expert model. The target sparse hybrid expert model can be deployed on devices for use.

[0057] In different scenarios, the target sparse hybrid expert model can process different input data and obtain corresponding output results.

[0058] Specifically, when processing graphic and text data, the graphic and text data to be processed is input into the target sparse hybrid expert model. The target sparse hybrid expert model performs graphic and text understanding, extracts key information, and other processing on the graphic and text data to be processed, and obtains the corresponding output results. For example, the output results can be response text or response images.

[0059] When processing videos, the video to be processed is input into the target sparse hybrid expert model. The target sparse hybrid expert model performs video content recognition, video editing, and other processing on the video to obtain the corresponding output results. For example, the output results can be videos with added text, edited videos, and long video summaries.

[0060] When processing audio, the audio to be processed is input into the target sparse mixture expert model. The target sparse mixture expert model performs speech recognition, audio editing, audio noise reduction, speech conversion and other processing on the audio to be processed, and obtains the corresponding output results. For example, the output results can be noise-reduced audio, speech-to-text text and audio with added sound effects.

[0061] When processing an image, the image to be processed is input into the target sparse mixture expert model. The target sparse mixture expert model performs image recognition, image editing, and image cropping on the image to be processed, and obtains the corresponding output results. For example, the output results can be the edited image, the cropped image, etc.

[0062] In this application, the hybrid expert model layers in the initial sparse hybrid expert model are grouped and divided into local blocks containing multiple hybrid expert model layers. Then, by comparing the output features of adjacent hybrid expert model layers in the local blocks, candidate layers with high redundancy are selected. Within the candidate layers, experts are pruned based on their expert importance scores to obtain the target sparse hybrid expert model. This application determines the feature similarity between hybrid expert model layers on a local block basis, identifies hybrid expert model layers with high repetition and functional redundancy, and performs pruning based on the feature similarity of adjacent hybrid expert model layers. Priority is given to retaining hybrid expert model layers and experts that contribute significantly to the overall output of the local block, maintaining the integrity of the feature transformation logic within the local block. The semantics and feature expressive power of the sparse hybrid expert model are not significantly attenuated after pruning, ensuring the accuracy of the pruned sparse hybrid expert model.

[0063] In one possible implementation, step S104 may include: S1041, based on the expert importance scores of the experts in the candidate layer, select the experts that need to be pruned from the candidate layer to obtain a candidate pruning expert set composed of the experts that need to be pruned.

[0064] In this embodiment, the experts in the candidate layer are sorted according to their importance scores, and those with an importance score greater than a preset value are extracted to obtain the target experts. The target experts are the experts that need to be pruned. The set of all target experts is the candidate pruning expert set.

[0065] In one approach, the methods for determining expert importance scores include: Multiple test samples from the sample set are input into the initial sparse mixture expert model to determine the experts activated when processing each test sample. The number of times each expert is activated when processing multiple test samples is counted, and the expert importance score is determined based on the number of times each expert is activated. Specifically, the expert importance scores corresponding to different activation frequencies are pre-stored.

[0066] In another way, such as Figure 3 As shown, the methods for determining expert importance scores include: S11, input the test samples from the sample set into the initial sparse hybrid expert model, and determine the initial routing score of each expert in each layer of the hybrid expert model.

[0067] In this embodiment, an initial routing score can be obtained for each expert when processing each test sample. If there are M test samples, expert q can obtain M initial routing scores.

[0068] S12, based on the initial routing scores of each expert, determine the aggregate routing score of each expert, wherein the aggregate routing score represents the probability that the expert is selected to process the test sample.

[0069] In this embodiment, according to the formula Calculate the expert's aggregated route score; where, The aggregated route score is the score of the e-th expert in the l-th hybrid expert model layer. Given the x-th test sample in the input test sample set, the normalized initial routing score of the e-th expert in the l-th hybrid expert model layer; The sample set for the test samples; This is the normalized initial routing score of the j-th expert in the l-th hybrid expert model layer when the x-th test sample is input; is the total number of experts in the l-th hybrid expert model layer. This represents the sum of all initial route scores for the e-th expert in the l-th hybrid expert model layer, given each test sample in the sample set.

[0070] S13, based on the initial routing score of the expert and the output vector of the expert, determine the aggregate activation strength of each expert, wherein the aggregate activation strength characterizes the output contribution of the expert.

[0071] In this embodiment, test samples are input into an initial sparse mixture expert model to obtain the output vectors of each expert. The output vectors of the experts are different for different test samples.

[0072] According to the formula Determine the aggregation activation strength of the experts. Among them, Let be the aggregation activation strength of the e-th expert in the l-th hybrid expert model layer; For the sample set; Given the x-th test sample in the input sample set, the normalized initial routing score of the e-th expert in the l-th hybrid expert model layer; This is the output vector of the e-th expert in the l-th hybrid expert model layer when the x-th test sample is input; For indicator functions, in It was 1 when it was established, and in The value is 0 when the condition is not met. These are preset parameters; for The L2 norm (vector magnitude).

[0073] S14, the aggregated routing score and the aggregated activation intensity of the experts are weighted and fused to determine the expert importance score of each expert.

[0074] In this embodiment, the aggregated route score is multiplied by a first weight to obtain a first value; the aggregated activation intensity is multiplied by a second weight to obtain a second value; and the first value is added to the second value to obtain the expert importance score.

[0075] Or, using Obtain the expert importance score. Among them, Let be the expert importance score of the e-th expert in the l-th hybrid expert model layer; The aggregated route score is the score of the e-th expert in the l-th hybrid expert model layer. Let be the aggregation activation strength of the e-th expert in the l-th hybrid expert model layer; These are preset parameters. Norm() represents the normalization function.

[0076] S1042, In the candidate layer, the experts in the candidate pruning expert set are pruned, and after pruning all the candidate layers, the target sparse hybrid expert model is obtained.

[0077] In this embodiment, in the j-th candidate layer of the n-th local block, the experts in the candidate pruning expert set are pruned once to obtain the j-th candidate layer in the n-th local block after pruning; the j+1 candidate layers in the n-th local block are pruned; after the candidate layers in all local blocks are pruned, the target sparse hybrid expert model is obtained, 1≤n≤N, where N is the total number of local blocks.

[0078] Alternatively, select experts with a preset step size from the candidate pruning expert set (denoted as candidate experts). Perform pruning attempts on these candidate experts in the candidate layer to obtain the pruned local block. Compare the output vector of the pruned local block with the output vector of the unpruned local block to determine the error. If the error is greater than a preset value, the candidate expert cannot be pruned, and no pruning is performed on the candidate experts in the candidate layer. If the error is less than or equal to the preset value, the candidate experts in the candidate layer can be pruned. Continue to select experts with a preset step size from the candidate pruning expert set, repeating the above operation until all experts in the candidate pruning expert set in the candidate layer have been pruned, finally obtaining the pruned candidate layer. Continue pruning other candidate layers in this local block. After the operation on one local block is completed, continue pruning another local block until all local blocks are pruned, obtaining the target sparse hybrid expert model.

[0079] Specifically, the implementation process of step S1042 may include: S21, In the candidate layer, the experts in the candidate pruning expert set are pruned to obtain the pruned candidate layer, and the local block where the pruned candidate layer is located is the pruned local block.

[0080] In this embodiment, according to the pruning step size Y, Y experts are selected from the candidate pruning expert set in ascending order of their expert importance scores; these Y experts form a subset. The pruning step size can be determined based on the total number of experts in the candidate layer. Specifically, the pruning step size corresponding to the total number of experts in the candidate layer is found to obtain the pruning step size Y corresponding to the current candidate layer.

[0081] Alternatively, the pruning step size Y can be the total number of experts in the candidate layer multiplied by a preset percentage, for example, the preset percentage can be 2.5% to 5%. If the total number of experts in the candidate layer is less than the preset number, then the pruning Y is set to 1, and there is no restriction here.

[0082] To preserve the basic expressive power of the candidate layer, a minimum number of experts to retain is set. If, after removing experts from a subset, the number of remaining experts in the candidate layer is less than the minimum number of experts to retain, then the experts in the subset are not pruned; alternatively, the pruning step size is reduced so that the number of remaining experts in the pruned candidate layer is greater than or equal to the minimum number of experts to retain. The minimum number of experts to retain can be determined based on the total number of experts in the unpruned candidate layer. For example, the minimum number of experts to retain is 25% to 3% of the total number of experts in the unpruned candidate layer, and there is no specific limit here.

[0083] For example, if the minimum number of experts to be retained in the candidate layer is 9; the number of experts in the candidate layer at the current moment is 11; and the number of experts in the subset determined at the current moment is 3, if after pruning 3 experts, the number of experts remaining in the candidate layer is less than 9, then the experts in the subset are not pruned, or 2 experts in the subset are pruned.

[0084] In this embodiment, after temporarily pruning the experts in the subset of the candidate layer, the pruned candidate layer at the current moment is obtained, and the local block where the candidate layer is located is the pruned local block.

[0085] Since the number of experts in the candidate layer is reduced after pruning, the routing scores of the remaining experts in the candidate layer are no longer accurate. Therefore, it is necessary to normalize the routing scores of the remaining experts in the candidate layer so that the pruned candidate layer can work.

[0086] Specifically, after obtaining the clipped candidate layers, step S1042 may further include: Obtain the initial routing score of the surviving experts in the candidate layer, wherein the surviving experts are the experts in the candidate layer needed to process the test samples; normalize the initial routing score of the surviving experts to obtain the current routing score of the surviving experts; write the current routing score of the surviving experts into the pruned candidate layer.

[0087] In this embodiment, the initial routing score is the routing score assigned by the router to each expert after the test samples are input into the initial sparse hybrid expert model.

[0088] Using formula The initial routing scores of the surviving experts are normalized. This represents the current routing score of the e-th expert in the l-th hybrid expert model layer when the x-th test sample is input. Given the x-th test sample in the input sample set, the normalized initial routing score of the e-th expert in the l-th hybrid expert model layer; The total number of surviving experts; Let be the normalized initial routing score of the j-th surviving expert in the l-th hybrid expert model layer when the x-th test sample is input.

[0089] In practical use, if some of the Top-K experts selected by the router have been deleted when the test samples are input into the initial sparse hybrid expert model, then experts are selected from the remaining experts in descending order of their routing scores to obtain surviving experts. The routing scores of the surviving experts are then normalized. Alternatively, for the candidate layer, the router re-selects experts from the candidate layer to process the test samples, and the selected experts are the surviving experts. Or, the value of Top-K is reduced, and the experts currently existing in the Top-K experts selected by the router are taken as surviving experts.

[0090] For example, if the test samples are input into the initial sparse hybrid expert model, the router selects Top-K experts including expert 1, expert 3, expert 4, and expert 5. After pruning the candidate layer, expert 4 is pruned, and the router now selects Top-K experts including expert 1, expert 3, and expert 5. Method 1: Select expert 8 with the highest initial routing score from the uncropped experts in the candidate layer and add it to the router's selected Top-K experts. Experts 1, 3, 8, and 5 are the surviving experts. Method 3: Experts 1, 3, and 5 are the surviving experts. Method 4: The router reselects 4 experts from the pruned candidate layer as surviving experts.

[0091] S22, obtain the reconstructed block output vector of the clipped local block and the initial block output vector of the local block when it is not clipped.

[0092] In this embodiment, test samples from the sample set are input into the initial sparse hybrid expert model to determine the initial block output vector of each local block, and the input feature vector of each local block is recorded.

[0093] After obtaining the cropped local block at the current time, the input feature vector of the recorded local block is input into the cropped local block to obtain the output vector of the cropped local block, which is denoted as the reconstruction block output vector.

[0094] S23, Based on the deviation between the reconstructed block output vector and the initial block output vector, determine the block reconstruction error of the local block.

[0095] In this embodiment, the block reconstruction error can also be determined using one or more of the following methods: mean cosine distance, mean square error, and relative entropy. If multiple methods are used, the errors determined by these methods can be weighted to obtain the final block reconstruction error. For example, the mean cosine distance is used to determine the first deviation between the reconstructed block output vector and the initial block output vector; the mean square error is used to determine the second deviation between the reconstructed block output vector and the initial block output vector; and the first and second deviations are weighted to obtain the block reconstruction error.

[0096] Optional, utilize Determine the block reconstruction error. Among them, This refers to block reconstruction error; For the test sample set; The initial block output vector for the b-th local block; The output vector of the reconstructed block after cropping the b-th local block; These are preset parameters.

[0097] S24, if the block reconstruction error is less than a preset threshold, the pruned candidate layer is written into the initial sparse hybrid expert model.

[0098] In this embodiment, the preset threshold can be set as needed and is not limited here. If the block reconstruction error is less than the preset threshold, it means that the output of the local block after pruning the expert is not much different from the output error between the pruning and the current pruning expert. Therefore, it is determined that the current pruning expert can be pruned. Thus, the pruned candidate layer is written into the initial sparse hybrid expert model.

[0099] S25, if the block reconstruction error is greater than or equal to the preset threshold, determine that the experts in the candidate pruning expert set in the initial sparse hybrid expert model will not be pruned.

[0100] In this embodiment, if the block reconstruction error is greater than or equal to a preset threshold, it indicates that the output error between the local block after pruning the expert and the pruning output is relatively large, and it is determined that the expert to be pruned cannot be pruned.

[0101] For each expert in a candidate layer, pruning is performed according to the pruning step size. After pruning an expert in the candidate layer once, the judgment in steps S22 to S25 above is executed once, and then the next pruning is performed according to the pruning step size, until all experts in the candidate layer that need to be pruned have been pruned, and then the pruning of the candidate layer ends. The pruning of another candidate layer in the same local block continues, until all candidate layers in the local block have been pruned, and then the candidate layers in the next local block are pruned, and so on, until all local blocks have been pruned, and the target sparse hybrid expert model is output.

[0102] When performing expert trimming, this application employs a progressive process within local blocks, involving trial trimming, recalibration, acceptance of trimming, or rollback of trimming. After each trimming, the error of the currently trimmed local block is verified, and the trimming is only written into the model structure when the error meets the requirements, thus ensuring the accuracy of the final trimmed model.

[0103] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0104] Corresponding to the pruning method of the sparse hybrid expert model described in the above embodiments, Figure 4 The diagram shows a structural block diagram of the pruning device for the sparse hybrid expert model provided in the embodiments of this application. For ease of explanation, only the parts related to the embodiments of this application are shown.

[0105] Reference Figure 4 The device 300 may include: a block division module 310, an importance determination module 320, a layer screening module 330, and a model trimming module 340.

[0106] The block partitioning module 310 is used to group and partition the hybrid expert model layers in the initial sparse hybrid expert model to obtain multiple local blocks, wherein a local block contains multiple and consecutive hybrid expert model layers. Importance determination module 320 is used to determine the layer importance of each hybrid expert model layer in the local block based on the similarity of the output features between adjacent hybrid expert model layers in the local block; The layer filtering module 330 is used to filter candidate layers from the hybrid expert model layers within the local block according to the importance of the layers, wherein the candidate layers are the hybrid expert model layers that are redundant in the local block; The model pruning module 340 is used to prune the experts in the candidate layer according to their expert importance scores to obtain the target sparse hybrid expert model.

[0107] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0108] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0109] This application also provides a terminal device, see [link to relevant documentation] Figure 5 The terminal device 400 may include: at least one processor 410, a memory 420, and a computer program stored in the memory 420 and executable on the at least one processor 410. When the processor 410 executes the computer program, it implements the steps in any of the above method embodiments, for example... Figure 1 Steps S101 to S104 in the illustrated embodiment. Alternatively, when the processor 410 executes the computer program, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 4 The block division module 310 to the model trimming module 340 are shown to have the following functions.

[0110] For example, a computer program may be divided into one or more modules / units, one or more of which are stored in memory 420 and executed by processor 410 to complete this application. The one or more modules / units may be a series of computer program segments capable of performing a specific function, which are used to describe the execution process of the computer program in terminal device 400.

[0111] Those skilled in the art will understand that Figure 5 This is merely an example of a terminal device and does not constitute a limitation on the terminal device. It may include more or fewer components than shown, or combine certain components, or different components, such as input / output devices, network access devices, buses, etc.

[0112] The processor 410 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0113] The memory 420 can be an internal storage unit of the terminal device or an external storage device, such as a plug-in hard drive, a smart media card (SMC), a secure digital (SD) card, or a flash card. The memory 420 is used to store the computer program and other programs and data required by the terminal device. The memory 420 can also be used to temporarily store data that has been output or will be output.

[0114] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0115] The pruning method for sparse hybrid expert models provided in this application can be applied to terminal devices such as computers, tablets, laptops, netbooks, and personal digital assistants (PDAs). This application does not impose any restrictions on the specific type of terminal device.

[0116] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0117] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0118] In the embodiments provided in this application, it should be understood that the disclosed terminal devices, apparatuses, and methods can be implemented in other ways. For example, the terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, apparatuses, or units, and may be electrical, mechanical, or other forms.

[0119] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0120] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0121] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by one or more processors, it can implement the steps of the various method embodiments described above.

[0122] Similarly, as a computer program product, when the computer program product is run on a terminal device, it enables the terminal device to implement the steps in the above-described method embodiments.

[0123] The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0124] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A pruning method for a sparse hybrid expert model, characterized in that, include: The hybrid expert model layers in the initial sparse hybrid expert model are divided into groups to obtain multiple local blocks, wherein a local block contains multiple consecutive hybrid expert model layers. Based on the similarity of the output features between adjacent hybrid expert model layers in the local block, the layer importance of each hybrid expert model layer in the local block is determined. Based on the importance of the layers, candidate layers are selected from the hybrid expert model layers within the local block, wherein the candidate layers are the hybrid expert model layers that are redundant in the local block; Based on the expert importance scores of the experts in the candidate layer, the experts in the candidate layer are pruned to obtain the target sparse hybrid expert model.

2. The method as described in claim 1, characterized in that, The determination of the layer importance of each hybrid expert model layer within the local block based on the similarity of the output features between adjacent hybrid expert model layers in the local block includes: The test samples in the sample set are input into the initial sparse hybrid expert model to obtain the layer output feature vectors of each layer of the hybrid expert model; Based on the layer output feature vector of the i-th hybrid expert model layer in the local block and the layer output feature vector of the target hybrid expert model layer, the similarity between the i-th hybrid expert model layer and the target hybrid expert model layer is determined, where 1≤i≤N, N is the total number of hybrid expert model layers in the local block, and the target hybrid expert model layer is the hybrid expert model layer in the local block that is adjacent to the i-th hybrid expert model layer. Based on the similarity, the layer importance of the i-th hybrid expert model layer is determined.

3. The method as described in claim 1, characterized in that, The step of pruning the experts in the candidate layer based on their expert importance scores to obtain the target sparse hybrid expert model includes: Based on the expert importance scores of the experts in the candidate layer, the experts that need to be pruned are selected from the candidate layer to obtain a candidate pruning expert set consisting of the experts that need to be pruned. In the candidate layer, the experts in the candidate pruning expert set are pruned. After pruning all the candidate layers, the target sparse hybrid expert model is obtained.

4. The method as described in claim 3, characterized in that, Before selecting experts to be pruned from the candidate layer based on their expert importance scores, and obtaining a candidate pruning expert set consisting of the experts to be pruned, the method further includes: The test samples in the sample set are input into the initial sparse hybrid expert model to determine the initial routing score of each expert in each layer of the hybrid expert model; Based on the initial routing scores of each expert, an aggregate routing score for each expert is determined, wherein the aggregate routing score represents the probability that the expert is selected to process the test sample; Based on the initial routing score of the expert and the output vector of the expert, the aggregate activation strength of each expert is determined, wherein the aggregate activation strength characterizes the output contribution of the expert; The aggregated routing score and aggregated activation intensity of the experts are weighted and fused to determine the expert importance score of each expert.

5. The method as described in claim 4, characterized in that, The step of cropping the experts in the candidate expert set in the candidate layer includes: In the candidate layer, the experts in the candidate pruning expert set are pruned to obtain the pruned candidate layer, and the local block where the pruned candidate layer is located is the pruned local block; Obtain the reconstructed block output vector of the cropped local block and the initial block output vector of the local block when it is not cropped; The block reconstruction error of the local block is determined based on the deviation between the reconstructed block output vector and the initial block output vector; If the block reconstruction error is less than a preset threshold, the pruned candidate layer is written into the initial sparse hybrid expert model.

6. The method as described in claim 5, characterized in that, After determining the block reconstruction error of the local block based on the deviation between the reconstructed block output vector and the initial block output vector, the method further includes: If the block reconstruction error is greater than or equal to the preset threshold, it is determined that experts in the candidate pruning expert set in the initial sparse hybrid expert model will not be pruned.

7. The method as described in claim 5, characterized in that, Before obtaining the reconstructed block output vector of the clipped local block and the initial block output vector of the unclipped local block, the method further includes: Obtain the initial routing score of the surviving experts in the candidate layer, wherein the surviving experts are the experts in the candidate layer needed to process the test samples; The initial routing score of the surviving expert is normalized to obtain the current routing score of the surviving expert; Write the current routing score of the surviving expert into the pruned candidate layer.

8. The method as described in claim 2, characterized in that, The target hybrid expert model layer includes the (i-1)th hybrid expert model layer and the (i+1)th hybrid expert model layer; The determination of the similarity between the i-th hybrid expert model layer and the target hybrid expert model layer based on the layer output feature vector of the i-th hybrid expert model layer in the local block and the layer output feature vector of the target hybrid expert model layer includes: Based on the layer output feature vector of the i-th hybrid expert model layer and the layer output feature vector of the (i-1)-th hybrid expert model layer in the local block, a first similarity between the i-th hybrid expert model layer and the (i-1)-th hybrid expert model layer is determined; Based on the layer output feature vector of the i-th hybrid expert model layer and the layer output feature vector of the (i+1)-th hybrid expert model layer in the local block, a second similarity between the i-th hybrid expert model layer and the (i+1)-th hybrid expert model layer is determined. Accordingly, based on the similarity, the layer importance of the i-th hybrid expert model layer is determined, including: The maximum value between the first similarity and the second similarity is determined to obtain the redundancy of the i-th hybrid expert model layer; Based on the redundancy, the layer importance of the i-th hybrid expert model layer is determined.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the pruning method for the sparse hybrid expert model as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the pruning method for the sparse hybrid expert model as described in any one of claims 1 to 8.