Data routing method and device for hybrid expert model

By using random routing matrix in the early stage of training of hybrid expert models to solve the load imbalance problem, the load balancing and stability of the model are achieved, thereby improving the training efficiency.

CN119990374APending Publication Date: 2025-05-13ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510072358.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The hybrid expert model is prone to load imbalance in the early stage of training, which causes some expert sub-models to occupy a large amount of video memory capacity and affect the training effect and convergence speed.

Method used

Provide a data routing method, by obtaining the original routing matrix of the mixed expert model, determine whether the random routing conditions are met, and if they are met, a random routing matrix is ​​generated, and the routing data is routed according to the matrix, so as to avoid the calculated weight being too concentrated on some expert submodels.

Benefits of technology

Through the use of random routing matrix, we ensure the load balancing in the early stage of model training and improve the stability of the model, thereby promoting the model to converge faster and improving training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990374A_ABST
    Figure CN119990374A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention provide a data routing method and apparatus for a hybrid expert model. The method comprises the steps of obtaining an original routing matrix of the hybrid expert model for to-be-routed data; each element in the original routing matrix is an original routing parameter of each expert sub-model in the hybrid expert model for the to-be-routed data; under the condition that the current training step number is not greater than the preset step number, namely in the initial stage of model training, routing is not directly carried out on data to be routed according to an original routing matrix any more, but a corresponding random routing matrix is determined according to the original routing matrix; each element in the random routing matrix is a random routing parameter corresponding to each original routing parameter; and then routing data to be routed according to the random routing matrix so as to determine which expert sub-model or several expert sub-models the data to be routed is input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present specification relate to the field of large model technology, and in particular, to a data routing method and device for a hybrid expert model. Background Art

[0002] Mixture of Experts (MOE) is a deep learning model composed of multiple sub-models; different sub-models, that is, different "experts" (hereinafter referred to as expert sub-models), are good at different fields or different types of tasks. Mixture of Experts can achieve higher computational efficiency and better performance by decomposing complex tasks into multiple sub-tasks and assigning them to different expert sub-models for processing.

[0003] During the training phase of the hybrid expert model, especially in the early stages of training, since the input data allocation strategy is also in a waiting-for-training state, load imbalance problems are prone to occur. That is, the input data is overly concentrated on certain expert sub-models, causing these expert sub-models to occupy a large amount of video memory capacity. It also causes the loss function to fluctuate greatly, affecting the training effect and convergence speed. Summary of the invention

[0004] One or more embodiments of the present specification provide a data routing method and device for a hybrid expert model to ensure load balancing and training efficiency during model training.

[0005] In a first aspect, one or more embodiments of this specification provide a data routing method of a hybrid expert model, including:

[0006] Obtaining an original routing matrix of the hybrid expert model for the data to be routed; each element in the original routing matrix is ​​respectively an original routing parameter of each expert sub-model in the hybrid expert model for the data to be routed;

[0007] Determine whether the hybrid expert model meets the random routing condition; the random routing condition includes: the current training step number is not greater than the preset step number;

[0008] In the case where the hybrid expert model satisfies the random routing condition, a corresponding random routing matrix is ​​determined according to the original routing matrix; each element in the random routing matrix is ​​a random routing parameter corresponding to each of the original routing parameters;

[0009] The data to be routed is routed according to the random routing matrix, so as to input the data to be routed into at least one of the expert sub-models.

[0010] In a possible implementation, the random routing condition further includes: a value of the random routing flag of the hybrid expert model is true.

[0011] In a possible implementation, determining a corresponding random routing matrix according to the original routing matrix includes:

[0012] Generate a random tensor with the same dimension as the original routing matrix;

[0013] The random routing matrix having the same mean and standard deviation as the original routing matrix is ​​generated according to the random tensor.

[0014] In a possible implementation, routing the to-be-routed data according to the random routing matrix includes:

[0015] Mixing the random routing matrix and the original routing matrix to obtain a mixed routing matrix;

[0016] The data to be routed is routed according to the hybrid routing matrix.

[0017] In a possible implementation, mixing the random routing matrix and the original routing matrix to obtain a mixed routing matrix includes:

[0018] Adding the product of the original routing matrix and the first mixing coefficient and the product of the random routing matrix and the second mixing coefficient to obtain the mixed routing matrix;

[0019] The first mixing coefficient is the ratio of the current training step number to the preset step number; and the sum of the first mixing coefficient and the second mixing coefficient is 1.

[0020] In a possible implementation, the data routing method further includes:

[0021] A routing loss function of the hybrid expert model is calculated according to the random routing matrix or the hybrid routing matrix.

[0022] In a second aspect, one or more embodiments of this specification provide a data routing device of a hybrid expert model, including:

[0023] The original data determination unit is used to obtain the original routing matrix of the hybrid expert model for the data to be routed; each element in the original routing matrix is ​​respectively the original routing parameter of each expert sub-model in the hybrid expert model for the data to be routed;

[0024] A condition judgment unit, used to judge whether the hybrid expert model meets the random routing condition; the random routing condition includes: the current training step number is not greater than the preset step number;

[0025] A random data determination unit, configured to determine a corresponding random routing matrix according to the original routing matrix when the hybrid expert model satisfies the random routing condition; each element in the random routing matrix is ​​a random routing parameter corresponding to each of the original routing parameters;

[0026] A data routing unit is used to route the data to be routed according to the random routing matrix when the hybrid expert model meets the random routing condition, so as to input the data to be routed into at least one of the expert sub-models.

[0027] In a possible implementation, the random routing condition further includes: a value of the random routing flag of the hybrid expert model is true.

[0028] In a possible implementation, the random data determination unit includes:

[0029] A random tensor generation unit, used to generate a random tensor with the same dimension as the original routing matrix;

[0030] The random routing generation unit is used to generate the random routing matrix having the same mean and standard deviation as the original routing matrix according to the random tensor.

[0031] In a possible implementation, the data routing unit includes:

[0032] A routing matrix mixing unit, used for mixing the random routing matrix and the original routing matrix to obtain a mixed routing matrix;

[0033] A first routing unit, configured to route the data to be routed according to the hybrid routing matrix when the hybrid expert model satisfies a random routing condition;

[0034] The second routing unit is used to route the data to be routed according to the original routing matrix when the hybrid expert model does not meet the random routing condition.

[0035] In a possible implementation, the routing matrix mixing unit is used to mix the random routing matrix and the original routing matrix to obtain a mixed routing matrix, including:

[0036] The routing matrix mixing unit is used to add the product of the original routing matrix and the first mixing coefficient and the product of the random routing matrix and the second mixing coefficient to obtain the mixed routing matrix;

[0037] The first mixing coefficient is the ratio of the current training step number to the preset step number; and the sum of the first mixing coefficient and the second mixing coefficient is 1.

[0038] In a possible implementation, the data routing device further includes:

[0039] The loss function calculation unit is used to calculate the routing loss function of the hybrid expert model according to the random routing matrix or the hybrid routing matrix.

[0040] In a third aspect, one or more embodiments of the present specification further provide an electronic device comprising a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory, and when the computer program is executed, the method of the first aspect described above is implemented.

[0041] In a fourth aspect, one or more embodiments of the present specification further provide a computer-readable storage medium, wherein the computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed, the method of the first aspect described above is implemented.

[0042] In summary, the data routing method and related device of the hybrid expert model provided by one or more embodiments of the present specification, when the current training step number of the hybrid expert model is not greater than the preset step number, that is, in the initial stage of model training, does not directly use the original routing matrix to route the data to be routed, but determines a random routing matrix based on the original routing matrix, and routes the data to be routed according to the random routing matrix, thereby utilizing the randomness of each random routing parameter in the random routing matrix to avoid the maximum value of the calculated weight always corresponding to one or several expert sub-models, so as to avoid the routing layer in the initial stage of model training from routing a large amount of data to be routed to one or several expert sub-models, thereby ensuring the load balancing of the hybrid expert model in the initial stage of model training, improving the stability of the model, and further helping the model to converge faster and improve training efficiency.

[0043] Secondly, in the above embodiment, by creating a random tensor and determining the corresponding random routing matrix based on the mean and standard deviation of each original routing parameter in the original routing matrix, it can ensure that each random routing parameter in the random routing matrix has a certain randomness to achieve load balancing, and it can also ensure that it has a certain correlation with the original routing matrix to avoid excessive deviation between the random routing matrix and the original routing matrix, which affects the model's learning and prediction of the current data to be routed.

[0044] Thirdly, in the above embodiment, the random routing matrix and the original routing matrix are mixed by linear interpolation. As the current number of training steps gradually increases during the training process, the coefficient of the original routing matrix gradually approaches 1, and the coefficient of the random routing matrix gradually approaches 0. Thus, in the obtained hybrid routing matrix, the proportion of the random routing matrix gradually decreases, while the proportion of the original routing matrix gradually increases. On this basis, routing is performed according to the hybrid routing matrix, so that the routing strategy of the model is also smoothly transitioned from the random routing strategy to the original routing strategy, which can not only ensure the load balancing in the initial stage of model training, but also avoid the problem of model instability caused by sudden changes in routing strategy during training, thereby improving the stability of the hybrid expert model, thereby helping the model to converge quickly and improving training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions of one or more embodiments of the present specification, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of one or more embodiments of the present specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0046] Figure 1 A flow chart of a data routing method of a hybrid expert model provided for one or more embodiments of this specification;

[0047] Figure 2 A logic block diagram of a hybrid expert model provided for one or more embodiments of this specification;

[0048] Figure 3 A logic block diagram of a hybrid expert model provided for one or more embodiments of this specification;

[0049] Figure 4 A flow chart of a data routing method of a hybrid expert model provided for one or more embodiments of this specification;

[0050] Figure 5 A structural block diagram of a data routing device of a hybrid expert model provided for one or more embodiments of this specification;

[0051] Figure 6 A structural block diagram of an electronic device provided for one or more embodiments of this specification. DETAILED DESCRIPTION

[0052] One or more embodiments of the present invention are further described in detail below through the accompanying drawings and examples. Through these descriptions, the features and advantages of one or more embodiments of the present invention will become clearer and more specific.

[0053] The word "exemplary" is used exclusively herein to mean "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise noted.

[0054] In addition, the technical features involved in different implementations of one or more embodiments of this specification described below can be combined with each other as long as they do not conflict with each other.

[0055] To facilitate understanding, the application scenarios of the technical solutions provided by one or more embodiments of this specification are first described below.

[0056] The hybrid expert model, or MOE model, combines multiple expert sub-models, each of which plays a role in its field of expertise, to jointly cope with complex and changing tasks, improving computing efficiency and task processing effects. The hybrid expert model is widely used in various scenarios such as natural language processing, image processing, and recommendation systems.

[0057] The logical architecture of the hybrid expert model mainly includes the expert layer and the routing layer, such as Figure 2 As shown in the figure, the expert layer is configured with multiple expert sub-models E1, E2, E3, etc.; the routing layer is used to determine to which expert sub-model the input data to be routed is routed (assigned). Specifically, for each data to be routed, the routing layer can calculate and predict the weights of each expert sub-model, and then route the data to be routed to one or several expert sub-models with the largest weights; Figure 2 As shown, assuming that it is known through calculation that the weight corresponding to the expert sub-model E2 is the largest, the data to be routed is allocated to the expert sub-model E2, that is, the data is processed by the expert sub-model E2.

[0058] The relevant calculation function or calculation model used to calculate the weight in the routing layer can be called the routing strategy of the routing layer; the adjustable parameters in the calculation function or calculation model can also be used as parameters to be trained, and iterative optimization is performed during the model training process to improve the routing strategy, so that the routing layer can better allocate each data to be routed to the appropriate expert sub-model, thereby improving the data processing effect of the hybrid expert model.

[0059] In the early stage of training the hybrid expert model, the parameters to be trained in the routing layer have not been trained yet, and the routing quality is poor. It is very easy to route a large amount of input data sets to one or several expert sub-models, resulting in overload of one or several expert sub-models, which in turn leads to an overall load imbalance of the hybrid expert model and even training interruption.

[0060] In view of this, one or more embodiments of this specification provide a data routing method and device for a hybrid expert model to avoid load imbalance in the initial stage of model training.

[0061] Figure 1 A data routing method of a hybrid expert model is provided in one embodiment of this specification. Figure 1 , the data routing method may include the following steps:

[0062] Step 102, obtaining the original routing matrix of the hybrid expert model for the data to be routed;

[0063] The elements in the original routing matrix are respectively the original routing parameters of the expert sub-models in the hybrid expert model for the data to be routed;

[0064] In some embodiments, such as Figure 2 As shown, in the routing layer of the hybrid expert model, the current input data to be routed can be calculated through the linear layer (linear_layer) to obtain the routing parameters corresponding to each expert sub-model, such as the logical value Logits, and then the routing parameters are normalized through the activation function layer (Softmax) to obtain the weights P1, P2, P3, etc. corresponding to each expert sub-model.

[0065] In some embodiments, the routing parameters corresponding to each expert sub-model can be centrally recorded in the form of a matrix. Figure 2 As shown, assuming that there are n expert sub-models E1~En in the hybrid expert model, the above original routing matrix L can be expressed as a matrix with 1 row and n columns:

[0066] [L(1)L(2)L(3)…L(n)];

[0067] Among them, any element L(i) (i=1, 2, 3, ..., n) in the original routing matrix L represents the original routing parameter corresponding to the i-th expert sub-model Ei calculated by the linear layer.

[0068] Step 104, determining whether the hybrid expert model satisfies a random routing condition; the random routing condition includes: the current training step number is not greater than a preset step number;

[0069] In some embodiments, the current number of training steps may be the current number of iterations of the hybrid expert model. Exemplarily, during the model training process, relevant data may be input into the model in batches, each batch may include multiple sample data; each time the model processes a batch of data, an iteration is completed, and the number of training steps increases by 1.

[0070] It should be noted that in addition to the model training process, the hybrid expert model can still be optimized through continuous iterations during the actual application process. Therefore, the above-mentioned current training steps can be the number of iterations in the model training process or the number of iterations in the model application process.

[0071] The load imbalance of the hybrid expert model usually occurs in the early stage of model training (also known as the startup stage WarmUp); after training to a certain extent, the relevant parameters in the routing layer are also optimized, which can basically meet the load balancing requirements. Therefore, in some embodiments, a preset number of steps s' can be set as the demarcation value of the initial training stage, that is, when the current training step number s is not greater than the preset number of steps s', that is, s≤s', the hybrid expert model is considered to be in the early stage of training, and the subsequent step 106 and the like are required to ensure the model load balance.

[0072] Step 106, when the hybrid expert model satisfies the random routing condition, determining a corresponding random routing matrix according to the original routing matrix;

[0073] The elements in the random routing matrix are random routing parameters corresponding to the original routing parameters.

[0074] In some embodiments, the corresponding random routing matrix is ​​determined according to the original routing matrix, which can at least ensure that the dimension of the random routing matrix is ​​the same as that of the original routing matrix.

[0075] Exemplarily, given that the original routing matrix L is expressed as [L(1) L(2) L(3) … L(n)], the random routing matrix Lr can be correspondingly expressed as:

[0076] [Lr(1) Lr(2) Lr(3) … Lr(n)];

[0077] Among them, any element Lr(i) (i=1, 2, 3, ..., n) in the random routing matrix Lr represents the random routing parameter corresponding to the i-th expert sub-model Ei.

[0078] In some embodiments, the random routing parameter may be a random number generated according to a random number generation algorithm.

[0079] Step 108: routing the data to be routed according to the random routing matrix, so as to input the data to be routed into at least one expert sub-model.

[0080] In the above step 108, routing the data to be routed according to the random routing matrix may include assigning the random routing matrix to the input parameters of the activation function layer, so that the activation function layer can calculate the weights of each expert sub-model in the hybrid expert model relative to the data to be routed according to the random routing matrix, and then the hybrid expert model can select one or several expert sub-models according to the size of the weight as the sub-model for processing the data to be routed.

[0081] In the above embodiment, when the current training step number is not greater than the preset step number (i.e., the model training process is in the early stage of training), the data to be routed is routed according to the random routing matrix, and the weights of each expert sub-model relative to the data to be routed are calculated, rather than directly calculating the weights according to the original routing matrix. The randomness of the random routing parameters in the random routing matrix is ​​utilized to avoid the maximum value of the calculated weight always corresponding to one or several expert sub-models. In this way, it is possible to avoid the routing layer from routing a large amount of data to one or several expert sub-models in the early stage of model training.

[0082] Therefore, the above embodiment can ensure the load balancing of the hybrid expert model in the early stage of model training, improve the stability of the model, and further help the model converge faster and improve training efficiency.

[0083] In some embodiments, when the hybrid expert model does not meet the random routing conditions, that is, the model training process has passed the initial training stage, the parameters to be trained related to the routing strategy in the routing layer have been adjusted and optimized to a certain extent, and can meet the load balancing requirements of the model to a certain extent. Therefore, the random routing matrix is ​​no longer calculated, and the routing data is directly routed according to the original routing matrix. For example, the original routing matrix is ​​directly assigned to the input parameters of the activation function layer, thereby realizing the routing strategy originally configured by the hybrid expert model.

[0084] It should be noted that, for the sake of ease of description, in this specification, the strategy for calculating the weights of each expert sub-model based on the original routing matrix and routing the routing data is called the original routing strategy, that is, the routing strategy originally configured and to be trained in the hybrid expert model; the strategy for calculating the weights of each expert sub-model based on the random routing matrix and routing the routing data is called the random routing strategy, that is, the routing strategy used to solve the load imbalance problem in the early stage of model training in some embodiments.

[0085] In some embodiments, such as Figure 3 As shown, the data routing method according to the embodiment of the present application is equivalent to adding a randomization layer between the linear layer and the activation function layer of the routing layer; the above steps 104 and 106 are implemented by the randomization layer, and different routing matrices are provided to the activation function layer for weight calculation in different cases, that is:

[0086] When the random routing conditions are met, the randomization layer provides the random routing matrix to the activation function layer. For example, the randomization layer can assign the random routing matrix to the input parameters of the activation function layer, so that the weights of each expert sub-model are calculated in the activation function layer according to the random routing matrix, and the routing data is routed based on the random routing strategy to avoid imbalance of model load in the initial stage of training.

[0087] When the random routing conditions are not met, the randomization layer provides the original routing matrix to the activation function layer. For example, the randomization layer can assign the original routing matrix to the input parameters of the activation function layer, thereby calculating the weights of each expert sub-model in the activation function layer according to the original routing matrix, and routing the routing data based on the original routing strategy, thereby achieving full adjustment and optimization of the relevant parameters to be trained in the original routing strategy.

[0088] In some embodiments, the random routing condition in the above step 104 may further include: the value of the random routing flag of the hybrid expert model is true.

[0089] Exemplarily, a random routing identifier can also be configured in the hybrid expert model; the random routing identifier can be Boolean data, that is, there are only two values ​​of true (True) and false (False). Relevant personnel can configure the value of the above random routing identifier according to actual needs, so as to indicate whether the hybrid expert model is allowed to perform random routing: when the value of the random routing identifier is true, routing can be performed according to the random routing matrix; when the value of the random routing identifier is false, step 106 can be skipped, and routing can be performed directly according to the original routing matrix.

[0090] In some embodiments, relevant personnel can configure the specific value of the preset number of steps according to training requirements. For example, in the case of a large number of parameters to be trained and a large amount of sample data, the preset number of steps can be set to a larger value to ensure that the relevant parameters of the routing layer are fully trained based on the random routing strategy, avoiding load imbalance in the subsequent processing of a large amount of sample data; conversely, in the case of a small number of parameters to be trained and a small amount of sample data, the preset number of steps can be set to a smaller value to avoid the stage of using the random routing strategy accounting for too large a proportion of the entire training process, affecting the model's effective learning of the sample data.

[0091] In some embodiments, to ensure normal execution of the above data routing, the random routing condition in the above step 104 may further include: the preset number of steps is greater than 0.

[0092] If the judgment in step 104 shows that the preset number of steps is less than or equal to 0, it means that there may be a configuration error in the model, so relevant prompt information can be used to remind relevant personnel to handle it, or the preset number of steps can be automatically changed to the default number of steps to ensure the smooth progress of the model training process.

[0093] In addition, in some embodiments, in order to ensure that the training process of the hybrid expert model proceeds normally, if the random routing condition includes multiple sub-conditions, such as simultaneously including the two sub-conditions of "the current training step number is not greater than the preset step number" and "the value of the random routing identifier of the hybrid expert model is true", the random routing condition can be regarded as being met only when it is determined that each sub-condition is met, and routing is performed according to the random routing matrix; and when any sub-condition is not met, it is regarded as not meeting the random routing condition, and routing is performed directly according to the original routing matrix.

[0094] In some embodiments, such as Figure 4 As shown, determining the corresponding random routing matrix according to the original routing matrix in the above step 106 may specifically include:

[0095] Step 1062, generating a random tensor with the same dimension as the original routing matrix;

[0096] The above random tensor, that is, a tensor with random values, can be generated by a related function, and the dimension of the random tensor is the same as that of the original routing matrix.

[0097] Given that the original routing matrix L is represented as [L(1)L(2)L(3)…L(n)], the random tensor T can also be represented in matrix form:

[0098] [T(1)T(2)T(3)…T(n)];

[0099] Among them, any element T(i) in the random tensor T is a random value.

[0100] Step 1064: Generate a random routing matrix having the same mean and standard deviation as the original routing matrix based on the random tensor.

[0101] In order to avoid excessive deviation between the random routing matrix and the original routing matrix, which would affect the model's learning and prediction of the current data to be routed, in some embodiments of the present application, the random routing matrix is ​​restricted by the mean and standard deviation of the original routing matrix.

[0102] Exemplarily, the mean L_mean and standard deviation L_std of each element in the original routing matrix L may be calculated; then, the random tensor T is scaled according to the mean L_mean and standard deviation L_std to obtain a random routing matrix Lr, as shown in the following formula:

[0103] Lr=T*L_std+L_mean.

[0104] That is to say, for any element Lr(i) in the random routing matrix Lr, the calculation formula is:

[0105] Lr(i)=T(i)*L_std+L_mean.

[0106] Through the above-mentioned step 1062, namely step 1064, it can be ensured that each random routing parameter in the random routing matrix has a certain randomness, and it can be ensured that it has a certain correlation with the original routing matrix, thereby avoiding excessive deviation between the random routing matrix and the original routing matrix, which affects the model's learning and prediction of the current data to be routed.

[0107] In some embodiments, such as Figure 4 As shown, in the above step 108, routing the routing data according to the random routing matrix may specifically include:

[0108] Step 1082, mixing the random routing matrix and the original routing matrix to obtain a mixed routing matrix;

[0109] Step 1084: routing the data to be routed according to the hybrid routing matrix.

[0110] Exemplarily, routing the data to be routed according to a hybrid routing matrix may specifically include assigning the hybrid routing matrix to the input parameters of the activation function layer, so that the activation function layer calculates the weights of each expert sub-model according to the hybrid routing matrix, and then determines one or several expert sub-models for processing the data to be routed according to the size of the weights.

[0111] Since the hybrid routing matrix is ​​obtained by mixing the original routing matrix and the random routing matrix, each hybrid routing parameter has a certain randomness and a certain correlation with the original routing matrix. Therefore, the weights of each expert sub-model are calculated according to the hybrid routing matrix, which can avoid the maximum value of the calculated weight always being concentrated on one or several expert sub-models to ensure the load balancing of the model, and avoid the routing parameters used to calculate the weights deviating too much from the original routing parameters, which affects the model's learning and prediction of the current data to be routed.

[0112] It should be noted that, for the sake of ease of description, corresponding to the above-mentioned original routing strategy and random routing strategy, in this specification, the strategy of calculating the weights of each expert sub-model based on the hybrid routing matrix and routing the routing data is called a hybrid routing strategy. This hybrid routing strategy is also a routing strategy used in some embodiments to solve the load imbalance problem in the initial stage of model training.

[0113] In some embodiments, the random routing matrix and the original routing matrix are mixed in the above step 1082, which may specifically include:

[0114] The product of the original routing matrix L and the first mixing coefficient α is added together with the product of the random routing matrix Lr and the second mixing coefficient β to obtain a hybrid routing matrix Lm;

[0115] The first mixing coefficient α is the ratio of the current training step number s to the preset step number s′; the sum of the first mixing coefficient α and the second mixing coefficient β is 1.

[0116] That is, the calculation formula of the hybrid routing matrix Lm can be expressed as:

[0117] Lm=α*L+β*Lr=(s / s')*L+(1-s / s')*Lr.

[0118] That is to say, for any element Lm(i) in the hybrid routing matrix Lm, the calculation formula is:

[0119] Lm(i)=α*L(i)+β*Lr(i)=(s / s')*L(i)+(1-s / s')*Lr(i).

[0120] Based on the calculation formula of the hybrid routing matrix, it can be seen that in the above embodiment, the random routing matrix and the original routing matrix are mixed by linear interpolation. The weights of each expert sub-model are calculated according to the hybrid routing matrix obtained by mixing, which can take into account the randomness of the relevant parameters and the correlation with the data to be routed, thereby taking into account the load balancing, learning efficiency and stability of the hybrid expert model.

[0121] Specifically, refer to Figure 3 In the logical framework shown, the input parameters of the randomization layer are the original routing matrix calculated by the linear layer, and the output parameters of the randomization layer (i.e., the input parameters of the activation function layer, i.e., the parameters used to calculate the weights) vary depending on the training process at different stages or nodes:

[0122] 1) In the stage of s≤s', the random routing condition is met, and the randomization layer takes the mixed routing matrix as its output parameter and outputs it to the activation function; specifically:

[0123] 1.1) At the node s=0, that is, before the training starts, the randomization layer actually has no input and no output. However, theoretically, based on the above calculation formula, it can be deduced that the closer s is to 0, the closer the first mixing coefficient α=s / s' is to 0, and the closer the second mixing coefficient β=1-α is to 1, the closer the corresponding mixing routing matrix Lm is to Lr, that is, the output parameters of the randomization layer are closer to the random routing matrix;

[0124] 1.2) In the stage of 0<s<s', as the training process proceeds, the current training step number s gradually increases, the first mixing coefficient α=s / s' gradually increases, that is, gradually approaches 1 from 0, and the second mixing coefficient β=1-α gradually decreases, that is, gradually approaches 0 from 1, so that in the mixed routing matrix output by the randomization layer, the proportion of the random routing matrix gradually decreases, while the proportion of the original routing matrix gradually increases, and the routing strategy of the model also transitions from the random routing strategy to the original routing strategy;

[0125] 1.3) At the node s=s', the first mixing coefficient α=s / s'=1, the second mixing coefficient β=1-α=0, and the corresponding mixed routing matrix Lm=L, that is, the output parameters of the randomization layer are the same as the original routing matrix, which is equivalent to the model's routing strategy just switching to the original routing strategy;

[0126] 2) When s>s', the random routing condition is not met, and the model training process has passed the initial training stage. There is no need to execute the above steps 106 and other steps. The randomization layer can directly use the original routing matrix as an output parameter and output it to the activation function, that is, the weights of each expert sub-model are calculated based on the original routing matrix. The routing strategies of the model in the training process after the initial training stage are all the original routing strategies originally configured by the model.

[0127] It can be seen that, compared with adopting a random routing strategy when the random routing condition (such as s≤s') is met, switching to the original routing strategy when the random routing condition (such as s>s') is not met, the above embodiment adopts a hybrid routing strategy for routing when the random routing condition is met, that is: as the current training step number s gradually approaches the preset step number s', the hybrid routing matrix gradually transitions from the random routing matrix to the original routing matrix, which makes the routing strategy of the hybrid expert model smoothly transition from the random routing strategy based on the random routing matrix to the original routing strategy based on the original routing matrix, thereby avoiding the problem of model instability caused by sudden changes in routing strategies during training, that is, improving the stability of the hybrid expert model, which helps the model to converge quickly and improve training efficiency.

[0128] In some embodiments, the above data routing method may further include the following steps:

[0129] According to the random routing matrix or the hybrid routing matrix, the routing loss function of the hybrid expert model is calculated.

[0130] During the model training process, one or more loss functions can be used to calculate the loss caused by different processing steps of the model to the final result, provide a basis for adjusting and optimizing the parameters to be trained, and realize the iteration of the model.

[0131] Exemplarily, for the routing processing of routing data by the routing layer of the hybrid expert model, the loss caused by the routing processing to the final result of the hybrid expert model can be calculated by constructing a corresponding routing loss function, thereby providing a basis for adjusting relevant routing parameters in the routing layer.

[0132] Specifically, in the above embodiment, in the early stage of model training, the loss function is calculated based on a random routing matrix or a hybrid routing matrix that helps load balancing; compared with calculating the loss function based on the original routing matrix, the above embodiment can reduce the fluctuation of the loss function, thereby helping the model to converge quickly and improve training efficiency.

[0133] In addition, the data routing method provided in the above embodiment is easy to implement, and does not require major modifications to the hybrid expert model or its routing strategy. It only needs to add a small amount of code to the relevant routing function or code, such as the judgment code of the random routing condition, the calculation code of the random routing matrix or the hybrid routing matrix, etc., to achieve a smooth and automatic transition of the routing strategy from the random routing strategy to the original routing strategy as the number of training steps increases. This avoids sudden changes in the routing strategy during the model training process and ensures the stability of the model. It also has no effect on the training stage after the initial training stage, and does not require relevant technical personnel to manually switch different routing strategies according to different training stages, thereby improving the degree of automation and training efficiency of the model training process.

[0134] Some embodiments of the present application also provide a data routing device for a hybrid expert model; the data routing device can be configured in the routing layer of the hybrid expert model to route the data to be routed input into the hybrid expert model and determine one or several expert sub-models to process the data to be routed.

[0135] Reference Figure 5 , the data routing device 300 may include:

[0136] The original data determination unit 301 is used to obtain an original routing matrix of the hybrid expert model for the data to be routed; each element in the original routing matrix is ​​an original routing parameter of each expert sub-model in the hybrid expert model for the data to be routed;

[0137] Exemplarily, the original data determination unit 301 can be regarded as Figure 3 The linear layer of the hybrid expert model is used to directly calculate the original routing matrix; alternatively, the original data determination unit 301 can be regarded as an input interface of the randomization layer, used to receive the original routing matrix calculated by the linear layer.

[0138] The condition judgment unit 302 is used to judge whether the hybrid expert model meets the random routing condition; the random routing condition includes: the current training step number is not greater than the preset step number;

[0139] A random data determination unit 303 is used to determine a corresponding random routing matrix according to the original routing matrix when the hybrid expert model satisfies the random routing condition; each element in the random routing matrix is ​​a random routing parameter corresponding to each of the original routing parameters;

[0140] Exemplarily, the condition judgment unit 302 and the random data determination unit 303 can be regarded as Figure 3 The randomization layers of the mixture of experts model shown.

[0141] The data routing unit 304 is used to route the data to be routed according to the random routing matrix when the hybrid expert model meets the random routing condition, so as to input the data to be routed into at least one expert sub-model.

[0142] In the above embodiment, when the current training step number of the hybrid expert model is not greater than the preset step number, that is, in the initial stage of model training, the original routing matrix is ​​not directly used to route the data to be routed, but a random routing matrix is ​​determined based on the original routing matrix, and the data to be routed is routed according to the random routing matrix, thereby utilizing the randomness of each random routing parameter in the random routing matrix to avoid the maximum value of the calculated weight always corresponding to one or several expert sub-models, so that it can be avoided that the routing layer in the initial stage of model training routes a large amount of data to be routed to one or several expert sub-models, thereby ensuring the load balancing of the hybrid expert model in the initial stage of model training, improving the stability of the model, and thus helping the model to converge faster and improve training efficiency.

[0143] In some embodiments, the data routing unit 304 is further configured to route the data to be routed according to the original routing matrix to input the data to be routed into at least one expert sub-model when the hybrid expert model does not satisfy the random routing condition.

[0144] When the above hybrid expert model does not meet the random routing conditions, that is, the hybrid expert model is not in the early stage of training or other situations where random routing is not allowed, weight calculation can be performed based on the original routing matrix according to the traditional training method, that is, routing can be performed based on the original routing strategy originally configured by the hybrid expert model, thereby training the relevant parameters to be trained in the original routing strategy to improve model performance.

[0145] Exemplarily, the data routing unit 304 can be regarded as Figure 3 The output interface of the randomization layer of the hybrid expert model, or the input interface of the activation function layer; accordingly, the data routing unit 304 is specifically configured as follows:

[0146] When the hybrid expert model meets the random routing conditions, the random routing matrix is ​​assigned to the output parameters of the randomization layer or the input parameters of the activation function layer, so that the activation function layer performs weight calculation according to the random routing matrix to implement the random routing strategy;

[0147] When the hybrid expert model does not meet the random routing conditions, the original routing matrix is ​​assigned to the output parameters of the randomization layer or the input parameters of the activation function layer, so that the activation function layer performs weight calculation according to the original routing matrix to implement the original routing strategy.

[0148] Exemplarily, the data routing unit 304 can also be regarded as Figure 3 The activation function layer of the hybrid expert model shown in FIG. 1 inputs different routing matrices into the activation function in different situations to implement weight calculation based on different routing matrices, that is, to implement different routing strategies.

[0149] In some embodiments, the random routing condition configured in the above-mentioned condition judgment unit 302 may include multiple sub-conditions, that is, in addition to the above-mentioned "the current training step number is not greater than the preset step number", it may also include "the value of the random routing identifier of the hybrid expert model is true" and the like.

[0150] By setting different sub-conditions, the execution of random routing strategies can be controlled in many aspects, which is more convenient to meet different training needs. For example, based on the sub-condition that "the value of the random routing identifier of the hybrid expert model is true", relevant personnel can change the value of the random routing identifier in the hybrid expert model at any time to quickly and easily change the randomness of the model routing strategy and start or stop random routing.

[0151] In some embodiments, in order to determine a random routing matrix corresponding to the original routing matrix, the random data determining unit 303 may specifically include:

[0152] A random tensor generation unit, used to generate a random tensor with the same dimension as the original routing matrix;

[0153] The random routing generation unit is used to generate the random routing matrix having the same mean and standard deviation as the original routing matrix according to the random tensor.

[0154] In some embodiments, the data routing unit 304 may specifically include:

[0155] A routing matrix mixing unit, used for mixing the random routing matrix and the original routing matrix to obtain a mixed routing matrix;

[0156] A first routing unit, configured to route the data to be routed according to the hybrid routing matrix when the hybrid expert model satisfies the random routing condition;

[0157] The second routing unit is configured to route the data to be routed according to the original routing matrix when the hybrid expert model does not satisfy the random routing condition.

[0158] In some embodiments, the routing matrix mixing unit can be considered together with the condition judgment unit 302 and the random data determination unit 303 as Figure 3 The randomization layers of the mixture of experts model shown.

[0159] In some embodiments, the first routing unit and the second routing unit can be regarded as Figure 3 The output interface of the randomization layer of the hybrid expert model shown can also be regarded as the input interface of the activation function layer.

[0160] Exemplarily, the first routing unit and the second routing unit can be implemented by a conditional assignment unit. According to different judgment results obtained by the conditional judgment unit, different routing matrices (hybrid routing matrix, original routing matrix, etc.) are assigned to the output parameters of the randomization layer (or the input parameters of the activation function layer), thereby realizing weight calculation according to different routing matrices, that is, realizing different routing strategies.

[0161] In some other embodiments, the first routing unit and the second routing unit can be regarded as Figure 3 Activation function layer of the hybrid expert model shown. Exemplarily, the first routing unit and the second routing unit can be implemented by a preset calculation unit configured with a preset activation function; the preset calculation unit is configured to input different routing matrices (hybrid routing matrix, original routing matrix, etc.) into the activation function for weight calculation according to different judgment results obtained by the condition judgment unit 302, thereby implementing different routing strategies.

[0162] In some embodiments, the routing matrix mixing unit is specifically configured to add the product of the original routing matrix and the first mixing coefficient, and the product of the random routing matrix and the second mixing coefficient to obtain the mixed routing matrix; wherein the first mixing coefficient is the ratio between the current training step number and the preset step number; and the sum of the first mixing coefficient and the second mixing coefficient is 1.

[0163] That is to say, the routing matrix mixing unit can mix the original routing matrix and the random routing matrix by linear interpolation to obtain a mixed routing matrix; the proportion of the original routing matrix and the random routing matrix in the mixed routing matrix can change with the change of the current training step number.

[0164] Compared with adopting a random routing strategy when the random routing condition is met, and switching to the original routing strategy when the random routing condition is not met, the data routing unit in the above embodiment adopts a hybrid routing strategy for routing when the random routing condition is met, that is, as the current training step number s gradually approaches the preset step number s', the hybrid routing matrix gradually transitions from the random routing matrix to the original routing matrix, so that the routing strategy of the hybrid expert model smoothly transitions from the random routing strategy based on the random routing matrix to the original routing strategy based on the original routing matrix, thereby avoiding the problem of model instability caused by sudden changes in routing strategies during training, that is, improving the stability of the hybrid expert model, which helps the model to converge quickly and improve training efficiency.

[0165] In some embodiments, the data routing device 300 may further include:

[0166] The loss function calculation unit is used to calculate the routing loss function of the hybrid expert model according to the random routing matrix or the hybrid routing matrix.

[0167] In the above embodiment, at the initial stage of model training, the loss function is calculated based on a random routing matrix or a hybrid routing matrix that helps load balancing; compared with calculating the loss function based on the original routing matrix, the above embodiment can reduce the fluctuation of the loss function, thereby helping the model to converge quickly and improve training efficiency.

[0168] See also Figure 6 , Figure 6 This is a structural block diagram of an electronic device provided in one or more embodiments of this specification. Figure 6 As shown, the electronic device 500 may include a processor 501 and a memory 502; the memory 502 may be coupled to the processor 501. It is worth noting that the Figure 6 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0169] In a possible implementation, the functions of the data routing device 300 may be integrated into the processor 501. The processor 501 may be configured to perform the following operations:

[0170] Obtaining an original routing matrix of the hybrid expert model for the data to be routed; each element in the original routing matrix is ​​respectively an original routing parameter of each expert sub-model in the hybrid expert model for the data to be routed;

[0171] Determine whether the hybrid expert model meets the random routing condition; the random routing condition includes: the current training step number is not greater than the preset step number;

[0172] In the case where the hybrid expert model satisfies the random routing condition, a corresponding random routing matrix is ​​determined according to the original routing matrix; each element in the random routing matrix is ​​a random routing parameter corresponding to each of the original routing parameters;

[0173] The data to be routed is routed according to the random routing matrix, so as to input the data to be routed into at least one of the expert sub-models.

[0174] In another possible implementation, the data routing device 300 may be configured separately from the processor 501 . For example, the data routing device 300 may be configured as a chip connected to the processor 501 , and the data routing process of the hybrid expert model may be implemented under the control of the processor 501 .

[0175] In addition, in some optional implementations, the electronic device 500 may also include: a communication module, an input unit, an audio processor, a display, a power supply, etc. It is worth noting that the electronic device 500 does not necessarily have to include Figure 6 In addition, the electronic device 500 may also include Figure 6 For components not shown, reference may be made to the prior art.

[0176] In some optional implementations, the processor 501 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor device and / or logic device, which receives input and controls the operation of various components of the electronic device 500.

[0177] The memory 502 may be, for example, a cache, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory or other suitable devices, or one or more thereof. The memory may store the above-mentioned information related to the data routing device 300, and may also store a program for executing the related information. The processor 501 may execute the program stored in the memory 502 to implement information storage or processing, etc.

[0178] The input unit can provide input to the processor 501. The input unit is, for example, a key or a touch input device. The power supply can be used to provide power to the electronic device 500. The display can be used to display display objects such as images and text. The display can be, for example, an LCD display, but is not limited thereto.

[0179] The memory 502 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be a memory that saves information even when the power is off, can be selectively erased, and is provided with more data, examples of which are sometimes referred to as EPROMs, etc. The memory 502 may also be some other type of device. The memory 502 includes a buffer memory (sometimes referred to as a buffer). The memory 502 may include an application / function storage unit for storing application programs and function programs or processes for executing the operation of the electronic device 500 through the processor 501.

[0180] The memory 502 may also include a data storage unit for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit of the memory 502 may include various drivers for the communication function of the computer device and / or for executing other functions of the computer device (such as a messaging application, a contact book application, etc.).

[0181] The communication module is a transmitter / receiver that sends and receives signals via an antenna. The communication module (transmitter / receiver) is coupled to the processor 501 to provide input signals and receive output signals, which may be the same as in a conventional mobile communication terminal.

[0182] Based on different communication technologies, multiple communication modules may be provided in the same computer device, such as a cellular network module, a Bluetooth module and / or a wireless local area network module, etc. The communication module (transmitter / receiver) is also coupled to a speaker and a microphone via an audio processor to provide an audio output via the speaker and receive an audio input from the microphone, thereby realizing a common telecommunication function. The audio processor may include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor is also coupled to the processor 501, so that recording can be performed on the machine through the microphone, and the sound stored on the machine can be played through the speaker.

[0183] One or more embodiments of this specification also provide a computer-readable storage medium capable of implementing all the steps in the data routing method of the hybrid expert model in the above embodiment, wherein a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, all the steps in the data routing method of the hybrid expert model in the above embodiment are implemented. The specific steps can be referred to the above embodiments, and will not be repeated here.

[0184] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative labor. The order of steps listed in the embodiments is only one way of executing the order of many steps and does not represent the only execution order. When the actual device or client product is executed, it can be executed in the order of the method shown in the embodiments or the drawings or in parallel (for example, in a parallel processor or multi-threaded processing environment).

[0185] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, devices (systems) or computer program products. Therefore, the embodiments of this specification may take the form of complete hardware embodiments, complete software embodiments or embodiments combining software and hardware. Moreover, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0186] One or more embodiments of the present specification are described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to one or more embodiments of the present specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0187] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0188] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1The steps for the functions specified in one or more boxes.

[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of one or more embodiments of this specification, rather than to limit them. Although one or more embodiments of this specification have been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features may be replaced by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of one or more embodiments of this specification, and they should all be included in the scope of the claims and the specification of one or more embodiments of this specification.

[0190] One or more embodiments of this specification are described above in conjunction with optional implementation methods, but these implementation methods are only exemplary and serve only as an illustration. On this basis, multiple replacements and improvements can be made to one or more embodiments of this specification, all of which fall within the scope of protection of one or more embodiments of this specification.

Claims

1. A data routing method of a hybrid expert model, characterized in that: include: Obtaining an original routing matrix of the hybrid expert model for the data to be routed; Each element in the original routing matrix is ​​respectively the original routing parameter of each expert sub-model in the hybrid expert model for the data to be routed; Determining whether the hybrid expert model satisfies a random routing condition; The random routing conditions include: the current training step number is not greater than the preset step number; In the case where the hybrid expert model satisfies the random routing condition, a corresponding random routing matrix is ​​determined according to the original routing matrix; each element in the random routing matrix is ​​a random routing parameter corresponding to each of the original routing parameters; The data to be routed is routed according to the random routing matrix, so as to input the data to be routed into at least one of the expert sub-models.

2. The method according to claim 1, characterized in that The random routing condition also includes: the value of the random routing flag of the hybrid expert model is true.

3. The method according to claim 1, characterized in that The determining a corresponding random routing matrix according to the original routing matrix includes: Generate a random tensor with the same dimension as the original routing matrix; The random routing matrix having the same mean and standard deviation as the original routing matrix is ​​generated according to the random tensor.

4. The method according to claim 1, characterized in that: The step of routing the data to be routed according to the random routing matrix includes: Mixing the random routing matrix and the original routing matrix to obtain a mixed routing matrix; The data to be routed is routed according to the hybrid routing matrix.

5. The method according to claim 4, characterized in that The step of mixing the random routing matrix and the original routing matrix to obtain a mixed routing matrix includes: Adding the product of the original routing matrix and the first mixing coefficient and the product of the random routing matrix and the second mixing coefficient to obtain the mixed routing matrix; The first mixing coefficient is the ratio of the current training step number to the preset step number; and the sum of the first mixing coefficient and the second mixing coefficient is 1.

6. The method according to claim 4, characterized in that Also includes: A routing loss function of the hybrid expert model is calculated according to the random routing matrix or the hybrid routing matrix.

7. A data routing device of a hybrid expert model, characterized in that: include: An original data determination unit, used for obtaining an original routing matrix of the hybrid expert model for the data to be routed; Each element in the original routing matrix is ​​respectively the original routing parameter of each expert sub-model in the hybrid expert model for the data to be routed; A condition judgment unit, used to judge whether the hybrid expert model meets the random routing condition; The random routing conditions include: the current training step number is not greater than the preset step number; A random data determination unit, configured to determine a corresponding random routing matrix according to the original routing matrix when the hybrid expert model satisfies the random routing condition; each element in the random routing matrix is ​​a random routing parameter corresponding to each of the original routing parameters; A data routing unit is used to route the data to be routed according to the random routing matrix when the hybrid expert model meets the random routing condition, so as to input the data to be routed into at least one of the expert sub-models.

8. The device according to claim 7, characterized in that The random routing condition also includes: the value of the random routing flag of the hybrid expert model is true.

9. The device according to claim 7, characterized in that The random data determination unit comprises: A random tensor generation unit, used to generate a random tensor with the same dimension as the original routing matrix; The random routing generation unit is used to generate the random routing matrix having the same mean and standard deviation as the original routing matrix according to the random tensor.

10. The device according to claim 7, characterized in that The data routing unit comprises: A routing matrix mixing unit, used for mixing the random routing matrix and the original routing matrix to obtain a mixed routing matrix; A first routing unit, configured to route the data to be routed according to the hybrid routing matrix when the hybrid expert model satisfies a random routing condition; The second routing unit is used to route the data to be routed according to the original routing matrix when the hybrid expert model does not meet the random routing condition.

11. The device according to claim 10, characterized in that The routing matrix mixing unit is used to mix the random routing matrix and the original routing matrix to obtain a mixed routing matrix, including: The routing matrix mixing unit is used to add the product of the original routing matrix and the first mixing coefficient and the product of the random routing matrix and the second mixing coefficient to obtain the mixed routing matrix; The first mixing coefficient is the ratio of the current training step number to the preset step number; and the sum of the first mixing coefficient and the second mixing coefficient is 1.

12. The device according to claim 10, characterized in that Also includes: The loss function calculation unit is used to calculate the routing loss function of the hybrid expert model according to the random routing matrix or the hybrid routing matrix.

13. An electronic device, characterized in that: The electronic device comprises: Memory for storing computer programs; A processor, configured to execute a computer program stored in the memory, and when the computer program is executed, implement the method described in any one of claims 1 to 6.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, which, when executed, implement the method described in any one of claims 1 to 6.

Citation Information

Cited By

  • Optimization method and device of hybrid expert system, computer equipment and readable storage medium

    CN120449952A