Efficient generation type task reasoning acceleration method based on hybrid expert network

By adopting a combination of hybrid expert network and multi-level gated network in the deep learning model, a multi-layer hybrid expert network architecture is formed, and only some experts are activated during inference is solved, which solves the problems of low computing resource efficiency and slow inference speed in the inference process, and achieves efficient inference acceleration effect.

CN120146184APending Publication Date: 2025-06-13NORTHEASTERN UNIV CHINA
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510207765.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the inference process, deep learning models face problems such as low computing resource efficiency, slow inference speed and computing bottlenecks, especially when processing large-scale data, which leads to high hardware resource consumption and increased inference latency.

Method used

Using an efficient generative task inference acceleration method based on hybrid expert network, a multi-layer hybrid expert network architecture is formed by constructing an n-layer Transformer decoder structure and a multi-level gated network, and only some experts are activated during inference to achieve sparse calculations.

Benefits of technology

It significantly reduces the amount of computing, improves the utilization efficiency of computing resources, effectively speeds up the inference process, avoids computing bottlenecks, and meets the needs of real-time and efficientness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146184A_ABST
    Figure CN120146184A_ABST
Patent Text Reader

Abstract

The efficient generative task reasoning acceleration method based on the hybrid expert network comprises the steps that context information of related generative tasks is obtained, context data is preprocessed, and it is ensured that the data is suitable for model input; constructing an efficient reasoning model based on the hybrid expert network, and determining a reasoning process of the efficient reasoning model based on the hybrid expert network; training an efficient reasoning model of the hybrid expert network, optimizing model parameters, and storing an optimal model structure; and generating a reasoning result by using the optimal model structure. In a model improved by the method, an expert network is composed of a plurality of independent experts, and each expert is responsible for processing different input characteristics. The gating network dynamically selects which experts participate in the calculation according to input characteristics, and assigns an expert to each input by calculating a probability distribution. In this way, the gating network and the expert network are closely matched, it is ensured that only the most suitable expert is used, and therefore the calculation efficiency and the model performance are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology and relates to an efficient generative task inference acceleration method based on a mixture of experts network. Background Art

[0002] With the continuous expansion of the scale of deep learning models, traditional fully connected neural networks face significant computational resource consumption problems when processing large-scale data. Expanding capacity is an obvious and effective way to improve performance. However, deeper models also bring greater inference costs, and the inference latency has a sublinear relationship with the model size. Under the same hardware conditions, the deeper the model, the slower the speed. Especially in tasks such as natural language processing and computer vision, the number of model parameters has increased exponentially, resulting in a significant increase in the computational burden during training and inference. This not only places higher requirements on hardware resources but also leads to slower inference speeds, unable to meet the requirements of real-time and efficiency. In addition, as the model scale expands, the computational bottleneck becomes more obvious. Especially under limited computational resources, it becomes particularly difficult to process huge model parameters and complex computational requirements. These problems not only increase the cost of model deployment but also limit its wide promotion in practical applications. Summary of the Invention

[0003] The purpose of the present invention is to provide an efficient generative task inference acceleration method based on a mixture of experts network, which solves the problems of low computational resource efficiency, slow inference speed, and computational bottleneck faced by deep learning models in the inference process in the prior art.

[0004] The present invention provides an efficient generative task inference acceleration method based on a mixture of experts network, including:

[0005] Step 1: Obtain the context information of relevant generative tasks, preprocess the context data to ensure that the data is suitable for model input;

[0006] Step 2: Construct an efficient inference model based on a mixture of experts network, and determine the inference process of the efficient inference model based on a mixture of experts network;

[0007] Step 3: Train the efficient inference model of the mixture of experts network, optimize the model parameters, and save the optimal model structure;

[0008] Step 4: Use the optimal model structure to generate inference results.

[0009] Further, the generative tasks include: text generation tasks, question answering tasks, text summarization tasks, machine translation tasks, and code generation tasks.

[0010] Further, the efficient inference model based on the mixture of experts network in step 2 includes: an n-layer Transformer decoder structure and a multi-level gating network; the n-layer Transformer decoder structure is an n-layer Decoder block, and the efficient inference model adopts a hierarchical architecture, which is divided into two parts: a basic layer and a mixture of experts layer;

[0011] The basic layer contains the first n - l layers of Decoder blocks, and the mixture of experts layer contains the last l layers of Decoder blocks and the multi-level gating network; the last l layers of Decoder blocks are divided into m mixture of experts sub-layers, and each mixture of experts sub-layer consists of l / m layers of Decoder blocks and a corresponding gating network; each mixture of experts sub-layer is implemented as an independent mixture of experts network layer, where each layer of Decoder block serves as an independent expert, with a total of l experts, and the experts in each mixture of experts sub-layer are in a parallel relationship, forming a multi-layer and multi-way parallel processing mechanism;

[0012] The multi-level gating network, corresponding to multiple mixture of experts sub-layers respectively, dynamically calculates the matching degree between the input token and the experts through learnable weights, selects the optimal expert path for each token, and realizes load balancing and calculation efficiency optimization; the output layer includes a layer normalization module, a linear transformation module, and a Softmax module.

[0013] Further, the Transformer decoder structure includes a self-attention sub-layer, a feed-forward neural network sub-layer, a residual connection, and layer normalization; the self-attention sub-layer re-represents the input sequence through the self-attention mechanism; the feed-forward neural network sub-layer further transforms the input sequence; the residual connection adds a direct connection to each sub-layer to facilitate information flow; and layer normalization normalizes the vector before output to ensure that the result is easy for subsequent processing.

[0014] Further, the gating network layer consists of an input layer, a hidden layer, and an output layer. The input layer comes from the feature vector output by the previous layer structure. The hidden layer contains multiple fully connected layers, which are used to extract the high-order features of the input and calculate the routing strategy. The output layer outputs a softmax probability distribution, and each expert corresponds to a probability value, indicating the possibility that the expert processes the input; the gating network outputs a sparse weight matrix to only activate some experts, thereby improving the inference efficiency while ensuring the accuracy.

[0015] Further, the inference process of the efficient inference model based on the mixture of experts network is specifically as follows:

[0016] (1) When a token is passed from the previous base layer to the first mixture-of-experts sub-layer, the token is fed into the gating network of the first mixture-of-experts sub-layer; the gating network determines which expert in the first mixture-of-experts sub-layer the token should be routed to according to the characteristics of the token;

[0017] (2) The Top-1 routing strategy is adopted, that is, only one expert is selected for routing each time, and the selected expert is the one with the highest score calculated by the gating network; this process also applies to the transfer of tokens between different mixture-of-experts sub-layers;

[0018] (3) The sparse calculation formula of the mixture-of-experts layer is:

[0019]

[0020] where y represents the output of the mixture-of-experts sub-layer, E(x) is the output of the expert network composed of l / m layer Decoder blocks, G(x) is the output of the gating network, N is the number of experts in the expert network, N = l / m; e i (x) is the output of the i-th expert, g i (x) is the output of the gating network at the i-th position. Both the expert network and the gating network are built based on fully connected neural networks and receive the same input;

[0021] (4) The output of the gating network is:

[0022] G(x) = softmax(TopK(g(x), k))

[0023] where the TopK(g(x), k) function only retains the first k terms of the original vector values and sets the other terms to -∞; after the subsequent softmax operation, all -∞ terms will become approximately zero; the hyperparameter k represents the number of experts to be routed, k = 1.

[0024] Furthermore, the loss function for training the efficient inference model based on the mixture-of-experts network in step 3 is the sum of the cross-entropy loss function and the expert balance loss function:

[0025] L s = L task + L bal

[0026] where L s is the total loss value of the training model, L task is the cross-entropy loss value of the generative task, L bal is the expert balance loss value;

[0027]

[0028] Among them, f i is the proportion of tokens assigned to expert i, P i is the proportion of router probabilities assigned to expert i, α is the hyperparameter, N is the number of experts in the expert network, N = l / m; p(x) represents the probability that the x-th token is routed to an expert, β is the batch, T represents the number of tokens in a batch, p i (x) represents the probability that the x-th token is routed to expert i; after training is completed, the optimal model structure will be saved for subsequent use in the validation set.

[0029] An efficient inference acceleration method for generative tasks based on a mixture of experts network has the following beneficial effects:

[0030] (1) Computational resource efficiency: Traditional deep learning models usually need to perform calculations on all input data through all layers, resulting in a high computational burden, especially when the model is large. The mixture of experts network processes tasks by assigning them to the most suitable experts instead of all experts. In the present invention, only one expert is routed each time, which is equivalent to only calculating one layer in several layers of the network, significantly reducing the amount of calculation and improving efficiency. This mechanism of allocating resources on demand enables the model to process large-scale data at a lower computational cost.

[0031] (2) Accelerating the inference process: Since each token is only routed to one most suitable expert instead of all experts participating in the calculation, the Top-1 routing strategy of the mixture of experts network can effectively accelerate the inference process, avoiding unnecessary calculations and delays. Especially when dealing with a large amount of data, the model can generate inference results more quickly, improving the performance of real-time applications.

[0032] (3) Avoiding computational bottlenecks caused by overly large models: For deep learning models with a large number of parameters, computational resources and memory consumption are often bottlenecks. The mixture of experts network avoids full-scale calculations by only activating a part of the experts, significantly reducing the computational burden of the model while ensuring the accuracy and effectiveness of the model.

[0033] (4) Flexibility and scalability of tasks. The task involved in the present invention is a generative task, which is applicable to a wide range of application scenarios, including various task types such as text summarization, machine translation, question answering systems, and dialogue generation. Due to the high flexibility and scalability of the present invention, it can be effectively applied to multiple tasks and flexibly adjusted according to different requirements to meet the specific requirements of different tasks. Whether in existing tasks or new tasks that may emerge in the future, the present invention can provide excellent performance through appropriate expansion and adjustment.

[0034] (5) By optimizing the model structure, the last l layers of the network are constructed into a multi-layer mixture-of-experts network architecture. Through a sparse activation strategy, only some experts are activated during inference, thus significantly reducing the computational burden, improving the utilization efficiency of computing resources, and effectively accelerating the inference process. This method can not only optimize the allocation of computing resources, but also maintain efficient inference, solve the computational bottleneck problem in large-scale model inference, and provide an efficient inference acceleration scheme for generative tasks. Description of the Drawings

[0035] Figure 1 is a flowchart of an efficient generative task inference acceleration method based on a mixture-of-experts network according to the present invention;

[0036] Figure 2 is a structural diagram of an efficient inference model based on a mixture-of-experts network according to the present invention. Detailed Embodiments

[0037] As Figure 1 shown, an efficient generative task inference acceleration method based on a mixture-of-experts network according to the present invention includes:

[0038] Step 1: Obtain the context information of relevant generative tasks, and preprocess the context data to ensure that the data is suitable for model input.

[0039] The method of the present invention is applicable to various generative tasks, has high flexibility and scalability, and can be effectively applied to a variety of task scenarios. There are many types of generative tasks, including: text generation tasks, question answering tasks, text summarization tasks, machine translation tasks, and code generation tasks, etc. Taking the dialogue generation task in text generation tasks as an example, the data of this task is selected and cleaned and preprocessed to ensure that the data format meets the model input requirements.

[0040] Step 2: Construct an efficient inference model based on a mixture-of-experts network, and determine the inference process of the efficient inference model based on a mixture-of-experts network.

[0041] As Figure 2 shown, the efficient inference model based on a mixture-of-experts network according to the present invention includes: an n-layer Transformer decoder structure and a multi-level gating network. The n-layer Transformer decoder structure is an n-layer Decoder block, such as Decoder block 1 to Decoder block n in the figure. The efficient inference model adopts a hierarchical architecture, which is divided into two parts: a basic layer and a mixture-of-experts layer as a whole.

[0042] The base layer contains the first n-l Decoder blocks, and the mixture-of-experts layer contains the last l Decoder blocks and a multi-level gating network; the last l Decoder blocks are divided into m mixture-of-experts sub-layers, and each mixture-of-experts sub-layer consists of l / m Decoder blocks and a corresponding gating network; each mixture-of-experts sub-layer is implemented as an independent mixture-of-experts network layer, where each Decoder block serves as an independent expert, with a total of l experts. The experts within each mixture-of-experts sub-layer are in a parallel relationship, forming a multi-layer and multi-path parallel processing mechanism;

[0043] The multi-level gating network, corresponding to multiple mixture-of-experts sub-layers respectively, dynamically calculates the matching degree between the input tokens and the experts through learnable weights, selects the optimal expert path for each token, and realizes load balancing and optimization of computing efficiency.

[0044] As Figure 2 shown, the Transformer decoder structure includes a self-attention sub-layer, a first residual connection and layer normalization, a feed-forward neural network sub-layer, and a second residual connection and layer normalization connected in sequence. The self-attention sub-layer re-represents the input sequence through the self-attention mechanism; the feed-forward neural network sub-layer further transforms the input sequence; the residual connection adds a direct connection to each sub-layer to facilitate information flow; and the layer normalization normalizes the vector before output to ensure that the result is easy to process subsequently.

[0045] The gating network layer consists of an input layer, a hidden layer, and an output layer. The input layer comes from the feature vector output by the previous layer structure. The hidden layer contains multiple fully connected layers, which are used to extract the high-order features of the input and calculate the routing strategy. The output layer outputs a softmax probability distribution, with each expert corresponding to a probability value, indicating the possibility of that expert processing the input; the gating network outputs a sparse weight matrix to only activate some experts, thereby improving the inference efficiency while ensuring the accuracy.

[0046] The output layer includes a layer normalization module, a linear transformation module, and a Softmax module.

[0047] In this model, the expert network consists of multiple independent experts, and each expert is responsible for processing different input features. The gating network dynamically selects which experts participate in the calculation according to the input features, and assigns experts to each input by calculating a probability distribution. In this way, the gating network and the expert network cooperate closely to ensure that only the most suitable experts are used, thereby improving the computing efficiency and the model performance.

[0048] During specific implementation, the inference process of the efficient inference model based on the mixture-of-experts network of the present invention is specifically as follows:

[0049] (1) When a token is passed from the previous base layer to the first mixture-of-experts sub-layer, the token is fed into the gating network of the first mixture-of-experts sub-layer; based on the characteristics of the token, the gating network determines which expert in the first mixture-of-experts sub-layer the token should be routed to.

[0050] (2) The Top-1 routing strategy is adopted, that is, only one expert is selected for routing each time, and the selected expert is the one with the highest score calculated by the gating network; this process also applies to the transfer of tokens between different mixture-of-experts sub-layers.

[0051] (3) The sparse calculation formula of the mixture-of-experts layer is:

[0052]

[0053] where y represents the output of the mixture-of-experts sub-layer, E(x) is the output of the expert network composed of l / m layer Decoder blocks, G(x) is the output of the gating network, N is the number of experts in the expert network, N = l / m; e i (x) is the output of the i-th expert, g i (x) is the output of the gating network at the i-th position. Both the expert network and the gating network are built based on fully connected neural networks and receive the same input.

[0054] (4) The output of the gating network is:

[0055] G(x) = softmax(TopK ( g(x), k ) )

[0056] where the TopK(g(x), k) function only retains the first k terms of the original vector values and sets the other terms to -∞; after the subsequent softmax operation, all -∞ terms will become approximately zero; the hyperparameter k represents the number of experts to be routed, k = 1.

[0057] Step 3: Train the efficient inference model of the mixture-of-experts network, optimize the model parameters, and save the optimal model structure.

[0058] Before applying the efficient inference model of the mixture-of-experts network, the model needs to be trained first so that the model can adapt to the inference mode of the mixture-of-experts and ensure that it can effectively perform specific tasks.

[0059] Specifically, when implementing, the loss function for training the efficient inference model based on the mixture-of-experts network is the sum of the cross-entropy loss function and the expert balance loss function:

[0060] L s = L task+L bal

[0061] where L s is the total loss value of the training model, L task is the cross-entropy loss value of the generative task, and L bal is the expert balance loss value;

[0062]

[0063] where f i is the proportion of tokens assigned to expert i, P i is the proportion of the router probability assigned to expert i, α is a hyperparameter, N is the number of experts in the expert network, N = l / m; p(x) represents the probability that the x-th token is routed to an expert, β is the batch, T represents the number of tokens in a batch, and p i (x) represents the probability that the x-th token is routed to expert i; After the training is completed, the optimal model structure will be saved for subsequent use in the validation set.

[0064] Step 4: Generate inference results using the optimal model structure.

[0065] When generating inference results using the optimal model structure, first load the optimal model saved during the training process. After the input data is passed into the model, through the processing of multiple mixture-of-experts network layers and gating networks, the model dynamically determines which experts each token should be routed to based on the input features and generates the corresponding output. In this way, the model can combine the advantages of each expert to produce efficient and accurate inference results.

[0066] The present invention provides a method for accelerating inference of generative tasks based on a mixture-of-experts network, aiming to solve problems such as waste of computing resources, slow inference speed, and computational bottlenecks existing in traditional deep learning models during the inference process. Through an innovative design of the model structure, several network layers are combined into a mixture-of-experts network layer, and a gating network is introduced to form a multi-layer mixture-of-experts network architecture. At the same time, a sparse activation strategy is adopted to activate only some experts during inference, significantly reducing the computational burden, improving the utilization efficiency of computing resources, and effectively accelerating the inference process. This method not only solves the bottleneck problem in large model inference but also optimizes the allocation of computing resources, maintains an efficient inference speed, provides a feasible acceleration scheme for generative tasks, has broad application potential, and performs excellently especially in practical scenarios with high requirements for real-time performance and computational performance.

[0067] The above are only the preferred embodiments of the present invention and are not intended to limit the idea of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An efficient generative task reasoning acceleration method based on a hybrid expert network, characterized in that: include: Step 1: Obtain context information of relevant generative tasks and preprocess the context data to ensure that the data is suitable for model input; Step 2: Construct an efficient reasoning model based on a hybrid expert network and determine the reasoning process of the efficient reasoning model based on a hybrid expert network; Step 3: Train the efficient reasoning model of the hybrid expert network, optimize the model parameters, and save the optimal model structure; Step 4: Generate inference results using the optimal model structure.

2. The efficient generative task reasoning acceleration method based on hybrid expert network as claimed in claim 1, characterized in that: The generative tasks include: text generation tasks, question answering tasks, text summarization tasks, machine translation tasks and code generation tasks.

3. The efficient generative task reasoning acceleration method based on hybrid expert network as claimed in claim 1, characterized in that: The efficient reasoning model based on the hybrid expert network in step 2 includes: an n-layer Transformer decoder structure, a multi-level gating network, and an output layer; the n-layer Transformer decoder structure is an n-layer Decoder block, and the efficient reasoning model adopts a layered architecture, which is divided into two parts: a basic layer and a hybrid expert layer; The base layer contains the first nl layers of Decoder blocks, and the hybrid expert layer contains the last l layers of Decoder blocks and a multi-level gating network; the last l layers of Decoder blocks are divided into m hybrid expert sub-layers, each of which consists of l / m layers of Decoder blocks and a corresponding gating network; each hybrid expert sub-layer is implemented as an independent hybrid expert network layer, in which each layer of Decoder blocks acts as an independent expert, containing a total of l experts, and the experts in each hybrid expert sub-layer are in parallel, forming a multi-layer and multi-path parallel processing mechanism; The multi-level gating network corresponds to multiple hybrid expert sub-layers, dynamically calculates the matching degree between the input token and the expert through learnable weights, selects the optimal expert path for each token, and realizes load balancing and computing efficiency optimization; The output layer includes a layer normalization module, a linear change module, and a Softmax module.

4. The efficient generative task reasoning acceleration method based on hybrid expert network as claimed in claim 3, characterized in that: The Transformer decoder structure includes a self-attention sublayer, a feedforward neural network sublayer, a residual connection and layer normalization; the self-attention sublayer re-represents the input sequence through the self-attention mechanism; the feedforward neural network sublayer further transforms the input sequence; the residual connection adds a direct connection to each sublayer to promote information flow; and the layer normalization normalizes the vector before output to ensure that the result is easy to process later.

5. The efficient generative task reasoning acceleration method based on hybrid expert network as claimed in claim 3, characterized in that: The gated network layer consists of an input layer, a hidden layer and an output layer. The input layer comes from the feature vector output by the previous layer structure. The hidden layer contains multiple fully connected layers for extracting high-order features of the input and calculating the routing strategy. The output layer outputs a softmax probability distribution. Each expert corresponds to a probability value, indicating the possibility of the expert processing the input. The gating network outputs a sparse weight matrix to activate only some experts, thereby improving inference efficiency while ensuring accuracy.

6. The efficient generative task reasoning acceleration method based on hybrid expert network as claimed in claim 5, characterized in that: The reasoning process of the efficient reasoning model based on the hybrid expert network is as follows: (1) When a token is passed from the previous base layer to the first hybrid expert sublayer, the token is sent to the gating network of the first hybrid expert sublayer; the gating network determines which expert in the first hybrid expert sublayer the token should be routed to based on the characteristics of the token; (2) Adopting the Top-1 routing strategy, that is, only one expert is selected for routing each time. The selected expert is the expert with the highest score calculated based on the gating network. This process is also applicable to the transfer of tokens between different hybrid expert sub-layers. (3) The sparse calculation formula of the hybrid expert layer is: Where y represents the output of the hybrid expert sublayer, E(x) is the output of the expert network composed of l / m layers of Decoder blocks, G(x) is the output of the gating network, and N is the number of experts in the expert network, N = l / m; e i (x) is the output of the i-th expert, g i (x) is the output of the i-th position of the gating network. Both the expert network and the gating network are built based on a fully connected neural network and receive the same input; (4) The output of the gating network is: G(x)=softmax(TopK ( g(x), k ) ) The TopK(g(x),k) function retains only the first k items of the original value of the vector and sets the other items to -∞. In the subsequent softmax operation, all -∞ items will become approximately zero. The hyperparameter k represents the number of experts to be routed, k=1.

7. The efficient generative task reasoning acceleration method based on hybrid expert network as claimed in claim 1, characterized in that: The loss function for training the efficient reasoning model based on the hybrid expert network in step 3 is the sum of the cross entropy loss function and the expert balance loss function: L s =L task +L bal Among them, L s is the total loss value of the training model, L task is the cross entropy loss value of the generative task, L bal is the expert equilibrium loss value; Among them, f i is the proportion of tokens allocated to expert i, P i is the ratio of the router probability assigned to expert i, α is a hyperparameter, N is the number of experts in the expert network, N = l / m; p(x) represents the probability that the xth token is routed to an expert, β is the batch, T represents the number of tokens in a batch, p i (x) represents the probability that the xth token is routed to expert i. After the training is completed, the optimal model structure will be saved for use in subsequent validation sets.

Citation Information

Cited By

  • Data processing method and system of low-energy-consumption large language model based on momentum mechanism and multiple types of experts

    CN120450054A

  • Distributed reasoning method and device based on hybrid expert architecture, equipment and medium

    CN120579646A

  • Text common sense reasoning method based on dynamic top-k selection expert model

    CN120875042A

  • Task action generation method and device, robot, electronic equipment and medium

    CN121468541A

  • Text error correction method and device based on multi-expert mixing mechanism, equipment and medium

    CN121659938A