Expert access prediction method and system suitable for hybrid expert architecture large language model
By using neural network predictors to predict expert access patterns in a hybrid expert architecture large language model, combined with active cache update and prefetch priority design, the problems of GPU memory limitation and limited PCIe bandwidth are solved, and the model inference speed and hardware resource utilization are improved.
Patent Information
- Application Number
- CN202510101866.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-22
AI Technical Summary
When the hybrid expert architecture large language model runs on personal computers, due to GPU memory limitations and PCIe bandwidth, the performance deteriorates during expert offloading and prefetching, and existing prediction and prefetching methods are difficult to meet the requirements in terms of accuracy and efficiency.
Using a neural network-based expert access predictor, a two-layer perceptron predictor is built by obtaining model structure and hardware feature information, to predict the experts that need to be activated on each layer, and to optimize the expert parameter loading and prefetching process through active cache update and prefetching priority design.
The performance and hardware resource utilization rate in the inference process of large language models are improved, and the prefetching effect is maximized by adjusting the prediction distance and dynamically adjusting the prefetching quantity, and the model inference speed is improved.
Smart Images

Figure CN120012935A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language model prediction, and in particular to an expert access prediction method and system suitable for a large language model with a hybrid expert architecture. Background Art
[0002] With the widespread application of large language models (LLMs) in natural language processing, content generation, and decision support, there is a growing demand to run LLMs on consumer-grade platforms for better privacy protection and responsiveness. However, since the huge memory requirements of LLMs often exceed the capacity of consumer-grade GPUs, this memory limitation seriously affects the performance and application of LLMs on personal computers.
[0003] Mixture-of-Experts (MoE) architectures provide an opportunity to address GPU memory limitations by partitioning the model into multiple experts and activating only a subset of them during inference. This allows most of the expert parameters to be offloaded to the host memory, and only the necessary experts are loaded into the GPU memory when needed. While this significantly reduces GPU memory requirements, expert offloading also brings severe performance degradation due to the limited PCIe bandwidth between host memory and GPU.
[0004] To address this problem, existing systems propose to cache frequently accessed expert parameters in GPU memory to minimize offloading overhead. However, the cache handles cache-missed experts in a passive manner. When the inference process encounters an expert missing from GPU memory, the calculation is blocked until the expert is obtained from the host memory. This leaves a high latency overhead on the critical path of inference, which has a significant impact on performance and utilization.
[0005] Another optimization method is to predict which experts will be used next and pre-fetch them in advance. In theory, this method can significantly improve efficiency because it can overlap the loading of expert parameters with the calculation process, thereby hiding the delay of data transmission. However, since the accuracy and advance amount of the prediction are difficult to grasp, this method often does not work well in practice. Too low prediction accuracy will lead to a large number of invalid pre-fetches, which not only wastes bandwidth resources, but may also interfere with the loading of experts that are really needed; and too small a prediction advance amount will result in insufficient time available for pre-fetching, making it impossible to complete pre-fetching before the expert is accessed. It is worth noting that these problems manifest themselves differently in different model architectures and application scenarios.
[0006] Therefore, we hope to have a solution that can accurately predict the experts to be visited and reasonably arrange the pre-fetching time, and can adapt to different model structures and hardware platforms. Specifically, we need an expert visit prediction mechanism for the hybrid expert architecture, which can effectively combine the model structure and runtime characteristics to improve the prediction accuracy, and dynamically adjust the pre-fetching strategy according to the hardware characteristics to maximize the pre-fetching effect and improve the model inference speed.
[0007] Patent application document CN118761472A discloses a hybrid expert model reasoning acceleration method, device, equipment, medium and program, wherein the method includes: loading target association parameters of the hybrid expert model; wherein the target association parameters include non-expert network parameters and expert network reference vectors of the hybrid expert model; predicting the target expert network of the network layer to be run according to the current layer output feature vector of the current running network layer of the hybrid expert model and the expert network reference vector of the network layer to be run; preloading the target expert network of the network layer to be run to the device storage running the hybrid expert model to perform the model reasoning process. However, this patent cannot completely solve the current technical problems, nor can it meet the needs of the present invention. Summary of the invention
[0008] In view of the defects in the prior art, the object of the present invention is to provide an expert interview prediction method and system suitable for a large language model of a hybrid expert architecture.
[0009] The expert interview prediction method applicable to the hybrid expert architecture large language model provided by the present invention includes:
[0010] Step 1: Obtain the structural information of the large language model of the hybrid expert architecture and the hardware characteristics of the computing platform, including the number of model layers, the number of experts in each layer, the number of experts that need to be activated for each word in each layer, the number of parameters of each expert, the PCIe link bandwidth, and the available GPU memory size;
[0011] Step 2: Build a neural network-based expert interview predictor P for each layer of the large language model i , the predictor is responsible for predicting the experts that need to be activated in the current layer based on the input before k layers, where P i is the predictor of the i-th layer, which consists of a two-layer perceptron with input h i-k is the input of the ikth layer of the large language model, and the output is e i is the probability of each expert in the i-th layer being activated, and k is the prediction distance pre-specified by the user;
[0012] Step 3: Collect historical data from the large language model inference process as training data. For each word, perform an inference process and record the input of each layer of the gating function and the actual selected expert.
[0013] Step 4: Use the collected training data to train the expert access predictors of each layer offline and persist them to disk;
[0014] Step 5: Start the inference service, load the non-expert parameters into the GPU memory, load the expert parameters into the CPU memory, and load the corresponding predictors for each layer of the large language model into the CPU memory;
[0015] Step 6: During online inference, the gate function input of each layer in the GPU inference process is collected into the CPU memory. After obtaining the input, the predictor corresponding to the collected input will calculate e on the CPU. i =P i (h i-k ), output the probability of each expert being used after k layers as the prediction result;
[0016] Step 7: According to the prediction results, select the same number of experts as the number of experts activated in the model for each layer, put them into the prefetch queue, and hand them over to the prefetcher to perform the prefetch operation. The prefetcher loads the expert parameters from the host memory to the GPU memory in the order of entry;
[0017] Step 8: When the reasoning reaches the predicted layer, the experts that were not pre-fetched due to prediction errors are put back into the pre-fetch queue and handed over to the pre-fetcher for pre-fetching;
[0018] Step 9: After the reasoning of a word is completed, the reasoning result is returned to the user, and the process returns to step 6 to start the reasoning of the next word.
[0019] Preferably, in step 4, the expert access predictors of each layer are trained offline, specifically: for each input in the training data, the output of the predictor is compared with the actual output, the loss function is calculated, and the predictor parameters are updated using back propagation combined with the gradient descent method, and the training data is traversed and the above process is executed repeatedly until the predictor converges.
[0020] Preferably, in step 6, the acquisition process and the prediction process will be performed on the CPU in parallel with the GPU inference process.
[0021] Preferably, in step 7, the expert parameter loading process is executed in parallel with the GPU inference process, and the number of pre-fetched experts is dynamically adjusted according to the predictor accuracy.
[0022] Preferably, in step 8, the re-prefetch caused by the prediction error enters the prefetch queue with high priority, and the unfinished erroneous prefetch tasks in the queue will be cleared.
[0023] The expert access prediction system applicable to the hybrid expert architecture large language model provided by the present invention includes:
[0024] Module M1: Obtain the structural information of the large language model of the hybrid expert architecture and the hardware characteristics of the computing platform, including the number of model layers, the number of experts in each layer, the number of experts that need to be activated for each word in each layer, the number of parameters of each expert, the PCIe link bandwidth, and the available GPU memory size;
[0025] Module M2: Build a neural network-based expert interview predictor P for each layer of the large language model i , the predictor is responsible for predicting the experts that need to be activated in the current layer based on the input before k layers, where P i is the predictor of the i-th layer, which consists of a two-layer perceptron with input h i-k is the input of the ikth layer of the large language model, and the output is e i is the probability of each expert in the i-th layer being activated, and k is the prediction distance pre-specified by the user;
[0026] Module M3: Collect historical data from the large language model inference process as training data. For each word, perform an inference process and record the input of each layer of the gating function and the actual selected expert.
[0027] Module M4: Use the collected training data to train the expert access predictors of each layer offline and persist them to disk;
[0028] Module M5: Start the inference service, load the non-expert parameters into the GPU memory, load the expert parameters into the CPU memory, and load the corresponding predictor for each layer of the large language model into the CPU memory;
[0029] Module M6: During online inference, the gate function input of each layer in the GPU inference process is collected into the CPU memory. After the predictor corresponding to the collected input obtains the input, it will calculate e on the CPU. i =P i (h i-k ), output the probability of each expert being used after k layers as the prediction result;
[0030] Module M7: According to the prediction results, select experts for each layer with the same number of experts as the model activation experts, put them into the pre-fetch queue, and hand them over to the pre-fetcher for pre-fetching. The pre-fetcher loads the expert parameters from the host memory to the GPU memory in the order of entry;
[0031] Module M8: When the reasoning reaches the predicted layer, the experts that were not pre-fetched due to prediction errors are put back into the pre-fetch queue and handed over to the pre-fetcher for pre-fetching.
[0032] Module M9: After the reasoning of a word is completed, the reasoning result is returned to the user and module M6 is triggered to start the reasoning of the next word.
[0033] Preferably, in module M4, the expert access predictors of each layer are trained offline, specifically: for each input in the training data, the output of the predictor is compared with the actual output, the loss function is calculated, and the predictor parameters are updated using back propagation combined with the gradient descent method, and the training data is traversed and the above process is executed repeatedly until the predictor converges.
[0034] Preferably, in module M6, the acquisition process and the prediction process will be performed on the CPU and executed in parallel with the GPU inference process.
[0035] Preferably, in module M7, the expert parameter loading process is executed in parallel with the GPU inference process, and the number of pre-fetched experts is dynamically adjusted according to the predictor accuracy.
[0036] Preferably, in module M8, the re-prefetching caused by the prediction error enters the prefetching queue with high priority, and the unfinished erroneous prefetching tasks in the queue will be cleared.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] (1) The present invention adopts active cache update, so that the large model reasoning process and the expert parameter loading process can be executed in parallel, reducing the parameter loading delay on the critical path and improving system performance and hardware resource utilization;
[0039] (2) The present invention can adjust the prediction strategy by modifying the size of the prediction distance k, thereby achieving an optimal balance between the prediction accuracy and the prefetching advance amount, thereby maximizing the prefetching effect;
[0040] (3) The present invention accurately predicts the expert access pattern through a neural network predictor. The neural network predictor can learn the expert access pattern from historical information, thereby maintaining a high prediction accuracy rate under long prediction distances and effectively improving the loading efficiency of expert parameters;
[0041] (4) The present invention is versatile and extensible, and can adapt to different model structures, hardware platforms, and application scenarios. It can improve the generalization ability of the predictor by collecting training data in actual application scenarios.
[0042] (5) The present invention adopts a pre-fetch priority design, which enables the system to quickly switch to extracting the correct expert, effectively reducing the performance loss caused by prediction errors. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0044] Figure 1 The expert interview prediction method of the present invention is applicable to a large language model of a hybrid expert architecture;
[0045] Figure 2 A schematic diagram of an embodiment scenario of this specification. DETAILED DESCRIPTION
[0046] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0047] Example 1
[0048] like Figure 1 The present invention provides an expert visit prediction method for a large language model with a hybrid expert architecture. The method can not only accurately predict the experts to be visited and reasonably arrange the pre-fetching time, but also flexibly adapt to different model structures and hardware platforms. The specific implementation process is:
[0049] Step 1: Obtain the structural information of the large language model of the hybrid expert architecture and the hardware feature information of the computing platform. This information includes but is not limited to the following important aspects: the total number of layers of the model, the number of experts contained in each layer, the number of experts that need to be activated for each input word in each layer, the total number of parameters contained in each expert, and the PCIe link bandwidth of the computing platform, the available GPU memory size and other key information.
[0050] Step 2: Build a neural network-based expert visit predictor P for each layer of the model i These predictors will predict the experts that need to be activated in the current layer based on the input information of the previous k layers. i is the predictor of the i-th layer, which consists of a two-layer perceptron with input h i-k is the input of the ikth layer of the large language model, and the output is e i is the probability of each expert in the i-th layer being activated; k is the prediction distance, and its specific value is pre-specified by the user according to actual needs.
[0051] Step 3: Collect the intermediate result data generated by the model during the reasoning process as the training data set. Specifically, for the reasoning process of each input word, record the input of the gate function in each layer and the expert that is actually activated in the end.
[0052] Step 4: Use the collected training data to train the expert access predictors of each layer offline and persist them to disk. Specifically, for each input in the training data, compare the output of the predictor with the actual output, calculate the loss function, and use backpropagation with the gradient descent method to update the predictor parameters. Repeat the training data and execute the above process until the predictor converges.
[0053] The loss function is a way to measure the gap between the model's predicted output and the actual label. Different loss functions may be selected for different tasks. For example:
[0054] Regression problem: The commonly used loss function is Mean Squared Error (MSE), and its formula is:
[0055]
[0056] Where N is the number of samples, y i is the true value, is the predicted value.
[0057] Classification problem: The commonly used loss function is cross-entropy loss. For the two-classification problem, the formula is:
[0058]
[0059] For multi-classification problems, Softmax combined with cross entropy loss is used.
[0060] Once the loss function is calculated, we can use the backpropagation algorithm to calculate the partial derivative of each parameter with respect to the loss function, i.e. the gradient. Then, based on this gradient information, the model parameters are adjusted by the gradient descent method to minimize the loss function. The update rule of gradient descent can be expressed as:
[0061]
[0062] Among them, θ j represents the jth parameter, α is the learning rate, which determines the size of the step of parameter update.
[0063] The steps of the back propagation algorithm are as follows:
[0064] Forward propagation: pass the input data through the network and calculate the predicted value;
[0065] Calculate loss: Calculate the difference between the predicted value and the true value using the loss function mentioned above;
[0066] Back propagation: Starting from the last layer, the error term (delta) of each node is calculated forward layer by layer, which reflects the contribution of the output of the node to the total error;
[0067] Calculate gradients: Use the error term to calculate the gradient of each weight;
[0068] Update weights: Update the weights in the network based on the calculated gradient and the pre-set learning rate.
[0069] The above process constitutes a complete training iteration (epoch), which usually needs to be repeated multiple times until the loss function converges to a satisfactory level or reaches the preset maximum number of iterations.
[0070] Step 5: Prepare to start the inference service, load the non-expert parameters into the GPU memory, and load the expert parameters into the CPU memory. At the same time, load the corresponding predictor for each layer into the CPU memory.
[0071] Step 6: During online inference, the gate function input of each layer in the GPU inference process is collected into the CPU memory. After receiving the input, the predictor corresponding to the collected input will calculate e on the CPU. i =P i (h i-k ), output the probability of each expert being used after the k-th layer as the prediction result.
[0072] Step 7: According to the prediction results, select the same number of experts as the model activated experts for each layer, put them into the prefetch queue, and hand them over to the prefetcher for prefetching. The prefetcher loads the expert parameters from the host memory to the GPU memory in the order of entry.
[0073] Step 8: When the reasoning reaches the predicted layer, the experts that were not pre-fetched due to prediction errors are put back into the pre-fetch queue and handed over to the pre-fetcher to perform the pre-fetch operation.
[0074] Step 9: After the reasoning of a word is completed, the reasoning result is returned to the user, and the process returns to step 6 to start the reasoning of the next word.
[0075] In step 1, the structural information of the large language model of the hybrid expert architecture and the hardware feature information of the computing platform are obtained. The model has 4 layers, each layer contains 3 experts, and each word needs to activate 2 experts. In this hardware platform, the memory space of the GPU allows 2 experts to be cached in each layer.
[0076] In step 2, a neural network-based expert predictor is built for each layer. For example, if the user pre-specifies the prediction distance 2, one of the predictors will predict the expert that needs to be activated in layer 2 based on the intermediate results of layer 0.
[0077] In step 3, historical data from the model inference process is collected as training data. For each word inference process, the input of each layer of the gating function and the actual selected expert are recorded.
[0078] In step 4, the expert predictors of each layer are trained offline using the collected training data. For example, the gated input of layer 0 and the expert activation of layer 2 are used as the input and output of the predictor that uses layer 0 to predict layer 2, respectively, and the predictor is trained until convergence.
[0079] In step 5, prepare to start the inference service, load the non-expert parameters of each layer into the GPU memory, and load the expert parameters, such as the experts E1, E2, and E3 of the second layer, into the CPU memory. At the same time, load the predictor, such as the predictor that uses the 0th layer to predict the 2nd layer, into the GPU memory. Figure 2 .
[0080] In step 6, during online inference, when the inference reaches the gate function of layer 0, the gate function input is collected into the CPU memory. The input will be sent to the predictor that uses layer 0 to predict layer 2 for prediction, and the prediction result is that layer 2 will activate E1 and E3.
[0081] In step 7, according to the prediction results, the prefetch tasks of E1 and E3 predicted to be activated in the second layer are placed in the prefetch queue. The prefetcher loads the expert parameters of E1 and E3 from the host memory to the expert cache in the GPU video memory in the order of enqueuing.
[0082] In step eight, the inference process is performed synchronously with the above loading process. When the inference reaches the second level, the experts actually activated are E1 and E2, of which E1 has been pre-fetched into the expert cache in the GPU, while E2 has not been loaded due to a prediction error. E2 will be placed in the pre-fetch queue as a high-priority task and will be pre-fetched by the pre-fetcher first.
[0083] In step nine, after E2 is loaded, the large language model will complete the remaining inference calculations, the inference results will be returned to the user, and the system will start inference for the next word.
[0084] Example 2
[0085] The present invention also provides an expert visit prediction system applicable to a large language model with a hybrid expert architecture. The expert visit prediction system applicable to a large language model with a hybrid expert architecture can be implemented by executing the process steps of the expert visit prediction method applicable to a large language model with a hybrid expert architecture, that is, a person skilled in the art can understand the expert visit prediction method applicable to a large language model with a hybrid expert architecture as a preferred implementation of the expert visit prediction system applicable to a large language model with a hybrid expert architecture.
[0086] The expert access prediction system for a large language model with a hybrid expert architecture provided by the present invention comprises: module M1: obtaining structural information of the large language model with a hybrid expert architecture and hardware feature information of a computing platform, including the number of model layers, the number of experts in each layer, the number of experts to be activated for each word in each layer, the number of parameters of each expert, the PCIe link bandwidth and the available GPU memory size; module M2: constructing a neural network-based expert access predictor P for each layer of the large language model i , the predictor is responsible for predicting the experts that need to be activated in the current layer based on the input before k layers, where P i is the predictor of the i-th layer, which consists of a two-layer perceptron with input h i-k is the input of the ikth layer of the large language model, and the output is e i is the probability of each expert in the i-th layer being activated, and k is the prediction distance specified in advance by the user; Module M3: collects historical data from the large language model inference process as training data, and records the input of each layer of the gating function and the actual selected expert for each word in the inference process; Module M4: uses the collected training data to train the expert access predictors of each layer offline and persists them to the disk; Module M5: starts the inference service, loads the non-expert parameters into the GPU memory, loads the expert parameters into the CPU memory, and loads the corresponding predictor for each layer of the large language model into the CPU memory; Module M6: during online inference, collects the gating function input of each layer in the GPU inference process into the CPU memory, and the predictor corresponding to the collected input will calculate e on the CPU after obtaining the input i =P i (h i-k), output the probability of each expert being used after k layers as the prediction result; Module M7: According to the prediction result, select experts with the same number as the model activated experts for each layer, put them into the prefetch queue, and hand them over to the prefetcher to perform the prefetch operation. The prefetcher loads the expert parameters from the host memory to the GPU memory in the order of entry; Module M8: When the reasoning reaches the predicted layer, the experts that were not prefetched due to prediction errors are put back into the prefetch queue and handed over to the prefetcher to perform the prefetch operation; Module M9: After the reasoning of a word is completed, the reasoning result is returned to the user, and module M6 is triggered to start the reasoning of the next word.
[0087] In module M4, the expert access predictors of each layer are trained offline. Specifically, for each input in the training data, the output of the predictor is compared with the actual output, the loss function is calculated, and the predictor parameters are updated using back propagation combined with the gradient descent method. The training data is traversed and the above process is executed repeatedly until the predictor converges.
[0088] In module M6, the acquisition and prediction processes will be performed on the CPU in parallel with the GPU inference process.
[0089] In module M7, the expert parameter loading process is executed in parallel with the GPU inference process, and the number of pre-fetched experts is dynamically adjusted according to the predictor accuracy.
[0090] In module M8, the re-prefetch caused by the prediction error enters the prefetch queue with high priority, and the unfinished erroneous prefetch tasks in the queue will be cleared.
[0091] Those skilled in the art know that, in addition to implementing the system, device and its various modules provided by the present invention in a purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Therefore, the system, device and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be considered as structures within the hardware component; the modules for implementing various functions can also be considered as both software programs for implementing the method and structures within the hardware component.
[0092] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. An expert interview prediction method suitable for a large language model with a hybrid expert architecture, characterized in that: include: Step 1: Obtain the structural information of the large language model of the hybrid expert architecture and the hardware characteristics of the computing platform, including the number of model layers, the number of experts in each layer, the number of experts that need to be activated for each word in each layer, the number of parameters of each expert, the PCIe link bandwidth, and the available GPU memory size; Step 2: Build a neural network-based expert interview predictor P for each layer of the large language model i , the predictor is responsible for predicting the experts that need to be activated in the current layer based on the input before k layers, where P i is the predictor of the i-th layer, which consists of a two-layer perceptron with input h i-k is the input of the ikth layer of the large language model, and the output is e i is the probability of each expert in the i-th layer being activated, and k is the prediction distance pre-specified by the user; Step 3: Collect historical data from the large language model inference process as training data. For each word, perform an inference process and record the input of each layer of the gating function and the actual selected expert. Step 4: Use the collected training data to train the expert access predictors of each layer offline and persist them to disk; Step 5: Start the inference service, load the non-expert parameters into the GPU memory, load the expert parameters into the CPU memory, and load the corresponding predictors for each layer of the large language model into the CPU memory; Step 6: During online inference, the gate function input of each layer in the GPU inference process is collected into the CPU memory. After obtaining the input, the predictor corresponding to the collected input will calculate e on the CPU. i =P i (h i-k ), output the probability of each expert being used after k layers as the prediction result; Step 7: According to the prediction results, select the same number of experts as the number of experts activated in the model for each layer, put them into the prefetch queue, and hand them over to the prefetcher to perform the prefetch operation. The prefetcher loads the expert parameters from the host memory to the GPU memory in the order of entry; Step 8: When the reasoning reaches the predicted layer, the experts that were not pre-fetched due to prediction errors are put back into the pre-fetch queue and handed over to the pre-fetcher for pre-fetching; Step 9: After the reasoning of a word is completed, the reasoning result is returned to the user, and the process returns to step 6 to start the reasoning of the next word.
2. The expert interview prediction method applicable to the hybrid expert architecture large language model according to claim 1, characterized in that: In step 4, the expert access predictors of each layer are trained offline. Specifically, for each input in the training data, the output of the predictor is compared with the actual output, the loss function is calculated, and the predictor parameters are updated using back propagation combined with the gradient descent method. The training data is traversed and the above process is executed repeatedly until the predictor converges.
3. The expert interview prediction method applicable to the hybrid expert architecture large language model according to claim 1, characterized in that: In step 6, the acquisition and prediction processes will be performed on the CPU in parallel with the GPU inference process.
4. The expert interview prediction method applicable to the hybrid expert architecture large language model according to claim 1, characterized in that: In step 7, the expert parameter loading process is executed in parallel with the GPU inference process, and the number of pre-fetched experts is dynamically adjusted according to the predictor accuracy.
5. The expert interview prediction method applicable to the hybrid expert architecture large language model according to claim 1, characterized in that: In step 8, the re-prefetch caused by the prediction error enters the prefetch queue with high priority, and the unfinished erroneous prefetch tasks in the queue are cleared.
6. An expert interview prediction system suitable for a large language model with a hybrid expert architecture, characterized in that: include: Module M1: Obtain the structural information of the large language model of the hybrid expert architecture and the hardware characteristics of the computing platform, including the number of model layers, the number of experts in each layer, the number of experts that need to be activated for each word in each layer, the number of parameters of each expert, the PCIe link bandwidth, and the available GPU memory size; Module M2: Build a neural network-based expert interview predictor P for each layer of the large language model i , the predictor is responsible for predicting the experts that need to be activated in the current layer based on the input before k layers, where P i is the predictor of the i-th layer, which consists of a two-layer perceptron with input h i-k is the input of the ikth layer of the large language model, and the output is e i is the probability of each expert in the i-th layer being activated, and k is the prediction distance pre-specified by the user; Module M3: Collect historical data from the large language model inference process as training data. For each word, perform an inference process and record the input of each layer of the gating function and the actual selected expert. Module M4: Use the collected training data to train the expert access predictors of each layer offline and persist them to disk; Module M5: Start the inference service, load the non-expert parameters into the GPU memory, load the expert parameters into the CPU memory, and load the corresponding predictor for each layer of the large language model into the CPU memory; Module M6: During online inference, the gate function input of each layer in the GPU inference process is collected into the CPU memory. After the predictor corresponding to the collected input obtains the input, it will calculate e on the CPU. i =P i (h i-k ), output the probability of each expert being used after k layers as the prediction result; Module M7: According to the prediction results, select experts for each layer with the same number of experts as the model activation experts, put them into the pre-fetch queue, and hand them over to the pre-fetcher for pre-fetching. The pre-fetcher loads the expert parameters from the host memory to the GPU memory in the order of entry; Module M8: When the reasoning reaches the predicted layer, the experts that were not pre-fetched due to prediction errors are put back into the pre-fetch queue and handed over to the pre-fetcher for pre-fetching. Module M9: After the reasoning of a word is completed, the reasoning result is returned to the user and module M6 is triggered to start the reasoning of the next word.
7. The expert interview prediction system applicable to the hybrid expert architecture large language model according to claim 6, characterized in that: In module M4, the expert access predictors of each layer are trained offline. Specifically, for each input in the training data, the output of the predictor is compared with the actual output, the loss function is calculated, and the predictor parameters are updated using back propagation combined with the gradient descent method. The training data is traversed and the above process is executed repeatedly until the predictor converges.
8. The expert interview prediction system applicable to the hybrid expert architecture large language model according to claim 6, characterized in that: In module M6, the acquisition and prediction processes will be performed on the CPU in parallel with the GPU inference process.
9. The expert interview prediction system applicable to the hybrid expert architecture large language model according to claim 6, characterized in that: In module M7, the expert parameter loading process is executed in parallel with the GPU inference process, and the number of pre-fetched experts is dynamically adjusted according to the predictor accuracy.
10. The expert interview prediction system applicable to the hybrid expert architecture large language model according to claim 6, characterized in that: In module M8, the re-prefetch caused by the prediction error enters the prefetch queue with high priority, and the unfinished erroneous prefetch tasks in the queue will be cleared.
Citation Information
Patent Citations
Hybrid expert model reasoning acceleration method, device, equipment, medium and program
CN118761472A
Hybrid expert model reasoning method
CN118863055A
Enhancing perceptual data using large language models in environment reconstruction systems and applications
CN119151006A
Hybrid expert model reasoning method and device
CN119312935A
USE OF LANGUAGE MODELS IN AUTONOMOUS AND SEMI-AUTOMATIC SYSTEMS AND APPLICATIONS
DE102024116258A1
Cited By
Memory limited device MoE large model reasoning optimization system and method based on dual prediction
CN120610905A
Compression and reasoning method of hybrid expert model, electronic equipment and medium
CN120806117A
Expert model preloading method and device, chip, electronic equipment, storage medium and computer program product
CN121116656A
Data processing method, terminal equipment and storage medium
CN121255115A
Heterogeneous reasoning acceleration method and system for hybrid expert model
CN121525859A