Hybrid expert model reasoning method based on cooperation of CPU and GPU

By adopting a coordinated inference method of CPU and GPU in the hybrid expert model, combining dynamic scheduling strategies and intelligent cache management, the storage requirements and resource balance problems of MoE on edge devices are solved, and efficient computing resource utilization and inference acceleration effects are achieved.

CN120235253APending Publication Date: 2025-07-01PEKING UNIV
View PDF 0 Cites 11 Cited by

Patent Information

Application Number
CN202510254307.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The hybrid expert model (MoE) has huge storage requirements on edge devices, especially on resource-constrained devices, which leads to the computing overhead caused by expert loading on demand becoming a bottleneck, and the workload balancing between the CPU and GPU is complex, resulting in unbalanced resource use.

Method used

Using a hybrid expert model inference method based on CPU and GPU collaboration, a collaborative execution system including a hybrid CPU-GPU expert scheduler, an impact-driven expert prefetching mechanism and score-based expert cache management is built. Through dynamic scheduling strategies and intelligent cache management, resource allocation and computing efficiency are optimized.

Benefits of technology

Effectively balance heterogeneous computing resource load, improve hardware utilization, reduce the transmission overhead caused by cache missing, realize the overlap between CPU computing and PCIe transmission during GPU execution, improve expert cache hit rate, and ensure stable and efficient inference acceleration on resource-constrained platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235253A_ABST
    Figure CN120235253A_ABST
Patent Text Reader

Abstract

The invention discloses a hybrid expert model reasoning method based on cooperation of a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit), and belongs to the field of deep learning. According to the method, a CPU-GPU computing framework of a hybrid expert model is constructed, heterogeneous computing resource loads are effectively balanced, and the hardware utilization rate is remarkably increased; an intelligent cache management mechanism based on dynamic priority scores is provided, high-demand experts are reserved preferentially, and the transmission overhead caused by cache missing is reduced; through pipeline parallel design for separating calculation and transmission tasks, CPU calculation and PCIe transmission are overlapped in the GPU execution period, and delay is effectively hidden. In addition, in combination with a multi-layer expert activation prediction prospective prefetching mechanism, the expert cache hit rate is improved. The method is compatible with hybrid expert models of different scales and structures, and stable and efficient reasoning acceleration is realized on a resource-limited heterogeneous platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning (machine learning), and particularly to a method for reasoning of a Mixture-of-Experts (MoE) model based on the cooperation of a Central Processing Unit (CPU) and a Graphics Processing Unit (GPU). Background Art

[0002] In recent years, with the wide application of large language models (LLMs) in natural language processing tasks, especially in fields such as human-computer dialogue and text generation, the excessive model computation has become an important issue. As an effective method to improve computational efficiency, the Mixture-of-Experts (MoE) model can reduce the consumption of computing resources without affecting the model performance. The MoE model distributes the input data to a subset of experts through a dynamic routing mechanism, thereby significantly improving the parameter capacity and processing ability of the LLM while maintaining the same computing requirements.

[0003] Although the MoE model has advantages in computational efficiency, its huge storage requirements still pose challenges in actual deployment, especially on edge devices with limited storage capacity. To solve this problem, the expert offloading technology has emerged. This technology stores the expert weights in a secondary storage device, such as memory or hard disk, and only loads the expert weights into the GPU video memory when needed. However, in the expert offloading scenario, the main bottleneck of the MoE model is the overhead caused by the on-demand loading of experts.

[0004] To further optimize the computational overhead, some research has explored reducing the offloading overhead through CPU-assisted computing. When the required expert weights are not cached in the GPU video memory, the CPU can directly read the weights from memory and perform the calculation, avoiding the significant latency of loading from memory to video memory.

[0005] However, the dynamic nature of the MoE model poses great challenges to the cooperative reasoning between the CPU and the GPU. Specifically, the expert activations in the MoE model usually change frequently, and the activated experts cannot be predicted during each decoding, which makes the workload balance between the CPU and the GPU complex. Static task allocation strategies cannot adapt to this real-time change, resulting in uneven resource utilization.

[0006] Most current solutions rely on fixed mapping strategies based on historical activation frequencies, but this approach fails to fully consider the dynamics and unpredictability in the MoE inference process, thus unable to achieve optimal allocation of CPU and GPU resources. Therefore, in response to the challenges brought by the dynamics in the MoE model inference process, there is an urgent need to develop more flexible task scheduling and resource allocation mechanisms to improve the running efficiency and resource utilization rate of the MoE model on edge devices. Summary of the Invention

[0007] In view of the problems existing in the above prior art, the present invention proposes a hybrid expert model inference method based on CPU-GPU cooperation.

[0008] The technical solution provided by the present invention is as follows:

[0009] A hybrid expert model inference method based on CPU-GPU cooperation, characterized in that a CPU-GPU computing framework of the hybrid expert model is constructed, and a cooperative execution system including a hybrid CPU-GPU expert scheduler, an influence-driven expert prefetching mechanism, and a score-based expert cache management is established, specifically including the following steps:

[0010] Step 1: Given an autoregressive hybrid expert model M, which includes L layers, and each layer has N experts E0, E1,... E N-1 , represent the input data as X, which is a tensor with the shape of batch size, sequence length, and hidden layer dimension. Denote the time for a single expert to process a load of i on the CPU platform as The time for processing a load of i on the GPU platform is denoted as The transmission time from the CPU platform to the GPU platform is denoted as T Trans ;

[0011] Step 2: According to the set expert cache ratio k, set the number of experts allocated to the video memory of each layer as t l , then t l = kN, and evenly distribute the experts of each layer to the GPU and CPU memories;

[0012] Step 3: After obtaining the activated experts by calling the gating function, the GPU gives priority to calculating the cached experts with high load, the CPU gives priority to calculating the uncached experts with low load, and the CPU-GPU transmission mechanism gives priority to moving the uncached experts with high load from the CPU to the GPU. Denote the time for the CPU platform to process the experts as T cpu (cpu_expert), and the time for the GPU platform to process the experts as T gpu (gpu_expert). The specific scheduling objectives are:

[0013]

[0014] Step 4: Based on the scheduling objective in Step 3, after the gating function gives the expert loads, set the GPU expert execution queue Q G and the CPU expert execution queue Q C . Meanwhile, record the total execution time T G of Q GPU , the total transfer time T PCIe from the CPU to the GPU of the experts, and the total execution time T C of Q CPU ;

[0015] Step 5: Divide all the activated experts into two lists, the cached experts L load and the uncached experts L unload , both sorted in descending order of load. Each time, select an expert and assign it to Q G or Q C . In a single selection, if the expert E i is assigned to Q G and the load is i, then the update of T GPU is as follows: If E i is cached: If E i is uncached: E i In the case that E PCIe is uncached, T PCIe needs to be updated to: T PCIe = T Trans + T i . If the expert E CPU is assigned to the CPU queue and the load is i, then the update of T CPU

[0016] G or the CPU expert execution queue Q C for model inference, establish a heterogeneous computing pipeline, create independent CUDA streams to handle GPU computing and the transfer of experts from the CPU to the GPU respectively, and make the CPU computing and PCIe transfer overlap during GPU computing.

[0017] Furthermore, in Step 4), if the model has N' shared experts E'0, E'1,... E' N′-1 , the time for the shared expert calculation needs to be added to T GPU :

[0018] Furthermore, after the expert E i is assigned in Step 5), the total execution time T GPU of the GPU queue and the total execution time T CPUMinimize the larger value of the two, that is, calculate the actual time after each allocation selection:

[0019] T m = max(T CPU , T GPU )

[0020] Each time, select the smaller allocation strategy for T m , and remove the selected experts from L until each expert in L is assigned to the GPU expert execution queue Q G or the CPU expert execution queue Q C .

[0021] Furthermore, in step 6), an expert prefetch mechanism is set up, that is, predict the expert activation patterns of the future m layers. Specifically, it includes: first, reuse the output of the gating function of the subsequent m layers to generate the routing score vector s of each layer of experts l , for the candidate experts of layer l, perform a simulated scheduling process to evaluate the impact of prefetching the experts of this layer on the overall latency, select the expert that reduces the estimated latency the most as the prefetch target, and load the target expert into the GPU video memory in advance through the PCIe channel.

[0022] Furthermore, in step 6), maintain the dynamic priority score S of each expert. After each execution of the gating function, obtain the routing scores of all experts and update them. The update formula is:

[0023] S = α · TopP(s) + (1 - α) · S

[0024] where α is the average coefficient, TopP(·) extracts the set of experts in the top p positions of the current routing scores. When layer l needs to load the uncached expert E i from the CPU to the GPU, and the number of cached experts reaches the threshold tl, trigger the cache replacement mechanism, and replace the expert with the lowest priority score with E according to the priority score S of the current cached experts i .

[0025] The beneficial effects of the present invention are as follows:

[0026] Through the hybrid CPU-GPU dynamic scheduling strategy, the present invention effectively balances the heterogeneous computing resource load and significantly improves the hardware utilization rate. Based on the intelligent cache management mechanism of dynamic priority scores, it preferentially retains high-demand experts and reduces the transmission overhead caused by cache misses. Through the pipeline parallel design that separates the computing and transmission tasks, the overlap of CPU computing and PCIe transmission during GPU execution is achieved, effectively hiding the latency. In addition, the forward-looking prefetch mechanism combined with multi-layer expert activation prediction improves the expert cache hit rate. The present invention is compatible with hybrid expert models of different scales and structures and realizes stable and efficient inference acceleration on resource-constrained heterogeneous platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a flowchart of the method provided by a specific embodiment of the present invention.

[0028] Figure 2 This is an example diagram of the hybrid scheduling strategy in a specific embodiment of the present invention.

[0029] Figure 3 This is a schematic diagram of the expert prefetching mechanism in a specific embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The present invention will be further described below in conjunction with the drawings through specific embodiments.

[0031] The present invention relates to a hybrid expert model inference method for CPU and GPU cooperation. As Figure 1 shown, taking the operation of the DeepSeek-V2-Lite-Chat model on an RTX 6000 graphics card as an example, it includes the following steps:

[0032] Step 1: Given the autoregressive hybrid expert model DeepSeek-V2-Lite-Chat, which contains 26 layers, and each layer has 64 experts E0, E1,... E 63 , and two shared experts E′0, E′1. Represent the input data as X, which is a tensor with the shape of (batch size, sequence length, hidden layer dimension). Before the inference starts, measure the performance of a single expert on the CPU and GPU platforms with different load quantities. Denote the time for a single expert to process a load of i on the CPU platform as Denote the time for a single expert to process a load of i on the GPU platform as Denote the transfer time from the CPU platform to the GPU platform as T Trans , and denote the time for a single shared expert to process a load of i on the GPU platform as

[0033] Step 2: Set the expert cache ratio to 25%, and the number of experts allocated to video memory for each layer is t l = 64k, that is, 16 experts are cached to the GPU for each layer, and 48 experts are loaded into memory.

[0034] Step 3: After obtaining the activated experts by calling the gating function, set the rules of the hybrid scheduling strategy: The GPU gives priority to calculating the cached experts with high loads. The CPU gives priority to calculating the uncached experts with low loads. The CPU-GPU transfer mechanism gives priority to moving the uncached experts with high loads from the CPU to the GPU. Denote the time for the CPU platform to process an expert as T cpu (cpu_expert), and denote the time for the GPU platform to process an expert as T gpu(gpu_expert), the specific scheduling objective is:

[0035]

[0036] Step 4: Based on the scheduling objective in Step 3, after the gating function gives the expert load, set the GPU expert execution queue Q G and the CPU expert execution queue Q C , and at the same time record the total execution time T G of Q GPU , the total transfer time T PCIe from the CPU to the GPU of the expert, and the total execution time T C of Q CPU . This model has two shared experts, and T GPU is updated to:

[0037]

[0038] Step 5: Divide all activated experts into two lists, the cached experts L load and the uncached experts L unload , both sorted in descending order of load. Each time, select an expert and assign it to Q G or Q C . To meet the rules of the hybrid scheduling strategy, the experts assigned to Q G should be selected from the beginning to the end of L load , or from the end to the beginning of L unload . The experts assigned to Q C should be selected from the end to the beginning of L unload . If L unload is empty, it can be selected from the end to the beginning of L load .

[0039] In a single selection, if the expert E i is assigned to Q G , and the load is i, then T GPU is updated to:

[0040] If E i is cached:

[0041]

[0042] If E i is not cached:

[0043]

[0044] E i is not cached, T PCle needs to be updated to

[0045] TPCle = T PCIe + T Trans

[0046] If expert E i is assigned to the CPU queue with a load of i, then the update of T CPU is as follows:

[0047]

[0048] Step 6: To achieve optimal scheduling, while satisfying the policy rules, make the larger value between the total execution time T i of the GPU queue and the total execution time T GPU of the CPU queue after expert E is assigned CPU minimal, that is, the actual calculated time is lower. Therefore, calculate the actual time after each allocation selection

[0049] T m = max(T CPU , T GPU )

[0050] Each time, select the allocation strategy with a smaller T m , and remove the selected expert from L until each expert in L is assigned to the GPU expert execution queue Q G or the CPU expert execution queue Q C .

[0051] Figure 2 A hybrid scheduling example is provided. The arrow direction in the figure is the selection order of the CPU queue, GPU queue, and PCIe queue. In step one, compare the three options of preferentially selecting expert D to execute in the GPU queue, preferentially selecting expert A to execute in the CPU queue, or preferentially selecting expert C to transfer in the PCIe queue, and select the one with the shortest time consumption, that is, select expert A to the CPU queue. Therefore, calculate expert A in the CPU and discard the options of expert C and expert D. Similarly in step two, after comparing the time consumption of selecting D to the GPU queue, C to the PCIe queue, or B to the CPU queue, select B to execute in the CPU and discard the options of C and D. In step three, after comparing the time consumption of selecting D to the GPU queue, C to the PCIe queue, or C to the CPU queue, select D to execute in the GPU. In step four, after comparing the three options, select C to transfer in the PCIe queue. In step five, after comparing the time consumption of E to the GPU queue, PCIe queue, or CPU queue, select to calculate in the CPU. In step six, after C finishes transferring from the PCIe, add it to the GPU queue for calculation. This strategy balances the execution time of the CPU and GPU to achieve better hardware utilization.

[0052] Step 7: According to the GPU expert execution queue Q obtained in step 6 Gor the CPU expert executes queue Q C Perform model inference. Establish a heterogeneous computing pipeline, apply a fine-grained parallelization strategy, create independent CUDA streams to separately process GPU computing and the expert transfer from CPU to GPU, and overlap CPU computing and PCIe transfer during GPU computing to improve the inference speed.

[0053] Step 8: Implement an impact-driven expert prefetch mechanism during model inference: As Figure 3 shown, predict the expert activation patterns of the next 3 layers based on the similarity features between the current layer's hidden state and the subsequent layers. First, the gating function 1, the MoE part of the first layer, and the non-MoE part of the second layer in the figure are the sequences to be executed soon. The present invention reuses the outputs of the gating functions of the next 3 layers of the first layer, i.e., the outputs of gating functions 2, 3, and 4, and generates the routing score vector s of each layer's experts through a simulator l . As shown in the figure, for the candidate experts of the third layer, perform a simulated scheduling process to evaluate the prefetch of the experts in the third layer that may be called but have not been cached to the GPU yet, and select the most efficient and important expert i as the prefetch target. While calculating the MoE part of the first layer and the non-MoE part of the second time, load the target expert i to the GPU video memory in advance through the PCIe channel, make full use of the idle bandwidth of PCIe during GPU computing, and ensure that the prefetch operation does not affect the computing tasks of the current layer. Similarly, when calculating to the second layer and executing the gating function 2, continue to reuse the gating functions of the next 3 layers of the second layer and implement the expert prefetch mechanism.

[0054] Step 9: Adopt a dynamic cache management strategy based on routing scores: Maintain the dynamic priority score S of each expert, obtain the routing scores of all experts after each gating function execution and update them, and the update formula is:

[0055] S = α · TopP(s) + (1 - α) · S

[0056] where α is the average coefficient, set to 0.5. The number of activated experts in this model is 6, so P is set to 12.

[0057] Step 10: When the l-th layer needs to load the uncached expert E i from the CPU to the GPU, and the number of cached experts reaches the threshold of 16, trigger the cache replacement mechanism, and replace the expert with the lowest priority score with E according to the current priority score S of the cached experts i .

[0058] Finally, it should be noted that the purpose of disclosing the embodiments is to help further understand the present invention. However, those skilled in the art can understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection claimed by the present invention shall be subject to the scope defined by the claims.

Claims

1. A hybrid expert model reasoning method based on CPU and GPU collaboration, specifically comprising the following steps: Step 1: Given an autoregressive mixture of experts model M, which consists of L layers, each layer has N experts E0, E1, …E N-1 , the input data is represented as X, which is a tensor with the shape of batch size, sequence length, and hidden layer dimension, and the time for a single expert to process load i on the CPU platform is recorded as The time for processing load i on the GPU platform is recorded as The transmission time from the CPU platform to the GPU platform is denoted as T Trans ; Step 2: According to the set expert cache ratio l, the number of experts allocated to each layer of video memory is t l , then t l = kN, evenly distribute the experts in each layer to the GPU and CPU memory; Step 3: After calling the gating function to get the activated expert, the GPU gives priority to the calculation of the high-load cached expert, and the CPU gives priority to the calculation of the low-load uncached expert. The CPU-GPU transfer mechanism gives priority to the movement of the high-load uncached expert from the CPU to the GPU. The time for the CPU platform to process the expert is recorded as T cpu (cpu_expert), the time for GPU platform to process the expert is recorded as T gpu (gpu_expert), the specific scheduling target is: Step 4: Based on the scheduling target in step 3, after the gating function gives the expert load, set the GPU expert execution queue Q G and CPU Expert Execution Queue Q C , while recording Q G Total execution time T GPU , the total transmission time from CPU to GPU of experts is T PCIe and Q C Total execution time T CPU ; Step 5: Divide all activated experts into cached experts L according to whether they are cached or not. load and uncached expert L unload The two lists are sorted from high to low according to the load, and each time an expert is selected and assigned to Q G or Q C , in a single choice, if expert E i Assigned to Q G , the load is i, then T GPU The update is: If E i Cached: If E i Not cached: E i In the case of no caching, T PCIe Need to be updated to: T PCIe =T PCIe +T Trans ; If expert E i Assigned to the CPU queue, the load is i, then T CPU The update is: Step 6: GPU expert execution queue Q obtained in step 5 G Or CPU Expert Execution Queue Q C Perform model inference, establish heterogeneous computing pipelines, and create independent CUDA streams to handle GPU computing and CPU-to-GPU expert transmission respectively, so that CPU computing and PCIe transmission can be overlapped during GPU computing.

2. The hybrid expert model reasoning method based on CPU and GPU collaboration as claimed in claim 1, characterized in that: In step 4), if the model has N′ shared experts E′0, E′1, …E′ N′-1 , then T GPU Plus the time of shared expert calculations:

3. The hybrid expert model reasoning method based on CPU and GPU collaboration as claimed in claim 1, characterized in that: In step 5), the expert E i After being allocated, the total execution time of the GPU queue is T GPU The total execution time T of the CPU queue CPU The larger of the two is minimized, i.e. the actual time after each allocation selection is calculated: T m =max(T CPU ,T GPU ) Each time you select T m Smaller allocation strategy, and remove the selected experts from L until each expert in L is assigned to the GPU expert execution queue Q G Or CPU Expert Execution Queue Q C .

4. The hybrid expert model reasoning method based on CPU and GPU collaboration as claimed in claim 1, characterized in that: In step 6), the expert pre-fetching mechanism is set, that is, predicting the expert activation mode of the future m layers, which specifically includes: firstly reusing the gating function output of the subsequent m layers to generate the routing score vector s of the experts in each layer l ,For the candidate experts of layer l, a simulation scheduling process is performed to evaluate the ,impact of pre-fetching the experts of this layer on the overall latency, and the ,expert that can reduce the estimated latency the most is selected as the ,pre-fetching target, and the target expert is pre-loaded into the GPU memory through the ,PCIe channel.

5. The hybrid expert model reasoning method based on CPU and GPU collaboration as claimed in claim 1, characterized in that: In step 6), the dynamic priority score S of each expert is maintained, and the routing scores of all experts are obtained and updated after each gating function is executed. The update formula is: S=α·TopP(s)+(1-α)·S Where α is the average coefficient, and TopP(·) extracts the expert set with the top p digits of the current routing score.

6. The hybrid expert model reasoning method based on CPU and GPU collaboration as claimed in claim 5, characterized in that: When the lth level needs to load the uncached expert E i From CPU to GPU, while the number of cached experts reaches the threshold t l When , the cache replacement mechanism is triggered. According to the priority score S of the current cache expert, the expert with the lowest priority score is replaced by E. i .

7. The hybrid expert model reasoning method based on CPU and GPU collaboration as claimed in claim 5, characterized in that: p is twice the number of activated experts.

Citation Information

Cited By

  • Processor system and operation method for large-scale deep learning

    CN120448132A

  • Large model batch reasoning and data flow optimization system oriented to MOE architecture

    CN120849141A

  • Mass inference and data flow optimization system for MOE architecture-oriented large model

    CN120849141B

  • Model loading and unloading method and electronic equipment

    CN120909807A

  • Model loading and unloading methods and electronic devices

    CN120909807B