Multi-expert model reasoning method, device and equipment and storage medium

By performing offline performance measurement, online request scheduling, batch division, and expert management in the multi-expert model inference system, the problem of excessive number of expert exchanges in the existing system is solved, and the inference efficiency is improved.

CN120196433APending Publication Date: 2025-06-24BEIHANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510244348.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing multi-expert model inference system is not efficient in request scheduling, expert management and memory management, resulting in too many expert exchanges, affecting the inference performance.

Method used

By measuring the performance of the expert model in the offline stage and generating the optimal resource allocation plan, the request scheduling module is used to schedule inference request requests, the batch division module divides request batches, and the expert management module dynamically manages expert loading to reduce the number of expert exchanges.

Benefits of technology

The number of expert exchanges during the inference process of multi-expert models is reduced, the inference efficiency is improved, and request scheduling, expert management and memory management are optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196433A_ABST
    Figure CN120196433A_ABST
Patent Text Reader

Abstract

The invention provides a multi-expert model reasoning method, device and equipment and a storage medium, and aims to reduce the number of times of expert model exchange among multi-level memories in a multi-expert model reasoning process so as to improve the reasoning efficiency of a multi-expert model. In the off-line stage, the system measures the performance of each expert model and obtains an optimal memory allocation scheme so as to reasonably allocate the memory to store the parameters and the reasoning intermediate quantity of the expert models. In the online stage, a reasoning request scheduling module schedules a reasoning request to a proper reasoning actuator queue to wait for reasoning. And the batch processing division module performs batch processing on the requests according to the expert performance and the available memory. When expert exchange is needed, the expert management module unloads the experts with the minimum use probability in the future and loads the needed experts. According to the method, the expert exchange frequency in the multi-expert model reasoning process is reduced, and the reasoning efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention relates to the technical field of expert model inference, and particularly to an expert model inference method, device, equipment, and storage medium. By utilizing heterogeneous computing power and multi-level storage resources, request scheduling and expert model management are performed in multi-expert model inference to achieve efficient multi-expert model inference. Background Art:

[0002] With the development of artificial intelligence technology, many models with excellent performance have emerged in different sub-tasks, such as code generation, language translation, etc. By combining these models as experts to form a multi-expert model and calling different experts according to different input requests, different types of tasks can be effectively processed, and compared with a single model, it has better performance.

[0003] Due to privacy and latency requirements, in many scenarios, such as intelligent manufacturing, smart grid, etc., it is required that these expert models be deployed on edge devices. However, the resources of edge devices are limited, and when the number of expert models is too large, they cannot all be loaded into the GPU memory. The current method is to unload some expert models to memory or hard disk, and unload the experts already loaded in the GPU memory when activating to release memory, and then load the activated experts into the GPU memory for calculation. This method is called "expert swapping". However, in one inference, the time ratio of expert swapping from hard disk to GPU memory can exceed 90%, and the time ratio of expert swapping from CPU memory to GPU memory will also exceed 60%, which leads to a performance bottleneck in inference. Therefore, reducing the number of expert swaps is crucial for improving the inference efficiency of multi-expert models.

[0004] The current multi-expert model inference system mainly focuses on the selection of replaced expert models, and usually adopts the least recently used (LRU) strategy to replace expert models to reduce the frequency of "expert swapping". However, this method does not consider the dependency relationship of expert models and the global usage probability in a task, resulting in less improvement in inference efficiency. In addition, during inference, the scheduling of inference requests and memory management will also affect inference efficiency.

[0005] To reduce the number of "expert switches" and improve the inference efficiency, this paper proposes an efficient inference method for multi-expert models, which improves the inference efficiency of multi-expert models by optimizing request scheduling, expert management, and memory management. In terms of request scheduling, requests are distributed across different inference executors, and requests processed by the same expert are combined for inference; in terms of expert management, experts are replaced during "expert switches" by considering the dependency relationships and usage probabilities of experts; in terms of memory management, a trade-off is made between the memory usage of expert model parameters and the intermediate quantities of batch processing during inference. Through the above methods, the number of expert switches is reduced, and the inference efficiency of multi-expert models is optimized.

[0006] In summary, although the existing multi-expert model inference systems have optimized the inference efficiency to a certain extent in expert replacement, their efficiency in request scheduling, expert management, and memory management is still not high. The present invention proposes an inference optimization method for multi-expert models, which reduces the number of "expert switches" and improves the inference efficiency of multi-expert models by fully optimizing request scheduling, expert management, and memory management. Summary of the Invention:

[0007] The present invention proposes an expert model inference method, device, equipment, and storage medium. In the offline stage, the performance of the expert model is measured and an optimal resource allocation plan is generated; in the online stage, inference requests are received, the execution order of the requests is managed through request scheduling, and the loading of experts is dynamically managed to efficiently execute the requests.

[0008] The technical solution of the present invention is as follows:

[0009] A multi-expert model inference method, characterized by comprising the following steps:

[0010] Step 1, using a performance measurement module, receiving an expert model and a routing rule, and generating configuration information by running the expert model on a small data set, including the inference latency of the expert model, the loading latency of the expert model, the memory occupancy, the usage probability of the expert model, the number of inference executors, and the memory allocation plan;

[0011] Step 2, using an executor creation module, creating multiple inference executors running on GPUs and CPUs according to the configuration information, and allocating memory for storing expert model parameters and intermediate inference quantities;

[0012] Step 3, using an expert initialization module, loading the expert model into the model pool during system initialization;

[0013] Step 4, using an inference request scheduling module, scheduling the input inference requests to the inference request queue of a suitable inference executor to minimize the total inference time and the additional inference time;

[0014] Step 5: Use the batch processing division module to divide a set of requests into different batches for batch processing. The number of batch processes is set to the number that minimizes the average inference latency under the condition of meeting the memory limit.

[0015] Step 6: Use the expert management model. When the required expert model is not in the memory pool, unload the existing expert and load the required expert. The unloaded expert is the one with the lowest probability of future use.

[0016] A multi-expert model inference system, device, or equipment, characterized by including:

[0017] (1) A performance measurement module, which is used to receive the expert model and routing rules, and generate configuration information by running the expert model on a small data set, including the inference latency of the expert model, the loading latency of the expert model, the memory occupancy, the usage probability of the expert model, the number of inference executors, and the memory allocation scheme.

[0018] (2) An executor creation module, which is used to create multiple inference executors running on the GPU and CPU according to the configuration information, and allocate the memory for storing the expert model parameters and inference intermediate quantities.

[0019] (3) An expert initialization module, which is used to load the expert model into the model pool during system initialization.

[0020] (4) An inference request scheduling module, which is used to schedule the input inference requests to the inference request queue of the appropriate inference executor to minimize the total inference time and the additional inference time.

[0021] (5) A batch processing division module: which is used to divide a set of requests into different batches for batch processing. The number of batch processes is set to the number that minimizes the average inference latency under the condition of meeting the memory limit.

[0022] (6) An expert management model: which is used to unload the existing expert and load the required expert when the required expert model is not in the memory pool. The unloaded expert is the one with the lowest probability of future use.

[0023] A multi-expert model inference storage medium, characterized in that the above multi-expert model inference method is included in the computer program stored in the storage medium.

[0024] The above-mentioned multi-expert model inference system, device, or equipment is characterized in that the performance and usage probability of all expert models are measured by offline running a small data set. The performance of the expert model includes the inference latency of the expert model, the loading latency of the expert model, and the memory occupancy of the expert model.

[0025] The described multi-expert model inference system, device, or equipment is characterized in that during the offline phase, the performance of different numbers of inference executors is measured to obtain the number of inference executors that optimizes the inference throughput. The inference executors can run on heterogeneous processors.

[0026] The described multi-expert model inference system, device, or equipment is characterized in that according to the cumulative distribution function characteristics of the expert model, by means of a decreasing measurement window, the throughput at both ends of the window is measured in sequence to obtain the optimal number of expert models to be loaded.

[0027] The described multi-expert model inference system, device, or equipment is characterized in that by predicting the execution time of inference requests on different inference executors and the expert exchange time, the inference requests are assigned to the request queue of the inference executor that minimizes the total execution time and the increased execution time.

[0028] The described multi-expert model inference system, device, or equipment is characterized in that by sensing the currently available memory of the system and combining with the inference performance of the expert model, a set of requests is divided into appropriate multiple batches of requests, and the batch processing quantity of each batch of requests is the quantity that minimizes the average inference latency under the condition of meeting the memory limit.

[0029] The described multi-expert model inference system, device, or equipment is characterized in that during expert exchange, a two-stage expert model unloading method is used to unload the loaded expert models. In the first stage, the subsequent expert models without pre-dependencies are unloaded in descending order of memory occupancy; in the second stage, other expert models are unloaded in descending order of usage probability until the memory space meets the requirements for loading the required expert models.

[0030] The technical effects of the present invention are as follows: The present invention provides a multi-expert model inference method, device, equipment, and storage medium, aiming to reduce the number of expert model exchanges between multiple levels of memory during the multi-expert model inference process to improve the inference efficiency of the multi-expert model. During the offline phase of the system, the performance of each expert model is measured, and the optimal memory allocation scheme is obtained to reasonably allocate memory to store the parameters and inference intermediate quantities of the expert models. During the online phase, the inference request scheduling module schedules the inference requests to the appropriate inference executor queues for waiting for inference. The batch processing division module performs batch processing on the requests according to the expert performance and available memory. The expert management module unloads the expert with the lowest future usage probability and loads the required expert when expert exchange is needed. The present invention reduces the number of expert exchanges during the multi-expert model inference process and improves the inference efficiency. BRIEF DESCRIPTION OF THE DRAWINGS:

[0031] Figure 1 It is a schematic diagram of the overall architecture of multi-expert model inference. DETAILED DESCRIPTION OF THE EMBODIMENTS:

[0032] The following further elaborates on the present invention in conjunction with the accompanying drawings ( Figure 1 ) and embodiments.

[0033] The present invention relates to an efficient multi-expert model inference system based on edge devices. By utilizing heterogeneous computing power and multi-level storage resources, request scheduling and expert model management are carried out in multi-expert model inference to achieve efficient multi-expert model inference.

[0034] The present invention proposes an efficient multi-expert model inference method. In the offline stage, the performance of experts is measured and an optimal resource allocation scheme is generated; in the online stage, inference requests are received, the execution order of requests is managed through request scheduling, and the loading of experts is dynamically managed to efficiently execute requests.

[0035] As Figure 1 shown, it is the overall architecture of multi-expert model inference. The system includes three stages: the offline stage, the system initialization stage, and the online operation stage.

[0036] In the offline stage, the performance measurement module measures the performance of experts according to the provided expert models and routing information, including the inference latency of experts, the expert exchange latency, the memory occupancy, and the expert usage probability. Moreover, the performance measurement module measures the optimal number of inference executors and the memory allocation scheme.

[0037] An inference executor refers to a process that performs inference on the GPU and CPU. By creating different numbers of inference executors, the inference performance is measured, and the number of inference executors that gives the best performance is selected as the optimal number of inference executors.

[0038] Regarding the memory occupancy, one is the occupancy for loading expert models. Allocating more memory to it can increase the number of loaded models; the other is the occupancy for intermediate inference quantities. Allocating more memory to it can increase the number of batch requests. In this article, different memory allocation methods are specified according to the characteristics of the processor's computing power.

[0039] On a processor with limited computing power, such as a CPU. The memory allocated to the intermediate inference quantities can satisfy the optimal number of batch requests, and the remaining memory is allocated to the loading of expert models.

[0040] On a processor with powerful computing capabilities, such as a GPU, the memory allocation is determined by measuring the optimal number of expert models to be loaded. First, the expert models are sorted in descending order of usage probability to obtain their cumulative distribution function. Then, a "window" is initialized from 0 on the cumulative distribution function, and the number of expert models to be loaded at both ends of the "window" is calculated, and the loading throughput at this number is obtained. After that, the "window" is decayed, and the decay rate is as shown in Equation 1, and the "window" is increased upward based on the original "window", and its throughput is tested again. The above process is repeated until it is found that the throughput decreases. At this time, within the last "window", a value is randomly selected, and the corresponding number of expert models to be loaded is calculated. Allocate memory that meets the loading quantity for the loading of expert models, and the remaining memory is used to store intermediate quantities during inference.

[0041]

[0042] In the system initialization phase, the executor creator creates an inference executor according to the previously generated configuration information. After that, the expert initialization module in each inference executor loads the expert model into the model pool according to the memory occupied by the expert model.

[0043] In the online phase, the incoming inference requests are first input into the inference request scheduling module, and the inference request scheduling module schedules the inference requests to the appropriate inference request queues. Each inference request queue corresponds to an inference executor. Then, the batch processing division module divides the requests in the request queue. Finally, the requests in the batch are executed for inference. During the execution, if the required expert model is in the model pool, the inference is directly performed; otherwise, the expert management module unloads the expert model with the lowest future usage probability and loads the required expert model to complete the inference.

[0044] During request scheduling, first, according to the measured expert performance, predict the increased inference time of the new request on each inference executor, including the inference execution time and the "expert exchange" time.

[0045] Then, select the request queue of the executor that makes the total inference time of the task the shortest and the increased inference time the shortest to join. Since the executors are parallel, the total inference time of the task is the time of the executor queue with the longest inference time. After the request joins the queue, if there are requests processed by the same expert in the queue, the request is scheduled behind these requests to form a group of requests to reduce the possible "expert exchange".

[0046] Finally, when a request is to be executed, the batch division module calculates the maximum number of batches that can be accommodated in the current memory based on the available memory space, and selects the smaller value between this number and the optimal batch number of the expert model as the maximum number of batches that can be executed currently. Based on the maximum number of batches that can be executed currently, a set of requests is divided into multiple batches of requests not exceeding this number.

[0047] When the required expert model is not in the model pool, the existing experts in the model pool need to be unloaded and the required expert needs to be loaded to complete the inference. Unloading the expert model is divided into two stages. In the first stage, the expert models without pre-order expert models are unloaded. Since the execution of some expert models depends on pre-order expert models, if their pre-order expert models are not in the memory pool, they will not be executed, resulting in memory waste. When unloading, they are unloaded in descending order of the memory space they use until there is enough memory to load the required expert model.

[0048] If there is still not enough memory left to load the required expert model after the above-mentioned model unloading is completed, then in the second stage, other expert models are unloaded. Other expert models are unloaded in ascending order of their usage probability until there is enough memory to load the required expert model.

[0049] When performing multi-expert model inference, the expert model and expert information are first input into the performance measurement module for offline measurement, and configuration information is generated. The configuration information includes the inference latency of the expert model, the loading latency of the expert model, the memory occupancy, the usage probability of the expert model, the number of inference executors, and the memory allocation scheme.

[0050] The inference latency, expert exchange latency, and memory occupancy of the expert model are obtained by running the expert model to record its running data.

[0051] The acquisition of the usage probability of the expert model can be divided into two cases: First, when the routing rule is uncertain, by performing inference on a small dataset and recording the number of times each expert model is used, the usage probability of the expert model can be obtained; Second, when the routing rule is determined, for example, in a task, if the data that the expert needs to process and the data distribution have been obtained, the usage probability of the expert model can be directly obtained through the data distribution.

[0052] The determination of the number of inference executors and the memory allocation scheme is obtained by running the inference process on a small dataset, obtaining its throughput, and selecting the configuration with the highest throughput according to the method described in the content of the present invention.

[0053] After obtaining the configuration information through the above measurements, the system starts to initialize. During the initialization phase, the executor creator creates multiple inference executors by creating multiple processes. Then, the expert initialization module loads the expert models into the memory in a polling manner according to the decreasing probability of using the expert models until the memory is full.

[0054] After the system is initialized, the system starts to run online. At this time, inference requests will be received. First, the inference request scheduling module first obtains the expert model required by this request according to the routing situation. Then, according to the performance information of this expert model, it predicts the increased execution time of this request in each inference executor. Finally, through the method described in the present invention content, the request is added to the request queue of the appropriate inference executor.

[0055] Inside the inference executor, the batch processing division module obtains multiple batches of requests to be processed according to the method described in the present invention content.

[0056] If the required expert model has been loaded into the model pool, the inference is directly completed using the required expert model. Otherwise, the expert management module loads the required expert model into the model pool according to the method described in the present invention content, and then completes the inference.

Claims

1. A multi-expert model reasoning method, characterized in that: The following steps are involved: Step 1: Utilize the performance measurement module to receive the expert model and routing rules, and generate configuration information by running the expert model on the small data set, including the inference delay of the expert model, the loading delay of the expert model, the memory usage, the usage probability of the expert model, the number of inference executors, and the memory allocation scheme; Step 2: Use the executor creation module to create multiple inference executors running on the GPU and CPU according to the configuration information, and allocate memory for storing expert model parameters and inference intermediate quantities; Step 3, using the expert initialization module, load the expert model into the model pool when the system is initialized; Step 4: Use the inference request scheduling module to schedule the input inference request to the inference request queue of the appropriate inference executor, so that the total inference time is the shortest and the increased inference time is the shortest; Step 5: Use the batch partitioning module to divide a group of requests into different batches for batch processing. The number of batches is set to the number that minimizes the average inference latency while meeting the memory limit. Step 6, using the expert management model, when the required expert model is not in the memory pool, the existing expert is unloaded and the required expert is loaded. The unloaded expert is the expert with the lowest probability of future use.

2. A multi-expert model reasoning system or device or apparatus, characterized in that: include: (1) A performance measurement module, which receives the expert model and routing rules and generates configuration information by running the expert model on a small data set, including the inference delay of the expert model, the loading delay of the expert model, the memory usage, the usage probability of the expert model, the number of inference executors, and the memory allocation scheme; (2) An executor creation module, which is used to create multiple inference executors running on the GPU and CPU according to the configuration information, and allocate memory for storing expert model parameters and inference intermediate quantities; (3) Expert initialization module, used to load the expert model into the model pool during system initialization; (4) an inference request scheduling module, which is used to schedule the input inference request to the inference request queue of the appropriate inference executor so that the total inference time is minimized and the increased inference time is minimized; (5) Batch partitioning module: used to divide a group of requests into different batches for batch processing. The number of batches is set to the number that minimizes the average inference latency while meeting the memory limit. (6) Expert management model: used to unload existing experts and load the required experts when the required expert model is not in the memory pool. The unloaded experts are the experts with the lowest probability of future use.

3. A multi-expert model reasoning storage medium, characterized in that: The multi-expert model reasoning method described in claim 1 is included in a computer program stored in a storage medium.

4. The multi-expert model reasoning system or device or apparatus according to claim 2, characterized in that The performance of all expert models and the probability of using expert models are measured by running a small data set offline. The performance of the expert model includes the inference delay of the expert model, the loading delay of the expert model, and the memory usage of the expert model.

5. The multi-expert model reasoning system or device or apparatus according to claim 2, characterized in that The performance of different numbers of inference executors is measured in the offline phase to obtain the number of inference executors that optimizes the inference throughput. The inference executors can run on heterogeneous processors.

6. The multi-expert model reasoning system or device or apparatus according to claim 2, characterized in that According to the characteristics of the cumulative distribution function used by the expert model, the throughput at both ends of the window is measured in turn through a decreasing measurement window to obtain the optimal expert model loading quantity.

7. The multi-expert model reasoning system or device or apparatus according to claim 2, characterized in that By predicting the execution time and expert exchange time of inference requests on different inference executors, the inference requests are allocated to the request queue of the inference executor that minimizes the total execution time and the increased execution time.

8. The multi-expert model reasoning system or device or apparatus according to claim 2, characterized in that By sensing the current available memory of the system and combining it with the reasoning performance of the expert model, a group of requests are allocated to appropriate batches of requests. The number of batches processed in each batch of requests is the number that minimizes the average reasoning delay while meeting the memory limit.

9. The multi-expert model reasoning system, device or apparatus according to claim 2, characterized in that During expert exchange, a two-stage expert model unloading method is used to unload the loaded expert models. In the first stage, the post-expert models without pre-dependencies are unloaded in descending order of memory usage; in the second stage, other expert models are unloaded in descending order of usage probability until the memory space is sufficient to load the required expert models.

Citation Information

Cited By

  • Information processing method and device, equipment, storage medium and program product

    CN121614274A