MoE model heterogeneous low-power scheduling method and system based on expert activated risk assessment
Patent Information
- Application Number
- CN202611061693.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-09-25
AI Technical Summary
[0011]本发明的目的在于针对上述现有技术的不足,提供一种基于专家激活风险评估的MoE 模型异构低功耗调度方法及系统,以解决现有 MoE 推理系统中专家低功耗控制缺乏前瞻性、简单休眠易造成误预测惩罚、PIM 兜底能力未被充分利用以及专家状态切换缺少反馈修正的问题,实现专家级细粒度低功耗控制,在保证推理延迟可控的同时降低 MoE 模型推理过程中的无效功耗
[0046]其一,本发明由于综合利用各专家在当前调度周期内被分配的词元 token 数量、历史激活信息、层间专家转移关系、请求级专家激活画像以及反馈修正信息,计算各专家在后续调度周期内的未来激活概率及其预测不确定性,而不是仅依据当前调度周期的专家冷热程度进行调度,因此能够提前识别可能被再次激活的专家,提高专家低功耗状态控制的前瞻性和准确性。
Smart Images

Figure CN122819489A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of large model inference acceleration and scheduling technology, specifically involving a heterogeneous low-power scheduling method and system for MoE models, which can be used for expert-level low-power state control, PIM fallback execution, and misprediction feedback correction for NPU-PIM heterogeneous inference platforms. Background Technology
[0002] As large language models and multimodal models continue to expand in scale, the number of model parameters, memory access overhead, and inference energy consumption increase significantly. To control the computational cost per inference while increasing model capacity, Mixture-of-Experts (MoE) models have gradually become an important structure in large models. MoE models typically include multiple expert modules and a gating network or router for selecting experts. For each input token, the router usually selects only a subset of experts to participate in the computation. Therefore, compared to dense models, MoE models exhibit structurally sparse activation characteristics. Within the same scheduling cycle, not all experts are invoked; the number of tokens allocated to different experts also varies significantly. Some experts are frequently activated over a longer period, while others are invoked with only a few tokens, or even not invoked at all for several scheduling cycles.
[0003] The aforementioned sparse activation characteristics provide space for low-power scheduling: on the one hand, experts with high token counts are generally more suitable for execution by NPUs, GPUs, or other high-throughput computing units; on the other hand, experts with low token counts or memory-sensitive experts can be handled by PIM, NDP, or near-memory computing units to reduce parameter movement and memory access overhead. Existing solutions related to MoE model inference and heterogeneous scheduling typically map different experts to different computing units based on the current number of expert tokens, current input data characteristics, or current hardware load, or select qualified experts to participate in computation through gating networks to improve throughput, reduce data movement overhead, or reduce unnecessary computation.
[0004] Patent document CN202510898422.3 discloses a data processing method and system for a low-energy, large-scale language model based on momentum mechanisms and multiple types of experts. It employs a hybrid expert model, acquires target data through a lazy loading mechanism, obtains the fitness scores of each expert network corresponding to the target data through a gating network, selects the expert network that meets the ranking requirements as the target network, acquires the output data of the target data in the target network, and performs a weighted summation of multiple output data through a combined network to obtain the final output. While this scheme, by setting multiple expert networks in the hybrid expert model and selecting the expert network participating in the calculation based on the fitness score of the target data, can reduce floating-point operations and computational memory overhead to some extent, improve computational efficiency, and reduce memory resource waste, it still has the following shortcomings.
[0005] First, this scheme focuses on expert selection, fitness ranking, and output combination based on current input data. It mainly focuses on the selection mechanism for experts to participate in the calculation and the optimization of computational efficiency in the inference process. It lacks the assessment of the future activation probability of experts in subsequent scheduling cycles and its prediction uncertainty, making it difficult to determine in advance whether an expert is suitable to enter a low-power state.
[0006] Second, because the scheme does not perform state division and migration management of experts from the perspective of expert-level fine-grained low-power control, it lacks a scheduling mechanism that distinguishes experts into multiple states such as high-performance execution, low-power execution, standby retention, and deep sleep. Therefore, it is difficult to further reduce the ineffective power consumption caused by experts that are not frequently activated.
[0007] Third, because the scheme does not fully consider the wake-up cost, misprediction penalty and inference delay impact that may result when experts enter a low-power state, it lacks a quantitative assessment of the risk of switching to a low-power state. This can easily lead to a situation where inactive experts are simply put to sleep and then suddenly activated, resulting in additional wake-up overhead and latency jitter.
[0008] Fourth, the heterogeneous inference scheme does not fully utilize the role of near-memory computing units such as PIM as a low-power fallback execution path in scenarios with uncertain predictions or a low number of tokens. Therefore, when an expert is suddenly activated by a small number of tokens, the system often needs to directly wake up high-power computing units, making it difficult to balance low power consumption and low latency.
[0009] Fifth, because the scheme lacks a feedback correction mechanism for actual operating results, it cannot dynamically correct future activation probabilities, low-power state risks, or state switching thresholds based on events such as sleepmiss, standby miss, PIM fallback event, or immediate NPU wake event. This results in poor scheduling stability and weak adaptability under different input requests and different expert activation modes.
[0010] Therefore, an expert-level low-power scheduling method is needed for the sparse activation characteristics of MoE. This method should be able to dynamically determine the expert's transition between NPU-active, PIM-active, standby, and sleep states based on the expert's future activation probability and the risk of low-power states. Furthermore, it should utilize PIM for low-power fallback execution when prediction is uncertain, thereby reducing inference energy consumption under controllable latency conditions. Summary of the Invention
[0011] The purpose of this invention is to address the shortcomings of the prior art by providing a heterogeneous low-power scheduling method and system for MoE models based on expert activation risk assessment. This addresses the problems in existing MoE inference systems, such as the lack of foresight in expert low-power control, the tendency of simple sleep mode to cause false prediction penalties, the underutilization of PIM fallback capabilities, and the lack of feedback correction during expert state switching. The invention achieves expert-level fine-grained low-power control, reducing invalid power consumption during MoE model inference while ensuring controllable inference latency.
[0012] The technical approach to achieving the objective of this invention is as follows: During the inference process, by collecting the token allocation quantity, historical activation status, inter-layer transition relationships, and request-level expert activation information of each expert, and based on this, predicting the future activation probability and prediction uncertainty of each expert in subsequent scheduling cycles, the forward-looking low-power control of experts in the MoE inference system is achieved; by further combining the expert wake-up cost, misprediction penalty, and PIM fallback execution cost, the risk of experts entering a low-power state is calculated, avoiding the misprediction penalty caused by directly putting an expert to sleep based solely on the current expert's popularity; by dividing experts into four states—NPU-active, PIM-active, standby, and sleep—and in activation scenarios where the prediction is uncertain or the token quantity is lower than the token threshold T_token_low, PIM is prioritized to execute the corresponding expert's calculation, improving the fallback capability of PIM; by feeding back and correcting the future activation probability, low-power state risk, and state switching threshold based on the deviation between the actual expert activation result and the prediction result, the ineffective power consumption during the MoE model inference process is reduced while ensuring controllable inference latency.
[0013] Based on the above ideas, the technical solution of the present invention includes:
[0014] 1. A heterogeneous low-power scheduling method based on a hybrid expert MoE model using expert activation risk assessment, characterized in that it includes:
[0015] S1) Obtain expert activation information during the inference process of the hybrid expert MoE model, which includes at least the number of tokens assigned to each expert in the current scheduling period;
[0016] S2) Based on the expert activation information, calculate the future activation probability of each expert in the subsequent scheduling cycle, and determine the prediction uncertainty corresponding to the future activation probability;
[0017] S3) Calculate the low-power state risk of each expert based on the future activation probability, prediction uncertainty, expert wake-up cost, misprediction penalty, and in-memory computing unit PIM fallback execution cost.
[0018] S4) Based on the future activation probability, prediction uncertainty and low power state risk, each expert is divided into any one of the following states: neural network processing unit active state (NPU-active), in-memory computing unit active state (PIM-active), standby state, and sleep state, and hybrid expert MoE inference is performed based on the state division results.
[0019] S5) Calculate the deviation between the hybrid expert MoE inference results and the prediction results, use the deviation to record sleep miss, standby miss, in-memory computing fallback event, and immediate NPU wake event, and correct the future activation probability, low power state risk or state switching threshold based on these events.
[0020] S6) Determine whether the current hybrid expert MoE inference task is complete:
[0021] If completed, output the hybrid expert model inference result and end the current scheduling process;
[0022] If not completed, proceed to the next scheduling cycle, that is, repeat S1) to S5 based on the correction result obtained in S5).
[0023] Furthermore, in S5), calculating the deviation between the hybrid expert MoE inference result and the prediction result involves comparing the actual activation state, actual token allocation quantity, and actual execution state of each expert inferred by the hybrid expert MoE in S4) with the predicted activation level, predicted high and low token thresholds in S2) and the predicted execution state in S4), respectively, to obtain the state or numerical deviation between the two, wherein:
[0024] The actual activation state is compared with the predicted activation level to obtain the deviation of the actual activation state.
[0025] The actual number of tokens allocated is compared with the low token threshold T_token_low and the high token threshold T_token_high to obtain the actual activation level.
[0026] The actual execution state is compared with the predicted execution state to obtain the execution state deviation.
[0027] Furthermore, in step S5), the event corresponding to the expert activation deviation record includes:
[0028] When an expert is classified as sleep in step S4), but is actually activated in a subsequent scheduling cycle, it is determined that there is an activation direction deviation and execution state deviation, and a sleep miss event is recorded.
[0029] When an expert is classified as standby in step S4), but is actually activated in a subsequent scheduling cycle, it is determined that there is an activation direction deviation and an execution state deviation, and a standby miss event is recorded.
[0030] When the future activation probability of an expert is lower than the hot expert threshold or the prediction uncertainty is higher than the uncertainty threshold, and the expert is activated by input with a word token number lower than the low word token threshold T_token_low in the subsequent scheduling cycle, and is executed by the in-memory computing unit PIM, the in-memory computing fallback event is recorded.
[0031] When an expert is not classified as NPU-active in step S4, but the actual number of tokens continues to increase in subsequent scheduling cycles, exceeding the high token threshold T_token_high, and is then migrated to NPU-active, an immediate NPU wake event is recorded.
[0032] Furthermore, in step S5), the future activation probability, low-power state risk, or state transition threshold are adjusted based on different recorded events. This adjustment is implemented by:
[0033] When a sleep miss event is recorded, increase the probability of the expert's future activation or low-power state risk, or decrease the priority of the expert entering a sleep state.
[0034] When a standby miss event is recorded, the subsequent prediction uncertainty of the expert is increased, the standby retention period is adjusted, or the priority of the expert entering the in-memory computing unit active state (PIM-active) is increased.
[0035] When the in-memory computation fallback event occurs frequently and the number of corresponding expert tokens continues to increase, increase the priority of migrating the expert to the neural network processing unit active state NPU-active, or adjust the in-memory computation PIM fallback execution ratio;
[0036] When recording an immediate NPU wake event, increase the priority of the expert's subsequent future activation probability, low-power state risk, or NPU-active state.
[0037] When an expert remains inactive for an extended period without any missed events, reduce the risk of the expert entering a low-power state, decrease the standby retention period, or increase the priority of the expert entering a sleep state.
[0038] 2. A heterogeneous low-power scheduling system based on the MoE model, characterized in that it comprises:
[0039] The expert activation information collection module is used to collect expert activation information during the inference process of the hybrid expert MoE model, including the number of tokens allocated to each expert in the current scheduling cycle, historical activation information, inter-layer routing information, or request-level expert activation information.
[0040] The expert future activation probability prediction module is used to calculate the future activation probability of each expert in the subsequent scheduling cycle based on the expert activation information, and to determine the prediction uncertainty corresponding to the future activation probability.
[0041] The low-power state risk assessment module is used to calculate the low-power state risk of each expert based on the future activation probability, prediction uncertainty, expert wake-up cost, misprediction penalty, and in-memory PIM fallback execution cost.
[0042] The four-state scheduling module is used to divide each expert into any one of the following states based on the future activation probability, prediction uncertainty and low power consumption risk: neural network processing unit active state (NPU-active), in-memory computing unit active state (PIM-active), standby state, and sleep state, and generate expert execution dispatch instructions based on the state division results.
[0043] The in-memory computing PIM fallback execution module is used to generate an in-memory computing PIM fallback execution request when the expert prediction is uncertain or the expert is activated by an input with a token count lower than the low token threshold T_token_low, and dispatch the corresponding expert's computing task to the in-memory computing unit PIM.
[0044] The feedback correction module is used to record sleep miss, standby miss, PIMfallback event, or immediate NPU wake event based on the deviation between the actual expert activation result and the prediction result, and to correct the future activation probability, low power state risk, or state switching threshold based on the events.
[0045] Compared with the prior art, the present invention has the following advantages:
[0046] Firstly, this invention comprehensively utilizes the number of tokens allocated to each expert in the current scheduling cycle, historical activation information, inter-layer expert transfer relationships, request-level expert activation profiles, and feedback correction information to calculate the future activation probability and prediction uncertainty of each expert in subsequent scheduling cycles, rather than scheduling based solely on the hotness or coldness of experts in the current scheduling cycle. Therefore, it can identify experts that may be reactivated in advance, improving the foresight and accuracy of expert low-power state control.
[0047] Secondly, since the present invention incorporates the future activation probability of experts, prediction uncertainty, expert wake-up cost, misprediction penalty, and the PIM fallback execution cost of in-memory computing units into the low-power state risk assessment, it can quantify the wake-up delay, additional energy consumption, and computation migration overhead that may occur after experts enter the sleep state, thereby reducing the problems of false sleep and frequent wake-up caused by directly putting experts to sleep based solely on fixed hot and cold thresholds.
[0048] Third, this invention divides each expert into an active state (NPU), an active state (PIM), a standby state, and a sleep state, and performs high-performance computing, low-power near-memory computing, fast recovery and retention, and deep low-power control respectively. Therefore, it can dynamically balance computing performance, recovery latency and energy consumption at the expert level, which has higher scheduling accuracy than uniformly implementing low-power control for the entire computing unit.
[0049] Fourth, this invention does not directly put experts with a future activation probability lower than the hot expert threshold and high prediction uncertainty into a dormant state. Instead, it puts them into a PIM-active or standby state. When the expert is activated by an input with a token number lower than the low token threshold T_token_low, the corresponding calculation is preferentially performed by the in-memory computing unit PIM. Therefore, PIM can be used as a low-power fallback path in prediction uncertainty scenarios, reducing unnecessary immediate wake-up of the neural network processing unit NPU and the resulting latency jitter and increased energy consumption.
[0050] Fifth, this invention compares the actual activation state of each expert, the actual number of tokens allocated, and the actual execution state with the prediction results, and records sleep miss, standby miss, PIM fallback event, and immediate NPU wake event based on the comparison results. This corrects future activation probabilities, prediction uncertainties, low-power state risks, or state switching thresholds, thus enabling it to adapt to different input requests and expert activation modes and improve the stability and adaptability of heterogeneous low-power scheduling.
[0051] Sixth, the future activation probability calculation weights, low term threshold T_token_low, high term threshold T_token_high, hot expert threshold, uncertainty threshold, risk threshold, and state preservation parameters in this invention can all be configured or updated according to the model structure, number of experts, scheduling cycle length, NPU / PIM hardware characteristics, and online feedback results. Therefore, it has good platform adaptability and scalability, and can be used in hybrid expert MoE type systems of different scales and NPU-PIM heterogeneous inference systems of different structures. Attached Figure Description
[0052] Figure 1This is a flowchart of the heterogeneous low-power scheduling method based on the hybrid expert MoE model of the present invention, which is based on expert activation risk assessment.
[0053] Figure 2 This is a schematic diagram illustrating the state transition and feedback correction process in the method of this invention.
[0054] Figure 3 This is the timing diagram of the PIM fallback execution of the in-memory computing unit based on the scheduling cycle in the method of this invention;
[0055] Figure 4 This is a block diagram of the heterogeneous low-power scheduling system based on the hybrid expert MoE model of the present invention, which is based on expert activation risk assessment. Detailed Implementation
[0056] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0057] Example 1: A heterogeneous low-power scheduling method based on a hybrid expert MoE model with expert activation risk assessment.
[0058] Terminology Explanation
[0059] To facilitate understanding of this invention, the main terms used in the specification are explained below. Unless otherwise stated, the following terms should not be construed as limiting the scope of protection of this invention.
[0060] MoE model refers to a neural network model that contains multiple expert modules and selects some experts to perform computations for the input token through a router or gating network.
[0061] Experts refer to the sub-networks, feedforward network modules, parameter blocks, or computation modules in the MoE model that can be selected by the router to perform computations.
[0062] The scheduling cycle refers to the basic time granularity at which one expert activation statistics, future activation probability prediction, prediction uncertainty calculation, low-power state risk assessment, expert state partitioning, and NPU / PIM execution dispatch are performed in this invention. The scheduling cycle can correspond to an expert routing and execution process for one MoE layer, one decoding step, or one micro-batch, depending on the specific implementation.
[0063] The token threshold T_token_low is a preset threshold used to determine whether an expert is in a low-token activation scenario. When the number of tokens allocated to an expert in the current or subsequent scheduling cycles is lower than the token threshold T_token_low, the expert is considered to meet the low-token activation condition. The token threshold T_token_low can be configured or updated based on the number of experts, micro-batch size, hardware execution latency, NPU / PIM energy consumption parameters, or online feedback results.
[0064] The token threshold T_token_high is a preset threshold used to determine whether an expert is suitable for migration to the NPU-active state. When the number of tokens allocated to an expert in the current scheduling cycle or multiple consecutive scheduling cycles is higher than the token threshold T_token_high, the expert is considered to meet the high token activation condition and can be prioritized for execution by the NPU.
[0065] Expert activation information refers to information reflecting the invocation of experts during the inference process, including one or more of the following: number of expert tokens, historical activation frequency, inter-layer transfer relationship, request-level activation profile, expert state retention time, or the interval between the most recent activation.
[0066] Future activation probability refers to the likelihood that an expert will be invoked in a subsequent scheduling cycle, as predicted based on expert activation information.
[0067] Prediction uncertainty refers to the reliability of future activation probability predictions, which can be obtained from historical activation fluctuations, the degree of divergence between different prediction components, or statistics on recent mispredicted events.
[0068] Low-power state risk refers to the cost of misprediction that may occur when an expert enters a low-power state, including wake-up latency, additional power consumption, computation migration overhead, or inference latency penalty.
[0069] PIM fallback execution refers to the process where, when an expert predicts a low probability of activation but has high uncertainty, or when an expert is suddenly activated by a small number of tokens, PIM is given priority to perform the expert's calculations to avoid immediately waking up the NPU or other high-power computing units.
[0070] A sleep miss refers to an event in which an expert is put into a sleep state but is actually activated in a subsequent scheduling cycle.
[0071] A standby miss refers to an event in which an expert, after being placed in a standby state, is actually activated in a subsequent scheduling cycle.
[0072] A PIM fallback event refers to an event in which PIM is used as a fallback mechanism when experts make uncertain predictions or when a small number of tokens are suddenly activated.
[0073] An immediate NPU wake event refers to an event in which an expert is migrated to the NPU-active state due to a continuous increase or sudden surge in the number of tokens.
[0074] Reference Figure 1 The implementation steps of this embodiment include the following:
[0075] Step 1: Obtain expert activation information during the inference process of the hybrid expert MoE model.
[0076] The input token is routed through the Router or Gate of the hybrid expert MoE model to obtain the expert routing results corresponding to each input token;
[0077] Collect expert activation information to obtain the expert routing results, and count the number of tokens allocated to each expert in the current scheduling period;
[0078] The system can acquire one or more of the following information: the number of times or frequency of activation for each expert within a preset historical window; the expert transfer relationship between adjacent MoE layers; the number of times each expert is repeatedly activated within the current input request; the previous state, current state, state holding time, and the interval between the most recent activations for each expert.
[0079] Step 2: Calculate the future activation probability of each expert and determine the corresponding prediction uncertainty.
[0080] 2.1) Based on the expert activation information, calculate the future activation probability of each expert in subsequent scheduling cycles:
[0081] 2.1.1) Statistical experts The number of tokens allocated in the current scheduling period. In conjunction with the historical heat of the previous scheduling cycle Computational experts During the scheduling period Historical popularity :
[0082] ,in, Historical update coefficient;
[0083] 2.1.2) Statistics of the first Experts in each MoE layer After being activated, the first Experts in each MoE layer Conditional probability of activation :
[0084] ;
[0085] 2.1.3) Based on the current scheduling cycle, the [number]th [item]... The routing weights or token proportions of activated experts in each MoE layer are used to weight and sum the transition probabilities between layers to obtain the expert's value. Inter-layer expert transfer prediction score :
[0086] ,
[0087] in, Indicates the number of times within the current scheduling period The set of experts activated in each MoE layer; Experts The percentage of routing weight, activation weight, or keyword token in the current scheduling cycle;
[0088] 2.1.4) Statistical analysis of the same input request across different scheduling cycles or different MoE layers for experts. The activation frequency, number of repeated activations, cumulative number of tokens, number of times the expert set is reused, and one or more of the most recent activation interval are considered. This information is normalized, and a weighted calculation is performed based on the weight of each piece of information to obtain the expert... Request level expert activation profile score ;
[0089] 2.1.5) Statistical experts The events that occur within the preset history window include sleep miss, standby miss, PIM fallback event, and immediate NPU wake event.
[0090] When a sleep miss occurs, increase the feedback correction term. .
[0091] When a standby miss occurs, increase the feedback correction item. .
[0092] When frequent in-memory computation fallback events (PIM fallback events) occur, and the actual number of tokens continues to increase, increase the feedback correction term. .
[0093] When an immediate NPU wake event occurs, increase the feedback correction term. .
[0094] When it is not activated and no miss event occurs within multiple consecutive scheduling cycles, reduce .
[0095] 2.1.6) Obtaining historical popularity among experts Inter-layer expert transfer prediction score Request expert to activate profile score and feedback correction items Then, this information is weighted and fused to obtain expert results. Probability of future activation in the next scheduling cycle:
[0096] ;
[0097] in, This represents the weighting parameter corresponding to the expert's historical popularity. This represents the weighting parameter corresponding to the inter-layer expert transfer prediction score. This indicates the weighting parameter corresponding to the request-level expert activation profile score. This indicates the weight parameters corresponding to the feedback correction items. These weight parameters can be configured or updated based on the model structure, the number of experts, the scheduling cycle length, or the online feedback results.
[0098] 2.2) Determine the prediction uncertainty corresponding to the future activation probability;
[0099] 2.2.1) Statistical experts Within a preset historical window, the degree of discrepancy between historical heat, inter-layer expert transfer prediction score, and request-level expert activation profile score is calculated based on activation fluctuations. The frequency of recent sleep misses, standby misses, and PIM fallback events is also counted. Based on one or more of this information, a prediction uncertainty index is obtained. The greater the activation fluctuation of experts within the preset historical window, the higher the prediction uncertainty index; the greater the difference between different prediction components, the higher the prediction uncertainty index; the higher the number or frequency of recent mispredicted events, the higher the prediction uncertainty index.
[0100] 2.2.2) Indicators of forecast uncertainty Compared with the preset uncertainty threshold By comparing the results, we can determine the level of forecast uncertainty.
[0101] when At that time, the experts made the judgment. Predictive uncertainty High;
[0102] when At that time, the experts made the judgment. Predictive uncertainty Lower.
[0103] Step 3: Calculate the low-power state risk for each expert.
[0104] The low-power state risk is a risk calculated by comprehensively considering the expert's future activation probability and prediction uncertainty, the wake-up cost after the expert enters sleep mode, the misprediction penalty, and the cost reduction that can be achieved by the in-memory computing unit PIM fallback execution; this embodiment uses... Experts The low-power state risk is calculated as follows:
[0105] 3.1) Obtaining Experts The information includes one or more of the following: wake-up latency, parameter reload overhead, cache recovery overhead, and additional energy consumption required to resume execution from a sleep state. This information is normalized, and then weighted according to the weights of each piece of information to obtain an expert... The cost of awakening experts ;
[0106] 3.2) Experts The scenario where a device enters a sleep state and is subsequently reactivated in a later scheduling cycle is considered a misprediction. The calculation of this misprediction scenario relative to the expert's prediction is then performed. The expert is calculated by normalizing one or more of the following factors when not in a dormant state: increased inference latency, state switching energy consumption, and computational migration overhead. Then, a weighted average is calculated based on the weights of each factor to obtain the expert's result. Misprediction penalty ;
[0107] 3.3) Determine whether the in-memory computing unit (PIM) currently has available computing resources, and whether the in-memory computing unit (PIM) can complete the expert evaluation within the preset execution delay. Correspondingly, based on one or more of the following information—the resource availability of the in-memory computing unit (PIM), the execution queue status, the expected execution latency, and the expected execution energy consumption—experts are obtained. The corresponding in-memory computing unit (PIM) provides fallback execution availability. ;
[0108] 3.4) Computational Expert The execution latency and power consumption generated during execution by the in-memory computing unit (PIM) are compared with those of the immediately wake-up neural network processing unit (NPU) or other high-power computing units. By comparing the wake-up latency, execution latency, and execution power consumption incurred during computation, the cost reduction achieved by using the in-memory compute unit (PIM) as a fallback execution compared to immediately waking up the high-power compute unit is obtained. ;
[0109] 3.5) Based on the future activation probabilities obtained above Forecast uncertainty The cost of awakening experts Misprediction penalty PIM (Power In-Memory) computing unit backstop for availability. The cost reduction achieved by the in-memory computing unit (PIM) as a fallback execution mechanism. Computational experts Low-power state risks of entering sleep mode :
[0110] ,
[0111] 3.6) Experts Low power state risk With risk threshold Compare them.
[0112] when At that time, the experts made the judgment. The low-power state poses a high risk, indicating that experts will... When placed in sleep mode, there is a higher probability of mispredictions and additional delays or energy consumption, according to experts. The risk conditions for entering a sleep state are not met;
[0113] when At that time, the experts made the judgment. The low-power state poses a low risk, indicating that experts will... The cost of mispredictions after putting the device into a sleep state is within an acceptable range, according to experts. The risk conditions for entering a sleep state are met.
[0114] Step 4: Divide the system into four states based on the future activation probability, prediction uncertainty, and low-power state risk, and perform hybrid expert MoE inference based on the division results.
[0115] The hybrid expert MoE inference refers to the process of dispatching expert computing tasks in the neural network processing unit (NPU-active) and the in-memory computing unit (PIM-active) to the NPU and PIM respectively, based on the expert routing results output by the router or gate network and the state division results of each expert. Standby and sleep control are performed on experts in the standby and sleep states respectively. Finally, the results of the completed expert computing tasks are aggregated according to the expert routing weights.
[0116] This step includes:
[0117] 4.1) Obtain thermal expert thresholds through offline calibration or online updates of the target hardware platform. Memory-based computing unit (PIM) execution threshold Sleep threshold Uncertainty threshold and risk threshold These are configurable thresholds;
[0118] 4.2) Based on the expert future activation probability obtained above Forecast uncertainty and low power state risks Each expert is classified into one of the following states according to the following rules: NPU-active (Neural Processing Unit Active), PIM-active (In-Memory Computing Unit Active), standby, and sleep:
[0119] like Then the expert enters the NPU-active state;
[0120] like Then the expert enters the PIM-active state;
[0121] like and Then the experts enter standby mode;
[0122] like , and Then the expert enters sleep mode;
[0123] 4.3) Reference Figure 2 Based on the state classification results of each expert, the state of the previous scheduling cycle, the actual number of tokens, and the state transition triggering conditions, the state of each expert is updated and transitioned to determine the execution state of each expert in the current scheduling cycle:
[0124] When the number of tokens assigned to an expert in the PIM-active state continues to increase and reaches the high token threshold T_token_high, or when its future activation probability reaches the hot expert threshold. When this happens, the expert will be transitioned from the PIM-active state to the NPU-active state.
[0125] When an expert in the NPU-active state experiences a decrease in popularity, a reduction in the number of assigned tokens, and no longer meets the criteria for NPU-active state, the expert is migrated from the NPU-active state to the PIM-active state.
[0126] When the probability of future activation of an expert in the PIM-active state decreases and the prediction uncertainty is higher than the uncertainty threshold, the expert is transferred from the PIM-active state to the standby state.
[0127] When an expert in the standby state experiences a standby miss event, or is activated by an input with a token count lower than the low token threshold T_token_low, the expert is transitioned from the standby state to the PIM-active state.
[0128] When an expert in the standby state is not activated for multiple consecutive scheduling cycles, and the recalculated future activation probability, prediction uncertainty, and low-power state risk meet the conditions for dividing the sleep state, the expert is transferred from the standby state to the sleep state.
[0129] When an expert in the sleep state experiences a sleep miss and is activated by an input with a token count lower than the low token threshold T_token_low, the expert is transitioned from the sleep state to the PIM-active state.
[0130] When an expert is effectively executed by the in-memory computing unit PIM, and the number of tokens assigned to the expert remains below the low token threshold T_token_low, the expert remains in the PIM-active state.
[0131] 4.4) Based on the expert's current state, generate corresponding expert execution dispatch instructions or low-power control instructions, and perform hybrid expert MoE inference according to the instructions:
[0132] When an expert is in the NPU-active state, an NPU execution dispatch instruction is generated to dispatch the computation task corresponding to that expert to the NPU for execution, while keeping the corresponding computation path, parameter loading path, and control path executable. This type of expert computation task can be a task where the number of tokens allocated in the current scheduling period exceeds the high token threshold T_token_high, or a task where the computational overhead percentage exceeds a preset computational overhead percentage threshold, or the expected execution latency exceeds a preset latency threshold. The computational overhead percentage can be the proportion of matrix computation overhead, cumulative computation overhead, or activation computation overhead in the expert's total execution overhead.
[0133] When an expert is in the PIM-active state, a PIM execution dispatch instruction is generated, dispatching the computation task corresponding to that expert to PIM execution. This type of expert computation task can be a task where the number of tokens allocated in the current scheduling period is less than the low token threshold T_token_low, or a task where the memory access overhead percentage is higher than a preset memory access percentage threshold. The memory access overhead percentage can be the proportion of parameter read overhead, cache access overhead, or data migration overhead in the expert's total execution overhead.
[0134] When an expert is in standby mode, a standby control command is generated to shut down the high-power computing paths corresponding to that expert, while retaining the status register, wake-up control path, or PIM fallback execution path, so that the expert can quickly resume or be executed by PIM when actually activated.
[0135] When an expert is in a sleep state, a sleep control command is generated to shut down or reduce at least a portion of the corresponding computation path, parameter access path, cache retention path, clock path, or power supply path, so that the expert enters a deep low-power state.
[0136] 4.5) For experts with a low probability of future activation but high predictive uncertainty, according to Figure 3 The in-memory computation unit (PIM) based on the scheduling cycle is used for fallback execution timing, as shown below:
[0137] When an expert is activated by an input with a token count below the low token threshold T_token_low, the four-state scheduling module initiates a PIM fallback execution request to the in-memory computing unit (PIM), prioritizing the assignment of the computation task corresponding to that expert to the PIM. The PIM executes the computation task for that expert and returns the result to the four-state scheduling module or the subsequent inference pipeline of the hybrid expert MoE model, thus avoiding the immediate activation of the neural network processing unit (NPU) due to a small number of token activations.
[0138] When the number of tokens assigned to an expert continues to increase in subsequent scheduling cycles, exceeding the high token threshold T_token_high, and the migration conditions are met, the four-state scheduling module initiates a migration request to the neural network processing unit (NPU) to switch the expert to the NPU-active state, and the NPU takes over the subsequent calculations for that expert.
[0139] Step 5: Calculate the deviation between the actual inference result and the prediction result of the hybrid expert MoE, record the feedback event and make corrections.
[0140] 5.1) After the current scheduling period ends, obtain the actual activation status, actual number of tokens allocated, and actual execution status of each expert, and make the following comparisons:
[0141] The actual activation state is compared with the predicted activation level to obtain the deviation of the actual activation state.
[0142] The actual number of tokens allocated is compared with the low token threshold T_token_low and the high token threshold T_token_high to obtain the expert's actual activation level;
[0143] The actual execution state is compared with the predicted execution state to obtain the execution state deviation.
[0144] 5.2) Reference Figure 2 and Figure 3 Based on the above deviation records and the actual state transition of each expert, record the corresponding events:
[0145] When an expert is classified as sleep in step 4, but is actually activated in a subsequent scheduling cycle, it is determined that the expert has an activation direction deviation and an execution state deviation, and a sleep miss event is recorded.
[0146] When an expert is classified as standby in step 4, but is actually activated in a subsequent scheduling cycle, it is determined that the expert has an activation direction deviation and an execution state deviation, and a standby miss event is recorded.
[0147] When the future activation probability of an expert is lower than the hot expert threshold or the prediction uncertainty is higher than the uncertainty threshold, and the expert is activated by input with a token number lower than the low token threshold T_token_low in the subsequent scheduling cycle and is executed by PIM, the PIM fallback event is recorded in memory.
[0148] When an expert is not classified as NPU-active in step 4, but their actual number of tokens continues to increase, surges, or exceeds the high token threshold T_token_high in subsequent scheduling cycles, and they are migrated to the NPU-active state, an immediate NPU wake event is recorded.
[0149] 5.3) Reference Figure 2 The state transition and feedback correction relationships shown are adjusted based on the recorded event type to adjust future activation probability, prediction uncertainty, low-power state risk, or state switching threshold:
[0150] When a sleep miss is recorded, the probability of the corresponding expert activating in the future or the risk of entering a low-power state are increased, or the priority of the expert entering a sleep state is reduced.
[0151] When a record is marked as a standby miss, the subsequent prediction uncertainty of the corresponding expert is increased, the standby status retention period is adjusted, or the priority for the expert to enter the PIM-active state is increased.
[0152] When a PIM fallback event occurs frequently within a preset history window and the number of tokens for the corresponding expert continues to increase, the priority for migrating the expert to the NPU-active state is increased, or the PIM fallback execution ratio is adjusted.
[0153] When recorded as an immediate NPU wake event, the priority of the corresponding expert in subsequent future activation probability, low-power state risk, or NPU-active state classification is increased;
[0154] If an expert is not activated and no miss event occurs within several consecutive scheduling cycles, reduce the risk of the expert's low-power state, shorten its standby state retention period, or increase the priority for the expert to enter the sleep state.
[0155] 5.4) The corrected feedback correction item, prediction parameter or state switching threshold shall be used in the next scheduling cycle.
[0156] Step 6: Determine whether the current hybrid expert MoE inference task is completed.
[0157] Whether the current hybrid expert MoE inference task is completed is determined based on whether all input tokens corresponding to the current inference task have completed expert routing, expert computation, and computation result aggregation, and whether there are still unfinished neural network processing unit (NPU) tasks, in-memory computing unit (PIM) tasks, or expert computation results that have not yet been returned. Its implementation includes:
[0158] When the hybrid expert MoE inference task is completed, the calculation results of each activated expert are aggregated according to the routing weights output by the Router or Gate to obtain the inference result of the hybrid expert MoE model, and the current scheduling process ends.
[0159] If the hybrid expert MoE inference task has not been completed, proceed to the next scheduling cycle, and based on the feedback correction results obtained in step 5, re-execute steps 1 to 5 until the current hybrid expert MoE inference task is completed.
[0160] Through the aforementioned cyclic scheduling method, the system can dynamically adjust the future activation probability of experts, prediction uncertainty, low-power state risk, and state partitioning results according to the changes in expert activation modes within different scheduling cycles. This reduces the ineffective energy consumption of experts that are not frequently activated while ensuring that inference latency is controllable.
[0161] It should be noted that the step numbers in this specification and claims are only for the purpose of clearly describing the embodiments of the present invention and facilitating understanding, and their order is not limited.
[0162] Example 2: Heterogeneous low-power scheduling system based on expert activation risk assessment using a hybrid expert MoE model.
[0163] Reference Figure 4This embodiment includes: an expert activation information acquisition module 1, an expert future activation probability prediction module 2, a low-power state risk assessment module 3, a four-state scheduling module 4, an in-memory computing PIM fallback execution module 5, and a feedback correction module 6. The four-state scheduling module 4 includes: a neural network activation state scheduling submodule 41, an in-memory computing activation state scheduling submodule 42, a standby state scheduling submodule 43, and a hibernation state scheduling submodule 44.
[0164] The working principle of the entire system is as follows:
[0165] The expert activation information acquisition module 1 is used to collect expert activation information during the inference process of the hybrid expert MoE model. This includes: collecting expert routing results and counting the number of tokens allocated to each expert within the current scheduling period. The expert routing results are obtained by routing the input tokens through the router or gate network of the hybrid expert MoE model; acquiring one or more of the following information: the number of activations or activation frequency of each expert within a preset historical window, the expert transition relationship between adjacent MoE layers, the number of repeated activations of each expert within the current input request, the previous state, current state, state holding time, and the interval between the most recent activations of each expert; and sending the collected expert activation information to the expert future activation probability prediction module 2.
[0166] The expert future activation probability prediction module 2 is used to calculate the future activation probability and corresponding prediction uncertainty of each expert based on the expert activation information such as historical popularity and inter-layer expert transfer prediction score transmitted by the expert activation information acquisition module 1, as well as the feedback correction information generated by the feedback correction module 6 in the previous scheduling cycle. The obtained future activation probability and prediction uncertainty are then sent to the low power state risk assessment module 3, the four-state scheduling module 4, and the feedback correction module 6, respectively.
[0167] The low-power state risk assessment module 3 is used to calculate the low-power state risk of each expert based on the future activation probability and prediction uncertainty output by the expert future activation probability prediction module 2, compare the low-power state risk with the risk threshold, determine whether each expert meets the conditions for entering the sleep state, and send the comparison result to the four-state scheduling module 4.
[0168] The four-state scheduling module 4 is used to perform state scheduling based on the future activation probability, prediction uncertainty, and low-power state risk of each expert. Specifically, it categorizes each expert into any one of the following states: Neural Processing Unit (NPU) active state, In-Memory Computing Unit (PIM) active state, Standby state, and Sleep state. It generates corresponding execution dispatch instructions or low-power control instructions and, based on the expert's current state, the actual number of allocated tokens, and prediction uncertainty, determines whether the expert meets the PIM fallback execution condition. If the condition is met, it sends PIM fallback execution trigger information to the PIM fallback execution module 5.
[0169] The neural network processing unit activation state scheduling submodule 41 is used to generate a neural network processing unit NPU execution dispatch instruction when an expert is classified as an NPU-active state, dispatch the computation task corresponding to the expert to the neural network processing unit NPU for execution, and keep the computation path, parameter loading path and control path corresponding to the expert in an executable state.
[0170] The in-memory computing unit activation state scheduling submodule 42 is used to generate an in-memory computing unit PIM execution dispatch instruction when an expert is classified as an in-memory computing unit active state PIM-active, and dispatch the computing task corresponding to the expert to the in-memory computing unit PIM for execution.
[0171] The standby state scheduling submodule 43 is used to generate standby control instructions when an expert is classified as standby, shut down some high-power computing paths corresponding to the expert, and retain the status register, wake-up control path or in-memory computing unit PIM fallback execution path; when an expert in standby state is actually activated in a subsequent scheduling cycle, and its actual token allocation number is lower than the low token threshold T_token_low, it is determined that the expert meets the PIM fallback execution condition;
[0172] The sleep state scheduling submodule 44 is used to generate sleep control instructions when an expert is classified as a sleep state, to shut down or reduce at least a portion of the corresponding computation path, parameter access path, cache retention path, clock path, or power supply path of the expert, so that the expert enters a deep low-power state; when an expert in a sleep state is actually activated in a subsequent scheduling cycle, and its actual token allocation number is lower than the low token threshold T_token_low, it is determined that the expert meets the PIM fallback execution condition.
[0173] The in-memory computation unit PIM fallback execution module 5 is used to perform fallback execution on experts with uncertain predictions or low-token activation experts according to the PIM fallback execution trigger information sent by the four-state scheduling module 4. After the in-memory computation unit PIM completes the computation task corresponding to the expert, it returns the computation result to the subsequent inference pipeline of the hybrid expert MoE model. Then, it aggregates the computation results of each expert according to the expert routing weights output by the router or gate network to obtain the actual inference result of the current scheduling cycle, and transmits it to the feedback correction module 6.
[0174] The feedback correction module 6 is used to compare the actual reasoning results of each expert in the current scheduling cycle with the prediction results output by the expert future activation probability prediction module 2. Based on the comparison results, it records sleep miss, standby miss, in-memory computation fallback event, or immediate NPU wake event. Based on the recorded events, it corrects parameters such as expert future activation probability and prediction uncertainty, and feeds the feedback correction information back to the expert future activation probability prediction module 2, low power state risk assessment module 3, and four-state scheduling module 4 for use in the next scheduling cycle.
[0175] It should be noted that the above functional modules can be implemented, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, they can be implemented, in whole or in part, as a program instruction product. A program instruction product includes one or a set of program instructions. When the program instructions are loaded and executed on a computer, all or part of the described process or function is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The program instructions can be stored in a computer-readable and writable storage medium or transferred from one computer-readable and writable storage medium to another.
[0176] In this embodiment, the direct coupling or communication connection between the modules can be achieved through indirect coupling or communication connection via interfaces, devices, or modules. The functional modules and sub-modules in this embodiment can dynamically reside within a single processing unit, or each module can exist physically independently, or two or more modules can dynamically reside within a single processing unit. When these dynamic components are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable and writable storage medium. This storage medium can be a memory, disk, or optical disc, etc.
[0177] The effectiveness of this invention can be further illustrated by the following simulation results:
[0178] I. Simulation Conditions
[0179] Thirty English natural language inputs were used as test samples, and expert activation trajectories were derived based on a hybrid expert MoE model containing 64 experts. The expert activation trajectories include an expert counting trajectory (CountTrace) and a token-level routing trajectory (TokenRouteTrace). The CountTrace records the number of tokens assigned to each expert within each scheduling period, while the TokenRouteTrace records the expert IDs assigned to the same token in different MoE layers.
[0180] Evaluation metrics include average latency, total energy consumption, average power consumption, average number of experts in PIM-active state, average number of experts in standby state, and average number of experts in sleep state. Average latency, total energy consumption, and average power consumption are normalized evaluation values from behavioral simulations.
[0181] II. Simulation Content
[0182] Under the above simulation conditions, behavioral-level low-power scheduling simulations were performed using the method of this invention and the all-NPU (all-neural network processing unit) and all-PIM (all-in-memory computing unit) execution strategies, respectively, and the average latency, total energy consumption, average power consumption and expert state distribution were obtained, as shown in Table 1.
[0183] Table 1. Performance comparison of the present invention and existing methods in low-power scheduling simulation.
[0184]
[0185] As shown in Table 1, compared with the all-neural network processing unit (all-NPU) execution strategy, the average latency of this invention is reduced from 0.2080 to 0.1569, a reduction of approximately 24.57%; total energy consumption is reduced from 790.9440 to 420.8942, a reduction of approximately 46.79%; and average power consumption is reduced from 21.5322 to 17.4961, a reduction of approximately 18.74%. Compared with the all-in-memory computing unit (all-PIM) execution strategy, the total energy consumption of this invention is reduced from 588.9840 to 420.8942, a reduction of approximately 28.54%; and average power consumption is reduced from 29.6798 to 17.4961, a reduction of approximately 41.05%. Although the average latency of this invention is higher than that of the all-in-memory computing unit (all-PIM) execution strategy, it is still lower than that of the all-neural network processing unit (all-NPU) execution strategy, indicating that the method of this invention can achieve a balance between inference latency and energy consumption. Furthermore, in this invention, an average of 29.0 experts are in the PIM-active state, 0.3944 experts are in the standby state, and 34.6056 experts are in the sleep state. Among them, experts in the sleep state account for approximately 54.07% of all 64 experts, indicating that this invention can enable experts who are not frequently activated to enter a low-power state based on the expert's future activation probability, prediction uncertainty, and low-power state risk, and utilize the in-memory computing unit PIM to process experts with low-term or prediction uncertainties.
[0186] The simulation results above show that the present invention can utilize the sparse activation characteristics of the hybrid expert MoE model, reduce the ineffective energy consumption in the model inference process through expert activation risk assessment, four-state scheduling, in-memory computing unit PIM fallback execution and feedback correction, and achieve low-power inference while maintaining acceptable inference latency.
[0187] The above description is merely a specific example of the present invention and does not constitute any limitation on the present invention. Obviously, those skilled in the art, after understanding the content and principles of the present invention, may make various modifications and changes in form and details without departing from the principles and structure of the present invention. However, these modifications and changes based on the ideas of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. A heterogeneous low-power scheduling method based on a hybrid expert MoE model using expert activation risk assessment, characterized in that, include: S1) Obtain expert activation information during the inference process of the hybrid expert MoE model, which includes at least the number of tokens assigned to each expert in the current scheduling period; S2) Based on the expert activation information, calculate the future activation probability of each expert in the subsequent scheduling cycle, and determine the prediction uncertainty corresponding to the future activation probability; S3) Calculate the low-power state risk of each expert based on the future activation probability, prediction uncertainty, expert wake-up cost, misprediction penalty, and in-memory computing unit PIM fallback execution cost. S4) Based on the future activation probability, prediction uncertainty and low power state risk, each expert is divided into any one of the following states: neural network processing unit active state (NPU-active), in-memory computing unit active state (PIM-active), standby state, and sleep state, and hybrid expert MoE inference is performed based on the state division results. S5) Calculate the deviation between the hybrid expert MoE inference results and the prediction results, use the deviation to record sleep miss, standby miss, in-memory computing fallback event, and immediate NPU wake event, and correct the future activation probability, low power state risk or state switching threshold based on these events. S6) Determine whether the current hybrid expert MoE inference task is complete: If completed, output the hybrid expert model inference result and end the current scheduling process; If not completed, proceed to the next scheduling cycle, that is, repeat S1) to S5 based on the correction result obtained in S5).
2. The method according to claim 1, characterized in that, In step S2), the future activation probability of each expert in subsequent scheduling cycles is calculated based on the expert activation information. This is done by calculating one or more of the following: expert historical popularity, inter-layer expert transfer prediction score, request-level expert activation profile score, and feedback correction term. ; in, Experts The probability of future activation in the next scheduling cycle Experts During the scheduling period Historical popularity Experts The number of tokens allocated during scheduling period t. This is the historical update coefficient. expert During the scheduling period Historical popularity; This indicates the inter-layer expert transfer prediction score. Experts The proportion of routing weight, activation weight, or term token within the current scheduling period. Indicates the first Layer experts After being activated, the first Layer experts The conditional probability of activation. Indicates the first Layer experts , Indicates the first Layer experts ; This indicates a request for an expert to activate the profile score. Indicates feedback correction items, This represents the weighting parameter corresponding to the expert's historical popularity. This represents the weighting parameter corresponding to the inter-layer expert transfer prediction score. This indicates the weighting parameter corresponding to the request-level expert activation profile score. This indicates the weight parameter corresponding to the feedback correction item. The weight parameter can be configured or updated according to the model structure, the number of experts, the scheduling cycle length, or the online feedback results.
3. The method according to claim 1, characterized in that, The determination of the prediction uncertainty corresponding to the future activation probability in S2) includes the following implementation: S2a) Statistical experts measure the activation fluctuations within a preset historical window, obtain the degree of divergence between different predicted components in the future activation probability, and count the number or frequency of recent mispredicted events. S2b) Based on one or more of the activation fluctuation, divergence degree, and the number or frequency of recent mispredicted events, a prediction uncertainty index is obtained; S2c) Compare the prediction uncertainty index with the uncertainty threshold: If the prediction uncertainty index is higher than the uncertainty threshold, then the prediction uncertainty is determined to be high. If the prediction uncertainty index is lower than or equal to the uncertainty threshold, then the prediction uncertainty is determined to be low.
4. The method according to claim 3, characterized in that: The occurrence of the recent mispredicted event includes at least one of the following: The occurrence of the sleep miss event indicates whether an expert is actually activated in subsequent scheduling cycles after being classified as a sleeper. The higher the number or frequency of this event within a preset historical window, the higher the prediction uncertainty index of the corresponding expert. The occurrence of the standby miss event indicates whether an expert is actually activated in subsequent scheduling cycles after being classified as a standby expert. The higher the number of times or frequency of this event occurs within a preset historical window, the higher the prediction uncertainty index of the corresponding expert. The occurrence of the in-memory computing fallback event (PIM) indicates the situation where the expert's prediction is uncertain or activated by input with a token count lower than the low token threshold (T_token_low), and the in-memory computing unit (PIM) executes the event. The higher the number of times or frequency of this event occurs within a preset historical window, the higher the corresponding expert's prediction uncertainty index.
5. The method according to claim 1, characterized in that, In step S3), the low-power state risk of each expert is calculated based on the future activation probability, prediction uncertainty, expert wake-up cost, misprediction penalty, and in-memory PIM fallback execution cost. The formula is as follows: ; in, This indicates the risk of expert e entering a sleep state. This indicates the probability of the expert being activated in the future. This represents the wake-up cost required for an expert to resume execution from a sleep state. Indicates forecast uncertainty. Indicates the penalty for misprediction. This indicates the availability of in-store PIM fallback execution. This indicates the cost reduction achieved by in-memory compute PIM fallback execution compared to immediately waking up high-power compute units.
6. The method according to claim 1, characterized in that: In step S4), the state is divided according to the future activation probability, prediction uncertainty and low power state risk, including comparing the expert future activation probability with the thermal expert threshold, the in-memory computing PIM execution threshold and the sleep threshold respectively. Compare the uncertainty of expert predictions with an uncertainty threshold; Compare the expert's low-power state risk with a risk threshold; In step S4), each expert is classified into any one of the following states: Neural Processing Unit Active (NPU-active), In-Memory Computing Unit Active (PIM-active), Standby, and Sleep, according to the following rules: If the future activation probability of an expert is higher than the hot expert threshold, then the expert is classified as the neural network processing unit activation state NPU-active. If the future activation probability of an expert is lower than the hot expert threshold but higher than the in-memory computing PIM execution threshold, then the expert is classified as the in-memory computing unit activation state PIM-active. If the future activation probability of an expert is lower than the in-memory PIM execution threshold and the prediction uncertainty is higher than the uncertainty threshold, then the expert is classified as standby. If an expert's future activation probability is lower than the sleep threshold, the prediction uncertainty is lower than the uncertainty threshold, and the low-power state risk is lower than the risk threshold, then the expert is classified as being in a sleep state.
7. The method according to claim 1, characterized in that, In S4), the hybrid expert MoE reasoning is performed on the divided state results, which involves different reasoning methods based on the different states in which the experts are located: When an expert is in sleep mode, at least one of the expert's corresponding computation path, parameter access path, cache retention path, clock path, or power supply path is turned off or reduced. When an expert is in standby mode, some computation paths are shut down, while the status register, wake-up control path, or in-memory computation PIM fallback execution path are retained. When an expert is in the in-memory computing unit active state (PIM-active), the in-memory computing unit (PIM) executes the expert's corresponding expert computing task. The expert computing task is either a task where the number of tokens allocated to the expert in the current scheduling period is lower than the low token threshold (T_token_low), or a task where the memory access overhead percentage is higher than a preset memory access percentage threshold. The memory access overhead percentage is the proportion of the parameter reading overhead, cache access overhead, or data migration overhead of the expert's corresponding computing task in the expert's total execution overhead. When an expert is in the NPU-active state, the NPU executes the expert's corresponding expert computation task. This expert computation task is either one where the number of tokens allocated to the expert in the current scheduling period exceeds a high token threshold T_token_high, or one where the computational overhead percentage exceeds a preset computational overhead percentage threshold or the expected execution delay exceeds a preset delay threshold. The computational overhead percentage is the proportion of the matrix computational overhead, cumulative computational overhead, or activation computational overhead of the expert's corresponding computation task within the expert's total execution overhead. For experts whose future activation probability is lower than the hot expert threshold and whose prediction uncertainty is higher than the uncertainty threshold, they are placed in the in-memory computing unit active state PIM-active or standby state; when the expert is activated by input with a token number lower than the low token threshold T_token_low in a subsequent scheduling cycle, the in-memory computing unit PIM will preferentially perform the corresponding calculation for the expert.
8. The method according to claim 1, characterized in that, In step S5), calculating the deviation between the hybrid expert MoE inference result and the prediction result involves comparing the actual activation state, actual token allocation quantity, and actual execution state of each expert inferred by the hybrid expert MoE in step S4) with the predicted activation level, predicted high and low token thresholds in step S2) and the predicted execution state in step S4), respectively, to obtain the state or numerical deviation between the two, where: The actual activation state is compared with the predicted activation level to obtain the deviation of the actual activation state. The actual number of tokens allocated is compared with the low token threshold T_token_low and the high token threshold T_token_high to obtain the actual activation level; The actual execution state is compared with the predicted execution state to obtain the execution state deviation.
9. The method according to claim 1, characterized in that, The S5) section, which records the corresponding events based on expert activation bias, includes: When an expert is classified as sleep in step S4), but is actually activated in a subsequent scheduling cycle, it is determined that there is an activation direction deviation and execution state deviation, and a sleep miss event is recorded. When an expert is classified as standby in step S4), but is actually activated in a subsequent scheduling cycle, it is determined that there is an activation direction deviation and an execution state deviation, and a standby miss event is recorded. When the future activation probability of an expert is lower than the hot expert threshold or the prediction uncertainty is higher than the uncertainty threshold, and the expert is activated by input with a word token number lower than the low word token threshold T_token_low in the subsequent scheduling cycle, and is executed by the in-memory computing unit PIM, the in-memory computing fallback event is recorded. When an expert is not classified as NPU-active in step S4, but the actual number of tokens continues to increase, surges, or exceeds the high token threshold T_token_high in subsequent scheduling cycles, and is migrated to NPU-active, an immediate NPU wake event is recorded.
10. The method according to claim 1, characterized in that, The process of adjusting the future activation probability, low-power state risk, or state transition threshold based on different recorded events in step S5 includes the following implementation: When a sleep miss event is recorded, increase the probability of the expert's future activation or low-power state risk, or reduce the priority of the expert entering a sleep state. When a standby miss event is recorded, the subsequent prediction uncertainty of the expert is increased, the standby retention period is adjusted, or the priority of the expert entering the in-memory computing unit active state (PIM-active) is increased. When the in-memory computation fallback event occurs frequently and the number of corresponding expert tokens continues to increase, increase the priority of migrating the expert to the neural network processing unit active state NPU-active, or adjust the in-memory computation PIM fallback execution ratio; When recording an immediate NPU wake event, increase the priority of the expert's subsequent future activation probability, low-power state risk, or NPU-active state. When an expert remains inactive for an extended period without any missed events, reduce the risk of the expert entering a low-power state, shorten the standby retention period, or increase the priority of the expert entering a sleep state.
11. A heterogeneous low-power scheduling system based on a hybrid expert MoE model using expert activation risk assessment, characterized in that, include: The expert activation information collection module is used to collect expert activation information during the inference process of the hybrid expert MoE model, including the number of tokens allocated to each expert in the current scheduling cycle, historical activation information, inter-layer routing information, or request-level expert activation information. The expert future activation probability prediction module is used to calculate the future activation probability of each expert in the subsequent scheduling cycle based on the expert activation information, and to determine the prediction uncertainty corresponding to the future activation probability. The low-power state risk assessment module is used to calculate the low-power state risk of each expert based on the future activation probability, prediction uncertainty, expert wake-up cost, misprediction penalty, and in-memory PIM fallback execution cost. The four-state scheduling module is used to divide each expert into any one of the following states based on the future activation probability, prediction uncertainty and low power consumption risk: neural network processing unit active state (NPU-active), in-memory computing unit active state (PIM-active), standby state, and sleep state, and generate expert execution dispatch instructions based on the state division results. The in-memory computing PIM fallback execution module is used to generate an in-memory computing PIM fallback execution request when the expert prediction is uncertain or the expert is activated by an input with a token count lower than the low token threshold T_token_low, and dispatch the corresponding expert's computing task to the in-memory computing unit PIM. The feedback correction module is used to record sleep miss, standby miss, in-memory computing fallback event, or immediate NPU wake event based on the deviation between the actual expert activation result and the prediction result, and to correct the future activation probability, low power state risk, or state switching threshold based on the event.
12. The system according to claim 11, characterized in that, The four-state scheduling module includes: The neural network activation state scheduling submodule is used to generate neural network processing unit execution dispatch instructions when an expert is classified as the neural network processing unit active state NPU-active, and dispatch the computation task corresponding to the expert to the neural network processing unit NPU for execution. The in-memory computing activation state scheduling submodule is used to generate an in-memory computing unit execution dispatch instruction when an expert is classified as an in-memory computing unit active state PIM-active, and dispatch the computing task corresponding to the expert to the in-memory computing unit PIM for execution. The standby state scheduling submodule is used to generate standby control instructions when an expert is assigned to the standby state, shut down some high-power computing paths corresponding to that expert, and retain the status register, wake-up control path or in-memory computing PIM fallback execution path. The sleep state scheduling submodule is used to generate sleep control instructions when an expert is assigned to a sleep state, and to shut down or reduce at least a portion of the corresponding computation path, parameter access path, cache retention path, clock path, or power supply path.
Citation Information
Patent Citations
Data processing method and system of low-energy-consumption large language model based on momentum mechanism and multiple types of experts
CN120450054A