Server core dynamic sleep scheduling system and method based on mixed load

CN122526754BActive Publication Date: 2026-09-29四川华鲲振宇智能科技有限责任公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610984376.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-29
Estimated Expiration
2046-07-03

AI Technical Summary

Technical Problem

当前企业广泛采用在线、离线混合负载部署模式,现有电源管理技术无法兼顾性能与能效,且难以适配异构CPU架构,亟需精细化、低开销、高适配的动态休眠调度方案

Benefits of technology

本发明提供了基于混合负载场景下的服务器核心动态休眠调度方案,通过非侵入式多维度采集、智能负载分类、精准空闲预测、差异化休眠决策及闭环性能保障等,可以在保证在线业务性能损失小于1%、采集开销低于1%的前提下,实现混合负载下精细化休眠调度,显著提升服务器能效,适配多种异构CPU架构。具体的,基于eBPF的非侵入式多维度采集,具有开销低、可扩展,支持超线程绑定采集的优点;设计的轻量级负载分类与延迟分级方案,具有分类准确率高,支持模型在线更新的特点;结合LSTM预测与多目标决策的休眠决策方法,可以适配异构CPU架构,权重可动态调整;以及提出了核心休眠与进程调度、软中断迁移的协同机制,含批量预唤醒策略;闭环式性能保障机制,可以确保在线业务性能损失小于1%;全局策略配置与异常处理机制,可以适配多种业务场景,保障系统稳定。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122526754B_ABST
    Figure CN122526754B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on the server core dynamic dormancy scheduling system and method under mixed load, belong to server dormancy scheduling technical field, including: multi-dimensional load characteristic perception module, for non-invasive load data acquisition;Load type and delay sensitivity identification module, for based on the data collected, through classification model, load type identification and delay tolerance classification are carried out to process;Core state decision module, for through the multi-objective decision based on performance, energy efficiency trade-off, in combination with timing prediction model, optimal dormancy state is independently decided for each core;Dynamic dormancy execution and collaborative scheduling module, for executing the dynamic switching of core dormancy state, process migration and batch pre-awakening;Closed-loop performance guarantee module, for monitoring the performance index of online business, establishes closed-loop performance feedback mechanism, adjusts dormancy strategy.The application can realize fine dormancy scheduling, improve server energy efficiency, and has high adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of server hibernation scheduling technology, and more specifically, to a dynamic hibernation scheduling system and method for server cores under mixed load. Background Technology

[0002] With the surge in data center energy consumption, CPU power consumption accounts for 40%-60% of total server power consumption. Modern CPUs support multiple hibernation levels (C1-C7), with deeper hibernation resulting in lower power consumption but longer wake-up latency. Currently, enterprises widely adopt online and offline hybrid load deployment models. Existing power management technologies cannot balance performance and energy efficiency and are difficult to adapt to heterogeneous CPU architectures. There is an urgent need for a refined, low-overhead, and highly adaptable dynamic hibernation scheduling solution.

[0003] The main drawbacks of existing technologies are: 1) Globalized hibernation strategy with low granularity: Existing technologies adopt a globally unified hibernation strategy, which cannot differentiate control based on the load characteristics of individual cores, does not adapt to hyper-threading and heterogeneous core characteristics, and results in serious energy waste. 2) Lack of load type differentiation and blind decision-making: Decisions are made solely based on CPU utilization and idle time, without distinguishing between online / offline load types, and without latency tolerance grading, which can easily affect online business response. 3) Uncontrollable wake-up latency and lack of performance guarantee: Low load prediction accuracy, passive wake-up can easily lead to latency spikes, and there is no batch wake-up strategy, often resulting in enterprises sacrificing energy efficiency for performance. 4) Poor coordination of mixed loads: Lack of deep coordination with process scheduling, failure to migrate soft interrupts before cores enter deep hibernation, making them susceptible to accidental wake-up. 5) Insufficient adaptation to heterogeneous architectures: Using uniform parameters for performance cores, energy efficiency cores, and different CPU architectures cannot leverage the energy efficiency advantages of the architecture. 6) Highly invasive data collection: Collection methods require code modification or module loading, resulting in high overhead and inability to achieve accurate multi-dimensional data collection.

[0004] Among the closest existing implementations, the main ones are the Linux kernel cpuidle subsystem and the Intel SpeedShift technology. The Linux kernel cpuidle subsystem selects a sleep state by monitoring kernel idle time, but its drawbacks include low prediction accuracy, lack of load classification, lack of support for heterogeneous adaptation, poor coordination, and limited data collection dimensions. The Intel SpeedShift technology controls sleep and frequency through hardware, but its disadvantages include globally uniform policies, lack of configurability, poor compatibility, lack of performance loop guarantees, and limited data collection methods. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a server core dynamic hibernation scheduling system and method based on mixed load, which realizes fine-grained hibernation scheduling under mixed load, improves server energy efficiency, and can be adapted to various heterogeneous CPU architectures.

[0006] The objective of this invention is achieved through the following approach: A server core dynamic hibernation scheduling system based on mixed load includes: a multi-dimensional load characteristic perception module, a load type and latency sensitivity identification module, a core state decision module, a dynamic hibernation execution and collaborative scheduling module, and a closed-loop performance guarantee module. The multi-dimensional load feature perception module is used for non-intrusive multi-dimensional load data acquisition. The load type and latency sensitivity identification module is used to identify the load type and classify the latency tolerance of processes based on the collected multi-dimensional load data using a pre-trained lightweight classification model. The core state decision module is used to make an independent decision on the optimal sleep state for each core through multi-objective decision-making based on performance and energy efficiency trade-offs, combined with a time-series prediction model, to adapt to heterogeneous CPU architecture. The dynamic hibernation execution and collaborative scheduling module is used to perform dynamic switching of core hibernation states, process migration, and batch pre-wake-up. The closed-loop performance assurance module is used to monitor the performance indicators of online services, establish a closed-loop performance feedback mechanism, and dynamically adjust the hibernation strategy.

[0007] Furthermore, the non-intrusive load data acquisition is specifically implemented using an eBPF module.

[0008] Furthermore, the classification model specifically adopts an incremental training model based on random forest or logistic regression, which supports online updates of training samples to adapt to load changes in different business scenarios; and the training features of the classification model include CPU usage fluctuation variance, IO wait ratio, sleep or wake-up frequency, packet sending frequency, and memory access pattern.

[0009] Furthermore, the load type identification and latency tolerance classification specifically include: in load type identification, processes are divided into four categories: low-latency sensitive online business, high-throughput tolerant offline computing, periodic scheduled tasks, and system basic services, to cover process types in mixed load scenarios; in latency tolerance classification, each type of process is assigned a latency tolerance level of 1 to 5, with the smaller the number, the more sensitive to latency, where level 1 is the highest sensitivity and level 5 is the lowest sensitivity.

[0010] Furthermore, the core state decision module includes an idle time prediction module, a heterogeneous core adaptation module, a multi-objective decision algorithm module, and a sleep state restriction module, and the time series prediction model includes an LSTM time series prediction model; The idle time prediction module is used to employ an LSTM time series prediction model, with a set time as the sliding window. The input sequence includes the idle percentage, interruption density, load gradient, and task arrival rate of the core consecutive set number of windows. The output is the core idle probability distribution within the future time range to obtain the expected idle duration with high confidence. When the prediction confidence is lower than the set value, it is downgraded to a conservative prediction based on exponential smoothing. The heterogeneous core adaptation module is used to pre-establish energy consumption and latency models for performance cores and energy efficiency cores respectively, and adapts to multiple architectures. Among them, the maximum allowed sleep depth of performance cores is limited to sleep state C3 to ensure that the wake-up latency is lower than the set time to meet the low latency requirements of online services. Energy efficiency cores are allowed to enter deep sleep states C6 or C7, and the wake-up latency is relaxed to within the set time to maximize energy consumption reduction. Binding sleep control is adopted for hyper-threaded sibling cores. The multi-objective decision algorithm module is used to construct a multi-objective decision function, dynamically calculate the comprehensive utility value of each sleep state, and select the sleep state with the highest utility value. The decision function is expressed as: Utility value = ω1 × power consumption reduction per unit time - ω2 × expected wake-up delay cost - ω3 × performance jitter risk value; where ω1, ω2, and ω3 are dynamically adjustable weight coefficients. In high-priority online business scenarios, ω2 and ω3 are increased to prioritize performance; in offline computing-intensive scenarios, ω1 is increased to prioritize energy consumption reduction. The hibernation state restriction module is used to allow the core to enter the hibernation state only when the predicted idle time is greater than the exit delay of the corresponding hibernation state.

[0011] Furthermore, the weighting coefficients are adjusted through the global strategy configuration module.

[0012] Furthermore, the dynamic hibernation execution and cooperative scheduling module is integrated with the Linux kernel's cpuidle subsystem and process scheduler. Therefore, the dynamic switching of the execution core's hibernation state, process migration, and batch pre-wake-up specifically include: In the dynamic switching of the core hibernation state, based on the output of the core state decision module, the interface of the cpuidle subsystem is called to control the core to enter the corresponding hibernation state of C1, C1E, C3, C6 or C7; before the core is about to enter C6 or above deep hibernation, the soft interrupt affinity of the core is migrated to other active cores and unnecessary timer events are turned off. In the coordinated scheduling of process migration, when a core is about to enter a deep sleep state, online business processes with a latency sensitivity level of less than or equal to 2 on that core are migrated to other cores in a shallow sleep or active state; when the online business load decreases, offline computing processes are migrated to energy efficiency cores to run; when the online business load increases, performance cores are woken up first. In terms of batch pre-wake optimization, when the LSTM model predicts a surge in business traffic within a set time period in the future, a batch pre-wake strategy is executed to wake up a group of cores in the order of performance cores first and energy efficiency cores last, so as to avoid latency spikes caused by core wake-up serialization when requests arrive in a concentrated manner.

[0013] Furthermore, the closed-loop performance assurance module is used to monitor the performance indicators of online services, establish a closed-loop performance feedback mechanism, and dynamically adjust the sleep strategy, specifically including: In terms of performance monitoring, the sampling period is set to continuously monitor the end-to-end latency, jitter, and throughput changes of online services, and track the P99 latency index to ensure the stability of online service performance. Regarding performance protection triggering, when the online service P99 latency exceeds the preset threshold within multiple consecutive sampling periods, the performance protection mechanism is triggered, which forcibly downgrades the sleep depth of the core involved by at least two levels and prohibits it from re-entering deep sleep within the next set time, while quickly waking up adjacent idle cores. In terms of strategy recovery, when the P99 latency of multiple consecutive sampling cycles is lower than the threshold set ratio, and the service throughput and error rate return to normal, the maximum allowable sleep depth of the core is gradually restored to the original value, and energy efficiency optimization is restarted, thereby achieving a dynamic balance between performance and power consumption.

[0014] Furthermore, the global policy configuration module serves as an auxiliary module of the system. It dynamically adjusts the weight coefficients ω1, ω2, ω3, data acquisition frequency, core maximum sleep depth, and performance alarm threshold according to different business scenarios, enabling the system to adapt to various mixed load deployment requirements.

[0015] A method for dynamic hibernation scheduling of server cores under mixed load includes: Step 1: Construct a server core dynamic hibernation scheduling system based on mixed load as described in any of the above steps; Step 2: Utilize this system to perform dynamic hibernation scheduling of server cores under mixed load scenarios.

[0016] The beneficial effects of this invention include: This invention provides a dynamic hibernation scheduling scheme for server cores under mixed load scenarios. Through non-intrusive multi-dimensional data collection, intelligent load classification, accurate idle prediction, differentiated hibernation decision-making, and closed-loop performance assurance, it achieves refined hibernation scheduling under mixed loads while ensuring online business performance loss is less than 1% and data collection overhead is less than 1%, significantly improving server energy efficiency and adapting to various heterogeneous CPU architectures. Specifically, the non-intrusive multi-dimensional data collection based on eBPF has the advantages of low overhead, scalability, and support for hyper-threaded data collection; the designed lightweight load classification and latency grading scheme features high classification accuracy and supports online model updates; the hibernation decision-making method combining LSTM prediction and multi-objective decision-making can adapt to heterogeneous CPU architectures, and the weights can be dynamically adjusted; a collaborative mechanism for core hibernation, process scheduling, and soft interrupt migration is proposed, including a batch pre-wake-up strategy; the closed-loop performance assurance mechanism ensures that online business performance loss is less than 1%; and the global policy configuration and exception handling mechanism can adapt to various business scenarios and ensure system stability. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the steps of the method in an embodiment of the present invention; Figure 2 This is a structural block diagram of the system according to an embodiment of the present invention. Detailed Implementation

[0019] All features disclosed in all embodiments of this specification, or steps in all methods or processes implied in the disclosure, may be combined and / or extended or replaced in any way, except for mutually exclusive features and / or steps.

[0020] As a first aspect of the present invention, in a preferred embodiment, such as Figure 2 As shown, a server core dynamic hibernation scheduling system based on mixed load is provided, including: a multi-dimensional load characteristic perception module, a load type and latency sensitivity identification module, a core state intelligent decision-making module, a dynamic hibernation execution and collaborative scheduling module, a closed-loop performance guarantee module, and an auxiliary module.

[0021] Conventional data acquisition schemes for the application scenarios of this invention mainly employ SystemTap and DTrace, which suffer from drawbacks such as complex deployment, high overhead, and poor compatibility. Therefore, in the specific implementation of the multi-dimensional load characteristic perception module of this invention, eBPF technology is used to achieve non-intrusive load data acquisition. This requires no modification to application code or kernel source code, can be loaded as an extension module, and the overall CPU utilization of the probe does not exceed 1%. Specifically, in terms of acquisition method, real-time data acquisition is achieved by hooking kernel functions schedule, tick_sched_timer, softirq, syscall, and workqueue execution points using eBPF probes. The acquisition frequency can be dynamically adjusted between 10ms and 1s, adaptively adjusting the acquisition granularity according to business load fluctuations. Based on this acquisition method, the system can output the following three dimensions of metrics: 1) Core-level metrics: CPU utilization, idle duration, soft and hard interrupt frequency, context switching frequency, cache miss interval, and process migration count for each physical CPU core; for hyper-threaded architectures, the sleep state of sibling cores is additionally acquired to achieve bound sleep control. 2) Process-level metrics: CPU usage percentage in user mode and kernel mode for each process, memory growth rate, I / O wait percentage, system call frequency, and thread blocking duration; also distinguishing the process type to provide data support for subsequent load classification. 3) Business-level metrics: Request latency (P50 / P95 / P99), queue backlog length, connection fluctuations, and service error rate metrics for online services, reflecting the real-time operational status of the business.

[0022] The alternative solutions for load classification in the application scenarios of this invention are generally SVM and lightweight neural network solutions, which have the disadvantages of low efficiency and high overhead, making it difficult to achieve both simultaneously. Therefore, in the specific implementation of the load type and latency sensitivity identification module, based on the collected multi-dimensional load data, a pre-trained lightweight classification model is used to identify the load type and classify the latency tolerance of processes. The specific implementation is as follows: The classification model adopts an incremental training model based on random forest or logistic regression, which supports online updates of training samples to adapt to load changes in different business scenarios; the training features include CPU usage fluctuation variance, IO wait ratio, sleep / wake frequency, packet sending frequency, and memory access pattern. The model's classification accuracy for online business and offline computing tasks is no less than 96%. In terms of load classification, processes are divided into four categories: low latency sensitive online business, high throughput tolerant offline computing, periodic scheduled tasks, and system basic services, covering the main process types in mixed load scenarios. In terms of latency tolerance grading, each type of process is assigned a latency tolerance level of 1 to 5. The smaller the number, the more sensitive to latency. Level 1 is the highest sensitivity (such as online business such as core transactions and real-time interfaces), and level 5 is the lowest sensitivity (such as offline tasks such as offline data backup and model training).

[0023] The alternative solutions for idle prediction in the application scenarios of this invention are mainly the ARIMA and Prophet schemes. Their drawbacks include low prediction accuracy and inability to adapt to mixed load fluctuations. Decision algorithms generally employ genetic algorithms and particle swarm optimization, but their real-time performance is poor, failing to meet millisecond-level decision-making requirements. Regarding heterogeneous adaptation, hardware-level power management is primarily used, but its disadvantages include non-configurability and limited compatibility. Therefore, in the specific implementation of the core state intelligent decision-making module, a multi-objective decision-making algorithm based on performance and energy efficiency trade-offs was designed. Combined with the LSTM time series prediction model, the optimal sleep state is independently decided for each core, adapting to heterogeneous CPU architecture. The specific implementation is as follows: In terms of idle time prediction, the LSTM time series prediction model is adopted, with a set time (e.g., 100ms) as the sliding window. The input sequence includes the idle percentage, interrupt density, load gradient and task arrival rate of the core for a consecutive set value (e.g., 30) window. The output is the core idle probability distribution within the future time range (e.g., 10ms to 1s) to obtain a high-confidence expected idle duration. When the prediction confidence is lower than the set value (e.g., 80%), it automatically degrades to a conservative prediction based on exponential smoothing to avoid decision-making errors. In terms of heterogeneous core adaptation, power consumption and latency models are pre-established for performance cores (P cores) and energy efficiency cores (E cores) to adapt to various architectures such as ARM big.LITTLE, x86 hyper-threading, and asymmetric heterogeneous CPUs. Specifically, the maximum allowed sleep depth for performance cores is limited to C3 to ensure wake-up latency is below a set time (e.g., 20 microseconds) to meet the low-latency requirements of online services. Energy efficiency cores are allowed to enter C6 or C7 deep sleep states, with wake-up latency relaxed to within a set time (e.g., 100 microseconds) to maximize energy efficiency. For hyper-threaded sibling cores, a bound sleep control is used to prevent one thread from waking up and forcing another thread out of sleep. In the design of the multi-objective decision algorithm, a multi-objective decision function is first constructed to dynamically calculate the comprehensive utility value of each sleep state, selecting the sleep state with the highest utility value. The decision function is expressed as: Utility value = ω1 × Power consumption reduction per unit time - ω2 × Expected wake-up latency cost - ω3 × Performance jitter risk value. Here, ω1, ω2, and ω3 are dynamically adjustable weight coefficients. In high-priority online business scenarios (such as e-commerce peak traffic), ω2 (wake-up latency cost weight) and ω3 (performance jitter risk weight) are automatically increased to prioritize performance. In offline computationally intensive scenarios (such as offline training), ω1 (power consumption reduction weight) is automatically increased to prioritize energy consumption reduction. The weight coefficients can be flexibly adjusted through the global policy configuration module. Regarding sleep state restrictions, a core is only allowed to enter a sleep state when the predicted idle duration is greater than the exit latency of the corresponding sleep state, avoiding performance loss due to untimely wake-up.

[0024] Regarding the specific implementation of the dynamic hibernation execution and collaborative scheduling module, it is deeply integrated with the Linux kernel's cpuidle subsystem and process scheduler to achieve dynamic switching of core hibernation states, process migration, and batch pre-wake-up. Specifically, in core hibernation execution, based on the output of the core state decision module, the cpuidle subsystem interface is called to control the core to enter the corresponding hibernation state (C1, C1E, C3, C6, or C7). Before a core is about to enter deep hibernation (C6 or higher), the system proactively migrates the soft interrupt affinity of that core to other active cores and disables unnecessary timer events to reduce the probability of being unexpectedly woken up during hibernation. In process collaborative scheduling, when a core is about to enter deep hibernation, online business processes with a latency sensitivity level ≤2 on that core are migrated to other cores in shallow hibernation or active states to avoid the impact of wake-up delays on online services. When the online business load decreases, offline computing processes are migrated to energy-efficient cores to fully utilize their low-power advantages. When the online business load increases, performance cores are woken up first to ensure service response speed. In terms of batch pre-wake optimization, when the LSTM model predicts a surge in business traffic within a set time (e.g., 200ms), a batch pre-wake strategy is executed, waking up a group of cores in the order of performance cores first and energy efficiency cores last, avoiding latency spikes caused by core wake-up serialization when requests arrive in a concentrated manner, and further ensuring the performance of online services.

[0025] The performance assurance scheme in this invention's application scenario mainly relies on static threshold triggering. Its drawback is its inability to adapt to dynamic loads, easily leading to energy efficiency or performance losses. Therefore, in terms of the closed-loop performance assurance module, real-time monitoring of online service performance indicators is implemented to establish a closed-loop performance feedback mechanism and dynamically adjust the sleep strategy. Specifically, the following is implemented: For performance monitoring, a set time (e.g., 50ms) is used as the sampling period to continuously monitor changes in end-to-end latency, jitter, and throughput of online services, with a focus on tracking the P99 latency indicator to ensure stable online service performance. For performance protection triggering, when the online service P99 latency exceeds a preset threshold (e.g., 15%) for multiple consecutive sampling periods (e.g., 3), the performance protection mechanism is triggered. This forces a reduction in the sleep depth of the core involved to at least two levels and prohibits re-entering deep sleep within the next set time (e.g., 1 second). Simultaneously, adjacent idle cores are quickly woken up to alleviate load pressure. In terms of strategy recovery, when the P99 latency of multiple consecutive (e.g., 10) sampling cycles is lower than the threshold set ratio (e.g., 80%), and the service throughput and error rate return to normal, the maximum allowable sleep depth of the core is gradually restored to the original value, and energy efficiency optimization is restarted to achieve a dynamic balance between performance and power consumption.

[0026] In terms of auxiliary modules, the global strategy configuration module is used as an auxiliary module of the system. Based on different business scenarios (such as e-commerce peak, offline training, edge nodes, cloud server overselling, etc.), the weight coefficients ω1, ω2, ω3, data collection frequency, core maximum sleep depth and performance alarm threshold are dynamically adjusted to enable the system to adapt to various mixed load deployment requirements and improve the system's flexibility and applicability.

[0027] Figure 2 It should be noted that the system of this embodiment of the invention is represented by a layered architecture consisting of a main process and a support layer. Figure 2 The five modules in the upper half form a complete decision-making loop in the order of data flow. The eBPF multi-dimensional acquisition module is a specific implementation of the multi-dimensional load characteristic perception module; the load type intelligent identification module is a specific implementation of the load type and latency sensitivity identification module; the core space prediction and hibernation decision module is a specific implementation of the core state decision module; and the coordinated scheduling and hibernation execution module is a specific implementation of the dynamic hibernation execution and cooperative scheduling module. The "all core modules" in the lower half represent the system's hardware execution carrier (i.e., P core / E core and hyper-threaded sibling cores), which specifically execute the control logic recorded in the dynamic hibernation execution and cooperative scheduling module and the heterogeneous core adaptation module. The global policy configuration module is configured via the configuration bus ( Figure 2 (Not shown in the image to avoid cluttered lines) The upper half of the module sends out the operating parameters. The exception handling module acts as a global monitoring sentinel. When the prediction confidence is insufficient or the performance index exceeds the limit, it triggers the degradation and recovery strategy. Specifically, it implements the performance protection triggering mechanism of the closed-loop performance guarantee module and the exponential smooth degradation mechanism in the idle time prediction module.

[0028] As a second aspect of the invention, in a preferred embodiment, such as Figure 1 As shown, a method for dynamic hibernation scheduling of server cores under mixed load is provided, including: Step 1: Construct a server core dynamic hibernation scheduling system based on mixed load as described in any of the above steps; Step two involves using this system to perform dynamic hibernation scheduling of server cores under mixed load scenarios, specifically including the following sub-steps: After the system starts, the eBPF module is used to collect three levels of indicators: core, process, and business. Based on the collected metrics, the load is classified and the latency tolerance is graded using the load type and latency sensitivity identification module. Next, the core state decision-making module enters the continuous optimization decision-making stage, and then the dynamic hibernation execution and collaborative scheduling module is used to re-execute the scheduling. After scheduling, the closed-loop performance assurance module monitors business performance and dynamically adjusts the hibernation strategy. Finally, the system checks whether the business performance is stable. If the result is negative, performance protection is triggered, the hibernation depth is reduced, and then the system returns to the continuous optimization decision-making stage for a new iteration. If the result is positive, the system returns to the scheduling stage for re-execution, and continues to monitor and adjust in a loop.

[0029] It should be noted that the method of the embodiments of the present invention, through the process of collection, classification, prediction, decision-making, execution and protection, combined with eBPF, LSTM and multi-objective decision technology, can ultimately achieve fine-grained sleep scheduling under mixed loads and adapt to a variety of heterogeneous CPU architectures.

[0030] The technical effects of the embodiments of the present invention are compared with those of existing solutions in Table 1 below: Table 1

[0031] The units described in the embodiments of the present invention can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0032] According to one aspect of the present invention, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described above.

[0033] In another aspect, embodiments of the present invention also provide a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.

Claims

1. A dynamic hibernation scheduling system for server cores under mixed load, characterized in that, include: The module includes a multi-dimensional load characteristic perception module, a load type and latency sensitivity identification module, a core state decision module, a dynamic sleep execution and collaborative scheduling module, and a closed-loop performance guarantee module. The multi-dimensional load feature perception module is used for non-intrusive multi-dimensional load data acquisition. The load type and latency sensitivity identification module is used to identify the load type and classify the latency tolerance of processes based on the collected multi-dimensional load data using a pre-trained lightweight classification model. The core state decision module is used to make an independent decision on the optimal sleep state for each core through multi-objective decision-making based on performance and energy efficiency trade-offs, combined with a time-series prediction model, to adapt to heterogeneous CPU architecture. Energy consumption and latency models are pre-established for performance cores and energy efficiency cores, adaptable to various architectures. The maximum allowed sleep depth for performance cores is limited to sleep state C3; energy efficiency cores are allowed to enter deep sleep states C6 or C7, with wake-up latency relaxed to within a set time. Binding sleep control is adopted for hyper-threaded sibling cores. A multi-objective decision function is constructed to dynamically calculate the comprehensive utility value of each sleep state, selecting the sleep state with the highest utility value. The decision function is expressed as: Utility value = ω1 × Power consumption reduction per unit time - ω2 × Expected wake-up latency cost - ω3 × Performance jitter risk value; where ω1, ω2, and ω3 are dynamically adjustable weight coefficients. In high-priority online business scenarios, ω2 and ω3 are increased to prioritize performance; in offline computationally intensive scenarios, ω1 is increased to prioritize energy consumption reduction. The dynamic hibernation execution and collaborative scheduling module is used to perform dynamic switching of core hibernation states, process migration, and batch pre-wake-up. The closed-loop performance assurance module is used to monitor the performance indicators of online services, establish a closed-loop performance feedback mechanism, and dynamically adjust the hibernation strategy.

2. The server core dynamic hibernation scheduling system based on mixed load as described in claim 1, characterized in that, The non-intrusive multi-dimensional load data acquisition is specifically implemented using the eBPF module.

3. The server core dynamic hibernation scheduling system based on mixed load as described in claim 1, characterized in that, The classification model specifically adopts an incremental training model based on random forest or logistic regression, which supports online updates of training samples to adapt to load changes in different business scenarios; and the training features of the classification model include CPU usage fluctuation variance, IO wait ratio, sleep or wake-up frequency, packet sending frequency, and memory access pattern.

4. The server core dynamic hibernation scheduling system based on mixed load as described in claim 1, characterized in that, The load type identification and latency tolerance classification specifically include: in load type identification, processes are divided into four categories: low latency sensitive online business, high throughput tolerant offline computing, periodic scheduled tasks, and system basic services, to cover process types in mixed load scenarios; in latency tolerance classification, each type of process is assigned a latency tolerance level of 1 to 5, with the smaller the number, the more sensitive to latency, where level 1 is the highest sensitivity and level 5 is the lowest sensitivity.

5. The server core dynamic hibernation scheduling system based on mixed load as described in claim 1, characterized in that, The core state decision module includes an idle time prediction module, a heterogeneous core adaptation module, a multi-objective decision algorithm module, and a sleep state restriction module; the time series prediction model includes an LSTM time series prediction model. The idle time prediction module is used to employ an LSTM time series prediction model, with a set time as the sliding window. The input sequence includes the idle percentage, interruption density, load gradient, and task arrival rate of the core consecutive set number of windows. The output is the core idle probability distribution within the future time range to obtain the expected idle duration with high confidence. When the prediction confidence is lower than the set value, it is downgraded to a conservative prediction based on exponential smoothing. The heterogeneous core adaptation module is used to pre-establish energy consumption and latency models for performance cores and energy efficiency cores respectively, and adapts to multiple architectures. Among them, the maximum allowed sleep depth of performance cores is limited to sleep state C3 to ensure that the wake-up latency is lower than the set time to meet the low latency requirements of online services. Energy efficiency cores are allowed to enter deep sleep states C6 or C7, and the wake-up latency is relaxed to within the set time to maximize energy consumption reduction. Binding sleep control is adopted for hyper-threaded sibling cores. The multi-objective decision algorithm module is used to construct a multi-objective decision function, dynamically calculate the comprehensive utility value of each sleep state, and select the sleep state with the highest utility value. The decision function is expressed as: Utility value = ω1 × power consumption reduction per unit time - ω2 × expected wake-up delay cost - ω3 × performance jitter risk value; where ω1, ω2, and ω3 are dynamically adjustable weight coefficients. In high-priority online business scenarios, ω2 and ω3 are increased to prioritize performance; in offline computing-intensive scenarios, ω1 is increased to prioritize energy consumption reduction. The hibernation state restriction module is used to allow the core to enter the hibernation state only when the predicted idle time is greater than the exit delay of the corresponding hibernation state.

6. The server core dynamic hibernation scheduling system based on mixed load as described in claim 5, characterized in that, The weighting coefficients are adjusted through the global strategy configuration module.

7. The server core dynamic hibernation scheduling system based on mixed load as described in claim 5, characterized in that, The dynamic hibernation execution and cooperative scheduling module is integrated with the Linux kernel's cpuidle subsystem and process scheduler. The dynamic switching of the execution core's hibernation state, process migration, and batch pre-wake-up specifically include: In the dynamic switching of the core hibernation state, based on the output of the core state decision module, the interface of the cpuidle subsystem is called to control the core to enter the corresponding hibernation state of C1, C1E, C3, C6 or C7; before the core is about to enter C6 or above deep hibernation, the soft interrupt affinity of the core is migrated to other active cores and unnecessary timer events are turned off. In the coordinated scheduling of process migration, when a core is about to enter a deep sleep state, online business processes with a latency sensitivity level of less than or equal to 2 on that core are migrated to other cores in a shallow sleep or active state; when the online business load decreases, offline computing processes are migrated to energy efficiency cores to run; when the online business load increases, performance cores are woken up first. In terms of batch pre-wake optimization, when the LSTM model predicts a surge in business traffic within a set time period in the future, a batch pre-wake strategy is executed to wake up a group of cores in the order of performance cores first and energy efficiency cores last, so as to avoid latency spikes caused by core wake-up serialization when requests arrive in a concentrated manner.

8. The server core dynamic hibernation scheduling system based on mixed load as described in claim 1, characterized in that, The closed-loop performance assurance module is used to monitor the performance indicators of online services, establish a closed-loop performance feedback mechanism, and dynamically adjust the sleep strategy, specifically including: In terms of performance monitoring, the sampling period is set to continuously monitor the end-to-end latency, jitter, and throughput changes of online services, and track the P99 latency index to ensure the stability of online service performance. Regarding performance protection triggering, when the online service P99 latency exceeds the preset threshold within multiple consecutive sampling periods, the performance protection mechanism is triggered, which forcibly downgrades the sleep depth of the core involved by at least two levels and prohibits it from re-entering deep sleep within the next set time, while quickly waking up adjacent idle cores. In terms of strategy recovery, when the P99 latency of multiple consecutive sampling cycles is lower than the threshold set ratio, and the service throughput and error rate return to normal, the maximum allowable sleep depth of the core is gradually restored to the original value, and energy efficiency optimization is restarted, thereby achieving a dynamic balance between performance and power consumption.

9. The server core dynamic hibernation scheduling system based on mixed load as described in claim 6, characterized in that, The global policy configuration module serves as an auxiliary module of the system. It dynamically adjusts the weight coefficients ω1, ω2, ω3, data collection frequency, maximum core sleep depth, and performance alarm threshold according to different business scenarios, enabling the system to adapt to various mixed load deployment requirements.

10. A method for dynamic hibernation scheduling of server cores under mixed load, characterized in that, include: Step 1: Construct a server core dynamic hibernation scheduling system based on any one of claims 1 to 9; Step two: Utilize the constructed system to perform dynamic hibernation scheduling of the server core under mixed load scenarios.

Citation Information

Patent Citations

  • Multi-core processor task migration and power consumption adjustment method and architecture based on performance monitoring mechanism

    CN115576664A

  • Workload adjusting method, device and equipment based on RISC-V, medium and product

    CN121326562A