An edge computing node computing power consumption linkage adaptive regulation method
Patent Information
- Application Number
- CN202611099103.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-09-25
AI Technical Summary
[0010]综上所述,现有边缘计算节点功耗管理技术因其非协同、非智能、非自适应的设计架构,已难以满足日益增长的对算力效率和功耗控制的双重高标准要求
1、通过构建多维度感知、精细化分级调控与闭环迭代优化深度融合的联动体系,在动态异构场景下实现算力效能与功耗控制的协同最优;具体而言,其一方面通过综合采集算力、资源、链路及状态等多类参数并采用改进型指数加权移动平均算法精准划分重载至空载的负载状态,再结合贝叶斯推理与玻尔兹曼分布动态按需分配异构算力资源,避免关键算力短缺及闲置空转以提升利用率,另一方面针对不同状态执行精细化的线性自适应动态电压频率调整与多级分层休眠策略,并配合唤醒窗口内按通信、存储、算力核心顺序的逐级预唤醒机制将整体唤醒延迟控制在毫秒级,从而在充分保障业务实时响应与稳定运行的同时,依托能耗-时延联合评估模型对加权系数、分配权重、分级阈值及调频参数进行反向迭代优化,使系统具备应对负载特征漂移的自适应更新能力,最终从整体上系统性达成算力资源高效利用与能效比显著提升的双重目标。
Smart Images

Figure CN122816897A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of edge computing technology, and in particular to a method for adaptive control of computing power and power consumption of edge computing nodes. Background Technology
[0002] With the evolution of computing paradigms, edge computing has become an important supplement to cloud computing. By offloading data processing and computing power to the network edge, it effectively reduces service latency and alleviates bandwidth pressure on the core network. However, as the types of services carried by edge nodes become increasingly complex, especially with the routine deployment of mixed workloads such as AI inference, video streaming processing, and multi-protocol data aggregation, the computing power requirements and power consumption of edge nodes are growing rapidly. Against this backdrop, how to maximize the utilization of computing resources and reduce overall energy consumption while ensuring the real-time performance and stability of services has become a core technical challenge facing the field of edge computing.
[0003] Currently, the main technical solutions for power consumption and resource management of edge computing nodes are as follows: Option 1: Power consumption control technology based on Dynamic Voltage-Frequency Scaling (DVFS). As a mainstream chip-level power management method, its core mechanism lies in dynamically adjusting the processor core's operating voltage and frequency based on the CPU's real-time utilization through the operating system or hardware interface to achieve a trade-off between power consumption and performance. Typical implementations include Intel SpeedStep, AMD PowerNow!, and the Linux kernel CPUFreq subsystem. Existing literature also contains numerous improved schemes based on proportional-integral-derivative (PID) controllers or preset operating modes (such as performance, balanced, and power-saving modes). These technologies reduce the processor's dynamic power consumption to a certain extent.
[0004] Option 2: Static resource scheduling technology based on load statistics and threshold determination. At the system-level resource management level, existing technologies typically collect node load metrics (mainly CPU utilization) periodically and use algorithms such as Exponentially Weighted Moving Average (EWMA) for smoothing and trend prediction. Based on the prediction results, the system matches preset static thresholds to trigger corresponding resource allocation strategies, such as adjusting the number of container instances, limiting process resource quotas, or managing task queues. Furthermore, to reduce static power consumption, some solutions employ global or coarse-grained hibernation strategies, putting the entire node or most hardware modules into hibernation mode when a node is determined to be under low load.
[0005] Option 3: Basic resource scheduling technology based on general container orchestration frameworks. In actual deployments, edge nodes often serve as remote worker nodes for cloud-native edge computing frameworks such as KubeEdge and OpenYurt. These mature frameworks provide basic scheduling capabilities such as node status reporting, container lifecycle management, and application orchestration. Currently, the industry's resource management of edge nodes largely relies on the native functions of these frameworks, and their scheduling logic is mainly based on the request and limit amounts of CPU and memory, without considering real-time power consumption as a core scheduling dimension.
[0006] However, as edge computing expands into dynamic, heterogeneous, and high-real-time business scenarios such as industrial automated production lines, advanced driver assistance systems, and multi-channel AI visual analysis, the aforementioned existing technical solutions have gradually revealed the following shortcomings: First, the single dimension of load perception leads to an imbalance in computing power allocation. Current load perception technologies heavily rely on a single CPU utilization metric, neglecting the actual load status of heterogeneous computing chips such as NPUs, GPUs, and FPGAs, as well as critical resources like storage I / O and network bandwidth in modern edge nodes. This limitation prevents the system from accurately grasping the true overall load of nodes. Furthermore, the coarse-grained resource allocation model based on static thresholds, due to its lag and rigid strategy, easily leads to a shortage of computing power for critical services during high loads and idle heterogeneous computing resources during low loads, resulting in suboptimal overall computing power utilization.
[0007] Secondly, the power consumption control mechanism is rigid, highlighting the contradiction between energy saving and stability. Traditional DVFS technology mostly adopts open-loop or fixed feedback control, lacking comprehensive awareness of global load trends and the status of heterogeneous modules. It is difficult to achieve fine-grained power consumption adaptation among multiple hardware modules, and its frequency adjustment granularity is coarse, with overall energy saving efficiency usually below 20%. In addition, the globally unified sleep strategy puts the system in an extreme state of "full sleep or full wake-up": frequent full sleep can introduce millisecond-level wake-up latency, causing service jitter or even interruption; and to avoid this risk, the system often abandons deep sleep, resulting in high static power consumption.
[0008] Third, there is a technological gap between edge computing frameworks and power management. As the de facto standard for edge node management, frameworks such as KubeEdge and OpenYurt focus on application lifecycle management and container scheduling. They lack a closed-loop logic that links with the underlying hardware power management, resulting in computing resource scheduling and power control operating on two separate technological planes, failing to create a synergistic effect. This directly leads to the inability of the upper-layer scheduling to drive the underlying hardware to perform fine-grained frequency reduction, core shutdown, or hibernation operations when there is excess computing power, resulting in significant wasted potential for energy efficiency optimization.
[0009] Fourth, the control parameters are fixed and lack dynamic adaptive capabilities. The control parameters of existing solutions (such as EWMA weighting coefficients, DVFS voltage-frequency mapping tables, and sleep determination thresholds) are typically fixed for a long operating cycle once set. Faced with vastly different load fluctuation patterns in various scenarios such as industrial production, vehicle sensing, and urban security, fixed parameters lack adaptive iteration capabilities. When the scenario changes or load characteristics drift, the system is prone to parameter mismatch, manifesting either as overly aggressive energy saving that compromises real-time performance, or overly conservative performance preservation that sacrifices energy-saving potential.
[0010] In summary, existing edge computing node power management technologies, due to their non-cooperative, non-intelligent, and non-adaptive design architecture, are struggling to meet the ever-increasing demands for both computing power efficiency and power consumption control. Therefore, providing a method for adaptively adjusting the computing power and power consumption of edge computing nodes, while ensuring the real-time performance and stability of dynamic heterogeneous business scenarios, and improving the utilization rate and energy efficiency ratio of computing resources, has become an urgent technical problem to be solved. Summary of the Invention
[0011] The technical problem to be solved by this invention is to provide a method for adaptive adjustment of computing power and power consumption of edge computing nodes, which improves the utilization rate and energy efficiency ratio of computing resources while ensuring the real-time performance and stability of dynamic heterogeneous business scenarios.
[0012] This invention provides a method for adaptive adjustment of computing power and power consumption of edge computing nodes, comprising the following steps: Step S1: Periodically collect the operating parameters of the edge nodes through the load acquisition plugin deployed on the edge computing framework to obtain multi-dimensional load data including computing power parameters, resource parameters, link parameters and status parameters; use the sliding time window algorithm to perform noise reduction preprocessing on the collected multi-dimensional load data, remove instantaneous peaks and pulsed abnormal data, and generate a standardized and time-series load feature dataset. Step S2: The preprocessed load feature dataset is fitted and calculated using an improved exponentially weighted moving average algorithm to predict the overall load change trend of the nodes in the next period, and the load growth rate and load fluctuation coefficient are output. Based on the overall load change trend of the nodes, the load growth rate, the load fluctuation coefficient and the measured average load, the node load status is divided into four levels: heavy load, medium load, light load and no load, to obtain the load status classification result. Step S3: Based on the load status classification results, the Boltzmann Bayesian dynamic computing power allocation algorithm is adopted. The load posterior probability of each hardware module of the edge node is updated in real time through Bayesian inference. The optimal computing power allocation weight of each hardware module is calculated by combining the Boltzmann distribution. This is used to dynamically allocate computing power resources to each hardware module of the edge node on demand, and obtain the dynamic computing power allocation result. Step S4: Based on the dynamic allocation result of computing power, an adaptive dynamic voltage and frequency adjustment algorithm is adopted to linearly and dynamically adjust the operating frequency and operating voltage of each hardware module according to the computing power allocation weight of each hardware module, the operating temperature of the equipment and the load fluctuation coefficient. Step S5: Based on the load status classification results, execute the multi-level sleep strategy of the hardware modules. The multi-level sleep strategy includes at least two different sleep depth levels. When the load rebounds based on the overall load change trend of the node, wake up the hibernating hardware modules in a predetermined order within a preset time window before the wake-up operation. Step S6: Construct an energy consumption-delay joint evaluation model, collect monitoring indicators for each preset control cycle and input them into the energy consumption-delay joint evaluation model for evaluation, and obtain the model evaluation results; based on the model evaluation results, iteratively optimize the weighting coefficients, Boltzmann-Bayes computing power allocation weights, load grading thresholds, and frequency adjustment parameters of the improved exponential weighted moving average algorithm.
[0013] Furthermore, in step S1, the computing power parameters include CPU utilization, NPU inference utilization, GPU parallel computing load, and number of computing cores working; the resource parameters include task queue length, memory utilization, storage read / write frequency, and cache utilization; the link parameters include network throughput, data packet transmission / reception rate, and communication channel occupancy status; the status parameters include device operating temperature and module power supply status; the periodic acquisition period is 100ms, and the sliding time window is a 5-frame sliding time window.
[0014] Furthermore, in step S2, the improved exponentially weighted moving average algorithm introduces a dynamic weighting factor, which adaptively adjusts the weighting coefficient according to the load fluctuation amplitude of the edge nodes. When the load fluctuation is severe, the weight of real-time data is increased, and the weight of historical data is retained when the load is in a steady state. The criteria for classifying the node load state are as follows: a measured average load greater than 70% is a heavy load state, 30% to 70% is a medium load state, 5% to 30% is a light load state, and less than 5% is an unloaded state.
[0015] Furthermore, step S3 specifically includes: Based on the load grading results, the baseline value for computing power allocation of each hardware module of the edge node is initialized. Under the heavy load state, each hardware module is allocated 100% of its rated computing power. Under the medium load state, the posterior probability of the load of each hardware module is updated through Bayesian inference according to the real-time load ratio of each hardware module, and the computing power allocation weight is dynamically calculated in combination with Boltzmann distribution. High-load hardware modules are given priority in computing power allocation, and idle hardware modules have their computing power quota gradually reduced. Under the light load state, 30% to 50% of redundant computing power is cut off and the backup computing power cores are shut down. Under the no-load state, only 15% of the basic computing power is reserved for system operation and maintenance and heartbeat detection. Then, the computing power resources of each hardware module of the edge node are dynamically allocated on demand to obtain the dynamic computing power allocation result.
[0016] Furthermore, in step S4, the adaptive dynamic voltage and frequency adjustment algorithm uses a linear dynamic adjustment method instead of fixed-level adjustment; under heavy load, the highest main frequency and standard operating voltage of each hardware module are locked and the power-saving strategy is turned off; under medium load, the main frequency and voltage are linearly reduced according to the computing power allocation weight; under light load, the main frequency is reduced to 60% of the rated main frequency and the operating voltage is reduced, and idle NPU cores and GPU cores are turned off; under no-load, the main frequency is reduced to 30% of the rated main frequency and only the minimum operating core of the system is retained.
[0017] Furthermore, in step S4, a temperature protection mechanism is also provided, which automatically reduces the operating frequency when the operating temperature of the equipment exceeds a preset temperature threshold.
[0018] Furthermore, in step S5, the multi-level sleep strategy includes: Light hibernation: Adapts to the light load state, shuts down idle industrial peripherals, redundant communication channels and backup storage interfaces, and retains core computing power and basic network services; Medium hibernation: Adapted for continuous light-load scenarios, it performs clock gating on idle computing power modules and cache modules, blocking clock signals; Deep sleep: Adapts to continuous idle state for more than a preset time, performs power gating on idle hardware modules, and cuts off the power supply to the modules; Cluster node hibernation: In cluster scenarios with multiple edge nodes, shut down idle standby nodes and retain only the master node to handle business.
[0019] Furthermore, in step S5, the predetermined sequence is to pre-wake up the communication module, storage module, and computing core in that order; the preset time window is 20ms, and the overall wake-up delay is controlled within 10ms.
[0020] Furthermore, step S6 specifically includes: Construct an energy consumption-latency joint evaluation model, collect monitoring indicators for each preset control cycle, including total node energy consumption, service processing latency, computing resource utilization rate, and sleep duration, input the collected monitoring indicators into the energy consumption-latency joint evaluation model for evaluation, and obtain the model evaluation results; Based on the model evaluation results of the energy consumption-latency joint evaluation model, the dynamic weighting coefficients of the improved exponential weighted moving average algorithm, the computing power allocation weights of the Boltzmann Bayes computing power allocation algorithm, the node load state partitioning threshold, and the frequency and voltage adjustment mapping parameters of the adaptive dynamic voltage and frequency adjustment algorithm are optimized in reverse iteratively. Through the edge node log statistics and parameter configuration capabilities of the edge computing framework, the optimal control parameters are automatically synchronized and updated to achieve scenario adaptive updating of control parameters.
[0021] Furthermore, the edge computing framework is the KubeEdge edge collaboration framework and / or the OpenYurt edge node management framework; the load acquisition plugin is a plugin that extends the functionality based on the native node status acquisition capability of the KubeEdge framework, or a plugin that extends the functionality based on the native node status acquisition capability of the OpenYurt framework. The dynamic allocation of computing power in step S3 relies on the edge container scheduling capability of the KubeEdge framework to dynamically adjust the container resource quota; the hibernation management in step S5 relies on the node scheduling capability of the OpenYurt framework to achieve module status management and cluster node hibernation.
[0022] The advantages of this invention are: 1. By constructing a linkage system that deeply integrates multi-dimensional perception, refined hierarchical control, and closed-loop iterative optimization, the system achieves optimal synergy between computing power efficiency and power consumption control in dynamic heterogeneous scenarios. Specifically, on the one hand, it comprehensively collects various parameters such as computing power, resources, links, and status, and uses an improved exponential weighted moving average algorithm to accurately classify the load states from heavy load to idle load. Then, it dynamically allocates heterogeneous computing power resources on demand using Bayesian inference and Boltzmann distribution to avoid critical computing power shortages and idle running, thereby improving utilization. On the other hand, it executes refined linear adaptive dynamic voltage and frequency adjustment and multi-level hierarchical sleep strategies for different states. Combined with a step-by-step pre-wake-up mechanism within the wake-up window according to the order of communication, storage, and computing power cores, the overall wake-up latency is controlled at the millisecond level. Thus, while fully ensuring real-time response and stable operation of services, it relies on the energy consumption-latency joint evaluation model to perform reverse iterative optimization of weighting coefficients, allocation weights, hierarchical thresholds, and frequency adjustment parameters, enabling the system to have adaptive update capabilities to cope with load characteristic drift. Ultimately, it systematically achieves the dual goals of efficient utilization of computing power resources and significant improvement in energy efficiency ratio.
[0023] 2. Collect four-dimensional load data (computing power, resources, links, and status), and use sliding time window noise reduction preprocessing to generate standardized time-series features, which can accurately depict the true overall load profile of edge nodes. Employ a Boltzmann-Bayes dynamic computing power allocation algorithm, which updates the posterior probability of the load of each hardware module in real time through Bayesian inference, and dynamically calculates the computing power allocation weights in combination with Boltzmann distribution. Differentiated on-demand allocation is performed on heterogeneous modules such as CPU, NPU, and GPU under different load conditions. Key modules are prioritized under high load, while redundant computing power is actively trimmed and backup cores are shut down under light load and no load, thereby ensuring that computing power resources are always accurately matched with real-time business needs and maximizing the overall computing power utilization of the system.
[0024] 3. By adaptively tracking load fluctuations through an improved exponentially weighted moving average algorithm, the node load is finely divided into four levels: heavy load, medium load, light load, and no load. Based on this, differentiated linear dynamic voltage and frequency adjustments are performed to achieve continuous and smooth adjustment of the main frequency and voltage, avoiding the coarse adjustment problems caused by fixed levels. At the same time, a multi-level sleep strategy including light sleep, medium sleep, deep sleep, and cluster node sleep is designed. When the load recovers, the communication module, storage module, and computing core are pre-wake up in a predetermined order within a 20ms preset window, with the overall wake-up latency controlled within 10ms. This hierarchical linkage design enables the system to accurately match the most suitable energy-saving measures under different load levels, maintaining the timeliness of business response while significantly reducing static and dynamic power consumption, and significantly improving the energy efficiency ratio.
[0025] 4. The load acquisition plugin is extended based on the native state acquisition capabilities of the KubeEdge or OpenYurt framework. Dynamic computing power allocation relies on KubeEdge's edge container scheduling capabilities to dynamically adjust container resource quotas, while hibernation management relies on OpenYurt's node scheduling capabilities to achieve module status and cluster node hibernation management. Through this deep integration design, the entire control process is fully embedded in the edge computing framework's control system, enabling computing power allocation decisions, voltage and frequency adjustment, and hibernation / wake-up operations to seamlessly coordinate with container orchestration and application scheduling, forming an efficient closed-loop link from upper-layer load perception to lower-layer hardware control. This simplifies system deployment complexity and enhances the real-time performance and reliability of control strategies.
[0026] 5. Construct an energy consumption-latency joint evaluation model, periodically collect key monitoring indicators such as total node energy consumption, business processing latency, computing resource utilization, and sleep duration for evaluation, and iteratively optimize the dynamic weighting coefficients of the improved exponential weighted moving average algorithm, the Boltzmann-Bayes computing power allocation weights, the load grading thresholds, and the frequency and voltage regulation mapping parameters for adaptive dynamic voltage and frequency adjustment based on the evaluation results. All updated optimal control parameters are automatically synchronized and take effect through the log statistics and parameter configuration capabilities of the edge computing framework without manual intervention. This closed-loop self-optimization mechanism enables the control strategy to continuously adapt to changes in load characteristics, always keeping the system near the optimal operating point, and ensuring that high computing efficiency and low power consumption can be stably balanced under different business scenarios.
[0027] 6. A sliding time window algorithm is used to preprocess the raw collected data for noise reduction, effectively eliminating instantaneous peaks and pulse-like abnormal data, avoiding the misleading influence of sudden interference on subsequent predictions. An improved exponential weighted moving average algorithm with dynamic weighting factors is introduced, which can adaptively adjust the weight ratio of historical data and real-time data according to the load fluctuation amplitude: when the load fluctuates drastically, the weight of real-time data is increased to quickly track changes, while the weight of historical data is retained when the load is in a steady state to smooth noise. The combination of the two makes the output load growth rate, fluctuation coefficient and final classification results more reliable and timely, providing solid data support for subsequent computing power allocation, frequency adjustment and hibernation decisions, and ensuring the effectiveness of the entire control strategy from the source.
[0028] 7. A linear dynamic adjustment method is adopted to continuously and steplessly adjust the operating frequency and voltage of each hardware module according to the computing power allocation weight, avoiding performance abrupt changes and oscillations during the gear switching process, making power consumption control more precise and smooth. At the same time, a temperature protection mechanism is specially set up. When the device operating temperature exceeds the preset threshold, the operating frequency is automatically reduced, effectively preventing chip overheating damage caused by high load or poor heat dissipation, significantly improving the long-term operational reliability and hardware lifespan of edge nodes in harsh environments such as industrial sites and outdoors.
[0029] 8. The defined multi-level sleep strategy not only includes four different sleep levels—light, medium, deep, and cluster node sleep—but also, when the load recovers and a wake-up is required, it pre-wakes up the communication module, storage module, and computing core in a predetermined order, and strictly limits the entire wake-up process to a preset time window of 20ms, with the overall wake-up latency controlled within 10ms. This ordered wake-up design fully considers the initialization time of different hardware modules and the order of business dependencies, prioritizing the restoration of network communication and storage paths before restoring the computing core, ensuring the shortest possible business interruption time. The extremely low wake-up latency allows the system to quickly respond to sudden tasks even after deep sleep, thus completely eliminating the inherent problems of "not daring to go into deep sleep and slow wake-up response" in traditional global sleep strategies.
[0030] 9. The energy consumption-latency joint evaluation model does not optimize energy consumption or latency in isolation, but rather collects multi-dimensional monitoring indicators such as total node energy consumption, business processing latency, computing resource utilization, and sleep duration for comprehensive evaluation. This ensures that the optimization objective balances energy saving and service quality. Based on the evaluation results, the improved EWMA weighting coefficient, Boltzmann-Bayes weight allocation, load grading threshold, and DVFS frequency and voltage modulation mapping parameters are iteratively optimized and automatically updated through the edge framework's logs and configuration capabilities. This closed-loop mechanism enables the system to continuously seek the optimal balance between "energy saving" and "performance" without manual intervention. Under long-term operation, it can adapt to changes in load characteristics of different business scenarios and always maximize overall benefits. Attached Figure Description
[0031] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0032] Figure 1 This is a flowchart of an edge computing node computing power and power consumption linkage adaptive control method according to the present invention. Detailed Implementation
[0033] The overall approach of the technical solution in this application is as follows: Addressing issues such as the mismatch between heterogeneous computing power and dynamic load at edge nodes, and the coarse and disconnected power consumption control from the scheduling framework, the solution first obtains a true load profile through multi-dimensional load collection and sliding window noise reduction. Then, an improved exponentially weighted moving average algorithm is used to adaptively track load trends. Combined with measured averages and fluctuation coefficients, nodes are subdivided into four levels: heavy load, medium load, light load, and no load, providing a precise basis for differentiated control. Based on this, two interconnected main lines of computing power allocation and power consumption control are executed in parallel according to the classification results. On the computing power side, a Boltzmann-Bayes dynamic algorithm is used to update the posterior probability of the load of each hardware module through Bayesian inference and calculate the optimal computing power weight, achieving on-demand allocation of heterogeneous resources. On the power consumption side, the main power allocation is adjusted based on weights, temperature, and fluctuation coefficients. The frequency and voltage are continuously and linearly adjusted. At the same time, a multi-level sleep strategy is implemented for light load and no-load scenarios. When the load recovers, the system is pre-wake-up in order of communication, storage and computing cores within a 20ms window to ensure that the wake-up latency is controlled within 10ms, thus balancing fast response and static power consumption reduction. The entire process uses an energy consumption-latency joint evaluation model to iteratively optimize the prediction weighting coefficient, computing power allocation weight, hierarchical threshold and frequency adjustment parameters in reverse iteratively based on indicators such as energy consumption, latency, utilization rate and sleep duration within the cycle. It also relies on the log and configuration capabilities of the edge framework to automatically update synchronously, forming a closed-loop self-evolution mechanism of "perception-hierarchy-allocation-frequency adjustment-sleep-evaluation-optimization", so that the system continuously adapts to changes in load characteristics and always remains stable in the optimal operating state of computing power efficiency and energy efficiency ratio.
[0034] Please refer to Figure 1 As shown, a preferred embodiment of the edge computing node computing power and power consumption linkage adaptive control method of the present invention includes the following steps: Step S1: Periodically collect the operating parameters of the edge nodes through the load acquisition plugin deployed on the edge computing framework to obtain multi-dimensional load data including computing power parameters, resource parameters, link parameters and status parameters; use the sliding time window algorithm to perform noise reduction preprocessing on the collected multi-dimensional load data, remove instantaneous peaks and pulsed abnormal data, and generate a standardized and time-series load feature dataset. In actual deployment, the load acquisition plugin is implemented by extending the NodeStatus reporting mechanism of the KubeEdge framework or the yurtctl node monitoring component of the OpenYurt framework. The acquisition cycle is set to 100ms to ensure timely capture of load fluctuations. The sliding time window algorithm uses a 5-frame window. Specifically, the window size is set to 5. Each time newly acquired data enters the window, the oldest frame of data is discarded, and the arithmetic mean of the 5 frames of data within the window is calculated as the effective load value at that moment. For instantaneous peak values and pulse-like abnormal data that exceed the historical mean ± 3 standard deviations, they are marked as invalid and replaced with the median value of other valid data within the window, thereby generating a standardized, time-series load characteristic dataset. The computing power parameters are obtained by reading the device file under / sys / class / hwmon / or calling the monitoring API provided by the NPU / GPU vendor; the resource parameters are obtained by parsing the / proc / stat, / proc / meminfo, and / proc / diskstats files of the Linux kernel.
[0035] Step S2: The preprocessed load feature dataset is fitted and calculated using an improved exponentially weighted moving average algorithm to predict the overall load change trend of the nodes in the next period, and the load growth rate and load fluctuation coefficient are output. Based on the overall load change trend of the nodes, the load growth rate, the load fluctuation coefficient and the measured average load, the node load status is divided into four levels: heavy load, medium load, light load and no load, to obtain the load status classification result. The prediction model of the improved exponentially weighted moving average algorithm is as follows: ; in, This is the predicted value for the next period. This is the measured value for the current period. This is the predicted value from the previous period. Dynamic weighting factor. Adaptive calculation formula used: ; in, We set β to 0.25, where β is the preset fluctuation response coefficient, ranging from 0.1 to 0.3. This is the load fluctuation coefficient within the current window (i.e., the ratio of the standard deviation to the mean of the data within the window). When the load fluctuates drastically ( )hour, The value is adaptively increased to 0.55~0.85 to increase the weight of real-time data and quickly track load changes; when the load is in steady state ( When the load growth rate approaches 0.25, it is used to retain the weight of historical data and smooth out noise. The load growth rate is obtained by calculating the average slope of the current predicted value compared to the predicted values of the previous three periods. The division of node load status adopts a two-level judgment: first, an initial judgment is made based on the measured average load. When the average is greater than 70%, it is considered heavy load; 30%~70% is considered medium load; 5%~30% is considered light load; and less than 5% is considered no load. Then, the load growth rate is combined with the load growth rate correction classification. If the load growth rate exceeds 20% for three consecutive periods, it is adjusted upward by one level (e.g., medium load is adjusted to heavy load).
[0036] Step S3: Based on the load status classification results, the Boltzmann Bayesian dynamic computing power allocation algorithm is adopted. The load posterior probability of each hardware module of the edge node is updated in real time through Bayesian inference. The optimal computing power allocation weight of each hardware module is calculated by combining the Boltzmann distribution. This is used to dynamically allocate computing power resources to each hardware module of the edge node on demand, and obtain the dynamic computing power allocation result. Suppose that the edge node has N hardware modules (including CPU cores, NPU, GPU, etc.), and at time t, the observed load state of each module is... (e.g., CPU utilization, NPU inference utilization). The Bayesian inference process is as follows: First, initialize the prior probabilities of each module. The distribution is uniform; then the likelihood function is calculated based on the current observations. The likelihood function is modeled using a Gaussian distribution, with its mean and variance dynamically updated from historical statistics; the posterior probability is then calculated. ; The computing power allocation weights of each module are calculated using the Boltzmann distribution. : ; Wherein, T is a temperature coefficient, ranging from 0.5 to 1.5, controlling the smoothness of weight allocation. Under heavy load conditions, T is set to 0.5, causing weights to favor high-load modules and concentrating computing power; under medium load conditions, T is set to 1.0 for balanced allocation; and under light load conditions, T is set to 1.5 for more even weight distribution. In a preferred embodiment, the dynamic allocation of computing power relies on the ResourceQuota and LimitRange capabilities of the KubeEdge framework, achieved by dynamically modifying the CPU and NPU resource quotas of the container group.
[0037] Step S4: Based on the dynamic allocation result of computing power, an adaptive dynamic voltage and frequency adjustment algorithm is adopted to linearly and dynamically adjust the operating frequency and operating voltage of each hardware module according to the computing power allocation weight of each hardware module, the operating temperature of the equipment and the load fluctuation coefficient. The linear dynamic adjustment method calculates the target operating frequency according to the following formula. : ; in, This is the rated maximum clock frequency of the hardware module. This is the minimum operating frequency allowed by the system (usually 30% of the rated main frequency). The computing power obtained in step S3 is assigned a weight, where γ is the temperature suppression coefficient. The temperature suppression coefficient γ is calculated as follows: ; in, The preset temperature protection threshold (preferably 75°C) is used. The preset temperature protection threshold (preferably 85℃) is used. This refers to the current physical operating temperature of the equipment. When the temperature exceeds 85°C, the frequency automatically drops to 50% of the calculated value. Target operating voltage. A linear mapping is then performed based on the chip's voltage-frequency curve to ensure that the voltage adjusts synchronously with the frequency, thereby reducing dynamic power consumption. In one embodiment, for an ARM architecture CPU core, linear frequency scaling is achieved by modifying / sys / devices / system / cpu / cpu* / cpufreq / scaling_governor to user-space control mode and directly writing to scaling_setspeed.
[0038] Step S5: Based on the load status classification results, execute the multi-level sleep strategy of the hardware modules. The multi-level sleep strategy includes at least two different sleep depth levels. When the load rebounds based on the overall load change trend of the node, wake up the hibernating hardware modules in a predetermined order within a preset time window before the wake-up operation. The entry and exit conditions for the multi-level sleep strategy are as follows: when the load status is rated as light load and lasts for more than 5 seconds, light sleep is triggered; when the light load status lasts for more than 30 seconds and there is no network data packet transmission or reception, medium sleep is triggered; when the idle status lasts for more than 60 seconds, deep sleep is triggered. In the step-by-step pre-wake-up mechanism, the preset time window is 20ms, and the specific wake-up sequence is as follows: from 0ms to 5ms after the load is determined to have rebounded, the power supply and clock of the communication module are restored, and the network connection is rebuilt; from 5ms to 10ms, the power supply of the storage module (such as SD card, eMMC controller) is restored and the file system is initialized; from 10ms to 20ms, the power supply and clock of the computing core are restored, and the register state is restored. Through this ordered wake-up design, the overall wake-up latency is controlled within 10ms (i.e., the time from the issuance of the wake-up command to the restoration of full computing power by the computing core).
[0039] Step S6: Construct an energy consumption-delay joint evaluation model, collect monitoring indicators for each preset control cycle and input them into the energy consumption-delay joint evaluation model for evaluation, and obtain the model evaluation results; based on the model evaluation results, iteratively optimize the weighting coefficients, Boltzmann-Bayes computing power allocation weights, load grading thresholds, and frequency adjustment parameters of the improved exponential weighted moving average algorithm.
[0040] The energy consumption-delay joint evaluation model uses a weighted summation method to construct the comprehensive cost function C: ; Where E represents the total energy consumption of the node. The preset baseline energy consumption value is 50% of the node's rated power consumption, and D is the service processing latency (including task queuing latency and processing latency). The preset baseline latency is 10ms, and U represents the computing resource utilization rate. The weighting coefficients are set to 0.4, 0.4, and 0.2 respectively. After one control cycle (preferably 60 seconds), the comprehensive cost C of that cycle is calculated. If C increases compared to the previous cycle, the gradient descent method is used to adjust the parameters to be optimized: for the EWMA weighting coefficients... The adjustment step size is ±0.02; for load grading thresholds (such as 70%, 30%, 5%), the adjustment step size is ±1%; for DVFS frequency modulation mapping parameters, the adjustment step size is ±2%. All updated parameters are automatically synchronized to each edge node through the edge computing framework's ConfigMap mechanism (KubeEdge) or the yurt-app-manager configuration distribution channel (OpenYurt), achieving hot updates without restarting.
[0041] In step S1, the computing power parameters include CPU utilization, NPU inference utilization, GPU parallel computing load, and number of computing cores working; the resource parameters include task queue length, memory utilization, storage read / write frequency, and cache utilization; the link parameters include network throughput, data packet transmission / reception rate, and communication channel occupancy status; the status parameters include device operating temperature and module power supply status; the periodic acquisition period is 100ms, and the sliding time window is a 5-frame sliding time window.
[0042] In step S2, the improved exponentially weighted moving average algorithm introduces a dynamic weighting factor, which adaptively adjusts the weighting coefficient according to the load fluctuation amplitude of the edge nodes. When the load fluctuation is severe, the weight of real-time data is increased, and the weight of historical data is retained when the load is in a steady state. The criteria for classifying the node load state are as follows: the measured average load is greater than 70% as a heavy load state, 30% to 70% as a medium load state, 5% to 30% as a light load state, and less than 5% as an unloaded state.
[0043] Step S3 specifically includes: Based on the load grading results, the baseline value for computing power allocation of each hardware module of the edge node is initialized. Under the heavy load state, each hardware module is allocated 100% of its rated computing power. Under the medium load state, the posterior probability of the load of each hardware module is updated through Bayesian inference according to the real-time load ratio of each hardware module, and the computing power allocation weight is dynamically calculated in combination with Boltzmann distribution. High-load hardware modules are given priority in computing power allocation, and idle hardware modules have their computing power quota gradually reduced. Under the light load state, 30% to 50% of redundant computing power is cut off and the backup computing power cores are shut down. Under the no-load state, only 15% of the basic computing power is reserved for system operation and maintenance and heartbeat detection. Then, the computing power resources of each hardware module of the edge node are dynamically allocated on demand to obtain the dynamic computing power allocation result.
[0044] In step S4, the adaptive dynamic voltage and frequency adjustment algorithm uses a linear dynamic adjustment method instead of a fixed-level adjustment; under heavy load, the highest main frequency and standard operating voltage of each hardware module are locked and the power-saving strategy is turned off; under medium load, the main frequency and voltage are linearly reduced according to the computing power allocation weight; under light load, the main frequency is reduced to 60% of the rated main frequency and the operating voltage is reduced, and idle NPU cores and GPU cores are turned off; under no-load, the main frequency is reduced to 30% of the rated main frequency and only the minimum operating core of the system is retained.
[0045] In step S4, a temperature protection mechanism is also provided, which automatically reduces the operating frequency when the operating temperature of the equipment exceeds the preset temperature threshold.
[0046] In step S5, the multi-level sleep strategy includes: Light hibernation: Adapts to the light load state, shuts down idle industrial peripherals, redundant communication channels and backup storage interfaces, and retains core computing power and basic network services; Medium hibernation: Adapted for continuous light-load scenarios, it performs clock gating on idle computing power modules and cache modules, blocking clock signals; Deep sleep: Adapts to continuous idle state for more than a preset time, performs power gating on idle hardware modules, and cuts off the power supply to the modules; Cluster node hibernation: In cluster scenarios with multiple edge nodes, shut down idle standby nodes and retain only the master node to handle business.
[0047] In step S5, the predetermined sequence is to pre-wake up the communication module, storage module, and computing core in that order; the preset time window is 20ms, and the overall wake-up delay is controlled within 10ms.
[0048] Step S6 specifically involves: Construct an energy consumption-latency joint evaluation model, collect monitoring indicators for each preset control cycle, including total node energy consumption, service processing latency, computing resource utilization rate, and sleep duration, input the collected monitoring indicators into the energy consumption-latency joint evaluation model for evaluation, and obtain the model evaluation results; Based on the model evaluation results of the energy consumption-latency joint evaluation model, the dynamic weighting coefficients of the improved exponential weighted moving average algorithm, the computing power allocation weights of the Boltzmann Bayes computing power allocation algorithm, the node load state partitioning threshold, and the frequency and voltage adjustment mapping parameters of the adaptive dynamic voltage and frequency adjustment algorithm are optimized in reverse iteratively. Through the edge node log statistics and parameter configuration capabilities of the edge computing framework, the optimal control parameters are automatically synchronized and updated to achieve scenario adaptive updating of control parameters.
[0049] The edge computing framework is the KubeEdge edge collaboration framework and / or the OpenYurt edge node management framework; the load acquisition plugin is a plugin that extends the functionality based on the native node status acquisition capability of the KubeEdge framework, or a plugin that extends the functionality based on the native node status acquisition capability of the OpenYurt framework. The dynamic allocation of computing power in step S3 relies on the edge container scheduling capability of the KubeEdge framework to dynamically adjust the container resource quota; the hibernation management in step S5 relies on the node scheduling capability of the OpenYurt framework to achieve module status management and cluster node hibernation.
[0050] This invention was deployed and tested on an AI quality inspection edge node in a smart factory. The node was configured with a 4-core ARM CPU, one NPU (4 TOPS computing power), and 2GB of memory. Before deploying this control method, the node's average daily power consumption was 15.2Wh, the average task processing latency was 45ms, the CPU utilization was only 23%, and the NPU utilization was consistently 0%, indicating a "high-performance idle" state. After deploying this control method, through multi-dimensional load perception, the characteristics of AI inference tasks were automatically identified, and the NPU computing power allocation weight was increased to 0.85. Simultaneously, the clock speed of irrelevant CPU cores was reduced to 60% of the rated frequency. After 30 days of operation, the average daily power consumption decreased to 8.7Wh (a reduction of 42.8%), the average task processing latency decreased to 12ms (a reduction of 73.3%), and the overall utilization rate of computing resources increased to 78%. In load surge tests, the entire process from deep sleep to full wake-up and recovery of computing power remained stable within 8ms, with no business timeouts or packet loss occurring.
[0051] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for adaptive adjustment of computing power and power consumption of edge computing nodes, characterized in that, Includes the following steps: Step S1: Periodically collect the operating parameters of the edge nodes through the load acquisition plugin deployed on the edge computing framework to obtain multi-dimensional load data including computing power parameters, resource parameters, link parameters and status parameters; use the sliding time window algorithm to perform noise reduction preprocessing on the collected multi-dimensional load data, remove instantaneous peaks and pulsed abnormal data, and generate a standardized and time-series load feature dataset. Step S2: The preprocessed load feature dataset is fitted and calculated using an improved exponentially weighted moving average algorithm to predict the overall load change trend of the nodes in the next period, and the load growth rate and load fluctuation coefficient are output. Based on the overall load change trend of the nodes, the load growth rate, the load fluctuation coefficient and the measured average load, the node load status is divided into four levels: heavy load, medium load, light load and no load, to obtain the load status classification result. Step S3: Based on the load status classification results, the Boltzmann Bayesian dynamic computing power allocation algorithm is adopted. The load posterior probability of each hardware module of the edge node is updated in real time through Bayesian inference. The optimal computing power allocation weight of each hardware module is calculated by combining the Boltzmann distribution. This is used to dynamically allocate computing power resources to each hardware module of the edge node on demand, and obtain the dynamic computing power allocation result. Step S4: Based on the dynamic allocation result of computing power, an adaptive dynamic voltage and frequency adjustment algorithm is adopted to linearly and dynamically adjust the operating frequency and operating voltage of each hardware module according to the computing power allocation weight of each hardware module, the operating temperature of the equipment and the load fluctuation coefficient. Step S5: Based on the load status classification results, execute the multi-level sleep strategy of the hardware modules. The multi-level sleep strategy includes at least two different sleep depth levels. When the load rebounds based on the overall load change trend of the node, wake up the hibernating hardware modules in a predetermined order within a preset time window before the wake-up operation. Step S6: Construct an energy consumption-delay joint evaluation model, collect monitoring indicators for each preset control cycle and input them into the energy consumption-delay joint evaluation model for evaluation, and obtain the model evaluation results; based on the model evaluation results, iteratively optimize the weighting coefficients, Boltzmann-Bayes computing power allocation weights, load grading thresholds, and frequency adjustment parameters of the improved exponential weighted moving average algorithm.
2. The edge computing node computing power and power consumption linkage adaptive control method as described in claim 1, characterized in that, In step S1, the computing power parameters include CPU utilization, NPU inference utilization, GPU parallel computing load, and number of computing cores working; the resource parameters include task queue length, memory utilization, storage read / write frequency, and cache utilization; the link parameters include network throughput, data packet transmission and reception rate, and communication channel occupancy status; and the status parameters include device operating temperature and module power supply status. The periodic acquisition period is 100ms, and the sliding time window is a 5-frame sliding time window.
3. The edge computing node computing power and power consumption linkage adaptive control method as described in claim 1, characterized in that, In step S2, the improved exponentially weighted moving average algorithm introduces a dynamic weighting factor, which adaptively adjusts the weighting coefficient according to the load fluctuation amplitude of the edge nodes. When the load fluctuation is severe, the weight of real-time data is increased, and the weight of historical data is retained when the load is in a steady state. The criteria for classifying the node load state are as follows: the measured average load is greater than 70% as a heavy load state, 30% to 70% as a medium load state, 5% to 30% as a light load state, and less than 5% as an unloaded state.
4. The edge computing node computing power and power consumption linkage adaptive control method as described in claim 1, characterized in that, Step S3 specifically includes: Based on the load grading results, the baseline value for computing power allocation of each hardware module of the edge node is initialized. Under the heavy load state, each hardware module is allocated 100% of its rated computing power. Under the medium load state, the posterior probability of the load of each hardware module is updated through Bayesian inference according to the real-time load ratio of each hardware module, and the computing power allocation weight is dynamically calculated in combination with Boltzmann distribution. High-load hardware modules are given priority in computing power allocation, and idle hardware modules have their computing power quota gradually reduced. Under the light load state, 30% to 50% of redundant computing power is cut off and the backup computing power cores are shut down. Under the no-load state, only 15% of the basic computing power is reserved for system operation and maintenance and heartbeat detection. Then, the computing power resources of each hardware module of the edge node are dynamically allocated on demand to obtain the dynamic computing power allocation result.
5. The edge computing node computing power and power consumption linkage adaptive control method as described in claim 1, characterized in that, In step S4, the adaptive dynamic voltage and frequency adjustment algorithm uses a linear dynamic adjustment method instead of a fixed-level adjustment; under heavy load, the highest main frequency and standard operating voltage of each hardware module are locked and the power-saving strategy is turned off; under medium load, the main frequency and voltage are linearly reduced according to the computing power allocation weight; under light load, the main frequency is reduced to 60% of the rated main frequency and the operating voltage is reduced, and idle NPU cores and GPU cores are turned off; under no-load, the main frequency is reduced to 30% of the rated main frequency and only the minimum operating core of the system is retained.
6. The edge computing node computing power and power consumption linkage adaptive control method as described in claim 1, characterized in that, In step S4, a temperature protection mechanism is also provided, which automatically reduces the operating frequency when the operating temperature of the equipment exceeds the preset temperature threshold.
7. The edge computing node computing power and power consumption linkage adaptive control method as described in claim 1, characterized in that, In step S5, the multi-level sleep strategy includes: Light hibernation: Adapts to the light load state, shuts down idle industrial peripherals, redundant communication channels and backup storage interfaces, and retains core computing power and basic network services; Medium hibernation: Adapted for continuous light-load scenarios, it performs clock gating on idle computing power modules and cache modules, blocking clock signals; Deep sleep: Adapts to continuous idle state for more than a preset time, performs power gating on idle hardware modules, and cuts off the power supply to the modules; Cluster node hibernation: In cluster scenarios with multiple edge nodes, shut down idle standby nodes and retain only the master node to handle business.
8. The edge computing node computing power and power consumption linkage adaptive control method as described in claim 1, characterized in that, In step S5, the predetermined sequence is to pre-wake up the communication module, storage module, and computing core in that order; the preset time window is 20ms, and the overall wake-up delay is controlled within 10ms.
9. The edge computing node computing power and power consumption linkage adaptive control method as described in claim 1, characterized in that, Step S6 specifically involves: Construct an energy consumption-latency joint evaluation model, collect monitoring indicators for each preset control cycle, including total node energy consumption, service processing latency, computing resource utilization rate, and sleep duration, input the collected monitoring indicators into the energy consumption-latency joint evaluation model for evaluation, and obtain the model evaluation results; Based on the model evaluation results of the energy consumption-latency joint evaluation model, the dynamic weighting coefficients of the improved exponential weighted moving average algorithm, the computing power allocation weights of the Boltzmann Bayes computing power allocation algorithm, the node load state partitioning threshold, and the frequency and voltage adjustment mapping parameters of the adaptive dynamic voltage and frequency adjustment algorithm are iteratively optimized in reverse. Through the edge node log statistics and parameter configuration capabilities of the edge computing framework, the optimal control parameters are automatically synchronized and updated to achieve scenario adaptive updating of control parameters.
10. The edge computing node computing power and power consumption linkage adaptive control method according to any one of claims 1 to 9, characterized in that, The edge computing framework is the KubeEdge edge collaboration framework and / or the OpenYurt edge node management framework; the load acquisition plugin is a plugin that extends the functionality based on the native node status acquisition capability of the KubeEdge framework, or a plugin that extends the functionality based on the native node status acquisition capability of the OpenYurt framework. The dynamic allocation of computing power in step S3 relies on the edge container scheduling capability of the KubeEdge framework to dynamically adjust the container resource quota; the hibernation management in step S5 relies on the node scheduling capability of the OpenYurt framework to achieve module status management and cluster node hibernation.