Low-carbon energy-saving operation method and system for edge nodes of Internet of Things

By building comprehensive energy efficiency evaluation indicators and deep Q networks, dynamically adjusting the operating frequency and mode of IoT edge nodes, the problem of insufficient energy efficiency in the existing technology is solved, and low-carbon energy saving and safe operation are achieved.

CN120475486APending Publication Date: 2025-08-12WANSIWEI (CHENGDU) TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510616999.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing energy-saving methods of IoT edge nodes are difficult to dynamically adapt to complex network states, heterogeneous task loads and environmental energy supply changes, resulting in insufficient energy efficiency or waste of resources.

Method used

By constructing comprehensive energy efficiency evaluation indicators, combining dynamic power consumption, carbon emissions, throughput and time-delay models, using the deep Q network learning state and action mapping relationship, generating action strategies for frequency adjustment, task scheduling and mode switching, dynamically adjusting the operating frequency and working mode of node equipment, and adaptively adjusting the model weight and learning rate.

Benefits of technology

It realizes low-carbon and energy-saving operation of node equipment in dynamic environments, improves energy efficiency and operation safety, reduces unnecessary energy consumption, and ensures service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475486A_ABST
    Figure CN120475486A_ABST
Patent Text Reader

Abstract

The invention relates to a low-carbon energy-saving operation method and system for an edge node of the Internet of Things, and relates to the technical field of data processing, and the method comprises the steps: constructing a state vector of space-time correlation through dynamic weighted fusion based on the state information of the edge node collected in real time; constructing a comprehensive energy efficiency evaluation index based on the state vector; taking an energy efficiency evaluation index as an optimization target, and generating an action strategy through a deep Q network learning state and action mapping relation; according to the action strategy, dynamically adjusting the operation frequency of the node equipment, the task queue priority and the node working mode; based on the deviation between the actual energy efficiency and the target value, adaptively adjusting the weight and learning rate of each model; and performing dynamic attenuation correction on the execution strategy of each edge node through each updated model to obtain a target strategy, and controlling node equipment to operate in a low-carbon and energy-saving manner based on the target strategy. According to the invention, the low-carbon energy-saving capability and the operation safety of node equipment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a low-carbon energy-saving operation method, system, electronic device, and non-transient computer-readable storage medium for an edge node of the Internet of Things. Background Art

[0002] With the rapid development of the Internet of Things (IoT), edge computing, as a key supporting architecture, is being widely deployed in various application scenarios, such as smart cities, industrial monitoring, and environmental sensing. Currently, common edge node energy-saving methods rely primarily on hardware optimization, time-slot wake-up mechanisms, task offloading strategies, and operating mode switching based on static thresholds.

[0003] However, existing practices are mostly based on static configuration or preset rules, which make it difficult to dynamically adapt to the complex network status, heterogeneous task loads and environmental energy supply changes faced by edge nodes during actual operation, easily leading to insufficient energy efficiency or waste of computing resources. Summary of the Invention

[0004] In response to the technical problems existing in the prior art, the present invention provides a low-carbon and energy-saving operation method, system, electronic device and non-transitory computer-readable storage medium for an edge node of the Internet of Things, which can significantly improve the low-carbon and energy-saving capabilities and operational safety of node devices while ensuring service quality.

[0005] The technical solution of the present invention to solve the above technical problems is as follows: The present invention provides a low-carbon and energy-saving operation method for an edge node of an Internet of Things, the method comprising: Based on the state information of edge nodes collected in real time, a spatiotemporal correlation state vector is constructed through dynamic weighted fusion; Based on the state vector, a comprehensive energy efficiency evaluation index is constructed in combination with a dynamic power consumption model, a carbon emission model, and throughput and latency constraints; Taking energy efficiency evaluation indicators as the optimization goal, the deep Q network learns the mapping relationship between state and action, and generates an action strategy including frequency regulation, task scheduling and mode switching; Dynamically adjust the operating frequency, task queue priority, and node working mode of the node device according to the action strategy; Adaptively adjust the weights and learning rates of each model based on the deviation between the actual energy efficiency and the target value; The execution strategy of each edge node is dynamically attenuated and corrected by the updated models to obtain a target strategy, and the node device is controlled to operate in a low-carbon and energy-saving manner based on the target strategy.

[0006] Optionally, the state vector is combined with a dynamic power consumption model, a carbon emission model, and throughput and latency constraints to construct a comprehensive energy efficiency evaluation index, including: Constructing a power consumption model that is positively correlated with the square of the computing load of the node device and the cube of the operating frequency of the node device; Building a carbon emission model based on the real-time power consumption of the node device and the ambient temperature of the environment in which it is located; Constructing a throughput model that reflects the coupling relationship between the computing efficiency of the node device and the network quality; Constructing a delay model that characterizes the combined effects of the operating frequency of the node device and the network status; The outputs of the power consumption model, the carbon emission model, the throughput model, and the delay model are weighted and combined to form an energy efficiency evaluation value.

[0007] Optionally, the comprehensive energy efficiency evaluation index is expressed as: ; in, It is a comprehensive energy efficiency evaluation indicator. is the power consumption model, It is a carbon emission model. is the throughput model, is the time delay model, are the first weight, the second weight, the third weight and the fourth weight respectively, is the normalization coefficient.

[0008] Optionally, the energy efficiency evaluation index is used as the optimization target, and a state-action mapping relationship is learned through a deep Q network to generate an action strategy including frequency regulation, task scheduling, and mode switching, including: Constructing an action space including dynamic adjustment of the operating frequency of the node device, task strategy scheduling and working mode switching; Constructing an immediate reward function whose value is negatively correlated with the comprehensive energy efficiency evaluation index and positively correlated with the stability of energy supply and the constraints of task delay; The instantaneous reward function is processed using a temporal difference algorithm, and a corresponding Q value is calculated to determine an action strategy for updating the action space.

[0009] Optionally, the Q value is expressed as: ; in, is the Q value of the state and action value function, is the mathematical expectation, is the immediate reward function, is the discount factor.

[0010] Optionally, dynamically adjusting the operating frequency, task queue priority, and node operating mode of the node device according to the action strategy includes: Determining a loss function of the deep Q network based on a deviation between a Q value currently output by the deep Q network and a target Q value; Obtaining updated policy parameters of the deep Q network according to the loss function and historical policy parameters of the deep Q network; updating the comprehensive energy efficiency evaluation value according to the updated strategy parameters to obtain an updated comprehensive energy efficiency evaluation value; According to the updated comprehensive energy efficiency evaluation value and the state vector, the optimal action for maximizing the Q value is determined, and the operating frequency, task queue priority and node working mode of the node device are dynamically adjusted based on the optimal action.

[0011] Optionally, the adaptive adjustment of the weights and learning rates of each model based on the deviation between the actual energy efficiency and the target value includes: Calculate the deviation between the actual comprehensive energy efficiency index and the target comprehensive energy efficiency index; performing an exponential processing on the deviation; The weights and learning rates of the models are adaptively adjusted according to the deviation amount after the exponential processing.

[0012] Optionally, dynamically attenuating and correcting the execution strategy of each edge node using each updated model to obtain a target strategy includes: Construct a multi-dimensional security boundary function; Calculate the corresponding safety threshold according to the multi-dimensional safety boundary function; Through the updated models, dynamic attenuation correction is performed according to the safety threshold and the operating frequency of the node device, and a strategy of correcting the operating frequency to the operating frequency adjustment value is determined as the target strategy.

[0013] Optionally, the method further includes: Based on historical operation data sampling, a Pareto frontier solution set of energy efficiency indicators and computing performance is constructed through a multi-objective optimization algorithm; Provide a visual interactive interface to map the energy efficiency preference coefficients input by the user into weight parameters of the multi-objective optimization model; According to real-time energy efficiency requirements and performance constraints, the optimal strategy combination is dynamically selected from the Pareto solution set and fed back to the control end.

[0014] The present invention also provides a low-carbon and energy-saving operation system for an edge node of the Internet of Things, the system comprising: The state vector module is used to construct a spatiotemporal correlated state vector through dynamic weighted fusion based on the state information of edge nodes collected in real time; An energy efficiency evaluation module, configured to construct a comprehensive energy efficiency evaluation index based on the state vector, in combination with a dynamic power consumption model, a carbon emission model, and throughput and latency constraints; The action strategy module is used to optimize the energy efficiency evaluation index and generate action strategies including frequency regulation, task scheduling and mode switching through deep Q network learning of the state-action mapping relationship; A dynamic adjustment module, used to dynamically adjust the operating frequency, task queue priority and node working mode of the node device according to the action strategy; The model update module is used to adaptively adjust the weights and learning rates of each model based on the deviation between the actual energy efficiency and the target value; The operation control module is used to dynamically attenuate and correct the execution strategy of each edge node through the updated models to obtain a target strategy, and control the low-carbon and energy-saving operation of the node device based on the target strategy.

[0015] In addition, to achieve the above-mentioned purpose, the present invention also proposes an electronic device, comprising: a memory for storing a computer software program; a processor for reading and executing the computer software program, thereby realizing a low-carbon and energy-saving operation method of an IoT edge node as described above.

[0016] In addition, to achieve the above-mentioned purpose, the present invention also proposes a non-transitory computer-readable storage medium, in which a computer software program is stored. When the computer software program is executed by a processor, it implements a low-carbon and energy-saving operation method of an Internet of Things edge node as described above.

[0017] The beneficial effects of the present invention are: (1) This invention constructs a comprehensive energy efficiency evaluation function, comprehensively considering power consumption, carbon emissions, throughput and delay, and truly realizes energy-saving optimization of "multi-objective balance". The reinforcement learning model can adjust the frequency, scheduling and mode in real time according to state changes, effectively reducing unnecessary energy consumption.

[0018] (2) This paper utilizes a deep Q-network (DQN) to implement learning and iteration of the state-action value function Q(S,A), enabling the system to continuously self-learn and optimize in a dynamic environment. The immediate reward function incorporates key performance indicators such as energy efficiency deviation and latency constraints into the evaluation, ensuring that the training objectives are consistent with the actual operation objectives.

[0019] (3) The present invention dynamically adjusts the model parameter weight ω by calculating the effect deviation Δ(t), continuously correcting and optimizing, and enabling the system to have adaptive evolution capabilities. This ensures that the strategy will not solidify during the long-term operation of the system, but will continue to optimize. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A flowchart of a low-carbon and energy-saving operation method for an edge node of the Internet of Things provided by the present invention; Figure 2 A schematic diagram of the structure of a low-carbon and energy-saving operation system for an edge node of the Internet of Things provided by the present invention; Figure 3 A schematic diagram of the hardware structure of a possible electronic device provided by the present invention; Figure 4 A schematic diagram of the hardware structure of a possible computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0022] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0023] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or illustration". Any embodiment of the present invention described as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed herein.

[0024] See also Figure 1 , provides a flowchart of a low-carbon and energy-saving operation method of an Internet of Things edge node of the present invention, comprising the following steps: Step 201: Based on the state information of edge nodes collected in real time, a spatiotemporal correlation state vector is constructed through dynamic weighted fusion.

[0025] In some embodiments, the system can collect key parameters such as the computing load, network status, energy supply and ambient temperature of the edge node in real time, and weightedly fuse them through dynamic weights to construct a state vector that can reflect the current operating status of the node and its spatiotemporal change characteristics, which is used for subsequent reinforcement learning decision input.

[0026] Step 202: Based on the state vector, combined with a dynamic power consumption model, a carbon emission model, throughput and latency constraints, a comprehensive energy efficiency evaluation index is constructed.

[0027] In some embodiments, step 202 may include: Constructing a power consumption model that is positively correlated with the square of the computing load of the node device and the cube of the operating frequency of the node device; Building a carbon emission model based on the real-time power consumption of the node device and the ambient temperature of the environment in which it is located; Constructing a throughput model that reflects the coupling relationship between the computing efficiency of the node device and the network quality; Constructing a delay model that characterizes the combined effects of the operating frequency of the node device and the network status; The outputs of the power consumption model, the carbon emission model, the throughput model, and the delay model are weighted and combined to form an energy efficiency evaluation value.

[0028] Among them, the comprehensive energy efficiency evaluation index can be expressed as: ; in, It is a comprehensive energy efficiency evaluation indicator. is the power consumption model, It is a carbon emission model. is the throughput model, is the time delay model, are the first weight, the second weight, the third weight and the fourth weight respectively, is the normalization coefficient.

[0029] In specific implementations, the comprehensive energy efficiency evaluation index calculated by the energy efficiency function is used to weigh multiple dimensions of edge nodes, such as energy consumption, carbon emissions, performance, and latency, to achieve a balanced optimization of low-carbon energy conservation and service quality. The overall structure of the energy efficiency function can be understood as: ; The numerator is the consumption cost, which reflects the consumption of node devices, including power consumption P(t) and carbon emissions C(t), which are weighted separately. The denominator is the operational quality, which reflects the performance of node devices, including throughput R(t) (a positive indicator) and task latency D(t) (a negative indicator), which are also weighted. Used to control dimensionality or numerical balance.

[0030] Specifically: ; in, This means that the higher the task load, the greater the amount of calculation, and the power consumption increases quadratically (nonlinearly). This means that the higher the operating frequency of the node device, the more the dynamic power consumption increases cubically, which is consistent with the CMOS power consumption model. Indicates static power consumption or basic power consumption.

[0031] ; in, Indicates that carbon emissions are linearly positively correlated with power consumption (as calculated based on the grid carbon factor). This means that when the ambient temperature is high, heat dissipation becomes more difficult, indirectly leading to more carbon emissions or cooling costs.

[0032] ; in, It means load × frequency = the amount of tasks that can be completed per unit time, that is, throughput. Indicates that poor network conditions (such as low bandwidth and high packet loss) will reduce effective throughput.

[0033] ; in, The higher the frequency, the faster the calculation and the smaller the calculation delay. Indicates poor network conditions, which will increase transmission or response delays.

[0034] It should also be noted that is the computational load, is the network status, is the ambient temperature, is the node device operating frequency, is the first delay coefficient, is the second delay coefficient, is the first strategy parameter, is the second strategy parameter.

[0035] The formula can be organized as follows: ; It is understandable that the larger the molecule, the higher the energy consumption and carbon emissions. The larger the denominator, the lower the energy efficiency. The larger the denominator, the higher the throughput and lower the latency. The smaller → the better the energy efficiency. This method hopes to minimize η(t) (which appears in the reward function Therefore, the overall optimization goal is to reduce energy consumption and carbon emissions, improve throughput, and reduce latency while meeting service requirements.

[0036] In summary, the trend expectations and effects of the values output by each model of the present invention are as follows: Reflects power consumption; the smaller the better. C(t) reflects the load environment; the smaller the better. R(t) reflects computing efficiency; the smaller the better. D(t) reflects service quality; the smaller the better. η(t) reflects the reference value of the reward function; the smaller the better.

[0037] Step 203: Taking the energy efficiency evaluation index as the optimization target, the state-action mapping relationship is learned through the deep Q network to generate an action strategy including frequency regulation, task scheduling and mode switching.

[0038] In some embodiments, step 203 may include: Constructing an action space including dynamic adjustment of the operating frequency of the node device, task strategy scheduling and working mode switching; Constructing an immediate reward function whose value is negatively correlated with the comprehensive energy efficiency evaluation index and positively correlated with the stability of energy supply and the constraints of task delay; The instantaneous reward function is processed using a temporal difference algorithm, and a corresponding Q value is calculated to determine an action strategy for updating the action space.

[0039] Among them, the Q value can be expressed as: ; in, is the Q value of the state and action value function, is the mathematical expectation, is the immediate reward function, is the discount factor.

[0040] In the specific implementation, It represents the expected total reward that can be obtained by taking at least one type of action in the action space A in the current state S and then acting according to the strategy in the environment. It is the expectation of uncertainty about future paths, that is, the weighted average of possible paths. is a discount factor that reduces the weight of future rewards (current benefits are more important than long-term benefits).

[0041] It is used to guide the reinforcement learning agent to select the optimal action (argmaxQ(S,A)) and is the basis for strategy optimization.

[0042] in, Specifically expressed as: ; in, It is a comprehensive energy efficiency evaluation index calculated by energy efficiency function. It is the energy supply, is the target value of energy supply, is the task delay, is the maximum allowable value of task delay, and the action space A includes adjusting the operating frequency of node devices , Scheduling Task Strategy , switch working mode .

[0043] This function comprehensively considers energy efficiency index, energy supply and demand difference, and task delay violation as the "score" of single-step behavior in reinforcement learning, and determines the reward function value at each moment through the time difference algorithm, for example wait.

[0044] Energy efficiency drivers , It is a comprehensive energy efficiency indicator. The smaller η(t) is, the lower the energy consumption and the better the performance → the greater the reward. is the corresponding weight.

[0045] Energy balance penalty , we hope that the current energy supply E(t) is close to the target value , preventing energy waste or depletion. It is a "soft constraint" that balances power supply and demand and can be used in dynamic power management scenarios (such as solar / battery power supply).

[0046] Delay breach penalty If the task delay Exceeding the maximum allowed value , triggering a penalty. This is the last line of defense for ensuring quality of service (QoS), allowing the model to achieve "energy saving" while not neglecting "real-time performance."

[0047] The action space defines the agent's optional "control behaviors" at each step. In this invention, it includes three dimensions, as follows:

[0048] These actions work together to transition the system’s state and affect all key variables in the reward function (η(t), D(t), E(t), etc.).

[0049] The core goal of reinforcement learning is to learn a strategy π(S) so that taking action A=π(S) in any state S can maximize the total expected reward: ; In summary, the present invention dynamically adjusts , the edge nodes can maintain high energy efficiency (low η(t)) and power supply and demand balance in dynamic environments. , do not overtime .

[0050] Step 204: Dynamically adjust the operating frequency, task queue priority, and node working mode of the node device according to the action strategy.

[0051] In some embodiments, step 204 may include: Determining a loss function of the deep Q network based on a deviation between a Q value currently output by the deep Q network and a target Q value; Obtaining updated policy parameters of the deep Q network according to the loss function and historical policy parameters of the deep Q network; updating the comprehensive energy efficiency evaluation value according to the updated strategy parameters to obtain an updated comprehensive energy efficiency evaluation value; According to the updated comprehensive energy efficiency evaluation value and the state vector, the optimal action for maximizing the Q value is determined, and the operating frequency, task queue priority and node working mode of the node device are dynamically adjusted based on the optimal action.

[0052] Among them, the optimal action can be expressed as , as follows: ; ; ; ; in, It is the optimal action. is the state vector, They are the first feature weight, the second feature weight, the third feature weight and the fourth feature weight, are the updated policy parameters, is the historical strategy parameter, is the loss function.

[0053] In the specific implementation, Under the current state vector S(t), the agent chooses the optimal action that maximizes the Q value This action is the optimal control behavior considered by the current strategy in the long term. It is output from the valuation of the Deep Q Network (DQN) and is the core decision-making basis in the strategy execution phase.

[0054] represents the system task pressure, N(t) represents the network delay / congestion, Reflects the available energy status, Affects carbon emissions and heat dissipation energy consumption.

[0055] Represents the parameter set of the deep Q network, and uses the gradient descent algorithm or its variants (such as Adam) to iteratively update the policy network so that the Q value output by the network is closer to the target.

[0056] Measure the square error between the current network output Q(S,A;θ) and the "target Q value". The target value consists of the immediate reward r and the maximum Q value of the next state, that is: ; Among them, r is the immediate reward of the current time step, which comes from the aforementioned reward function. is a discount factor that controls the importance of future rewards. is the next state and all possible actions. are the parameters of the target network (used to stabilize training), and Q(S,A;θ) is the current policy network's estimate of Q. This loss function is the core of the standard DQN (Deep Q Network), used to drive the policy network to gradually learn to approach the optimal behavior.

[0057] From the above, it can be seen that the process of determining the optimal strategy of the present invention is as follows: (1) Input state vector S(t) to the DQN network; (2) Output the Q value of all actions and select the action corresponding to the maximum value ; (3) Execution , obtain the reward r(t) and observe the next state S′(t+1); (4) Constructing TD target ; (5) Calculate the loss L(θ) and update the network parameters θ; (6) Repeat the above process and gradually approach the optimal strategy.

[0058] Step 205: Based on the deviation between the actual energy efficiency and the target value, the weight and learning rate of each model are adaptively adjusted.

[0059] In some embodiments, step 205 may include: Calculate the deviation between the actual comprehensive energy efficiency index and the target comprehensive energy efficiency index; performing an exponential processing on the deviation; The weights and learning rates of the models are adaptively adjusted according to the deviation amount after the exponential processing.

[0060] Among them, the bias value and the updated weight are expressed as: ; ; in, is the comprehensive energy efficiency index actually calculated at the current time step, is the expected or target energy efficiency indicator, It is the effect deviation, is the updated weight, is the historical weight, is the learning rate.

[0061] In the specific implementation, Usually it is an energy efficiency benchmark value set by the system designer (such as a low-carbon energy saving threshold). It indicates the deviation between the current operating result of the system and the expectation, that is, a direct measure of "whether the optimization is effective". This value serves as a "feedback signal" to control the fine-tuning intensity of the next strategy. It can be any weight parameter in the energy efficiency function, and can also be extended to the dynamic update of multiple parameters. is the learning rate (or adjustment coefficient), which controls the sensitivity of the response to deviations.

[0062] like is very small, indicating that the actual energy efficiency is close to the target. , maintain the current strategy; if A larger value indicates a significant deviation → the index factor tends to decrease → the impact of the current weight is weakened to push the strategy towards a better direction.

[0063] Exponential decay is more "gentle" than linear decline, but more "punitive"; it ensures that the direction of weight change is always in the direction of reducing deviation; and it can achieve gradual adaptive optimization over time series.

[0064] This mechanism is primarily used in scenarios such as dynamic model self-tuning, system environment changes, and enhanced reward feedback for reinforcement learning. It dynamically adjusts weight ratios based on actual results to enhance system adaptability. When external conditions such as network status, temperature, and load change, the model weights are adaptively reconfigured. By adjusting energy efficiency weights, the strategy converges more quickly to the target strategy of "high efficiency, stability, and low carbon."

[0065] It can be set as a moving average energy efficiency baseline in long-term operation. , each of which may adopt the adjustment mechanism individually or in concert; It is recommended to set it to a smaller value (such as 0.01 to 0.1) to avoid parameter oscillation; upper and lower limits can be set to prevent the parameter from "exponential disappearance" or "explosion": for example .

[0066] Step 206: Dynamically attenuate and modify the execution strategy of each edge node using the updated models to obtain a target strategy, and control the low-carbon and energy-saving operation of the node device based on the target strategy.

[0067] In some embodiments, step 206 may include: Construct a multi-dimensional security boundary function; Calculate the corresponding safety threshold according to the multi-dimensional safety boundary function; Through the updated models, dynamic attenuation correction is performed according to the safety threshold and the operating frequency of the node device, and a strategy of correcting the operating frequency to the operating frequency adjustment value is determined as the target strategy.

[0068] The safety threshold and operating frequency adjustment value can be expressed as: ; ; in, is the safety threshold of the safety boundary function, It is the operating frequency adjustment value of the node device under safety protection. is the attenuation coefficient.

[0069] Specifically, is the energy supply at the current moment, is the current ambient temperature, is the current system computing load, is the upper limit of energy supply (such as maximum battery capacity or power supply limit), is the upper temperature limit (such as the chip safety temperature threshold), is the upper load limit (the maximum load capacity of the equipment), They are the first adjustment coefficient, the second adjustment coefficient and the third adjustment coefficient, which are used to adjust the importance or unit difference of different dimensions.

[0070] The safety margin function B(t) represents the "minimum margin" of the current system's operating distance from the safety limit. This means that when any dimension approaches its maximum safety value (i.e., the margin decreases), the entire system is considered to be approaching an unsafe state. This design approach, where "the weakest link determines the health of the system," embodies the security principle of protecting against worst-case scenarios.

[0071] is the operating frequency of the node device under the original decision, This is equivalent to the frequency reduction value after the safety protection mechanism target strategy is activated. is the attenuation coefficient, which controls the sensitivity of the frequency reduction amplitude to B(t).

[0072] When B(t) is large, the system is far away from the upper limits. , the system runs close to the original strategy; when B(t) is small, the system is close to the safety critical point, Decrease sharply, reduce This proactively mitigates energy consumption, load, or heat. This is similar to an exponential backoff mechanism, preventing system crashes or downtime due to over-limit conditions.

[0073] In summary, the mechanism of the present invention can reduce the risk of hardware failure caused by high temperature (especially in summer or passive cooling scenarios), extend the survival time in power-constrained scenarios (such as battery edge nodes), avoid task accumulation or deadlock, and buffer load growth by pre-downgrading.

[0074] It should be noted that It can be set based on the actual safety margin experience of the hardware (such as 0.1 to 1), the larger the value, the more sensitive it is. Frequency reduction may cause delay growth, and its impact should be evaluated in conjunction with the delay function D(t). It is possible to consider setting a lower limit threshold for B(t), and trigger protection only when B(t) is less than the lower limit threshold. For multi-core processor systems, Calculate by core Achieve local thermal protection.

[0075] In some embodiments, the present invention may further include: Based on historical operating data sampling, a Pareto frontier solution set of energy efficiency indicators and computing performance is constructed through a multi-objective optimization algorithm; Provide a visual interactive interface to map the energy efficiency preference coefficients input by the user into weight parameters of the multi-objective optimization model; According to real-time energy efficiency requirements and performance constraints, the optimal strategy combination is dynamically selected from the Pareto solution set and fed back to the control end.

[0076] In specific implementation, based on historical operation data sampling, a multi-objective optimization algorithm can be used to construct a Pareto frontier solution set between energy efficiency and computing performance. The user's energy efficiency preferences can be received through a visual interactive interface and converted into optimization weights. Combined with real-time operation requirements, the optimal strategy combination can be dynamically selected from the solution set and fed back to the control end to achieve personalized, low-carbon and efficient operation control.

[0077] See also Figure 2 , Figure 2This is a structural schematic diagram of a low-carbon and energy-saving operation system for an Internet of Things edge node provided by the present invention.

[0078] like Figure 2 As shown, an embodiment of the present invention proposes a low-carbon energy-saving operation system for an IoT edge node, including: The state vector module 301 is used to construct a spatiotemporal correlated state vector through dynamic weighted fusion based on the state information of edge nodes collected in real time; An energy efficiency evaluation module 302 is configured to construct a comprehensive energy efficiency evaluation index based on the state vector, in combination with a dynamic power consumption model, a carbon emission model, and throughput and latency constraints; The action strategy module 303 is used to use the energy efficiency evaluation index as the optimization target, learn the state-action mapping relationship through the deep Q network, and generate an action strategy including frequency regulation, task scheduling and mode switching; A dynamic adjustment module 304 is used to dynamically adjust the operating frequency, task queue priority and node working mode of the node device according to the action strategy; Model update module 305, used to adaptively adjust the weights and learning rates of each model based on the deviation between the actual energy efficiency and the target value; The operation control module 306 is configured to dynamically attenuate and modify the execution strategy of each edge node using the updated models to obtain a target strategy, and control the low-carbon and energy-saving operation of the node device based on the target strategy.

[0079] It should be noted that the specific implementation methods and beneficial effects of the above modules 301-306 can be found in the above description of steps 201-206, which will not be repeated here.

[0080] See also Figure 3 , Figure 3 Schematic diagram of an embodiment of an electronic device provided by an embodiment of the present invention. Figure 3 As shown, an embodiment of the present invention provides an electronic device 400, including a memory 410, a processor 420, and a computer program 411 stored in the memory 410 and executable on the processor 420. When the processor 420 executes the computer program 411, the following steps are implemented: Based on the real-time collected state information of edge nodes, a spatiotemporal correlation state vector is constructed through dynamic weighted fusion; Based on the state vector, a comprehensive energy efficiency evaluation index is constructed in combination with a dynamic power consumption model, a carbon emission model, and throughput and latency constraints; Taking energy efficiency evaluation indicators as the optimization goal, the deep Q network learns the mapping relationship between state and action, and generates an action strategy including frequency regulation, task scheduling and mode switching; Dynamically adjust the operating frequency, task queue priority, and node working mode of the node device according to the action strategy; Adaptively adjust the weights and learning rates of each model based on the deviation between the actual energy efficiency and the target value; The execution strategy of each edge node is dynamically attenuated and corrected by the updated models to obtain a target strategy, and the node device is controlled to operate in a low-carbon and energy-saving manner based on the target strategy.

[0081] See also Figure 4 , Figure 4 Schematic diagram of an embodiment of a computer-readable storage medium provided in an embodiment of the present invention. Figure 4 As shown, this embodiment provides a computer-readable storage medium 500 on which a computer program 411 is stored. When the computer program 411 is executed by a processor, the following steps are implemented: Based on the real-time collected state information of edge nodes, a spatiotemporal correlation state vector is constructed through dynamic weighted fusion; Based on the state vector, a comprehensive energy efficiency evaluation index is constructed in combination with a dynamic power consumption model, a carbon emission model, and throughput and latency constraints; Taking energy efficiency evaluation indicators as the optimization goal, the deep Q network learns the mapping relationship between state and action, and generates an action strategy including frequency regulation, task scheduling and mode switching; Dynamically adjust the operating frequency, task queue priority, and node working mode of the node device according to the action strategy; Adaptively adjust the weights and learning rates of each model based on the deviation between the actual energy efficiency and the target value; The execution strategy of each edge node is dynamically attenuated and corrected by the updated models to obtain a target strategy, and the node device is controlled to operate in a low-carbon and energy-saving manner based on the target strategy.

[0082] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0083] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0084] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A system that specifies the functions of a box or boxes.

[0085] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction system that is implemented in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0087] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0088] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A low-carbon and energy-saving operation method for an edge node of the Internet of Things, characterized in that: The method comprises: Based on the real-time collected state information of edge nodes, a spatiotemporal correlation state vector is constructed through dynamic weighted fusion; Based on the state vector, a comprehensive energy efficiency evaluation index is constructed in combination with a dynamic power consumption model, a carbon emission model, and throughput and latency constraints; Taking energy efficiency evaluation indicators as the optimization goal, the deep Q network learns the mapping relationship between state and action, and generates an action strategy including frequency regulation, task scheduling and mode switching; Dynamically adjust the operating frequency, task queue priority, and node working mode of the node device according to the action strategy; Adaptively adjust the weights and learning rates of each model based on the deviation between the actual energy efficiency and the target value; The execution strategy of each edge node is dynamically attenuated and corrected by the updated models to obtain a target strategy, and the node device is controlled to operate in a low-carbon and energy-saving manner based on the target strategy.

2. The low-carbon and energy-saving operation method of the Internet of Things edge node according to claim 1, characterized in that: Based on the state vector, combined with the dynamic power consumption model, carbon emission model, throughput and latency constraints, a comprehensive energy efficiency evaluation index is constructed, including: Constructing a power consumption model that is positively correlated with the square of the computing load of the node device and the cube of the operating frequency of the node device; Building a carbon emission model based on the real-time power consumption of the node device and the ambient temperature of the environment in which it is located; Constructing a throughput model that reflects the coupling relationship between the computing efficiency of the node device and the network quality; Constructing a delay model that characterizes the combined effects of the operating frequency of the node device and the network status; The outputs of the power consumption model, the carbon emission model, the throughput model, and the delay model are weighted and combined to form an energy efficiency evaluation value.

3. The low-carbon and energy-saving operation method of the Internet of Things edge node according to claim 2, characterized in that: The comprehensive energy efficiency evaluation index is expressed as: ; in, It is a comprehensive energy efficiency evaluation indicator. is the power consumption model, It is a carbon emission model. is the throughput model, is the time delay model, are the first weight, the second weight, the third weight and the fourth weight respectively, is the normalization coefficient.

4. The low-carbon and energy-saving operation method of the Internet of Things edge node according to claim 3 is characterized in that: The energy efficiency evaluation index is used as the optimization target, and the state-action mapping relationship is learned through the deep Q network to generate an action strategy including frequency regulation, task scheduling and mode switching, including: Constructing an action space including dynamic adjustment of the operating frequency of the node device, task strategy scheduling and working mode switching; Constructing an immediate reward function whose value is negatively correlated with the comprehensive energy efficiency evaluation index and positively correlated with the stability of energy supply and the constraints of task delay; The instantaneous reward function is processed using a temporal difference algorithm, and a corresponding Q value is calculated to determine an action strategy for updating the action space.

5. The low-carbon and energy-saving operation method of the Internet of Things edge node according to claim 4, characterized in that: The Q value is expressed as: ; in, is the Q value of the state and action value function, is the mathematical expectation, is the immediate reward function, is the discount factor.

6. The low-carbon and energy-saving operation method of the Internet of Things edge node according to claim 5, characterized in that: The dynamically adjusting the operating frequency, task queue priority, and node working mode of the node device according to the action strategy includes: Determining a loss function of the deep Q network based on a deviation between a Q value currently output by the deep Q network and a target Q value; Obtaining updated policy parameters of the deep Q network according to the loss function and historical policy parameters of the deep Q network; updating the comprehensive energy efficiency evaluation value according to the updated strategy parameters to obtain an updated comprehensive energy efficiency evaluation value; According to the updated comprehensive energy efficiency evaluation value and the state vector, the optimal action for maximizing the Q value is determined, and the operating frequency, task queue priority and node working mode of the node device are dynamically adjusted based on the optimal action.

7. The low-carbon and energy-saving operation method of the Internet of Things edge node according to claim 6, characterized in that: The adaptive adjustment of the weights and learning rates of each model based on the deviation between the actual energy efficiency and the target value includes: Calculate the deviation between the actual comprehensive energy efficiency index and the target comprehensive energy efficiency index; performing an exponential processing on the deviation; The weights and learning rates of the models are adaptively adjusted according to the deviation amount after the exponential processing.

8. The low-carbon and energy-saving operation method of the Internet of Things edge node according to claim 7, characterized in that: The dynamically attenuated correction of the execution strategy of each edge node by the updated models to obtain the target strategy includes: Construct a multi-dimensional security boundary function; Calculate the corresponding safety threshold according to the multi-dimensional safety boundary function; Through the updated models, dynamic attenuation correction is performed according to the safety threshold and the operating frequency of the node device, and a strategy of correcting the operating frequency to the operating frequency adjustment value is determined as the target strategy.

9. The low-carbon and energy-saving operation method of the Internet of Things edge node according to claim 8, characterized in that: The method further comprises: Based on historical operating data sampling, a Pareto frontier solution set of energy efficiency indicators and computing performance is constructed through a multi-objective optimization algorithm; Provide a visual interactive interface to map the energy efficiency preference coefficients input by the user into weight parameters of the multi-objective optimization model; According to real-time energy efficiency requirements and performance constraints, the optimal strategy combination is dynamically selected from the Pareto solution set and fed back to the control end.

10. A low-carbon energy-saving operation system for edge nodes of the Internet of Things, characterized in that: The system comprises: The state vector module is used to construct a spatiotemporal correlated state vector through dynamic weighted fusion based on the state information of edge nodes collected in real time; An energy efficiency evaluation module, configured to construct a comprehensive energy efficiency evaluation index based on the state vector, in combination with a dynamic power consumption model, a carbon emission model, and throughput and latency constraints; The action strategy module is used to optimize the energy efficiency evaluation index and generate action strategies including frequency regulation, task scheduling and mode switching through deep Q network learning of the state-action mapping relationship; A dynamic adjustment module, used to dynamically adjust the operating frequency, task queue priority and node working mode of the node device according to the action strategy; The model update module is used to adaptively adjust the weights and learning rates of each model based on the deviation between the actual energy efficiency and the target value; The operation control module is used to dynamically attenuate and correct the execution strategy of each edge node through the updated models to obtain a target strategy, and control the low-carbon and energy-saving operation of the node device based on the target strategy.

Citation Information

Cited By

  • Smart watch control method and device based on Internet of Things

    CN120928739A

  • Energy-saving method and energy-saving system based on electric power big data

    CN121071529A

  • An energy-saving method and system based on power big data

    CN121071529B