An EHES power management method and system based on energy neutrality perception
Through the energy-neutral perception EHES power management method, combined with the Actor-Critic model and DVFS technology, the voltage and frequency of task execution are adjusted, and the energy imbalance problem of energy collection embedded equipment is solved, the energy neutral and stable operation of the system is achieved, and the energy utilization efficiency is improved.
Patent Information
- Application Number
- CN202411279060.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-09-12
AI Technical Summary
When dealing with energy-harvesting embedded devices, the existing power management strategies fail to fully consider the real-time energy supply capacity of the equipment, resulting in insufficient or excessive energy supply, resulting in intermittent operation of the system or waste of energy, making it difficult to achieve energy-neutral operation.
The EHES power management method based on energy neutral perception is adopted. By defining the system energy model and energy neutral perception model, combining the Actor-Critic model and DVFS technology, the voltage and frequency of task execution are adjusted to achieve the balance of energy consumption and collection, and energy management is optimized using reinforcement learning methods.
The energy neutral operation of the embedded system is realized, ensuring the stable operation of the system in the dynamic energy changes, improving energy utilization efficiency, and ensuring the permanent operation of the system.
Smart Images

Figure BDA0005041088580000041 
Figure BDA0005041088580000042 
Figure BDA0005041088580000051
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power management of embedded systems, and particularly to an EHES power management method and system based on energy neutrality awareness. Background Art
[0002] In recent years, small embedded devices have been used more and more widely and have become increasingly important in our lives, covering fields such as smart cities, health detection, industrial monitoring, agricultural Internet of Things, and wearable devices. As an important component of embedded devices, the battery has always been one of its main drivers. Currently, some embedded devices are deployed in inaccessible areas, where replacement and maintenance are difficult, and the battery has a limited lifespan and high maintenance costs, thus restricting the persistent operation of the devices. To solve this problem, obtaining energy from the environment (such as solar energy, mechanical energy, kinetic energy, etc.) is a feasible solution. Energy harvesting technology converts environmental energy into electrical energy through an energy harvester to power small embedded devices, thereby achieving the permanent operation of small embedded devices. This embedded system that harvests energy from the environment and stores the energy in an energy storage unit is defined as an energy harvesting embedded system, which consists of traditional embedded system components and an energy management component. In an embedded system using energy harvesting for power supply, an important challenge is to balance the energy consumption and energy harvesting of the embedded device so that the system can operate continuously without relying on an external power supply. This mode is called energy neutral operation.
[0003] The dynamic characteristics of the energy harvesting process are manifested as non-linear changes, which is particularly significant in scenarios such as solar energy harvesting. The output power of a solar energy harvesting system is affected by environmental factors such as seasonal light intensity fluctuations and cloud cover, resulting in a non-linear dynamic fluctuation characteristic of energy supply. This non-linear dynamic change in the energy harvested by the energy harvester may cause the system to temporarily fail due to insufficient energy supply or energy waste due to excessive energy. Traditional power management strategies, such as Dynamic Voltage and Frequency Scaling (DVFS) and Dynamic Power Management (DPM), show their limitations when dealing with energy harvesting embedded devices. These strategies mainly focus on reducing energy consumption by adjusting voltage and frequency, but do not fully consider the real-time energy supply capacity of the device, and are insufficient in adapting to the dynamics of energy harvesting, easily causing energy waste or intermittent operation of the system. Summary of the Invention
[0004] In view of this, the present invention provides an EHES power management method and system based on energy-neutral awareness, which fully considers the real-time energy supply capacity of the device, senses the energy-neutral state according to the energy-neutral distance, and balances the energy consumption and collection of the embedded system by adjusting the voltage and frequency of task execution, thereby shortening the energy-neutral distance and achieving energy-neutral operation.
[0005] The technical solution adopted by the embodiment of the present invention to solve its technical problems is:
[0006] An EHES power management method based on energy-neutral awareness, including:
[0007] Step S1, define a system energy model, which consists of an energy harvesting model, an energy storage model, a power management model, and an energy consumption model; define a task sequence Γ = {τ1, τ2, …, τ i}, where τ i is the i-th executed task, and the requirement vector R i of each task τ i = {T ai , C i , D i , T i , W i , <f i , U i >}, where T ai represents the start time of the task, C i represents the execution duration of the task, D i represents the deadline of the task, T i represents the task period, <f i , U i > represents the frequency / voltage pair that can be executed when the CPU is working. Assume that the CPU has N discrete execution frequencies, then f i ∈ {f1, f2, …, f N}, and a frequency f i corresponds to a voltage U i ; W i is the load of the task, representing the number of clock cycles required for task execution;
[0008] Step S2, define an energy-neutral awareness model <E d (i), E s (i), EN th , E dead , E max >, where E d (i) is the energy-neutral distance before the execution of task τ i , and E s (i) is the energy-neutral distance before the execution of task τ iRemaining energy before execution, EN th Is the energy neutral value, E dead Is the minimum energy for system death, E max Is the maximum stored energy;
[0009] Step S3, system initialization, obtain the initial states such as the task sequence to be scheduled and executed and the current battery energy level of the system, and initialize the network parameters of the Actor network and the Critic network;
[0010] Step S4, define task τ i Energy neutral distance E before execution d (i), the energy harvesting situation E i-1 during the execution of task τ harvest (i - 1) and the task τ i in the upcoming task sequence Γ form state S(i), S(i) = [E d (i), E harvest (i - 1), τ i ;
[0011] Step S5, send the state S(i) to the agent, and the agent outputs action A(i) according to the state S(i), A(i) = [f i , U i , i ∈ [0, N], where A(i) is the frequency f i required to execute the task τ i corresponding to S(i) and voltage U i ;
[0012] Step S6, the system uses the DVFS technology to adjust the frequency and voltage of the currently executing task according to <f i , U i > corresponding to A(i), so that the system realizes energy - neutral operation;
[0013] Step S7, after executing action A(i), give an immediate reward R(i) according to whether the task τ i is successfully executed, the current energy - neutral state and the voltage - frequency of task execution;
[0014] Step S8, use the Critic network to estimate the state value of the current state and calculate the TD error δ(i);
[0015] Step S9, use the stochastic gradient descent method to update the Critic network parameters with β as the Critic learning rate to minimize the prediction loss of the Critic network, and realize the update of the Critic network, where the prediction loss function L C = (δ(i)) 2 ;
[0016] Step S10, use the stochastic gradient descent method to update the Actor network parameters with α as the Actor learning rate, minimize the prediction loss of the Actor network, and achieve the update of the Actor network. Among them, the prediction loss function L of the Actor network A = -π θ (S(i), A(i))δ(i);
[0017] Step S11, update the environmental state, and obtain the environmental state before executing task τ at time slot i + 1, that is, S(i) = S(i + 1), S(i + 1) = [E i+1 (i + 1), E d (i), τ harvest (i), τ i+1 , and further execute the next task τ based on the updated environmental state i+1 ;
[0018] Step S12, repeatedly execute Step S4 - Step S11 until all tasks in the task sequence Γ are executed
[0019] Preferably, in the system energy model in Step S1:
[0020] Energy harvesting model: Define the energy harvested within the time interval [t1, t2] as E harvest as:
[0021]
[0022] where P h (t) is the power of energy harvesting at time t;
[0023] Energy storage model: Define the energy E stored in the storage unit at time t storage as:
[0024] E storage (t + 1) = E storage (t) + E harvest (t) - E c (t)
[0025] When the stored energy exceeds the storage capacity of the energy storage unit, define the energy stored at the current time slot t as:
[0026] E storage (t) = min(E storage (t), E max )
[0027] where E max is the maximum energy that the system can store;
[0028] Power management model: The dynamic power consumption P of the processor executing a task at frequency f d is defined as:
[0029] P d = C × V 2 × f
[0030] Where C is the capacitance, V is the circuit voltage, and f is the CPU frequency when executing the task;
[0031] Energy consumption model: The energy consumption E required to complete each task τ i is c as follows:
[0032] E c = P d × C i
[0033] Where the execution duration C of task τ i is calculated as: i
[0034] C i = W i × 1 / f i
[0035] Where W i represents the total number of clock cycles required to execute task τ i and 1 / f i represents the time length of each clock cycle at frequency f i .
[0036] Preferably, in step S2, based on the energy neutrality awareness model <E d (i), E s (i), EN th , E dead , E max >, the energy collected by the embedded device within a period T is E harvest , and the energy consumed is E c , and the energy-neutral operation is expressed as:
[0037]
[0038] where E0 is the initial energy level of the storage unit;
[0039] The formula for the energy neutrality value EN th is expressed as:
[0040]
[0041] E d (i) represents Es (i) The relative distance from the energy neutral value EN th is expressed by the formula:
[0042] E d (i) = E s (i) - EN th .
[0043] Preferably, the functional expression of the immediate reward R(i) in the step S7 is:
[0044]
[0045] where T ai + C i ≤ D i indicates that the task is completed before the deadline, and at this time R(i) is a positive reward. T ai + C i > D i indicates that the task misses the deadline and fails to execute, and at this time R(i) is a negative reward;
[0046] In the formula, is the energy neutral distance reward function, and the range of the reward value is (0, is the task execution frequency and voltage reward function. The agent gives <f i , U i > for task execution according to the state. DVFS adjusts the voltage and frequency to shorten the energy neutral distance and achieve energy neutrality. The range of the reward value is (0, 1).
[0047] Preferably, the calculation formula of the TD error δ(i) in the step S8 is expressed as:
[0048] δ(i) = R(i) + γV w (S(i + 1)) - V w (S(i))
[0049] In the formula, γ is the discount factor, and V w (S(i)) is the state value function with w as a parameter, representing the approximate value of the state S(i). V w (S(i + 1)) is the state value function with w as a parameter, representing the approximate value of the state S(i + 1).
[0050] As can be seen from the above technical solutions, the EHES power management method and system based on energy-neutral awareness provided by the embodiments of the present invention enable the system to sense the energy-neutral operation during the execution of the previous task, calculate the energy-neutral distance, and use the DVFS technology to adjust the frequency and voltage of the task being executed, thereby achieving energy-neutral operation during the execution of the task. This allows the system to use the collected energy at an appropriate rate, ensuring the permanent operation of the energy-harvesting embedded system, improving energy utilization efficiency, and ensuring the stable operation and energy-efficiency optimization of the system in the dynamic changes of energy harvesting. Brief Description of the Drawings
[0051] Figure 1 It is the overall flowchart of the EHES power management method based on energy-neutral awareness. Detailed Embodiments
[0052] The following further elaborates on the technical solutions and technical effects of the present invention in conjunction with the drawings of the present invention.
[0053] When implementing the energy-aware power management method based on reinforcement learning and DVFS, while ensuring the success rate of task operation, the battery energy-neutral operation is considered, that is, the battery energy level is maintained between E max and E dead for a period of time, which can prevent the system from experiencing temporary death due to insufficient energy or energy waste due to excessive energy, enabling the embedded system to operate permanently.
[0054] The specific implementation process of the power management method for the embedded system based on energy-neutral awareness is as follows:
[0055] Step S1, define the system energy model, which consists of an energy harvesting model, an energy storage model, a power management model, and an energy consumption model; define the task sequence Γ = {τ1, τ2, …, τ i}, where τ i is the i-th executed task, and the requirement vector R i of each task τ i = {T ai , C i , D i , T i , W i , <f i , U i >}, where T ai represents the start time of the task, C i represents the task execution duration, D i represents the task deadline, T i represents the task period, <f i , U i>Denoted as the frequency / voltage pairs that can be executed when the CPU is working. Assume that the CPU has N discrete execution frequencies, then f i ∈{f1,f2,…,f N}, and a frequency f i corresponds to a voltage U i ; W i is the load of the task, indicating the number of clock cycles required for task execution;
[0056] Step S2, define the energy-neutral awareness model <E d (i),E s (i),EN th ,E dead ,E max >, where E d (i) is the energy-neutral distance before task τ i execution, E s (i) is the remaining energy before task τ i execution, EN th is the energy-neutral value, E dead is the minimum energy for system death, E max is the maximum energy for storage;
[0057] Step S3, system initialization, obtain the initial states such as the task sequence to be scheduled and executed and the current battery energy level of the system, where E harvest (0) = 0, E storage (0) = E max / 2; initialize the network parameters of the Actor network and the Critic network;
[0058] Step S4, define the energy-neutral distance E i (i) before task τ d execution, the energy harvesting situation E i-1 (i - 1) during task τ harvest execution and the task τ i in the upcoming task sequence Γ to form the state S(i), S(i) = [E d (i),E harvest (i - 1),τ i ;
[0059] Step S5, input the state S(i) into the agent, and the agent outputs the action A(i) according to the state S(i), A(i) = [f i ,U i , i ∈ [0,N], where A(i) is the frequency f i and voltage U i required to execute the task τ i corresponding to S(i);
[0060] Step S6, the system uses the DVFS technology to adjust the frequency and voltage of the currently executing task according to <f i ,U i > corresponding to A(i), so that the system realizes energy-neutral operation;
[0061] Step S7, after executing action A(i), give an immediate reward R(i) according to whether the task τ i is successfully executed, the current energy-neutral state and the voltage frequency of the task execution;
[0062] Step S8, use the Critic network to estimate the state value of the current state and calculate the TD error δ(i);
[0063] Step S9, use the stochastic gradient descent method to update the Critic network parameters with β as the Critic learning rate to minimize the prediction loss of the Critic network, realizing the update of the Critic network. The Critic will use the new network parameters to help the Actor calculate the optimal value of the state. Among them, the prediction loss function L of the Critic network C = (δ(i)) 2 ;
[0064] Step S10, use the stochastic gradient descent method to update the Actor network parameters with α as the Actor learning rate to minimize the prediction loss of the Actor network, realizing the update of the Actor network. Among them, the prediction loss function L of the Actor network A = -π θ (S(i), A(i))δ(i);
[0065] Step S11, update the environmental state, and obtain the environmental state before executing the task τ i+1 in the (i + 1)-th time slot, that is, S(i) = S(i + 1), S(i + 1) = [E d (i + 1), E harvest (i), τ i+1 . And further execute the next task τ i+1 based on the updated environmental state; E d (i + 1) is the energy-neutral distance before the execution of the task τ i+1 , E harvest (i) is the energy harvesting situation during the execution of the task τ i , τ i+1 is the next task to be executed;
[0066] Step S12, repeat Steps S4 - S11 until all tasks in the task sequence Γ are executed. Until all tasks in the task sequence Γ are executed.
[0067] In the system energy model in step S1:
[0068] Energy harvesting model: The role of the energy harvesting model is to efficiently harvest various forms of energy in the environment and convert it into electrical energy as an energy source. However, energy harvesting is random. The harvested energy is stored in the energy storage module to provide energy for subsequent devices. Define the energy harvested within the time interval [t1, t2] as E harvest as:
[0069]
[0070] where, P h (t) is the power of energy harvesting at time t;
[0071] Energy storage model: The energy storage model is a device that dynamically manages energy storage, ensuring that energy is available during peak demand and storing it when there is an excess of energy supply. The energy E stored in the storage unit at time t storage is defined as:
[0072] E storage (t + 1) = E storage (t) + E harvest (t) - E c (t) (2)
[0073] When the stored energy exceeds the storage capacity of the energy storage unit, the energy stored in the current time slot t is defined as:
[0074] E storage (t) = min(E storage (t), E max ) (3)
[0075] where, E max is the maximum energy that the system can store;
[0076] Power management model: The power management model is equipped with a frequency / voltage regulator to adjust and output a stable voltage to adapt to the voltage requirements during different task executions and achieve energy-neutral operation. The dynamic power consumption P of the processor executing a task at frequency f d is defined as:
[0077] P d = C × V 2 × f (4)
[0078] In the formula, C is the capacitance, V is the circuit voltage, and f is the CPU frequency when executing the task;
[0079] Energy consumption model: The energy consumption model is mainly the energy consumption of tasks running in real time in the device. Task τi The demand vector is R i ={T ai , C i , D i , T i , W i , <f i , U i >,}, to complete each task τ i The energy consumption E c required is:
[0080] E c = P d × C i (5)
[0081] In the formula, for task τ i the execution duration C i has the following calculation formula:
[0082] C i = W i × 1 / f i (6)
[0083] In the formula, W i represents the total number of clock cycles required to execute task τ i , and 1 / f i represents the time length of each clock cycle at frequency f i .
[0084] In step S2, based on the energy-neutral awareness model <E d (i), E s (i), EN th , E dead , E max >, energy-neutral operation means that the embedded device achieves a balance between energy consumption and energy harvesting, enabling the system to operate continuously without relying on an external power supply. Energy-neutral awareness senses the energy-neutral state based on the energy-neutral distance before the execution of task τ i , that is, the smaller the absolute value of the energy-neutral distance, the better the energy-neutral state; conversely, the larger the absolute value of the energy-neutral distance, the worse the energy-neutral state. The power management strategy adjusts the voltage and frequency of the current task τ i execution according to the energy-neutral distance before the execution of task τ th , that is, the relative distance between the remaining energy level and the energy-neutral value EN i through the DVFS technology, so that the embedded system achieves a balance between energy consumption and energy harvesting, thereby shortening the energy-neutral distance, realizing energy-neutral operation, optimizing the energy utilization efficiency, and enabling the embedded device to operate continuously.
[0085] The energy collected by the embedded device within a period of time T is E harvest The energy consumed is E c The energy-neutral operation is expressed as:
[0086]
[0087] where E0 is the initial energy level of the storage unit;
[0088] The energy-neutral value EN th is expressed by the formula:
[0089]
[0090] E d (i) represents the relative distance between E s (i) and the energy-neutral value EN th The formula is expressed as:
[0091] E d (i) = E s (i) - EN th (9)
[0092] The functional expression of the immediate reward R(i) in step S7 is:
[0093]
[0094] where T ai + C i ≤ D i indicates that the task is completed before the deadline, and at this time R(i) is a positive reward. T ai + C i > D i indicates that the task misses the deadline and fails to execute, and at this time R(i) is a negative reward;
[0095] In the formula, is the energy-neutral distance reward function, which is specifically manifested as: the smaller the energy-neutral distance, the closer the current energy level of the system is to the energy-neutral value. Therefore, as the energy-neutral distance becomes smaller, the obtained reward value should be higher, and the range of the reward value is is the task execution frequency and voltage reward function. The agent gives <f i , U i > according to the state. DVFS adjusts the voltage and frequency to shorten the energy-neutral distance and achieve energy neutrality. Therefore, the more reasonable the <f i , U i > given, the higher the obtained reward value, and the range of the reward value is (0, 1).
[0096] In step S8, the calculation formula of TD error δ(i) is expressed as:
[0097] δ(i) = R(i) + γV w (S(i + 1)) - V w (S(i))(11)
[0098] In the formula, γ is the discount factor, and V w (S(i)) is the state value function with w as the parameter, representing the approximate value of state S(i), and V w (S(i)) is the state value function with w as the parameter, representing the approximate value of state S(i + 1).
[0099] Through the above solution, the system senses the energy-neutral operation during the execution of the previous task, calculates the energy-neutral distance, and uses the DVFS technology to adjust the frequency and voltage of the task being executed, so as to achieve energy-neutral operation during the execution of this task, enabling the system to use the collected energy at an appropriate rate and ensuring the permanent operation of the energy-harvesting embedded system.
[0100] Furthermore, the present invention also provides an EHES power management system based on energy-neutral perception for implementing Figure 1 and the aforementioned method.
[0101] Compared with the background technology, the present invention has the following beneficial effects:
[0102] The present invention combines the reinforcement learning method with the power management technology, proposes an EHES power management method based on energy-neutral perception, realizes energy-neutral operation, and maintains the permanent operation of the embedded system. The present invention has the following two beneficial effects:
[0103] (1) It provides a solution for the energy-harvesting embedded system based on energy-neutral perception, laying a foundation for the dynamic and adaptive power management strategy.
[0104] (2) By combining the Actor-Critic model with the DVFS technology, it realizes the adjustment of the frequency and voltage levels of the scheduled tasks, enabling the system to use the collected energy at an appropriate rate, thereby achieving energy-neutral operation and ensuring the permanent operation of the energy-harvesting embedded system.
[0105] The present invention provides a solution for the energy-harvesting embedded system based on energy-neutral perception, senses the energy-neutral state according to the energy-neutral distance, and makes the energy consumption and collection of the embedded system reach balance by adjusting the voltage and frequency of task execution, thereby shortening the energy-neutral distance, realizing energy-neutral operation, improving the energy utilization efficiency, and ensuring the stable operation and energy efficiency optimization of the system in the dynamic changes of energy harvesting.
[0106] The above-disclosed is only the preferred embodiment of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. An EHES power management method based on energy neutrality perception, characterized in that, Including: Step S1, define the system energy model, which consists of an energy harvesting model, an energy storage model, a power management model, and an energy consumption model; define the task sequence Γ = {τ1, τ2, …, τ i}, where τ i is the i-th executed task, and the requirement vector R i of each task τ i = {T ai , C i , D i , T i , W i , <f i , U i >}, where T ai represents the start time of the task, C i represents the execution duration of the task, D i represents the deadline of the task, T i represents the task period, <f i , U i > represents the frequency / voltage pair that can be executed when the CPU is working. Assume that the CPU has N discrete execution frequencies, then f i ∈ {f1, f2, …, f N}, and a frequency f i corresponds to a voltage U i ; W i is the load of the task, indicating the number of clock cycles required for task execution; Step S2, define the energy-neutral perception model <E d (i), E s (i), EN th , E dead , E max >>, where E d (i) is the energy-neutral distance before the execution of task τ i , E s (i) is the remaining energy before the execution of task τ i , EN th is the energy-neutral value, E dead is the minimum energy for system death, E max is the maximum energy for storage; Step S3: System initialization, obtaining initial states such as the task sequence to be scheduled and executed and the current battery energy level of the system, and initializing the network parameters of the Actor network and the Critic network; Step S4, define task τ i Execute the previous energy-neutral distance E d (i), task τ i-1 Energy harvesting situation E during execution harvest (i - 1) and the task τ in the upcoming task sequence Γ i constitute state S(i), S(i) = [E d (i), E harvest (i - 1), τ i ; Step S5, send the state S(i) to the agent, and the agent outputs an action A(i) according to the state S(i), A(i) = [f i , U i , i ∈ [0, N], where A(i) is the frequency f i and voltage U i required to execute the task τ i ; Step S6, the system uses the DVFS (Dynamic Voltage and Frequency Scaling) technology to adjust the frequency and voltage of the currently executing task according to <f i ,U i > so that the system can achieve energy-neutral operation; Step S7, after performing action A(i), give an immediate reward R(i) according to whether the task τ i is successfully executed, the current energy-neutral state, and the voltage frequency of task execution; Step S8: Using the Critic network to estimate the state value of the current state and calculating the TD error δ(i) according to the immediate reward R(i); Step S9: Use the stochastic gradient descent method to update the Critic network parameters with β as the Critic learning rate to minimize the prediction loss of the Critic network, thereby realizing the update of the Critic network. Among them, the prediction loss function L of the Critic network C =(δ(i)) 2 ; Step S10, use the stochastic gradient descent method to update the Actor network parameters with α as the Actor learning rate, minimize the prediction loss of the Actor network, and implement the update of the Actor network. Among them, the prediction loss function L of the Actor network A = -π θ (S(i), A(i))δ(i); Step S11, update the environmental state and obtain the task τ to be executed in the (i + 1)-th time slot i+1 The previous environmental state, i.e., S(i) = S(i + 1), S(i + 1) = [E d (i + 1), E harvest (i), τ i+1 , and further execute the next task τ based on the updated environmental state i+1 ; Step S12: Repeatedly execute Step S4 - Step S11 until all tasks in the task sequence Γ are completed.
2. The EHES power management method based on energy neutrality perception according to claim 1, wherein In the system energy model in Step S1: Energy harvesting model: Define the energy collected within the time interval [t1, t2] as E harvest as follows: Among them, P h (t) is the power of energy harvesting at time t; Energy storage model: The energy E stored in the storage unit at time t+1 storage (t + 1) is defined as: E storage (t + 1)= E storage (t)+ E harvest (t)- E c (t) When the stored energy exceeds the storage capacity of the energy storage unit, the energy stored at the current moment t is defined as: E storage (t) = min(E storage (t), E max ) Among them, E max is the maximum energy that the system can store; Power management model: The dynamic power consumption P of the processor executing a task at frequency f d is defined as: P d = C × V 2 × f In the formula, C is the capacitance, V is the circuit voltage, and f is the CPU frequency when executing tasks; Energy consumption model: Energy consumption E required to complete each task τ i is c as follows: E c = P d × C i In the formula, the execution duration C i of task τ i is calculated by the following formula: C i = W i × 1 / f i Where, W i represents the total number of clock cycles required to execute task τ i and 1 / f i represents the time length of each clock cycle at frequency f i 3. The method for power management of EHES based on energy neutrality awareness as described in claim 2, characterized in that In the step S2, based on the energy-neutral perception model <E d (i), E s (i), EN th , E dead , E max >>, the energy collected by the embedded device within a period T is E harvest , and the consumed energy is E c . The energy-neutral operation is expressed as: where E0 is the initial energy level of the storage unit; Energy neutral value EN th The formula is expressed as: E d (i) represents E s (i)'s relative distance from the energy neutral value EN th is expressed by the formula as follows: E d (i) = E s (i) - EN th 。 4. The EHES power management method based on energy neutrality perception according to claim 3, wherein The functional expression of the immediate reward R(i) in Step S7 is: where T ai + C i ≤ D i means that the task is completed before the deadline, and at this time R(i) is a positive reward, T ai + C i > D i means that the task misses the deadline and fails to execute, and at this time R(i) is a negative reward; In the formula, is the energy-neutral distance reward function, and the range of the reward value is (0, is the task execution frequency and voltage reward function. The agent gives the <f i , U i > according to the state. DVFS shortens the energy-neutral distance and achieves energy neutrality by adjusting the voltage and frequency. The range of the reward value is (0, 1).
5. The energy-neutrality-aware EHES power management method according to claim 4, wherein The calculation formula of the TD error δ(i) in Step S8 is expressed as: δ(i) = R(i) + γV w (S(i + 1)) - V w (S(i)) where γ is the discount factor, and V w (S(i)) is the state value function parameterized by w, representing the approximate value of state S(i), and V w (S(i + 1)) is the state value function parameterized by w, representing the approximate value of state S(i + 1).
6. The energy-neutrality-aware EHES power management method according to claim 5, characterized in that In the step S3, in the initial state, E harvest (0) = 0, E storage (0) = E max / 2.
7. An EHES power management system based on energy-neutral awareness, characterized in that, For implementing the method according to any one of claims 1 - 6.
Citation Information
Patent Citations
Sensor node energy supplement and data acquisition method based on deep reinforcement learning
CN115835350A
Energy management method for energy collection wireless sensor
CN117412367A