EHES task scheduling method based on relaxation time perception and DQN
Through a task scheduling method based on relaxation time perception and deep reinforcement learning, the EHES task is decomposed into subtask sets, and the task execution order is dynamically adjusted, which solves the problem of EHES task scheduling failure and realizes efficient and sustainable operation of the system.
Patent Information
- Application Number
- CN202510608179.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-09-26
AI Technical Summary
During the task scheduling process, EHES fails to perceive the slack time in a timely manner, which greatly increases the probability of task scheduling failure. Especially when multiple tasks are approaching their deadlines at the same time, it is difficult to achieve sustainable operation of the system.
A task scheduling method based on relaxation time perception and deep reinforcement learning is adopted. By decomposing tasks into subtask sets, the reward function is designed in combination with energy neutrality theory, the task execution order is dynamically adjusted, and the scheduling strategy is optimized using the deep Q network to improve the task success rate.
The successful execution rate of EHES task scheduling is improved, energy waste and system downtime are reduced, and the sustainable operation of the system is ensured.
Smart Images

Figure BDA0005398824300000042 
Figure BDA0005398824300000043 
Figure BDA0005398824300000051
Abstract
Description
Technical Field
[0001] The present invention relates to the field of task scheduling of energy harvesting embedded systems, and in particular to an EHES task scheduling method based on slack time perception and DQN. Background Art
[0002] The Internet of Things (IoT) has driven the widespread adoption of embedded devices in areas such as industrial control, smart homes, healthcare, environmental monitoring, and smart cities. Embedded devices are typically battery-powered. This simple and easy-to-implement approach, however, limits the device's operating life. Once the battery is depleted, the device ceases operation, requiring replacement or recharging. This approach not only increases maintenance costs but also poses potential ecological risks to the environment. The rapid development of energy harvesting technology has provided a solution for achieving sustainable, permanent operation of battery-powered embedded systems. Energy harvesting technology utilizes renewable energy sources in the environment, such as solar energy, thermal energy, and mechanical vibration, to provide a continuous energy supply to embedded devices, significantly reducing reliance on traditional batteries and extending the device's lifespan. Embedded systems powered by energy harvesting technology are called energy harvesting embedded systems (EHES).
[0003] The unique characteristics of EHES compared to battery-powered embedded systems lie in the unlimited and dynamic nature of harvested energy. 1) Unlimited Energy: The upper limit of usable energy in battery-powered embedded systems is the maximum capacity of the battery, while the energy source of EHES is unlimited. This results in battery-powered embedded systems primarily relying on energy-saving technologies to extend device lifespan. However, due to the introduction of energy harvesting technology, the usable energy of EHES is unlimited, though it may be limited at certain moments. Therefore, the core of EHES energy utilization is no longer energy conservation, but rather the rational control of the energy utilization efficiency of each unit to maintain permanent system operation. 2) Dynamic Variability: Energy harvesting efficiency is primarily influenced by the environment, resulting in large fluctuations in the energy harvesting rate and difficulty in maintaining stable output. On the one hand, when the energy harvesting rate exceeds the consumption rate, reaching the battery capacity limit will result in energy waste. For example, the ALAP (As Late As Possible) scheduling strategy delays task execution as much as possible, resulting in excess energy waste. On the other hand, when the harvesting rate is insufficient to support device operation, the battery discharges to support device operation. If the battery is depleted, the system is forced to stop until sufficient energy is harvested and restarted, causing intermittent system switching. For example, the ASAP (As Soon As Possible) scheduling strategy executes tasks immediately once the stored energy meets the task start requirement, which can easily lead to unexpected system interruptions due to insufficient energy later. Therefore, the scheduling strategy needs to dynamically optimize task execution based on the collected energy to control the battery energy level within a reasonable range.
[0004] Affected by these two characteristics of EHES, the slack time of tasks in the task scheduling process of EHES is scattered. Especially when multiple tasks are approaching their deadlines at the same time, these slack times cannot be perceived in time, resulting in a significant increase in the probability of task scheduling failure. Summary of the Invention
[0005] In view of this, the present invention provides an EHES task scheduling method based on slack time perception and DQN. By integrating deep reinforcement learning and divide-and-conquer thinking, the system tasks are decomposed into multiple subtask sets. The slack time perception of the subtask sets is realized based on the battery energy neutral operation method. The slack time is fully utilized to dynamically adjust the task execution order, thereby improving the task scheduling completion rate while ensuring the sustainable operation of the system.
[0006] The technical solution adopted by the embodiment of the present invention to solve the technical problem is:
[0007] An EHES task scheduling method based on slack time perception and DQN includes:
[0008] Step S1, defining the system energy model: the energy model of the energy harvesting embedded system EHES consists of three parts: energy harvesting model, energy storage model and energy consumption model;
[0009] Step S2, subtask set calculation: tasks with the same sub-period in the task sequence are grouped into a subtask set, and J subtask sets Γ are obtained. j ;
[0010] Step S3, initializing the environment: obtaining the initial state S0 and creating an experience replay buffer;
[0011] Step S4, subtask set relaxation time RT j Perception: Perception of the subtask set Γ at time slot t j Relaxation time RT j ,determine the action space in combination with the battery energy level of the current time slot;
[0012] Step S5, define the available energy E of the current time slot t storage (t) and the subtask set Γ j Relaxation time RT j The state S(t) constituting time slot t: S(t) = {E storage (t),RT j};
[0013] Step S6, the state S(t) is sent to the agent, the agent explores the environment, and gives an action A(t) according to the state S(t): A(t) = {ξ j *|ξ j *∈γ}, where ξ j * is the set of currently executable tasks, γ={Γ,0} represents the total action space, Γ={Γ j} represents the subtask set, and 0 represents idle;
[0014] Step S7, after executing action A(t), the environment returns to the next state s t ', and gives an immediate reward R based on the battery energy level and task completion: R = R1-R2;
[0015] Step S8: The current interactive experience (s t ,a t ,s t ',r t ) is added to the experience replay buffer;
[0016] Step S9, randomly extracting small batches of samples from the experience replay buffer for training the Q network;
[0017] Step S10, regularly copying the parameters of the Q network to the target network to maintain the stability of the target network;
[0018] Step S11, subtask set relaxation time RT jUpdate: Run task τ in the action space at time slot t i Afterwards, we need to j The slack time is used for synchronization update;
[0019] Step S12, update the environment state: obtain the time slot t+1 to execute the task τ i+1 The previous environment state, and further perform the next task τ based on the updated environment state i+1 ;
[0020] Step S13, repeating steps S4 to S12 to process subtask set Γ j+1 , until all tasks in the task sequence are completed.
[0021] The task modeling of the energy consumption model in step S1 includes:
[0022] The task is represented as a two-tuple G = (∑, Γ), where ∑ represents the task set,
[0023] i)Σ={τ1,τ2,…,τ i ,…τ n} indicates that there are n tasks that need to be scheduled for execution, τ i For the i-th executed task, each task τ i The demand vector ∈Σ is R i ={N i ,K i ,C i ,T i ,D i ,E i}, where N i Represents the task τ i Name, K i Indicates the start time of the task, C i represents the worst-case task execution time, T i Indicates the task cycle, D i Indicates the relative deadline of the task, E i Indicates task execution C i The energy consumption of each time unit is defined as H. The LCM represents the least common multiple of all task periods in the task set. The super period can be calculated by the following formula:
[0024] H=LCM(T1,T2,…,T i ,…T n )
[0025] ii) Subtask set Γ j The demand vector R j Expressed as in Represented as subtask set Γ j Sub-period, RT j is the slack time of the subtask set;
[0026] If a task is scheduled to execute in time slot t, then task τ is executed in time slot t i Energy consumption E consume (t) is expressed as:
[0027]
[0028] Among them, S(τ i ,t) indicates whether task τ is executed in time slot t i The decision variable is defined as:
[0029]
[0030] The subtask set relaxation time perception process in step S4 includes:
[0031] Calculate each Γ j Relaxation time RT j , all the j =0 subtask set Γ j Merge into the set of tasks to be scheduledξ j *Inside; from the set of tasks to be scheduledξ j *Filter out tasks that are unschedulable and whose energy requirements are higher than the current remaining battery energy τ i ;
[0032] If the set of tasks to be scheduled is j * is not empty, then the action space is updated to the set of tasks to be scheduled ξ j *;
[0033] If the set of tasks to be scheduled is j * is empty and the remaining energy is higher than E max / 2, then the action space is updated to the set of schedulable tasks in the total task set Σ;
[0034] Otherwise, update the action space to γ = {Γ, 0}.
[0035] In the reward function R in step S7, R1 and R2 are expressed as:
[0036]
[0037] Where, E max and E min They are the upper and lower limits of battery energy level respectively.
[0038] For each sample in step S9, the target Q value is calculated using the Bellman equation, which is:
[0039] y=r+γmaxa'Q(s',a';θ)
[0040] Where θ is the parameter of the target network and γ is the discount factor.
[0041] The subtask set slack time updating process in step S11 includes:
[0042] If the task scheduled to be executed in the current time slot t belongs to the subtask set Γ j , then the slack time RT of the subtask set is j remains unchanged; otherwise, if the task scheduled to be executed in the current time slot t does not belong to the subtask set Γ j , then the slack time RT of the subtask set is j Minus 1; the initial slack time RT of the subtask set j The calculation formula is:
[0043] RT j =T s j -C s j
[0044] in, is a sub-period, is the subtask set Γ j The total execution time of all tasks in
[0045] It can be seen from the above technical solutions that the EHES task scheduling method based on relaxation time perception and DQN provided by the embodiment of the present invention is based on deep reinforcement learning combined with the divide-and-conquer idea to decompose all tasks into subtask sets. The action space is determined by the relaxation time constraints and energy constraints of the subtask sets, and the range of schedulable tasks is adjusted. DQN gives the tasks to be scheduled for execution in the corresponding action space, thereby increasing the probability of successful execution of the tasks. In addition, the present invention adopts a reward function based on energy efficiency optimization, aiming to reduce energy waste and system downtime in the scheduling process. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a flowchart of the EHES task scheduling method based on slack time perception and DQN of the present invention.
[0047] Figure 2 The algorithm structure diagram of the EHES task scheduling method based on slack time perception and DQN.
[0048] Figure 3The DQN decision process diagram of the EHES task scheduling method based on slack time perception and DQN. DETAILED DESCRIPTION
[0049] The technical solutions and technical effects of the present invention are further described in detail below with reference to the accompanying drawings of the present invention.
[0050] refer to Figure 1-Figure 3 As shown, the present invention provides an EHES task scheduling method based on slack time perception and DQN, which can improve the scheduling completion rate of the energy harvesting embedded system EHES execution task set. The specific implementation steps include:
[0051] Step S1, defining the system energy model: The energy model of the energy harvesting embedded system EHES consists of three parts: energy harvesting model, energy storage model and energy consumption model. Among them, the energy consumption model is mainly based on the computing unit and mainly describes the task energy consumption of the task set running in real time in the embedded system;
[0052] Step S2, subtask set calculation: tasks with the same sub-period in the task sequence are grouped into a subtask set Γ j , there are J subtask sets, j∈[1,J];
[0053] Step S3, initializing the environment: obtaining the initial state S0 and creating an experience replay buffer;
[0054] Step S4, subtask set relaxation time RT j Perception: Perception of the subtask set Γ at time slot t j Relaxation time RT j ,determine the action space in combination with the battery energy level of the current time slot;
[0055] Step S5, define the available energy E of the current time slot t storage (t) and the subtask set Γ j Relaxation time RT j The state S(t) constituting time slot t: S(t) = {E storage (t),RT j};
[0056] Step S6, the state S(t) is sent to the agent, the agent explores the environment, and gives an action A(t) according to the state S(t): A(t) = {ξ j *|ξ j *∈γ}, where ξ j * is the set of currently executable tasks, γ={Γ,0} represents the total action space, where Γ represents the set of subtasks and 0 represents idle;
[0057] Step S7, after executing action A(t), the environment returns to the next state s t ', and gives an immediate reward R based on the battery energy level and task completion: R = R1-R2;
[0058] Step S8: The current interactive experience (s t ,a t ,s t ',r t ) is added to the experience replay buffer;
[0059] Step S9: randomly extract small batches of samples from the experience replay buffer to train the Q network; (a small batch size refers to a batch size below a preset threshold)
[0060] Step S10, regularly copying the parameters of the Q network to the target network to maintain the stability of the target network;
[0061] Step S11, subtask set relaxation time RT j Update: Run task τ in the action space at time slot t i Afterwards, we need to j The slack time is used for synchronization update;
[0062] Step S12, update the environment state: obtain the time slot t+1 to execute the task τ i+1 The previous environment state, and further perform the next task τ based on the updated environment state i+1 ;
[0063] Step S13, repeating steps S4 to S12 to process subtask set Γ j+1 , until all tasks in the task sequence are completed.
[0064] The task modeling of the energy consumption model in step S1 includes:
[0065] The task is represented as a two-tuple G = (Σ, Γ), where Σ represents the task set and Γ = {Γ1, Γ2, …, Γ j ,…Γ J}, Γ represents the set of subtask sets, and the set of subtask sets refers to the subset of the task set, that is,
[0066] i)Σ={τ1,τ2,…,τ i ,…τ n} indicates that there are n tasks that need to be scheduled for execution, τ i For the i-th executed task, each task τ i The demand vector ∈Σ is R i ={N i ,Ki ,C i ,T i ,D i ,E i}, where N i Represents the task τ i Name, K i Indicates the start time of the task, C i represents the worst-case task execution time, T i Indicates the task cycle, D i Indicates the relative deadline of the task, E i Indicates task execution C i The energy consumption of each time unit is defined as H. The LCM represents the least common multiple of all task periods in the task set. The super period can be calculated by the following formula:
[0067] H=LCM(T1,T2,…,T i ,…T n ) (1)
[0068] ii) Subtask set Γ j The demand vector R j Expressed as in Represented as subtask set Γ j The sub-period of the task, that is, the super-period of all tasks in the sub-task set, RT j is the slack time of the subtask set. Slack time refers to the amount of time a task can be delayed without affecting the on-time completion of the entire schedule;
[0069] If a task is scheduled to execute in time slot t, then task τ is executed in time slot t i Energy consumption E consume (t) is expressed as:
[0070]
[0071] Among them, S(τ i ,t) indicates whether task τ is executed in time slot t i The decision variable is defined as:
[0072]
[0073] The subtask set relaxation time perception process in step S4 includes:
[0074] Calculate each Γ j Relaxation time RT j , all the j =0 subtask set Γ j Merge into the set of tasks to be scheduledξ j*Inside; from the set of tasks to be scheduledξ j *Filter out tasks that are unschedulable and whose energy requirements are higher than the current remaining battery energy τ i ;
[0075] If the set of tasks to be scheduled is j * is not empty, then the action space is updated to the set of tasks to be scheduled ξ j *;
[0076] If the set of tasks to be scheduled is j * is empty and the remaining energy is higher than E max / 2, then the action space is updated to the set of schedulable tasks in the total task set ∑;
[0077] Otherwise, update the action space to Υ = {Γ, 0}.
[0078] R1 and R2 in the reward function R in step S7 are expressed as:
[0079]
[0080]
[0081] Where, E max and E min They are the upper and lower limits of battery energy level respectively.
[0082] It can be seen that R1 is related to the basic execution status of the task; if in time slot t, the task does not miss the deadline and the action is not idle, the reward is 1, if the action is idle, the reward is -0.15; if the task misses the deadline and fails to execute, the reward is -50.
[0083] R2(e) is closely related to the energy level of the system; the battery energy level is maintained at E max and E min Between, the system can avoid entering the two unfavorable states of energy shortage and energy waste, and ultimately achieve the goal of reducing energy waste and increasing the number of successful task runs; it is defined here that if the remaining energy of the system is closer to the energy neutral value, the longer the system service life, so the system will receive additional rewards in this case.
[0084] For each sample in step S9, the target Q value is calculated using the Bellman equation, which is:
[0085] y=r +γmaxa'Q(s',a';θ) (6)
[0086] Where θ is the parameter of the target network and γ is the discount factor used to weigh the importance of future rewards.
[0087] The subtask set slack time update process in step S11 includes:
[0088] If the task scheduled to be executed in the current time slot t belongs to the subtask set Γ j , then the slack time RT of the subtask set is j remains unchanged; otherwise, if the task scheduled to be executed in the current time slot t does not belong to the subtask set Γ j , then the slack time RT of the subtask set is j Minus 1; the initial slack time RT of the subtask set j The calculation formula is:
[0089] RT j =T s j -C s j (7)
[0090] in, is a sub-period, is the subtask set Γ j The total execution time of all tasks in
[0091] This paper combines the reinforcement learning DQN method with the task scheduling method and proposes an EHES task scheduling method based on slack time perception and DQN to maximize the task success rate and maintain the permanent operation of the embedded system. The present invention has the following three beneficial effects:
[0092] (1) Based on the divide-and-conquer approach, the task is decomposed into multiple subtask sets and the corresponding task model is constructed. By sensing the slack time of each subtask set in real time, the task execution order is reasonably selected according to the preset scheduling strategy, thereby effectively improving the success rate of task execution.
[0093] (2) By integrating the DQN algorithm from deep reinforcement learning, we innovatively propose a reinforcement learning method that dynamically updates the action space. At each time step, the action space is dynamically adjusted in real time based on the slack time of the subtask set and the change in battery energy level. The optimal action is selected through an optimization strategy to achieve efficient operation of the system.
[0094] (3) Based on the energy neutrality theory (ENO), a reinforcement learning reward function is designed to quantify the deviation between the battery energy level and the energy neutral value, thereby motivating the agent to maintain a stable energy level and avoid energy excess or deficiency, thereby improving system performance.
[0095] The present invention provides an EHES task scheduling method based on slack time perception and DQN. The method first designs an action space update method that considers both time constraints and energy constraints. Based on deep reinforcement learning combined with the divide-and-conquer concept, all tasks are decomposed into subtask sets. By analyzing the slack time of the subtask sets and dynamically adjusting the range of schedulable tasks based on this, the probability of successful task execution is increased. The reward function based on energy efficiency optimization of the present invention is intended to reduce energy waste and system downtime during the scheduling process, ensuring that the energy harvesting embedded system works permanently.
[0096] The above disclosure is only a preferred embodiment of the present invention, and it is certainly not intended to limit the scope of the present invention. A person skilled in the art can understand that all or part of the processes of the above embodiment and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.
Claims
1. An EHES task scheduling method based on slack time perception and DQN, characterized in that: include: Step S1, defining the system energy model: the energy model of the energy harvesting embedded system EHES consists of three parts: energy harvesting model, energy storage model and energy consumption model; Step S2, subtask set calculation: tasks with the same sub-period in the task sequence are grouped into a subtask set, and J subtask sets Γ are obtained. j ; Step S3, initializing the environment: obtaining the initial state S0 and creating an experience replay buffer; Step S4, subtask set relaxation time RT j Perception: Perception of the subtask set Γ at time slot t j Relaxation time RT j ,determine the action space in combination with the battery energy level of the current time slot; Step S5, define the available energy E of the current time slot t storage (t) and the subtask set Γ j Relaxation time RT j The state S(t) constituting time slot t: S(t) = {E storage (t),RT j }; Step S6: Send the state S(t) to the agent, which explores the environment and takes action A(t) based on the state S(t): where ξ j * is the set of currently executable tasks, represents the total action space, Γ={Γ j } represents the subtask set, and 0 represents idle; Step S7, after executing action A(t), the environment returns to the next state s t ', and gives an immediate reward R based on the battery energy level and task completion: R = R1-R2; Step S8: The current interactive experience (s t ,a t ,s t ',r t ) is added to the experience replay buffer; Step S9, randomly extracting small batches of samples from the experience replay buffer for training the Q network; Step S10, regularly copying the parameters of the Q network to the target network to maintain the stability of the target network; Step S11, subtask set relaxation time RT j Update: Run task τ in the action space at time slot t i Afterwards, we need to j The slack time is used for synchronization update; Step S12, update the environment state: obtain the time slot t+1 to execute the task τ i+1 The previous environment state, and further perform the next task τ based on the updated environment state i+1 ; Step S13, repeating steps S4 to S12 to process subtask set Γ j+1 , until all tasks in the task sequence are completed.
2. The EHES task scheduling method based on slack time perception and DQN as claimed in claim 1, characterized in that: The task modeling of the energy consumption model in step S1 includes: The task is represented as a two-tuple G = (∑, Γ), where Σ represents the task set, i)∑={τ1,τ2,…,τ i ,…τ n } indicates that there are n tasks that need to be scheduled for execution, τ i For the i-th executed task, each task τ i ∈∑ The demand vector is R i ={N i ,K i ,C i ,T i ,D i ,E i }, where N i Represents the task τ i Name, K i Indicates the start time of the task, C i represents the worst-case task execution time, T i Indicates the task cycle, D i Indicates the relative deadline of the task, E i Indicates task execution C i The energy consumption of each time unit is defined as H. The LCM represents the least common multiple of all task periods in the task set. The super period can be calculated by the following formula: H=LCM(T1,T2,…,T i ,…T n ) ii) Subtask set Γ j The demand vector R j Expressed as in Represented as subtask set Γ j Sub-period, RT j is the slack time of the subtask set; If a task is scheduled to execute in time slot t, then task τ is executed in time slot t i Energy consumption E consume (t) is expressed as: Among them, S(τ i ,t) indicates whether task τ is executed in time slot t i The decision variable is defined as:
3. The EHES task scheduling method based on slack time perception and DQN as claimed in claim 2, characterized in that: The subtask set relaxation time perception process in step S4 includes: Calculate each Γ j Relaxation time RT j , all the j =0 subtask set Γ j Merge into the set of tasks to be scheduledξ j *Inside; from the set of tasks to be scheduledξ j *Filter out tasks that are unschedulable and whose energy requirements are higher than the current remaining battery energy τ i ; If the set of tasks to be scheduled is j * is not empty, then the action space is updated to the set of tasks to be scheduled ξ j *; If the set of tasks to be scheduled is j * is empty and the remaining energy is higher than E max / 2, then the action space is updated to the set of schedulable tasks in the total task set ∑; Otherwise, update the action space to 4. The EHES task scheduling method based on slack time perception and DQN as claimed in claim 3, characterized in that: In the reward function R in step S7, R1 and R2 are expressed as: Where, E max and E min They are the upper and lower limits of battery energy level respectively.
5. The EHES task scheduling method based on slack time perception and DQN as claimed in claim 4, characterized in that: For each sample in step S9, the target Q value is calculated using the Bellman equation, which is: y=r+γmaxa'Q(s',a';θ) Where θ is the parameter of the target network and γ is the discount factor.
6. The EHES task scheduling method based on slack time perception and DQN as claimed in claim 5, characterized in that: The subtask set slack time updating process in step S11 includes: If the task scheduled to be executed in the current time slot t belongs to the subtask set Γ j , then the slack time RT of the subtask set is j remains unchanged; otherwise, if the task scheduled to be executed in the current time slot t does not belong to the subtask set Γ j , then the slack time RT of the subtask set is j Minus 1; the initial slack time RT of the subtask set j The calculation formula is: RT j =T s j -C s j in, is a sub-period, is the subtask set Γ j The total execution time of all tasks in