Method for task offloading and resource allocation in energy harvesting MEC system
By decoupling task offloading and resource allocation in energy harvesting MEC systems, and combining deep reinforcement learning and adaptive genetic algorithms, the problem of limited computing power and battery capacity in dynamic MEC systems is solved, achieving efficient energy management and task offloading for terminal devices.
Patent Information
- Application Number
- CN202310212011.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-07
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2043-03-07
AI Technical Summary
In dynamic energy harvesting MEC systems, existing technologies struggle to effectively address the issues of insufficient computing power and limited battery capacity in terminal devices, especially when devices are distributed in remote or hazardous environments. Traditional algorithms require frequent iterations and are costly, while deep reinforcement learning-based methods have failed to effectively impose long-term performance constraints.
This paper proposes a method to decouple the task unloading and resource allocation problem based on Lyapunov optimization theory. By combining deep reinforcement learning and an improved adaptive genetic algorithm, a task unloading strategy and resource allocation scheme are designed. The action space and reward function are defined through Markov decision process, and the resource allocation is optimized using the adaptive genetic algorithm to achieve optimal unloading decision and allocation.
Under long-term queue stability constraints, the execution time and total energy consumption cost of terminal devices are reduced, and the stability and adaptability of the system are improved. Simulation results show superior performance.
Smart Images

Figure CN116209084B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mobile communication technology and relates to a method for task offloading and resource allocation in an energy harvesting MEC system. Background Technology
[0002] With the rapid development of mobile communication and IoT technologies, the number of smart terminals and data traffic have exploded. Driven by technologies such as artificial intelligence, machine learning, and edge intelligence, emerging applications including virtual reality / augmented reality, autonomous driving, smart cities, and smart factories are constantly emerging. However, terminal devices are limited by manufacturing processes and costs, resulting in severely restricted computing, storage, and battery capacities, making it difficult to meet the processing demands of these emerging applications. Mobile edge computing (MEC) allows terminal devices to offload computing tasks to the cloud, alleviating the problem of limited device resources. However, the limited battery capacity of terminal devices and insufficient edge computing power remain unresolved, making it difficult to meet the processing needs of these emerging applications.
[0003] In certain scenarios, such as when devices are located in remote or hazardous environments, it is difficult to obtain a continuous battery power supply from the traditional power grid. Therefore, to meet the long-term battery life requirements of terminal devices, energy harvesting (EH) technology is typically used to enable devices to obtain energy from the environment to support communication and task processing. This technology has become an important means of achieving green mobile communication. Combining EH and MEC technologies can effectively solve the problems of insufficient device computing power and limited battery capacity, supporting computationally intensive and latency-sensitive applications through green mobile communication. This is of great significance for constructing green and energy-efficient MEC systems.
[0004] Combining EH technology with MEC task offloading technology, the formulation of task offloading strategies and resource allocation in a green communication manner has attracted widespread attention from many scholars. Some major achievements include: (1) An online computation offloading algorithm in MEC edge environments with time-varying channels and task arrival (Reference: Bi S., Huang L., Wang H., et al. Lyapunov-guided deep reinforcement learning for stable online computation offloading in mobile-edge computing networks[J].IEEE Transactions on Wireless Communications, 2021.): This algorithm studies multi-user MEC networks with random task arrival. Under the constraints of long-term task queue stability and average power, an online computation offloading algorithm based on Lyapunov is designed to maximize the network data processing capability. (2) Computation Offloading and Resource Allocation Scheme in Energy Harvesting MEC Systems: GCN-DDPG Algorithm (References: Chen J., Wu Z. Dynamic Computation Offloading With Energy Harvesting Devices: A Graph-Based Deep Reinforcement Learning Approach[C] / / 2021IEEE Communications Letters.IEEE,2021. Kashyap P K., Kumar S., Jaiswal A. Deep Learning Based Offloading Scheme for IoT Networks Towards GreenComputing[C] / / 2019IEEE International Conference on Industrial Internet(ICII).IEEE,2019:22-27.): This algorithm proposes a centralized DDPG-based reinforcement learning algorithm to address the computation offloading and resource allocation problem of energy harvesting devices. It is used to learn the decisions of mobile devices, including the offloading ratio, local computing power, and uplink transmission power.(3) Computation Offloading in Heterogeneous Mobile Edge Computing with Energy Harvesting: A Non-Cooperative Computation Offloading Game Theory Algorithm (Reference: Zhang T., Chen W. Computation Offloading in Heterogeneous Mobile Edge Computing With Energy Harvesting[J].2021.): This algorithm studies the computation offloading problem from multiple users to multiple MECs in heterogeneous MEC systems with energy harvesting from the perspective of game theory, and establishes an M / G / 1 queue model to minimize the latency of all devices.
[0005] In MEC systems with energy harvesting, the dynamic nature of energy harvesting, the randomness of task arrival, and the real-time changes in network channel states pose significant challenges to task offloading and resource allocation. Traditional algorithmic solutions often require numerous numerical iterations to produce satisfactory solutions, and once the system state changes, it is necessary to frequently resolve complex optimization problems, which is too costly in highly dynamic MEC systems. On the other hand, algorithms based on deep reinforcement learning can adapt to the dynamic changes of the system. In energy harvesting MEC systems, system stability and computational performance are equally important, such as the stability of the task queue and the energy queue. In existing research, most deep reinforcement learning-based methods do not impose long-term performance constraints. In particular, the energy coupling between time slots after the introduction of energy harvesting will greatly affect the offloading scheme, bringing more challenges. Therefore, designing appropriate task offloading and resource allocation strategies in dynamic MECs with energy harvesting is of significant research value. Summary of the Invention
[0006] In view of this, in order to minimize the execution time and total energy consumption cost of terminal devices completing tasks and to ensure queue stability, this invention proposes a task offloading and resource allocation method in an energy harvesting MEC system, specifically including the following steps:
[0007] Based on an MEC system consisting of multiple terminal devices with EH functionality and a base station equipped with an edge server, a task queue model, a task computation model, and an energy harvesting model are established respectively.
[0008] Based on the dynamic energy harvesting, random task arrival, and real-time channel changes of the MEC system, a long-term stochastic optimization problem in the sense of time average is established according to the task queue model, task calculation model, and energy harvesting model, in order to minimize the execution time and total energy consumption cost of the terminal device to complete the task.
[0009] The optimization problem is decoupled into an unloading decision subproblem and a resource allocation subproblem within each defined time slot using Lyapunov optimization theory.
[0010] By using deep reinforcement learning, we model a Markov decision process, define the action space, state space, and reward function to solve the unloading decision subproblem and obtain the optimal unloading strategy.
[0011] An adaptive genetic algorithm is used to solve the resource allocation subproblem through crossover, mutation, and selection operations to obtain the optimal resource allocation scheme.
[0012] The beneficial effects of this invention are:
[0013] This invention considers the dynamics of energy harvesting, the randomness of task generation, and the real-time changes in channel conditions within an energy harvesting MEC system. To adapt to the system's dynamics and minimize the total system cost under long-term queue stability constraints, a long-term stochastic optimization problem is modeled. Using Lyapunov stochastic optimization theory, this problem is decoupled into a task offloading decision subproblem and a resource allocation subproblem within each defined time slot. A task offloading and resource allocation scheme combining reinforcement learning and an adaptive genetic algorithm is designed. For the offloading decision subproblem, a deep reinforcement learning-based algorithm is used, defining the algorithm's state space, action space, and reward function according to the dynamic MEC system to obtain the optimal task offloading strategy. For the resource allocation subproblem, an improved adaptive genetic algorithm is used, with adaptive parameters designed based on the algorithm's execution process to improve the algorithm's global search capability and convergence speed. The four main processes in the improved adaptive genetic algorithm are population initialization, mutation, crossover, and selection to obtain the optimal resource allocation. Simulation results show that this scheme has good performance in stabilizing the queue and satisfying system dynamics, demonstrating advantages over existing schemes. Attached Figure Description
[0014] Figure 1 This is a flowchart of a task offloading and resource allocation method in an energy harvesting MEC system according to an embodiment of the present invention;
[0015] Figure 2 A model of an MEC system with energy harvesting;
[0016] Figure 3 This forms the framework for the joint computing offloading and resource allocation scheme in this invention;
[0017] Figure 4 The task queue length under different control parameters V;
[0018] Figure 5 The total cost is given under different control parameters V. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] This invention proposes a method for task offloading and resource allocation in an energy harvesting MEC system, such as... Figure 1 As shown, it includes the following steps:
[0021] Based on an MEC system consisting of multiple terminal devices with EH functionality and a base station equipped with an edge server, a task queue model, a task computation model, and an energy harvesting model are established respectively.
[0022] Based on the dynamic energy harvesting, random task arrival, and real-time channel changes of the MEC system, a long-term stochastic optimization problem in the sense of time average is established according to the task queue model, task calculation model, and energy harvesting model, in order to minimize the execution time and total energy consumption cost of the terminal device to complete the task.
[0023] The optimization problem is decoupled into an unloading decision subproblem and a resource allocation subproblem within each defined time slot using Lyapunov optimization theory.
[0024] By using deep reinforcement learning, we model a Markov decision process, define the action space, state space, and reward function to solve the unloading decision subproblem and obtain the optimal unloading strategy.
[0025] An adaptive genetic algorithm is used to solve the resource allocation subproblem through crossover, mutation, and selection operations to obtain the optimal resource allocation scheme.
[0026] This embodiment describes the present invention from four aspects: system model, problem description, algorithm design, simulation results and analysis.
[0027] I. System Model
[0028] Consider an EH-MEC system consisting of multiple terminal devices with energy harvesting capabilities and a base station equipped with an edge server, such as... Figure 2As shown. The set of terminal devices is represented by M = {1,2,...,m,...,M}. Each terminal device can collect energy from the environment for computing and communication, and the collected energy is stored in a battery. The system is divided into several time slots, with time slot indices represented by T = {1,2,...,t...,T}. The length of each time slot is δ. In each time slot, a centralized training and distributed execution approach is adopted. The base station is responsible for collecting the status information of all terminal devices, including task queue status, energy queue status, channel status, etc., to train the model for offloading decisions and resource allocation. Finally, the terminal devices execute the decisions.
[0029] 1. Task Queue Model
[0030] Define the task generated by terminal device m in time slot t using I. m (t)={Q m (t),b m (t),U m (t),τ m (t)} represents, where Q is a constant. m (t) represents the number of tasks (bits) in the task queue of terminal device m in time slot t. m (t) represents the actual amount of tasks processed by terminal device m at time t, U m (t) represents the number of CPU cycles required to process a unit task, τ m (t) represents the latency tolerance threshold of the terminal device. Task arrivals are random and independently identically distributed, assumed to follow a parameter of... Poisson distribution, the workload generated by terminal device m in time slot t is represented by A. m (t) represents the unloading strategy adopted by the terminal device, which is binary unloading, and the unloading variable α. m (t)∈{0,1} represents the unloading decision of terminal device m, when α m When (t) = 1, it indicates that the task is offloaded to the edge server for execution; when α... m When (t) = 0, it indicates that the task will be executed locally. The actual amount of task processed satisfies b. m (t)=min{b max Q m (t)}, where b max This represents the maximum number of tasks that can be processed; therefore, the number of tasks offloaded from the terminal device m to the edge server for execution is represented as:
[0031]
[0032] The amount of work performed locally by terminal device m can be represented as:
[0033]
[0034] The dynamic update of the task queue of terminal device m is represented as follows:
[0035] Q m (t+1)=Q m (t)-b m (t)+A m (t)
[0036] Due to the randomness of task arrival, the task queue will also change over time. Therefore, in order to ensure the stability of the task queue, the following constraints apply:
[0037]
[0038] 2. Energy Harvesting Model
[0039] In an EH-MEC system, terminal devices have rechargeable batteries to store energy harvested from the environment, defined as B. m (t) represents the remaining energy in the battery at time slot t, e m (t) represents the energy collected within time slot t. This indicates the energy consumed when the task is processed locally. This represents the energy consumed during the task transmission process; the total energy consumed is... The dynamic update of the battery power queue of terminal device m is then represented as:
[0040]
[0041] To prevent over-discharge of the terminal device battery, the following constraints should be met:
[0042]
[0043] Where E min and E max These represent the maximum and minimum battery discharge energy, respectively. Furthermore, to ensure the battery life of the terminal device, the energy in the battery during time slot t must be greater than the energy required by the terminal device, satisfying the following constraints:
[0044]
[0045] 3. Communication model
[0046] The communication system uses 5G technology with orthogonal channels. The base station allocates bandwidth to the terminal devices, and all terminal devices will share the entire channel bandwidth B. Then, the uplink transmission rate achievable between terminal device m and the base station is:
[0047]
[0048] Where β m(t) represents the proportion of uplink bandwidth allocated to terminal device m, h m (t) represents the channel gain between terminal device m and base station. It is assumed that the channel gain is quasi-static (i.e., constant) in each time slot, but varies in different time slots. m (t) represents the transmission power of the terminal device, σ 2 Indicates noise power.
[0049] 4. Task Calculation Model
[0050] 1) Local computing model
[0051] When the task is processed locally, the amount of computation required is: Local computing power is In the EH-MEC system, assuming all terminal devices support dynamic voltage and frequency adjustment technology, this technology dynamically adjusts the chip's operating frequency and voltage according to the different computing needs of the applications running on the chip, thereby achieving energy saving. The local computing latency is then expressed as:
[0052]
[0053] Local computing power consumption is:
[0054]
[0055] Among them κ m It is the effective capacitance coefficient of the terminal device chip architecture.
[0056] 2) Unloading calculation model
[0057] When the task is unloaded, the amount of task to be unloaded is: The computing resources allocated by the server to terminal device m are: The unloading calculation involves three processes: 1) Task upload; 2) Server execution of the task; 3) Server returning the task execution result to the terminal device. The transmission latency of the task upload process is:
[0058]
[0059] The corresponding transmission energy consumption is:
[0060]
[0061] After receiving the task, the edge server makes a reasonable allocation of computing resources based on its own computing resources to compute the unloaded task. At this time, the computing latency is:
[0062]
[0063] The time consumed by the edge server in processing the task is:
[0064]
[0065] The size of the task's execution result is negligible compared to the input task size. Therefore, the result return time and energy consumption are also negligible. Thus, the total latency spent processing the task in time slot t is:
[0066]
[0067] The total energy consumption for task processing in time slot t is:
[0068]
[0069] Therefore, the execution time and total energy consumption cost incurred by the terminal device to complete the task can be expressed as:
[0070]
[0071] Where γ1 and γ2 represent weighting factors used to balance latency and energy consumption.
[0072] II. Problem Description
[0073] 1. Optimize the problem description
[0074] To minimize the total cost of the system to complete the task under the constraints of queue stability and limited computational and communication resources, matrix A is defined. t ={α m {α1(t), α2(t), ..., α} = {α1(t), α2(t), ..., α} m (t)} represents the set of unloading decisions, B represents the set of computing resource allocations for the server. t ={β m (t)}={β1(t),β2(t),...,β m If (t)} represents the set of sub-channel allocation decisions, then it can be modeled as a long-term stochastic optimization problem in the sense of time average:
[0075]
[0076] Wherein, C1 represents the constraint on the offloading decision variable; C2 represents the constraint on the channel allocation variable; C3 and C4 represent the constraint on the server computing power, indicating that the total computing power allocated to the edge server cannot exceed the maximum value; C5 and C6 represent the constraints on latency and energy, ensuring that the total execution time does not exceed the maximum tolerable latency and that the battery energy is not depleted after each computing task; C7 represents the stability constraint of the task queue.
[0077] Among them, A t ={α m {α1(t), α2(t), ..., α} = {α1(t), α2(t), ..., α} m (t)},B t ={β m (t)}={β1(t),β2(t),...,β m (t)} and Represent the terminal device task offloading decision set, bandwidth allocation set, and server computing resource allocation set, respectively; C m (t) represents the execution time and total energy consumption cost of the terminal device m in completing the task; α m (t) represents the unloading decision variable for terminal device m; β m (t) represents the proportion of uplink bandwidth allocated to terminal device m; This represents the computing resources allocated by the server to terminal device m; f m s ax This indicates the server's maximum computing resources; B represents the total energy consumption of terminal device m in time slot t; m (t) represents the remaining battery power in terminal device m in time slot t; e m (t) represents the energy collected by terminal device m in time slot t; τ represents the total latency spent processing the task in time slot t; m Q represents the latency tolerance threshold of terminal device m; m (t) represents the number of tasks (bits) in the task queue of terminal device m in time slot t; T represents the system running time; M represents the number of terminal devices; It indicates a desire for the expected value.
[0078] 2. Optimize problem transformation
[0079] Analysis reveals that problem P is a non-convex mixed-integer nonlinear programming (MINLP) problem, where the task unloading strategy and resource allocation strategy are coupled in each time slot. To decouple the problem, we employ Lyapunov optimization theory, constructing a Lyapunov quadratic function based on the task queue and energy queue; determining the Lyapunov drift function by controlling the Lyapunov quadratic function; determining the Lyapunov drift plus penalty function based on the Lyapunov drift function; and minimizing the Lyapunov drift plus penalty function to determine when to make task unloading decisions and resource allocations based on the observed state of the task queue. This transforms the decision problem in continuous time slots into determining two sub-problems within each time slot.
[0080] To jointly control the task queue and energy queue, a joint queue Z(t) is defined as {Q(t), B(t)}, where Q(t) = {Q...} m B(t)} represents the task queue, and B(t) = {B} m Let (t)} represent the energy queue, therefore the Lyapunov quadratic function is defined as:
[0081]
[0082] When t = 0, L(Z(t)) = 0. The more the task queue accumulates, the larger L(Z(t)) becomes; conversely, the smaller L(Z(t)) becomes, the smaller it becomes. Therefore, the task queue backlog can be reduced by controlling the value of L(Z(t)). The Lyapunov drift function is defined as:
[0083]
[0084] To minimize the total cost for terminal devices to complete tasks while maintaining a stable joint queue, a drift plus penalty function is defined as follows:
[0085]
[0086] Where V>0 is a parameter that measures the penalty, which is achieved by minimizing Δ. V Z(t) can guarantee the stability of the joint queue while minimizing the total cost for the terminal devices to complete the task. Therefore, the following will derive Δ V The upper bound of Z(t) can be obtained according to the triangle inequality [].
[0087]
[0088] From the above inequality, we can obtain the following for all terminal devices m:
[0089]
[0090] Substituting the above equation into the Lyapunov drift function, we get:
[0091]
[0092] in They are b m (t),A m (t), e m The upper bound of (t) is therefore expressed as:
[0093]
[0094] Based on Lyapunov's expectation minimization theory, task unloading decisions and resource allocation are made when the state of the task queue is observed. Definition:
[0095]
[0096] Therefore, the problem can be minimized within each time slot:
[0097]
[0098] Among them, H(A) t B t ,F t ) represents the cost function, A t ={α m {α1(t), α2(t), ..., α} = {α1(t), α2(t), ..., α} m (t)},B t ={β m (t)}={β1(t),β2(t),...,β m (t)} and These represent the terminal device task offloading decision set, the bandwidth allocation set, and the server computing resource allocation set, respectively; V>0 is a parameter that measures the penalty.
[0099] III. Algorithm Design
[0100] 1. Optimize problem transformation
[0101] Optimization problem P1 is an optimization problem within a fixed time slot, involving unloading decision variables A with discrete integer values. t and B with continuous values t ,F t The system contains both continuous and discrete values and is highly dynamic. As the dimension of the variables increases, the computational complexity of the entire system increases significantly, making it difficult to solve such highly complex dynamic problems using traditional optimization algorithms. On the other hand, to solve the P1 problem within time slot t, it is necessary to determine the joint queue Z(t) = {Q(t), B(t)} and the channel gain {h} within that time slot. m The state of (t)} determines the task unloading decision and resource allocation. Once the task unloading decision is determined, the resource allocation scheme can be solved using a heuristic algorithm. Therefore, this invention designs a joint computational unloading and resource allocation scheme based on deep reinforcement learning and an improved adaptive genetic algorithm through a multi-algorithm combination heuristic search. The algorithm framework is as follows: Figure 3 As shown.
[0102] 2. Optimize problem transformation
[0103] For the optimization problem P1, obtaining the task offloading decision and resource allocation strategy based on changes in the joint queue and channel state is an NP-hard problem. However, once the task offloading decision A is determined... t The P1 problem can be simplified to the resource allocation subproblem without integer variables. Based on the resource optimization results, the optimal unloading decision (A) can be obtained. t ) * :
[0104]
[0105] For the unloading decision subproblem P2, considering the dynamic characteristics of the system, an unloading strategy algorithm based on deep reinforcement learning will be adopted to obtain the unloading decision through interactive learning with the environment. The problem is modeled as a Markov Decision Process (MDP), which mainly includes the following three elements:
[0106] 1) State space: is the set of all possible states of the system, including changes in channel conditions at each time slot, as well as changes in energy queues and task queues. Therefore, the state space is defined as:
[0107] s t ={h m (t),Q m (t),B m (t)}
[0108] 2) Action Space: This is the set of all possible actions that the agent can perform. The agent chooses different offloading decisions based on different rewards according to the current system state, hoping to obtain a greater reward. Therefore, the action space is defined as:
[0109] a t ={α m (t)}
[0110] 3) Reward Function: This refers to the reward given to the agent by the system environment after the agent executes the offloading decision. The higher the weighted sum of latency and energy consumption of the terminal device executing the task, the lower the reward. If the constraints are not met after executing the offloading decision, a negative reward is returned, representing a penalty to the agent. The agent's goal is to maximize the reward obtained after executing the offloading decision. The objective function of this invention is to minimize the total cost for the terminal device to complete the task. Therefore, we define the reward function as follows:
[0111]
[0112] C0 and C1 are positive constants, and their values are greater than H(A). t B t ,F t The theoretical upper limit of ).
[0113] 1. Resource allocation module based on adaptive genetic algorithm
[0114] For solving problem P2, the deep reinforcement learning algorithm outputs the task offloading decision (A). t ) * Therefore, solving the resource allocation subproblem can be expressed as:
[0115]
[0116] To effectively solve the resource allocation subproblem P3, and in the exploration and development of balanced reinforcement learning algorithms, the traditional adaptive genetic algorithm was improved by designing an adaptive scaling factor mutation strategy and an adaptive crossover factor increase strategy. To evaluate the effectiveness of individuals in the algorithm, the fitness function is defined as follows:
[0117]
[0118] The larger the value of this function, the better the individual's adaptability, and the easier it is to be preserved in the next generation.
[0119] The improved adaptive genetic algorithm has four main steps: population initialization, mutation, crossover, and selection, as detailed below.
[0120] 1) Population initialization: This refers to initializing a population of size NP, where x represents the individuals in the population, which is the solution, expressed as:
[0121]
[0122] Each chromosome of an individual is a solution to an optimization problem, expressed as:
[0123]
[0124] 2) Mutation operation: After population initialization, a new generation of solutions is generated through mutation. The mutation operation generates the k-th generation solution depending on the crossover probability F. k The crossover probability will affect the algorithm's global search capability. When F k A larger F value is beneficial for maintaining population diversity and global search; a smaller F value... k To improve convergence speed, the algorithm needs to meet the requirements of different stages based on its progress. Therefore, the following adaptive mutation probability was designed:
[0125]
[0126] Among them, F k F represents the scaling factor for the k-th generation. max F represents the maximum scaling factor.min Let k represent the minimum scaling factor, and k represent the current iteration number of the population. max This represents the maximum number of iterations in the population.
[0127] During the search process, the algorithm should initially maintain a large F. k To ensure population diversity and global search capability, and to avoid premature convergence due to getting trapped in local optima, F increases with the number of iterations. k The number of individuals should be gradually reduced to ensure that the superior individuals found in the early stages are not destroyed, thus guaranteeing the probability of finding the globally optimal solution.
[0128] 3) Crossover operation: To obtain better individuals, crossover operation is required. The crossover process needs to be configured with a reasonable crossover probability, CR. k This will affect global search capabilities and convergence speed. CR k A larger CR value is beneficial for providing a faster convergence speed for the algorithm. k When the value is small, the search process becomes slow or even stalls. Therefore, the adaptive crossover probability is set as follows:
[0129]
[0130] Among them, CR k CR represents the crossover factor in the k-th generation. max CR represents the maximum crossover factor. min Let k denote the minimum crossover factor, and k represent the current generation number of the population. max This represents the maximum number of iterations in the population.
[0131] 4) Selection operation: The newly generated individual is compared with the target individual. If the fitness value of the new individual is greater than or equal to the fitness value of the target individual, the new individual will replace the corresponding target individual and enter the next generation. Otherwise, the target individual will enter the next generation.
[0132] IV. Simulation Results and Analysis
[0133] This invention mainly analyzes the feasibility and effectiveness of the designed algorithm. First, it introduces the setup of the simulation environment, and then illustrates the feasibility and effectiveness of the algorithm by examining the impact of different parameters on the designed algorithm.
[0134] 1. Simulation parameter settings
[0135] Consider a scenario with a single base station and multiple terminal devices. The number of terminal devices M = 20, the total time slot length T = 2000, and the length of each time slot δ = 10ms. Users arrive randomly in each time slot. What is the average arrival rate of the task? The maximum computing power of a server with channel bandwidth B = 10MHz Maximum computing power of terminal devices Unit density U of task processing m (t) = 1000 cycles / bit, capacitance coefficient κ of the terminal device m =10 -28 The maximum energy that the terminal device can collect is 0.2 mJ, and the maximum discharge capacity of the terminal device is... Maximum discharge capacity
[0136] Figure 4 and Figure 5 This demonstrates the impact of different control parameters V on the task queue and total cost. In the algorithm, parameter V is mainly used to measure the total cost of the system and the stability of the task queue. Figure 4 The display shows the change in task queue length under different parameters V. As V increases, the task queue length increases. Figure 5 The diagram shows the change in total cost under different control parameters V. As V increases, the system cost gradually decreases. This is because as V increases, the EH-MEC system becomes more cost-conscious, and the proposed scheme dynamically adjusts the unloading decision to reduce overall cost.
[0137] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for task offloading and resource allocation in an energy harvesting MEC system, the method comprising: Comprising the following steps: Based on the MEC system composed of a plurality of terminal devices with EH function and a base station equipped with an edge server, a task queue model, a task computing model and an energy collection model are established respectively; wherein EH is energy collection; Based on the dynamic energy collection, random task arrival and real-time channel change of the MEC system, a long-term random optimization problem in the sense of time average is established according to the task queue model, the task computing model and the energy collection model, so as to minimize the total cost of the execution time and energy consumption of the terminal device to complete the task; The long-term random optimization problem in the sense of time average: Constraint conditions: C1: a m (t) e {0, 1} wherein A t = {a m (t)} = {a1(t), a2(t),..., a m (t)}, B t = {b m (t)} = {b1(t), b2(t),..., b m (t)} and denote the set of terminal device task offloading decisions, the set of bandwidth allocations and the set of server computing resource allocations, respectively; C m (t) denotes the execution time and total energy cost of terminal device m to complete the tasks; a m (t) denotes the offloading decision variable of terminal device m; b m (t) denotes the proportion of uplink bandwidth allocated to terminal device m; denotes the computing resource allocated to terminal device m by the server; denotes the maximum computing resource of the server; denotes the total energy consumption of terminal device m at time slot t; B m (t) denotes the remaining battery power of terminal device m at time slot t; e m (t) denotes the energy harvested by terminal device m at time slot t; denotes the total latency spent by terminal device m in processing tasks at time slot t; τ m denotes the latency tolerance threshold of terminal device m; Q m (t) denotes the amount of tasks (in bits) in the task queue of terminal device m at time slot t; T denotes the system running time; M denotes the number of terminal devices; denotes the expectation; The optimization problem is decoupled into an offloading decision sub-problem and a resource allocation sub-problem in each determined time slot by using Lyapunov optimization theory; The sub-problem in each determined time slot is represented as: Constraint conditions: C1: a m (t) e {0, 1} where H(A t ,B t ,F t ) represents the cost function; V > 0 is a parameter that measures the penalty; The offloading decision sub-problem is represented as: wherein (A t ) * denotes the optimal offloading decision at time slot t; The resource allocation sub-problem is represented as: s.t.C2-C7 The offloading decision sub-problem is solved by using deep reinforcement learning to model a Markov decision process, define an action space, a state space and a reward function, and obtain an optimal offloading strategy; Solving the decoupled offloading decision sub-problem includes modeling the offloading decision sub-problem into a Markov decision process by using a deep reinforcement learning algorithm; constructing a state space according to the state of the channel condition in each time slot, the state of the energy queue and the state of the task queue; determining an action space according to the different offloading decisions selected by an intelligent agent based on different rewards according to the current system state; and constructing a reward function according to the reward fed back by the current system to the intelligent agent after executing the offloading decision. State space is: s t = {h m (t), Q m (t), B m (t)}; Action space is: a t = {a m (t)}; The reward function is: where h m (t) denotes the channel gain between the terminal device m and the base station, H(A t ,B t ,F t ) denotes the cost function, C0and C1are normal numbers; The resource allocation sub-problem is solved by using an adaptive genetic algorithm to perform crossover, mutation and selection operations to obtain an optimal resource allocation scheme. 2.The method for task offloading and resource allocation in an energy harvesting MEC system of claim 1, wherein, Decoupling the long-term random optimization problem into an offloading decision sub-problem and a resource allocation sub-problem in each determined time slot by using Lyapunov random optimization theory includes constructing a Lyapunov quadratic function according to the task queue and the energy queue; determining a Lyapunov drift function by controlling the Lyapunov quadratic function; determining a Lyapunov drift plus penalty function according to the Lyapunov drift function; and determining the task offloading decision and resource allocation when the state of the task queue is observed by minimizing the Lyapunov drift plus penalty function. 3.The method for task offloading and resource allocation in an energy harvesting MEC system of claim 1, wherein, Solving the decoupled resource allocation sub-problem includes initializing a population by using an adaptive genetic algorithm, generating a mutation vector according to an adaptive mutation factor; generating a crossover vector according to an adaptive crossover factor; comparing the newly generated resource allocation individual and the target resource allocation individual, and selecting the corresponding resource allocation individual to enter the next generation iteration until the final resource allocation individual is determined. 4.The method for task offloading and resource allocation in an energy harvesting MEC system of claim 3, wherein, The adaptive mutation factor is: where F k represents the scaling factor for the kth generation, F max represents the maximum scaling factor, F min represents the minimum scaling factor, k represents the current iteration number of the population, k max represents the maximum iteration number of the population. 5.The method for task offloading and resource allocation in an energy harvesting MEC system of claim 3, wherein, The adaptive crossover factor is: where CR k represents the crossover factor of the kth generation, CR max represents the maximum crossover factor, CR min represents the minimum crossover factor, k represents the current iteration number of the population, k max represents the maximum iteration number of the population.
Citation Information
Patent Citations
Distributed task unloading and computing resource management method based on energy collection
CN113114733A
Distributed task offloading and computing resource management method based on energy harvesting
WO2022199036A1