Task Offloading and Resource Allocation Method for Hybrid Energy WPT-MEC System
Through the hybrid energy WPT-MEC system, combined with wireless power transmission and green energy harvesting technology, the deep deterministic strategy gradient algorithm is used to optimize task offloading and resource allocation, and the problem of insufficient optimization capabilities in the WPT-MEC system is solved, stable power supply and efficient calculation of equipment are achieved, and energy costs and carbon emissions are reduced.
Patent Information
- Application Number
- CN202510677972.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-26
AI Technical Summary
There are limitations of energy supply mode and insufficient optimization capabilities in dynamic environments in the existing WPT-MEC system, which leads to the inability to operate stably under limited energy conditions, especially in remote areas and dynamic environments to calculate tasks that are prone to interruption.
The hybrid energy WPT-MEC system is adopted, combining wireless power transmission and green energy harvesting technology, and optimizes task offloading and resource allocation through a deep deterministic strategy gradient algorithm, and uses renewable energy to prioritize power supply and switch to grid energy. The task offloading strategy is dynamically adjusted to minimize long-term average grid energy consumption.
It realizes stable power supply and efficient computing of equipment, reduces dependence on traditional power grids, reduces energy costs and carbon emissions, improves the computing efficiency and stability of the system, and adapts to long-term operation in complex environments.
Smart Images

Figure CN120201498B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of mobile edge computing, and particularly relates to a task offloading and resource allocation method based on a hybrid energy WPT-MEC system. Background Art
[0002] With the rapid development of Internet of Things (IoT) and artificial intelligence (AI) technologies, wireless devices are facing increasingly complex computing requirements. Computationally intensive tasks such as high-definition video processing and real-time image recognition pose higher demands on the computing power and energy efficiency of wireless devices. To address the dual limitations of computing power and energy of wireless devices, mobile edge computing (MEC) technology has emerged. MEC technology sinks computing resources and storage resources to the network edge, close to wireless devices, allowing devices to offload some or all of their computing tasks to be executed on edge servers, significantly reducing local energy consumption and improving response speed. For example, in the field of intelligent transportation, wireless devices on vehicles can offload computing tasks such as real-time traffic condition analysis and autonomous driving decision-making to the MEC server on the roadside, avoiding decision-making delays caused by limited local computing resources of the vehicle and ensuring driving safety. In industrial IoT, sensors on the production line can offload complex real-time data analysis tasks to nearby edge servers to improve production efficiency. However, MEC technology does not fundamentally solve the problem of limited battery energy of wireless devices. For devices in remote areas or difficult-to-access locations (such as environmental monitoring sensors, deep-sea detectors, etc.), it is unrealistic to frequently replace or charge the battery, which severely limits the long-term operating ability of the devices. Therefore, how to ensure the continuous operation of wireless devices under limited energy has become an urgent challenge to be solved.
[0003] To further solve the device battery life problem, wireless power transfer (WPT) technology has been introduced into the MEC network. In the MEC network, the base station (BS) can wirelessly charge surrounding wireless devices through WPT technology, breaking through the physical limitations of wired charging and ensuring the continuous operation of the devices. For example, in industrial IoT, a large number of sensors and actuators on the production line can obtain electrical energy from a nearby base station through WPT technology to ensure the stable operation of the devices and improve production efficiency. In a smart city, widely distributed monitoring devices can achieve continuous power supply through WPT technology, avoiding data interruption caused by battery depletion. However, existing WPT-MEC systems still have the following key defects:
[0004] Limitations of the energy supply mode: Currently, MEC base stations mainly rely on grid power supply. However, high transmission losses occur during the power transmission process, especially during long-distance transmission. At the same time, relying on grid power supply means high carbon emissions, which does not conform to the globally increasing emphasis on the concept of green environmental protection. Moreover, the problem of insufficient grid coverage in remote areas is prominent. Although renewable energy sources such as solar and wind energy have been proposed for energy supply, the initial installation process is complex, requiring professional technology and equipment, with high costs, and regular maintenance is needed. These factors have hindered large-scale application.
[0005] Insufficient optimization ability in a dynamic environment: Renewable energy sources (such as solar and wind energy) are strongly volatile due to weather changes (such as solar energy being affected by day and night alternation and weather changes, and wind energy being restricted by wind speed and wind direction changes), resulting in unstable energy supply for the base station. Traditional optimization methods (such as static resource allocation, traditional optimization algorithms, and heuristic algorithms) have high computational loads in dynamic scenarios and cannot make real-time decisions for a large number of users. It is difficult to balance energy availability and the computational requirements of wireless devices, and it is easy for computational tasks to be interrupted due to lack of sufficient energy support or for wireless devices to malfunction due to insufficient energy, seriously affecting the performance and reliability of the entire mobile edge computing network.
[0006] In view of the above problems, the present invention proposes a task offloading and resource allocation method based on a hybrid energy WPT-MEC system. Summary of the Invention
[0007] The purpose of the present invention is to provide a task offloading and resource allocation method based on a hybrid energy WPT-MEC system, aiming to solve the problems raised in the above background technology.
[0008] The purpose of the present invention is achieved through the following technical solutions:
[0009] A task offloading and resource allocation method based on a hybrid energy WPT-MEC system, assuming that the hybrid energy WPT-MEC system includes a base station and N wireless devices. The system runs according to the time line and, in each time slot, the wireless device i has a randomly arriving task , where respectively represent the data volume, the number of CPU cycles required to process 1 bit of the task, and the maximum tolerable delay for executing the task;
[0010] The method includes the following steps:
[0011] Step 1: Establish a system model;
[0012] By establishing an energy harvesting model, a communication model, and an energy consumption model, comprehensively describe the dynamic behavior of the system;
[0013] Step 2: Formulate the optimization problem;
[0014] By analyzing the renewable energy storage and consumption, and the grid energy replenishment mechanism, with the goal of minimizing the long-term average grid energy consumption, jointly optimize the task offloading ratio, the bandwidth ratio allocated to wireless devices i and the local computing frequency, and set constraints to transform the actual scenario requirements into a mathematical optimization problem;
[0015] Step 3: Transform the optimization problem into a Markov decision process;
[0016] By defining the state space including the task offloading amount, channel gain, and battery level , the action space including the task offloading decision, channel allocation ratio, and local computing frequency , and the reward function set as a negative target value , and perform min-max normalization on the data before training to transform the optimization problem into a form suitable for solving by the deep reinforcement learning algorithm;
[0017] Step 4: Implement the action space reduction strategy;
[0018] By fixing the task offloading decision, decompose the resource allocation sub-problem into two sub-constraints, and use the convex optimization method to solve the optimal resource allocation scheme, reducing the original high-dimensional action space to an action space N only containing the task offloading decision;
[0019] Step 5: Perform data normalization;
[0020] Use the min-max normalization method to process the state observations and map them to a unified scale range;
[0021] Step 6: Construct an improved deep reinforcement learning network;
[0022] Combine the deep deterministic policy gradient algorithm with the cross-entropy method to form the enhanced deep deterministic policy gradient algorithm. The deep deterministic policy gradient algorithm provides a stable policy optimization foundation, and the cross-entropy method improves the algorithm's exploration ability and convergence speed by generating diverse samples and selecting elite samples;
[0023] Step 7: Generate and output the optimal policy;
[0024] Use the trained deep reinforcement learning network to dynamically generate the best task offloading and resource allocation schemes based on the real-time system state.
[0025] Furthermore, the energy harvesting model is used to capture the dynamic changes of green energy; assuming the radio frequency transmission power is , the received power is expressed as follows:
[0026] ;
[0027] Among them, is the RF received power; and are the RF signal transmitting and receiving antenna gains respectively; is the RF signal wavelength; d is the distance between the wireless device and the base station;
[0028] Renewable energy collected by the base station obeys a normal distribution with a mean of and a variance of . At time slot t , the renewable energy is randomly sampled from the normal distribution.
[0029] Furthermore, the communication model uses the Rayleigh channel fading model to characterize the communication process between the wireless device and the base station. The channel gain is expressed as follows:
[0030] ;
[0031] Among them, is the communication channel gain; is the communication antenna gain; is the carrier frequency; is the wireless device i and the distance to the base station; α is the path loss exponent; is an independent random channel fading factor, following an exponential distribution;
[0032] The transmission rate of the wireless device is expressed as:
[0033] ;
[0034] Among them, is the task transmission rate of the wireless device i ; is the bandwidth ratio allocated to the wireless device i ; is the communication bandwidth; is the transmission power; is the noise power.
[0035] Furthermore, the energy consumption model is used to quantify the energy consumption of the wireless device and the base station. The base station energy consumption includes the energy consumption of the MEC server for assisted calculation and the energy consumption of replenishing the power of the wireless device through the WPT technology;
[0036] Assume the wireless device iThe task offloading ratio is , and the energy consumption of the MEC server is expressed as follows:
[0037] ;
[0038] Among them, is the energy consumption of the MEC server; is the energy consumption required for the MEC server to run one CPU; N is the number of wireless devices; i is the index of the wireless device; is the data volume; is the number of CPU cycles required to process 1 bit of task;
[0039] Each time slot t The base station replenishes time slots for wireless devices t -1 consumption, and the WPT energy consumption of the base station is expressed as follows:
[0040] ;
[0041] Among them, is the WPT energy consumption of the base station; is the loss factor of radio frequency energy; is the energy consumption of the wireless device at time slot t- 1, including local computing energy consumption and offloading transmission energy consumption of the task.
[0042] Furthermore, the specific process of step 2 is as follows:
[0043] Renewable energy is stored in a battery with limited capacity. Assuming that the battery level at time slot t is , then the evolution formula over time is expressed as follows:
[0044] ;
[0045] Among them, is the battery level at time slot t+ 1; is the WPT energy consumption of the base station; is the energy consumption of the MEC server; is the time slot t when the renewable energy;
[0046] When the green energy is exhausted, the grid energy is used as backup energy to continue to provide continuous energy support for the entire system. The corresponding grid energy consumption is expressed as follows:
[0047] ;
[0048] Among them, is the power grid energy consumption;
[0049] The offloading ratio of the joint optimization task and the bandwidth ratio i allocated to the wireless device and the local computing frequency to minimize the long - term average power grid energy, the optimization objective is expressed as follows:
[0050] ;
[0051] where, T is the length of the time line; t is the index of the time slot; meanwhile, it satisfies the maximum execution delay constraint of the task, the sum of the channel allocation ratios is less than 1, the local computing frequency does not exceed its maximum capacity value, and the battery level does not exceed the maximum battery capacity constraint.
[0052] Furthermore, the deep deterministic policy gradient algorithm includes an environment, an Actor network, a Critic network, and an experience replay buffer;
[0053] The Actor network is a three - layer neural network that obtains the mapping relationship from state to action according to the deterministic policy with the goal of maximizing the expected total discounted reward, approximated as the Q - function; its parameters are updated using the gradient ascent method:
[0054] ;
[0055] where, is the gradient of the objective function with respect to the parameter ; M is the size of the experience batch; m is the index sampled from the experience buffer; is Q the gradient of the network with respect to the parameter ; is the gradient of the policy at the given state with respect to the parameter
[0056] The structure of the target Actor network is the same as that of the Actor network, and the policy and parameters of the target Actor network are represented as and ;
[0057] The Critic network is a three - layer neural network used to evaluate the quality of actions and update the network parameters , and according to the Bellman equation, the current Q - value is expressed as follows:
[0058] ;
[0059] Among them, is the discount factor; E[ ] is the expected value of the random variable inside the brackets; is the state under the policy the value corresponding to the generated action Q ;
[0060] The structure of the target Critic network is the same as that of the Critic network, and its parameters are represented as ;
[0061] Use the target Actor network to predict the action of the next state, and use the target Critic network to calculate the target Q value y , and update the network parameters:
[0062] ;
[0063] Among them, is the loss function; is Q the network's value estimation of the state and the action .
[0064] Furthermore, the specific process of the cross-entropy method is as follows:
[0065] First, add Gaussian noise to the current Actor network parameters to generate candidate samples; then, interact the candidate samples with the environment and record the corresponding rewards, and select the elite samples; finally, use the average value of the elite sample parameters to update the Actor network parameters to accelerate the convergence speed of the DDPG algorithm.
[0066] Compared with the prior art, the beneficial effects of the present invention are:
[0067] 1. Energy supply optimization: The present invention combines wireless power transfer (WPT) technology and green energy harvesting technology, and adopts a green-priority hybrid energy supply mode. The base station preferentially uses renewable energy such as solar energy and wind energy, switches to grid energy when insufficient, and provides stable power for wireless devices within the coverage area through WPT technology. This mode avoids long-distance power transmission losses, solves the power supply problem in remote areas, realizes local energy acquisition and flexible distribution, and significantly reduces the dependence on the traditional power grid.
[0068] 2. Enhancement of device performance: By virtue of the efficient utilization and stable supply of green energy, the present invention extends the battery life of wireless devices and reduces the dependence on the traditional power grid. Meanwhile, by introducing mobile edge computing (MEC) technology, wireless devices can offload complex computing tasks to the MEC server of the base station, reducing the local computing burden, improving the overall system computing efficiency, and ensuring the stable operation of the device in complex environments.
[0069] 3. Optimization of algorithm decision-making: The present invention designs an online decision-making mechanism based on the enhanced deep deterministic policy gradient (EDDPG) algorithm. This mechanism transforms complex system problems into Markov decision processes, and uses deep neural networks to quickly analyze and process a large amount of historical data and real-time information, dynamically optimizing task offloading and resource allocation strategies. This not only simplifies the problem-solving process, reduces the computing resource requirements, but also can quickly adapt to the fluctuations of renewable energy, ensuring the efficient and stable operation of the system.
[0070] Improvement of comprehensive benefits: The present invention reduces the dependence on the traditional power grid and lowers the energy cost. By maximizing the utilization of green energy, it reduces carbon emissions and environmental pollution, achieving a win-win situation in economic and environmental benefits. In addition, by simplifying the energy supply architecture and reducing the energy harvesting component requirements of individual wireless devices, the present invention not only reduces the frequency and difficulty of device maintenance, but also improves the system stability and reliability, providing strong support for the long-term stable operation of Internet of Things devices. Description of the Drawings
[0071] Figure 1 is the flowchart of the method of the present invention.
[0072] Figure 2 is an example of the application scenario of the present invention.
[0073] Figure 3 is the performance comparison chart of the EDDPG algorithm and multiple classic DRL algorithms (DDPG, SAC, TD3). Detailed Description of the Invention
[0074] For a clearer understanding of the technical features, objectives, and beneficial effects of the present invention, the technical solutions of the present invention are described in detail below, but it should not be construed as a limitation on the scope of implementation of the present invention.
[0075] The present invention provides a method for task offloading and resource allocation based on a hybrid energy WPT-MEC system, and its flowchart is as Figure 1 shown. The hybrid energy WPT-MEC system consists of a system modeling module, an optimization problem formulation module, an MDP conversion module, an action space reduction module, a normalization processing module, an improved deep reinforcement learning module, and an output optimal result module.
[0076] Assume that the hybrid energy WPT-MEC system consists of a base station and N wireless devices. The system operates according to a timeline . In each time slot, a wireless device i will has a randomly arriving task , where represent the data volume, the number of CPU cycles required to process 1 bit of the task, and the maximum tolerable delay for task execution, respectively.
[0077] The task offloading and resource allocation method specifically includes the following steps:
[0078] Step 1: Establish a system model;
[0079] In the hybrid energy WPT-MEC system, the system modeling module comprehensively describes the dynamic behavior of the system by establishing an energy harvesting model, a communication model, and an energy consumption model.
[0080] 1. Energy harvesting model;
[0081] The energy harvesting model is used to capture the dynamic changes of green energy, considering the radio frequency energy harvested by the wireless device and the renewable energy collected by the base station. Assume the radio frequency transmission power is , then the received power is expressed as follows:
[0082] ;
[0083] where is the radio frequency received power; and are the radio frequency signal transmission and reception antenna gains, respectively; is the radio frequency signal wavelength; d is the distance between the wireless device and the base station.
[0084] In addition, according to the harvesting trajectory of real-world renewable energy, the renewable energy collected by the base station follows a normal distribution with a mean of and a variance of t . The renewable energy at time slot
[0085] is randomly sampled from the normal distribution.
[0086] 2. Communication model;
[0087] ;
[0088] where is the communication channel gain; is the communication antenna gain; is the carrier frequency; is the wireless device i distance from the base station; α is the path loss exponent; is an independent random channel fading factor, following an exponential distribution. Based on this, the transmission rate of the wireless device can be expressed as:
[0089] ;
[0090] where, is the task transmission rate of the wireless device i ; is the bandwidth ratio allocated to the wireless device i ; is the communication bandwidth; is the transmission power; is the noise power.
[0091] 3. Energy consumption model;
[0092] The energy consumption model quantifies the energy consumption of the wireless device and the base station. The base station energy consumption includes the energy consumption of the MEC server for assisting in computing and the energy consumption of replenishing the power of the wireless device through the WPT technology.
[0093] Assume that the task offloading ratio of the wireless device i is , the MEC server energy consumption is expressed as follows:
[0094] ;
[0095] where, is the MEC server energy consumption; is the energy consumption required for the MEC server to run one CPU; N is the number of wireless devices; i is the index of the wireless device; is the data volume; is the number of CPU cycles required to process 1 bit of task.
[0096] To ensure the stable operation of the wireless device, in each time slot t the base station replenishes the wireless device with the consumption of time slot t -1, then the WPT energy consumption of the base station is expressed as follows:
[0097] ;
[0098] where, is the WPT energy consumption of the base station; is the loss factor of radio frequency energy, calculated based on the relationship between the transmitted and received power of radio frequency energy; is the energy consumption of the wireless device at time slot t- 1, including the local computing energy consumption and the offloading transmission energy consumption of the task.
[0099] The above model provides an accurate system information basis for the subsequent optimization problem formulation and algorithm design, ensuring the feasibility and effectiveness of the algorithm in practical applications.
[0100] Step 2: Formulate the optimization problem;
[0101] In the hybrid energy WPT-MEC system, the optimization problem formulation module transforms the requirements in the actual application scenario into a mathematical optimization problem. Renewable energy is usually stored in a battery with a limited capacity. Assume that the battery level at time slot t is , then the evolution formula over time is expressed as follows:
[0102] ;
[0103] where, is the battery level at time slot t+ 1; is the WPT energy consumption of the base station; is the energy consumption of the MEC server; is the time slot t when the renewable energy;
[0104] When the green energy is exhausted, the grid energy is used as backup energy to continue to provide continuous energy support for the entire system. The corresponding grid energy consumption is expressed as follows:
[0105] ;
[0106] where, is the grid energy consumption;
[0107] The present invention aims to achieve a green and energy-saving network environment, and the problem is defined as minimizing the long-term average grid energy consumption. To this end, the offloading ratio of the task , the bandwidth ratio i allocated to the wireless device and the local computing frequency are jointly optimized to minimize the long-term average grid energy. The optimization objective is expressed as follows:
[0108] ;
[0109] where, T is the length of the time line; tis the index of the time slot; simultaneously satisfying the maximum execution delay constraint of the task, the sum of the channel allocation ratios being less than 1, the local computing frequency not exceeding its maximum capacity value, and the battery level not exceeding the maximum battery capacity constraint condition.
[0110] This process abstracts complex practical problems into computable mathematical models, laying a theoretical foundation for the application of deep reinforcement learning algorithms and also clarifying the direction and boundaries for subsequent algorithm design.
[0111] Step 3: Convert the optimization problem into a Markov decision process (MDP);
[0112] In the hybrid energy WPT-MEC system, the MDP conversion module converts the optimization problem into a Markov decision process (MDP), specifically defining the state space, action space, and reward function.
[0113] State space : Includes system information such as the amount of task offloading, channel gain, and battery level.
[0114] Action space : Includes optimization variables such as task offloading decisions, the bandwidth ratio allocated to wireless devices i and local computing frequency.
[0115] Reward function : Since the goal of deep reinforcement learning is to maximize the value of the reward function, the reward function is the negative value of the goal.
[0116] In addition, min-max normalization is used to process the data before training to ensure the training effect.
[0117] Through this transformation, the complex optimization problem becomes a form suitable for solving by deep reinforcement learning algorithms, enabling the algorithm to gradually learn the optimal strategy by interacting with the environment, achieving efficient allocation and utilization of system resources.
[0118] Step 4: Implement the action space reduction strategy;
[0119] In the hybrid energy WPT-MEC system, the action space reduction module decomposes the resource allocation sub-problem into two sub-constraints by fixing the task offloading decision and uses convex optimization methods to solve the optimal resource allocation scheme. Through this process, the original high-dimensional action space (including task offloading decisions, CPU frequency, communication resource allocation, etc.) is reduced to an action space N with only task offloading decisions, significantly reducing the computational complexity and training difficulty of the algorithm, improving the efficiency of the algorithm, and enhancing its practicality in actual applications.
[0120] Step 5: Perform data normalization processing;
[0121] In the hybrid energy WPT - MEC system, the normalization processing module normalizes the state observation values, and uses the min - max normalization method to map them into a unified scale range. This eliminates the influence of data scale differences on model training, improves the training efficiency and stability of the algorithm, not only accelerates the convergence of the neural network, but also enhances the adaptability and robustness of the algorithm in different scenarios.
[0122] Step 6: Construct an improved deep reinforcement learning (DRL) network;
[0123] In the hybrid energy WPT - MEC system, the improved deep reinforcement learning module combines the deep deterministic policy gradient (DDPG) algorithm with the cross - entropy method (CEM) to form the enhanced deep deterministic policy gradient (EDDPG) algorithm, achieving efficient optimization of task offloading and resource allocation.
[0124] 1. DDPG algorithm;
[0125] The DDPG algorithm consists of an environment, an Actor network, a Critic network, and an experience replay buffer, providing a stable foundation for policy optimization.
[0126] The environment is an external entity that interacts with the agent in the deep reinforcement learning system, responsible for simulating the actual scenario where the agent is located, and providing state observations, reward signals, and termination conditions.
[0127] The experience buffer follows the first - in - first - out principle and is initialized as an empty buffer with a limited capacity. At each time slot, the agent of deep reinforcement learning deposits the experience tuple into the experience buffer. Then, it samples uniformly from the experience buffer for network parameter update.
[0128] The Actor network is a three - layer neural network, which obtains the mapping relationship from state to action according to the deterministic policy, and the goal is to maximize the expected total discounted reward, which can be approximated as the Q function; its parameters are updated using the gradient ascent method:
[0129] ;
[0130] where, is the gradient of the objective function with respect to the parameter ; M is the size of the experience batch; m is the index sampled from the experience buffer; is Q the gradient of the network with respect to the parameter for the given state and the policy At the parameter gradient. In addition, the structure of the target Actor network is the same as that of the Actor network, and the policy and parameters of the target Actor network are represented as and .
[0131] The Critic network is a three-layer neural network used to evaluate the quality of actions and update network parameters . According to the Bellman equation, the current Q value is expressed as follows:
[0132] ;
[0133] where is the discount factor; E[ ] is the expected value of the random variable inside the brackets; is the state under the policy generated by the action corresponding to Q value. Similarly, the structure of the target Critic network is the same as that of the Critic network, and its parameters are represented as .
[0134] Use the target Actor network to predict the action of the next state, and use the target Critic network to calculate the target Q value y , and update the network parameters:
[0135] ;
[0136] where is the loss function; is Q the network's value estimate of the state and the action .
[0137] Different from the deep Q network, the target network of DDPG uses the soft update method to partially update the parameters.
[0138] 2. Cross-entropy method (CEM);
[0139] First, to generate candidate samples, Gaussian noise is added to the current Actor network parameters to generate new network parameters, thereby increasing sample diversity and promoting exploration of the parameter space; then, the candidate samples are interacted with the environment and the corresponding rewards are recorded to screen out elite samples; finally, the average value of the elite sample parameters is used to update the Actor network parameters to accelerate the convergence speed of the DDPG algorithm.
[0140] The CEM enhancement mechanism significantly improves the exploration ability and convergence speed of the algorithm by generating diverse samples and selecting elite samples. It continuously optimizes the network parameters with a large amount of learning historical data to achieve the rapid convergence of the algorithm.
[0141] Step 7: Generate and output the optimal policy;
[0142] In the hybrid energy WPT - MEC system, the optimal result output module uses the trained deep reinforcement learning network to dynamically generate the best task offloading and resource allocation scheme according to the real - time system state (such as task requirements, network state, device energy state, etc.), realizing the efficient operation of the green - priority hybrid energy WPT - MEC system. This module can not only efficiently handle dynamic and fluctuating task requirements and network states, but also achieve the optimal allocation of resources in complex multi - user scenarios, significantly improving the system performance and user experience.
[0143] In the embodiments of the present invention, Figure 2 An application scenario example of the present invention is shown. In the figure, there is a base station (BS) that collects green energy such as solar energy and wind energy through direct environmental energy harvesting (EH) components (such as solar panels, wind turbines), and also collaboratively uses grid energy for power supply. There are multiple wireless devices distributed around the base station, and these wireless devices can indirectly receive green energy from the base station through wireless power transfer (WPT). In terms of data processing, the computing tasks of the wireless devices can either be executed locally (local computing as marked in the figure) or partially offloaded to the mobile edge computing (MEC) server equipped with the base station for processing. Through the data transmission represented by solid arrows and the energy transmission represented by dashed arrows, the interaction of data and energy is realized, and thus efficient and flexible data processing is achieved.
[0144] The following describes the specific implementation of the present invention in detail in combination with specific embodiments.
[0145] Embodiment 1: This embodiment verifies the energy consumption optimization ability of the enhanced deep deterministic policy gradient (EDDPG) algorithm in the green - priority hybrid energy WPT - MEC system through numerical simulation and analysis.
[0146] To verify the applicability of the EDDPG algorithm in the WPT - MEC system, it is compared with multiple benchmark algorithms. The results are shown in Table 1:
[0147] Table 1 Influence of the number of devices and the amount of tasks on the average electrical energy (J)
[0148]
[0149] From the data on the impact of the number of devices and the amount of tasks on the average electrical energy (J, the optimization objective) given in Table 1, it can be seen that the EDDPG algorithm can dynamically adjust the task offloading and resource allocation strategies when the number of devices and the amount of tasks increase. For example, when the amount of tasks is 300 Kbit, the grid energy consumption of the EDDPG algorithm is reduced by 78% compared with greedy local computing, by 61% compared with greedy offloading computing, and by 28% compared with random offloading. Thus, compared with the random offloading strategy, the EDDPG algorithm shows higher stability in terms of energy consumption and performance. By preferentially using green energy and optimizing the use of grid energy, it significantly reduces the grid energy consumption, reduces the energy cost and environmental pollution. This fully demonstrates the superiority and applicability of the EDDPG algorithm in energy consumption optimization and adaptability, and it is applicable to large-scale Internet of Things scenarios.
[0150] The EDDPG algorithm is compared with multiple classical DRL algorithms (DDPG, SAC, TD3) to prove that the improvement of the EDDPG algorithm has improved its convergence speed. As Figure 3 can be seen, as the number of time steps increases, the EDDPG algorithm quickly converges to an average reward of about -5.2 after about 30,000 time steps, while algorithms such as DDPG, SAC, and TD3 converge relatively slowly, and the average rewards finally reached are generally lower than that of the EDDPG algorithm. Thus, the EDDPG algorithm is superior to other DRL algorithms in terms of convergence speed and reward, which means that the EDDPG algorithm not only has strong global search ability, but also can use elite samples to improve its stability, and is more suitable for application in the WPT-MEC system.
[0151] The above experimental results show that in the face of a large number of users and a huge amount of tasks, the EDDPG algorithm can provide optimal task offloading and resource allocation solutions for multiple users in real time, significantly improving the computational efficiency and energy utilization rate of the system.
[0152] The above is only the preferred implementation manner of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent.
Claims
1. A task offloading and resource allocation method for a hybrid energy WPT-MEC system, characterized in that Assume that the hybrid energy WPT-MEC system includes a base station and N wireless devices, and the system operates according to a timeline . In each time slot, the wireless device i has a randomly arriving task , where represent the data volume, the number of CPU cycles required to process 1 bit of the task, and the maximum tolerable delay for executing the task, respectively; The method includes the following steps: Step 1: Establish a system model; By establishing an energy harvesting model, a communication model, and an energy consumption model, comprehensively describe the dynamic behavior of the system; Step 2: Formulate an optimization problem; By analyzing the renewable energy storage and consumption, and the grid energy replenishment mechanism, with the goal of minimizing the long-term average grid energy consumption, jointly optimize the task offloading ratio, the bandwidth ratio allocated to wireless devices i and the local computing frequency, and set constraints to transform the actual scenario requirements into a mathematical optimization problem; Step 3: Transform the optimization problem into a Markov decision process; By defining a state space that includes the task offloading volume, channel gain, and battery level , an action space that includes task offloading decisions, channel allocation ratios, and local computing frequencies , and a reward function set to a target negative value , and performing min-max normalization on the data before training to transform the optimization problem into a form suitable for solution by deep reinforcement learning algorithms; Step 4: Implement an action space reduction strategy; By fixing the task offloading decision, the resource allocation sub-problem is decomposed into two sub-constraints, and the convex optimization method is used to solve the optimal resource allocation scheme, reducing the original high-dimensional action space to an action space with only task offloading decisions N dimensional action space; Step 5: Perform data normalization processing; Use the min-max normalization method to process the state observation values and map them to a unified scale range; Step 6: Construct an improved deep reinforcement learning network; Combine the deep deterministic policy gradient algorithm with the cross-entropy method to form an enhanced deep deterministic policy gradient algorithm. The deep deterministic policy gradient algorithm provides a stable policy optimization basis, and the cross-entropy method improves the algorithm's exploration ability and convergence speed by generating diverse samples and selecting elite samples; Step 7: Generate and output the optimal policy; Use the trained deep reinforcement learning network to dynamically generate the best task offloading and resource allocation scheme according to the real-time system state.
2. The task offloading and resource allocation method for the hybrid energy WPT-MEC system according to claim 1, wherein The energy harvesting model is used to capture the dynamic changes of green energy; assuming that the radio frequency transmission power is , the received power is expressed as follows: ; Wherein, is the radio frequency received power; and are the radio frequency signal transmitting and receiving antenna gains respectively; is the radio frequency signal wavelength; d is the distance between the wireless device and the base station; Renewable energy collected by the base station obeys a normal distribution with a mean of and a variance of At time slot t the renewable energy is randomly sampled from the normal distribution.
3. The task offloading and resource allocation method for the hybrid energy WPT-MEC system according to claim 1, characterized in that, The communication model uses the Rayleigh channel fading model to characterize the communication process between the wireless device and the base station. The channel gain is expressed as follows: ; Among them, is the communication channel gain; is the communication antenna gain; is the carrier frequency; is the wireless device i and the distance between the base station; α is the path loss exponent; is an independent random channel fading factor, following an exponential distribution; The transmission rate of the wireless device is expressed as: ; wherein, is the task transmission rate of the wireless device i ; is the bandwidth ratio allocated to the wireless device i ; is the communication bandwidth; is the transmission power; is the noise power.
4. The task offloading and resource allocation method for the hybrid energy WPT-MEC system according to claim 1, wherein, The energy consumption model is used to quantify the energy consumption of the wireless device and the base station. The base station energy consumption includes the energy consumption of the MEC server for assisting in computing and the energy consumption of replenishing the power of the wireless device through the WPT technology; Assume that the task offloading ratio of the wireless device i is , and the energy consumption of the MEC server is expressed as follows: ; Among them, is the energy consumption of the MEC server; is the energy consumption required for the MEC server to run one CPU; N is the number of wireless devices; i is the index of the wireless device; is the data volume; is the number of CPU cycles required to process 1 bit of task; Each time slot t The base station replenishes time slots for wireless devices t -1 consumption, the WPT energy consumption of the base station is expressed as follows: ; Among them, is the WPT energy consumption of the base station; is the loss factor of radio frequency energy; is the energy consumption of the wireless device at time slot t- 1, including the local computing energy consumption and the offloading transmission energy consumption of the task.
5. The task offloading and resource allocation method for the hybrid energy WPT-MEC system according to claim 4, wherein The specific process of Step 2 is as follows: Renewable energy is stored in a battery with a limited capacity. Assuming that the battery level at time slot t is , its evolution formula over time is expressed as follows: ; Among them, is the battery level of the battery in time slot t+ 1; is the WPT energy consumption of the base station; is the energy consumption of the MEC server; is the time slot t and the renewable energy at that time; When the green energy is exhausted, the grid energy is used as backup energy to continue to provide continuous energy support for the entire system. The corresponding grid energy consumption is expressed as follows: ; Among them, is the power grid energy consumption; The offloading ratio of the joint optimization task , the bandwidth ratio allocated to the wireless device i , and the local computing frequency to minimize the long-term average grid energy, the optimization objective is expressed as follows: ; Among them, T is the time line length; t is the index of the time slot; at the same time, it satisfies the maximum execution delay constraint of the task, the sum of the channel allocation ratios is less than 1, the local computing frequency does not exceed its maximum capacity value, and the battery level does not exceed the maximum battery capacity constraint condition.
6. The task offloading and resource allocation method for the hybrid energy WPT-MEC system according to claim 1, wherein The deep deterministic policy gradient algorithm includes an environment, an Actor network, a Critic network, and an experience replay buffer; The Actor network is a three-layer neural network that, based on a deterministic policy obtains the mapping relationship from states to actions, with the goal of maximizing the expected total discounted reward, approximated as the Q function; its parameters are updated using the gradient ascent method: ; Among them, is the gradient of the objective function with respect to the parameter ; M is the size of the experience batch; m is the index sampled from the experience buffer; is Q the gradient of the network with respect to the parameter ; is the gradient of the given state , policy with respect to the parameter The structure of the target Actor network is the same as that of the Actor network. The policy and parameters of the target Actor network are represented as and ; The Critic network is a three-layer neural network used to evaluate the quality of actions and update network parameters , according to the Bellman equation, the current Q value is expressed as follows: ; where, is the discount factor; E[ ] is the expected value of the random variable inside the brackets; is the state at which the policy generates the action corresponding to the Q value; The structure of the target Critic network is the same as that of the Critic network, and its parameters are denoted as ; Predict the action of the next state using the target Actor network, and use the target Critic network to calculate the target Q value y , and update the network parameters: ; Among them, is the loss function; is Q the value estimation of the network for the state and the action .
7. The task offloading and resource allocation method for the hybrid energy WPT-MEC system according to claim 6, characterized in that The specific process of the cross-entropy method is as follows: First, add Gaussian noise to the current Actor network parameters to generate candidate samples; then, interact the candidate samples with the environment and record the corresponding rewards, and select the elite samples; finally, update the Actor network parameters using the average value of the elite sample parameters to accelerate the convergence speed of the DDPG algorithm.
Citation Information
Patent Citations
Mobile edge computing resource allocation method and system for dynamic user random access
CN117793805A
Automatic driving resource arrangement and task unloading method based on model segmentation
CN119603653A