Task unloading and resource allocation method based on hybrid energy WPT-MEC system

By adopting a WPT-MEC system with hybrid energy in mobile edge computing systems, combining wireless power transmission and green energy harvesting technology, dynamically optimized task offloading and resource allocation, the problems of instability in energy supply and insufficient optimization capabilities in the existing systems are solved, and efficient and stable energy management and computing performance improvements are achieved.

CN120201498AActive Publication Date: 2025-06-24JILIN UNIVERSITY

Patent Information

Application Number
CN202510677972.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

The existing mobile edge computing (MEC) systems have limitations in the energy supply mode, and insufficient optimization capabilities in dynamic environments lead to instability in energy supply and affect the performance and reliability of the entire network.

Method used

We adopt a WPT-MEC system based on hybrid energy, combining wireless power transmission (WPT) technology and green energy harvesting technology, prioritize the use of renewable energy (such as solar energy, wind energy), switch to grid energy when insufficient, and dynamically optimize task offloading and resource allocation strategies through deep reinforcement learning algorithms.

Benefits of technology

It has achieved optimization of energy supply, reduced dependence on traditional power grids, extended the battery life of wireless equipment, improved the system's computing efficiency and energy utilization rate, and ensured the stable operation of the equipment in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201498A_ABST
    Figure CN120201498A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of mobile edge computing, and provides a hybrid energy WPT-MEC system-based task unloading and resource allocation method, which comprises the following steps of: establishing a system model; an optimization problem with the goal of minimizing the long-term average power grid energy consumption is drawn up and converted into a Markov decision process; implementing an action space reduction strategy, and executing data normalization processing; constructing an enhanced depth deterministic policy gradient (EDDPG) algorithm; and generating and outputting an optimal strategy according to a real-time system state by using the trained network. According to the method, the wireless power transmission (WPT) technology, the green energy harvesting technology and the mobile edge computing technology are combined, so that the energy use duration and the computing capacity of the wireless equipment are improved. The online decision-making mechanism based on the EDDPG algorithm can dynamically optimize task unloading and resource allocation strategies, adapts to a dynamic environment, reduces energy cost and environmental pollution, and provides reliable technical support for large-scale deployment of the Internet of Things.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of mobile edge computing, and particularly relates to a task offloading and resource allocation method based on a hybrid energy WPT-MEC system. Background Art

[0002] With the rapid development of Internet of Things (IoT) and artificial intelligence (AI) technologies, wireless devices are facing increasingly complex computing requirements. Computationally intensive tasks such as high-definition video processing and real-time image recognition pose higher demands on the computing power and energy efficiency of wireless devices. To address the dual limitations of computing power and energy of wireless devices, mobile edge computing (MEC) technology has emerged. MEC technology sinks computing resources and storage resources to the network edge, close to wireless devices, allowing devices to offload some or all of their computing tasks to edge servers for execution, significantly reducing local energy consumption and improving response speed. For example, in the field of intelligent transportation, wireless devices on vehicles can offload computing tasks such as real-time traffic condition analysis and autonomous driving decision-making to roadside MEC servers, avoiding decision-making delays caused by limited local computing resources and ensuring driving safety. In industrial IoT, sensors on production lines can offload complex real-time data analysis tasks to nearby edge servers, improving production efficiency. However, MEC technology does not fundamentally solve the problem of limited battery energy of wireless devices. For devices in remote areas or difficult-to-access locations (such as environmental monitoring sensors, deep-sea detectors, etc.), it is unrealistic to frequently replace or charge the battery, which severely limits the long-term operating ability of the devices. Therefore, how to ensure the continuous operation of wireless devices under limited energy has become an urgent challenge to be solved.

[0003] To further solve the device battery life problem, wireless power transfer (WPT) technology has been introduced into the MEC network. In the MEC network, a base station (BS) can wirelessly charge surrounding wireless devices through WPT technology, breaking through the physical limitations of wired charging and ensuring the continuous operation of the devices. For example, in industrial IoT, a large number of sensors and actuators on production lines can obtain electrical energy from nearby base stations through WPT technology, ensuring the stable operation of the devices and improving production efficiency. In smart cities, widely distributed monitoring devices can achieve continuous power supply through WPT technology, avoiding data interruption caused by battery depletion. However, existing WPT-MEC systems still have the following key defects: Limitations of the energy supply mode: Currently, MEC base stations mainly rely on grid power supply. However, high transmission losses occur during the power transmission process, especially during long-distance transmission. At the same time, relying on grid power supply means high carbon emissions, which does not conform to the globally increasing emphasis on green environmental protection concepts, and the problem of insufficient grid coverage in remote areas is prominent. Although renewable energy sources such as solar energy and wind energy have been proposed for energy supply, the initial installation process is complex, requiring professional technology and equipment, with high costs, and regular maintenance is needed. These factors have hindered large-scale applications.

[0004] Insufficient optimization ability in a dynamic environment: Renewable energy sources (such as solar energy and wind energy) are strongly volatile due to weather changes (such as solar energy being affected by day-night alternation and weather changes, and wind energy being restricted by wind speed and wind direction changes), resulting in unstable energy supply for base stations. Traditional optimization methods (such as static resource allocation, traditional optimization algorithms, and heuristic algorithms) have high computational loads in dynamic scenarios, cannot make real-time decisions for a large number of users, and are difficult to balance energy availability and the computational requirements of wireless devices, easily causing computational tasks to be interrupted due to lack of sufficient energy support or wireless devices being unable to operate normally due to insufficient energy, seriously affecting the performance and reliability of the entire mobile edge computing network.

[0005] In response to the above problems, the present invention proposes a task offloading and resource allocation method based on a hybrid energy WPT-MEC system. Summary of the Invention

[0006] The purpose of the present invention is to provide a task offloading and resource allocation method based on a hybrid energy WPT-MEC system, aiming to solve the problems raised in the above background technology.

[0007] The purpose of the present invention is achieved through the following technical solutions: A task offloading and resource allocation method based on a hybrid energy WPT-MEC system, assuming that the hybrid energy WPT-MEC system includes a base station and N wireless devices, and the system operates according to the time line and, in each time slot, the wireless device i has a randomly arriving task , where respectively represent the data volume, the number of CPU cycles required to process 1 bit of the task, and the maximum tolerable delay for executing the task; The method includes the following steps: Step 1: Establish a system model; Comprehensively describe the dynamic behavior of the system by establishing an energy harvesting model, a communication model, and an energy consumption model; Step 2: Formulate an optimization problem; By analyzing the renewable energy storage and consumption, and the grid energy replenishment mechanism, with the goal of minimizing the long-term average grid energy consumption, jointly optimize the task offloading ratio, the bandwidth ratio allocated to wireless devices i and the local computing frequency, and set constraints to transform the actual scenario requirements into a mathematical optimization problem; Step 3: Transform the optimization problem into a Markov decision process; By defining a state space that includes the task offloading amount, channel gain, and battery level , an action space that includes task offloading decisions, channel allocation ratios, and local computing frequencies , and a reward function set as a negative target , and perform min-max normalization on the data before training to transform the optimization problem into a form suitable for solving by deep reinforcement learning algorithms; Step 4: Implement an action space reduction strategy; By fixing the task offloading decision, decompose the resource allocation sub-problem into two sub-constraints, and use convex optimization methods to solve the optimal resource allocation scheme, reducing the original high-dimensional action space to an action space N with only task offloading decisions; Step 5: Perform data normalization processing; Use the min-max normalization method to process the state observations and map them to a unified scale range; Step 6: Construct an improved deep reinforcement learning network; Combine the deep deterministic policy gradient algorithm with the cross-entropy method to form an enhanced deep deterministic policy gradient algorithm. The deep deterministic policy gradient algorithm provides a stable policy optimization foundation, and the cross-entropy method improves the algorithm's exploration ability and convergence speed by generating diverse samples and selecting elite samples; Step 7: Generate and output the optimal policy; Use the trained deep reinforcement learning network to dynamically generate the best task offloading and resource allocation schemes based on the real-time system state.

[0008] Furthermore, the energy harvesting model is used to capture the dynamic changes of green energy; assuming the radio frequency transmit power is , the received power is expressed as follows: ; where is the radio frequency received power; and are the radio frequency signal transmit and receive antenna gains respectively; is the radio frequency signal wavelength; d is the distance between the wireless device and the base station; Renewable energy collected by the base station The mean is , the variance is The normal distribution of t Renewable energy Randomly sample from a normal distribution.

[0009] Furthermore, the communication model uses a Rayleigh channel fading model to describe the communication process between the wireless device and the base station, and the channel gain is expressed as follows: ; in, is the communication channel gain; is the communication antenna gain; is the carrier frequency; For wireless devices i Distance from the base station; α is the path loss index; is an independent random channel fading factor, following an exponential distribution; The transmission rate of a wireless device is expressed as: ; in, For wireless devices i The task transfer rate; Assigned to wireless devices i Bandwidth ratio; is the communication bandwidth; is the transmission power; is the noise power.

[0010] Furthermore, the energy consumption model is used to quantify the energy consumption of wireless devices and base stations. The energy consumption of base stations includes the energy consumption of MEC server auxiliary computing and the energy consumption of replenishing power to wireless devices through WPT technology. Assuming wireless device i The task offloading ratio is , the energy consumption of MEC server is expressed as follows: ; in, is the energy consumption of MEC server; The energy consumption required to run a CPU for the MEC server; N is the number of wireless devices; i An index of wireless devices; is the amount of data; The number of CPU cycles required to process 1 bit of task; Each time slot t Base stations fill time slots for wireless devices tThe consumption of -1, the WPT energy consumption of the base station is expressed as follows: ; Among them, is the WPT energy consumption of the base station; is the loss factor of radio frequency energy; is the energy consumption of the wireless device at time slot t- 1, including the local computing energy consumption and the offloading transmission energy consumption of the task.

[0011] Furthermore, the specific process of step 2 is as follows: Renewable energy is stored in a battery with limited capacity. Assuming that the battery level at time slot t is , then the evolution formula over time is expressed as follows: ; Among them, is the battery level at time slot t+ 1; is the WPT energy consumption of the base station; is the energy consumption of the MEC server; is the renewable energy at time slot t ; When the green energy is exhausted, the grid energy is used as backup energy to continue to provide continuous energy support for the entire system. The corresponding grid energy consumption is expressed as follows: ; Among them, is the grid energy consumption; Jointly optimize the offloading ratio of the task , the bandwidth ratio i allocated to the wireless device and the local computing frequency to minimize the long-term average grid energy. The optimization objective is expressed as follows: ; Among them, T is the length of the time line; t is the index of the time slot; at the same time, it satisfies the maximum execution delay constraint of the task, the sum of the channel allocation ratios is less than 1, the local computing frequency does not exceed its maximum capacity value, and the battery level does not exceed the maximum battery capacity constraint condition.

[0012] Furthermore, the deep deterministic policy gradient algorithm includes an environment, an Actor network, a Critic network, and an experience replay buffer; The Actor network is a three-layer neural network, based on the deterministic policy Obtain the mapping relationship from states to actions, with the goal of maximizing the expected total discounted reward, approximated by the Q-function; its parameters are updated using the gradient ascent method: ; where, is the gradient of the objective function with respect to the parameter ; M is the size of the experience batch; m is the index sampled from the experience buffer; is Q the gradient of the network with respect to the parameter ; is the given state , and the policy is the gradient at the parameter ; The structure of the target Actor network is the same as that of the Actor network, and the policy and parameters of the target Actor network are represented as and ; The Critic network is a three-layer neural network used to evaluate the quality of actions and update the network parameters . According to the Bellman equation, the current Q-value is expressed as follows: ; where, is the discount factor; E[ ] is the expected value of the random variable inside the parentheses; is the state under which the action generated by the policy corresponds to the Q value; The structure of the target Critic network is the same as that of the Critic network, and its parameters are represented as ; Use the target Actor network to predict the action of the next state, use the target Critic network to calculate the target Q-value y , and update the network parameters: ; where, is the loss function; is Q the value estimation of the network for the state and the action .

[0013] Furthermore, the specific process of the cross-entropy method is as follows: First, Gaussian noise is added to the current Actor network parameters to generate candidate samples; then, the candidate samples are interacted with the environment and the corresponding rewards are recorded, and the elite samples are screened out; finally, the Actor network parameters are updated using the average value of the elite sample parameters to accelerate the convergence speed of the DDPG algorithm.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Energy supply optimization: The present invention combines wireless power transfer (WPT) technology and green energy harvesting technology, and adopts a green-priority hybrid energy supply mode. The base station preferentially utilizes renewable energy such as solar energy and wind energy, switches to grid energy when insufficient, and provides stable power for wireless devices within the coverage area through WPT technology. This mode avoids long-distance power transmission losses, solves the power supply problem in remote areas, realizes local energy acquisition and flexible distribution, and significantly reduces the dependence on the traditional power grid.

[0015] 2. Device performance enhancement: By virtue of the efficient utilization and stable supply of green energy, the present invention extends the battery life of wireless devices and reduces the dependence on the traditional power grid. At the same time, by introducing mobile edge computing (MEC) technology, wireless devices can offload complex computing tasks to the MEC server of the base station, reducing the local computing burden, improving the overall system computing efficiency, and ensuring the stable operation of the devices in complex environments.

[0016] 3. Algorithm decision optimization: The present invention designs an online decision-making mechanism based on the enhanced deep deterministic policy gradient (EDDPG) algorithm. This mechanism transforms complex system problems into Markov decision processes, and uses deep neural networks to quickly analyze and process a large amount of historical data and real-time information, dynamically optimizing task offloading and resource allocation strategies. This not only simplifies the problem-solving process, reduces the computing resource requirements, but also can quickly adapt to the fluctuations of renewable energy, ensuring the efficient and stable operation of the system.

[0017] Comprehensive benefit improvement: The present invention reduces the dependence on the traditional power grid and lowers the energy cost. By maximizing the utilization of green energy, carbon emissions are reduced and environmental pollution is alleviated, achieving a win-win situation in economic and environmental benefits. In addition, by simplifying the energy supply architecture, the present invention reduces the energy harvesting component requirements of individual wireless devices, not only reducing the frequency and difficulty of device maintenance, but also improving the system stability and reliability, providing strong support for the long-term stable operation of Internet of Things devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flowchart of the method of the present invention.

[0019] Figure 2 is an application scenario example of the present invention.

[0020] Figure 3 It is a performance comparison chart of the EDDPG algorithm and various classic DRL algorithms (DDPG, SAC, TD3). Detailed implementation manners

[0021] For a clearer understanding of the technical features, objectives, and beneficial effects of the present invention, the technical solution of the present invention will be described in detail below, but it should not be construed as a limitation on the implementable scope of the present invention.

[0022] The present invention provides a task offloading and resource allocation method based on a hybrid energy WPT-MEC system, and its flowchart is as Figure 1 shown. The hybrid energy WPT-MEC system consists of a system modeling module, an optimization problem formulation module, an MDP conversion module, an action space reduction module, a normalization processing module, an improved deep reinforcement learning module, and an output optimal result module.

[0023] Assume that the hybrid energy WPT-MEC system includes a base station and N wireless devices, and the system operates according to the time line and, at each time slot, the wireless device i will has a randomly arriving task , where respectively represent the data volume, the number of CPU cycles required to process 1 bit of the task, and the maximum tolerable delay for executing the task.

[0024] The task offloading and resource allocation method specifically includes the following steps: Step 1: Establish a system model; In the hybrid energy WPT-MEC system, the system modeling module comprehensively describes the dynamic behavior of the system by establishing an energy harvesting model, a communication model, and an energy consumption model.

[0025] 1. Energy harvesting model; The energy harvesting model is used to capture the dynamic changes of green energy, considering the radio frequency energy harvested by the wireless device and the renewable energy collected by the base station. Assume that the radio frequency transmission power is , then the received power is expressed as follows: ; where is the radio frequency received power; and are the radio frequency signal transmission and reception antenna gains respectively; is the radio frequency signal wavelength; d is the distance between the wireless device and the base station.

[0026] In addition, according to the harvesting trajectory of renewable energy in the real world, the renewable energy collected by the base station Follow a normal distribution with a mean of and a variance of . At time slot t , the renewable energy is randomly sampled from the normal distribution.

[0027] 2. Communication model; The communication model uses the Rayleigh channel fading model to characterize the communication process between wireless devices and the base station. The channel gain includes large-scale and small-scale attenuation effects and is expressed as follows: ; where is the communication channel gain; is the communication antenna gain; is the carrier frequency; is the distance between the wireless device i and the base station; α is the path loss exponent; is an independent random channel fading factor that follows an exponential distribution. Based on this, the transmission rate of the wireless device can be expressed as: ; where is the task transmission rate of the wireless device i ; is the bandwidth ratio allocated to the wireless device i ; is the communication bandwidth; is the transmission power; is the noise power.

[0028] 3. Energy consumption model; The energy consumption model quantifies the energy consumption of wireless devices and the base station. The base station energy consumption includes the energy consumption of the MEC server for assisting in computing and the energy consumption of replenishing the power of wireless devices through the WPT technology.

[0029] Assume that the task offloading ratio of the wireless device i is . The MEC server energy consumption is expressed as follows: ; where is the MEC server energy consumption; is the energy consumption required for the MEC server to run one CPU; N is the number of wireless devices; i is the index of the wireless device; is the data volume; is the number of CPU cycles required to process 1 bit of task.

[0030] To ensure the stable operation of wireless devices, each time slot t The base station replenishes time slots for wireless devices t -1 consumption, the WPT energy consumption of the base station is expressed as follows: ; Among them, is the WPT energy consumption of the base station; is the loss factor of radio frequency energy, calculated according to the relationship between the transmission and reception power of radio frequency energy; is the energy consumption of the wireless device in time slot t- 1, including local computing energy consumption and offloading transmission energy consumption of tasks.

[0031] The above model provides an accurate system information basis for subsequent optimization problem formulation and algorithm design, ensuring the feasibility and effectiveness of the algorithm in practical applications.

[0032] Step 2: Formulate the optimization problem; In the hybrid energy WPT-MEC system, the optimization problem formulation module converts the requirements in the actual application scenario into a mathematical optimization problem. Renewable energy is usually stored in a battery with limited capacity. Assuming that the battery level at time slot t is , then the evolution formula over time is expressed as follows: ; Among them, is the battery level at time slot t+ 1; is the WPT energy consumption of the base station; is the energy consumption of the MEC server; is the renewable energy at time slot t ; When the green energy is exhausted, the grid energy serves as backup energy to continue to provide continuous energy support for the entire system. The corresponding grid energy consumption is expressed as follows: ; Among them, is the grid energy consumption; The present invention aims to achieve a green and energy-saving network environment, and the problem is defined as minimizing the long-term average grid energy consumption. To this end, the offloading ratio of tasks , the bandwidth ratio i allocated to wireless device and the local computing frequency are jointly optimized to minimize the long-term average grid energy. The optimization objective is expressed as follows: ; Among them, Tis the timeline length; t is the index of the time slot; meanwhile, it satisfies the maximum execution delay constraint of the task, the sum of the channel allocation ratios is less than 1, the local computing frequency does not exceed its maximum capacity value, and the battery level does not exceed the maximum battery capacity constraint.

[0033] This process abstracts complex practical problems into computable mathematical models, laying a theoretical foundation for the application of deep reinforcement learning algorithms and clarifying the direction and boundaries for subsequent algorithm design.

[0034] Step 3: Convert the optimization problem into a Markov decision process (MDP); In the hybrid energy WPT-MEC system, the MDP conversion module converts the optimization problem into a Markov decision process (MDP), specifically defining the state space, action space, and reward function.

[0035] State space : Includes system information such as the amount of task offloading, channel gain, and battery level.

[0036] Action space : Includes optimization variables such as task offloading decisions, the bandwidth ratio allocated to wireless devices i and local computing frequency.

[0037] Reward function : Since the goal of deep reinforcement learning is to maximize the reward function value, the reward function is the negative value of the goal.

[0038] In addition, min-max normalization is used to process the data before training to ensure the training effect.

[0039] Through this conversion, the complex optimization problem becomes a form suitable for solving by deep reinforcement learning algorithms, enabling the algorithm to gradually learn the optimal policy by interacting with the environment, achieving efficient allocation and utilization of system resources.

[0040] Step 4: Implement the action space reduction strategy; In the hybrid energy WPT-MEC system, the action space reduction module decomposes the resource allocation sub-problem into two sub-constraints by fixing the task offloading decision and uses convex optimization methods to solve the optimal resource allocation scheme. Through this process, the original high-dimensional action space (including task offloading decisions, CPU frequency, communication resource allocation, etc.) is reduced to an N action space with only task offloading decisions, significantly reducing the computational complexity and training difficulty of the algorithm, improving the efficiency of the algorithm, and enhancing its practicality in actual applications.

[0041] Step 5: Perform data normalization processing; In the hybrid energy WPT-MEC system, the normalization processing module normalizes the state observation values and maps them to a unified scale range using the min-max normalization method. This eliminates the influence of data scale differences on model training, improves the training efficiency and stability of the algorithm, not only accelerates the convergence of the neural network, but also enhances the adaptability and robustness of the algorithm in different scenarios.

[0042] Step 6: Construct an improved deep reinforcement learning (DRL) network; In the hybrid energy WPT-MEC system, the improved deep reinforcement learning module combines the deep deterministic policy gradient (DDPG) algorithm with the cross-entropy method (CEM) to form the enhanced deep deterministic policy gradient (EDDPG) algorithm, achieving efficient optimization of task offloading and resource allocation.

[0043] 1. DDPG algorithm; The DDPG algorithm consists of an environment, an Actor network, a Critic network, and an experience replay buffer, providing a stable foundation for policy optimization.

[0044] The environment is an external entity that interacts with the agent in the deep reinforcement learning system, responsible for simulating the actual scenario where the agent is located, and providing state observations, reward signals, and termination conditions.

[0045] The experience buffer follows the first-in-first-out principle and is initialized as an empty buffer with a limited capacity. At each time slot, the agent of the deep reinforcement learning deposits the experience tuple into the experience buffer. Then, it uniformly samples from the experience buffer for network parameter update.

[0046] The Actor network is a three-layer neural network that obtains the mapping relationship from state to action according to the deterministic policy The goal is to maximize the expected total discounted reward, which can be approximated as the Q function; its parameters are updated using the gradient ascent method: ; where, is the gradient of the objective function with respect to the parameter ; M is the size of the experience batch; m is the index of the experience buffer sampling; is Q the gradient of the network with respect to the parameter For a given state the policy at the parameter The gradient. In addition, the structure of the target Actor network is the same as that of the Actor network, and the policy and parameters of the target Actor network are represented as and .

[0047] The Critic network is a three-layer neural network used to evaluate the quality of actions and update network parameters . According to the Bellman equation, the current Q value is expressed as follows: ; where is the discount factor; E[ ] is the expected value of the random variable inside the brackets; is the state under which the action corresponding to the policy generates the Q value. Similarly, the structure of the target Critic network is the same as that of the Critic network, and its parameters are represented as .

[0048] Use the target Actor network to predict the action of the next state, and use the target Critic network to calculate the target Q value y , and update the network parameters: ; where is the loss function; is the Q network's value estimate for the state and the action .

[0049] Different from the deep Q network, the target network of DDPG uses the soft update method to partially update the parameters.

[0050] 2. Cross-entropy method (CEM); First, to generate candidate samples, Gaussian noise is added to the current Actor network parameters to generate new network parameters, thereby increasing sample diversity and promoting exploration of the parameter space; then, the candidate samples are interacted with the environment and the corresponding rewards are recorded to screen out elite samples; finally, the average value of the elite sample parameters is used to update the Actor network parameters, accelerating the convergence speed of the DDPG algorithm.

[0051] The CEM enhancement mechanism significantly improves the exploration ability and convergence speed of the algorithm by generating diverse samples and selecting elite samples, and continuously optimizes the network parameters with a large amount of learning history data to achieve fast convergence of the algorithm.

[0052] Step 7: Generate and output the optimal policy; In the hybrid energy WPT-MEC system, the output optimal result module utilizes the trained deep reinforcement learning network to dynamically generate the best task offloading and resource allocation schemes based on the real-time system states (such as task requirements, network states, device energy states, etc.), achieving the efficient operation of the green-priority hybrid energy WPT-MEC system. This module can not only efficiently handle the dynamically fluctuating task requirements and network states, but also achieve the optimal allocation of resources in complex multi-user scenarios, significantly improving the system performance and user experience.

[0053] In the embodiment of the present invention, Figure 2 An application scenario example of the present invention is shown. In the figure, there is a base station (BS) that collects green energy such as solar energy and wind energy through direct environmental energy harvesting (EH) components (such as solar panels and wind turbines), and also cooperatively uses grid energy for power supply. There are multiple wireless devices distributed around the base station, and these wireless devices can indirectly receive green energy from the base station through wireless power transfer (WPT). In terms of data processing, the computing tasks of the wireless devices can either be executed locally (such as the local computing marked in the figure) or partially offloaded to the mobile edge computing (MEC) server equipped in the base station for processing. Through the data transmission represented by solid arrows and the energy transmission represented by dashed arrows, the interaction of data and energy is achieved, and thus efficient and flexible data processing is achieved.

[0054] The following describes the specific implementation of the present invention in detail in combination with specific embodiments.

[0055] Embodiment 1: This embodiment verifies the energy consumption optimization ability of the enhanced deep deterministic policy gradient (EDDPG) algorithm in the green-priority hybrid energy WPT-MEC system through numerical simulation and analysis.

[0056] To verify the applicability of the EDDPG algorithm in the WPT-MEC system, it is compared with multiple benchmark algorithms. The results are shown in Table 1: Table 1 Influence of the number of devices and the amount of tasks on the average electrical energy (J)

[0057] From the data on the impact of the number of devices and the amount of tasks on the average electrical energy (J, the optimization objective) given in Table 1, it can be seen that the EDDPG algorithm can dynamically adjust the task offloading and resource allocation strategies when the number of devices and the amount of tasks increase. For example, when the amount of tasks is 300 Kbit, the grid energy consumption of the EDDPG algorithm is reduced by 78% compared with greedy local computing, 61% compared with greedy offloading computing, and 28% compared with random offloading. Thus, compared with the random offloading strategy, the EDDPG algorithm shows higher stability in terms of energy consumption and performance. By preferentially utilizing green energy and optimizing the use of grid energy, it significantly reduces the grid energy consumption, reduces the energy cost and environmental pollution. This fully demonstrates that the EDDPG algorithm has superiority and applicability in energy consumption optimization and adaptability, and is suitable for large-scale Internet of Things scenarios.

[0058] The EDDPG algorithm is compared with a variety of classic DRL algorithms (DDPG, SAC, TD3) to prove that the improvement of the EDDPG algorithm has improved its convergence speed. From Figure 3 it can be seen that as the number of time steps increases, the EDDPG algorithm quickly converges to an average reward of about -5.2 after about 30,000 time steps, while algorithms such as DDPG, SAC, and TD3 converge relatively slowly, and the average rewards finally achieved are generally lower than that of the EDDPG algorithm. Thus, the EDDPG algorithm is superior to other DRL algorithms in terms of convergence speed and reward, which means that the EDDPG algorithm not only has strong global search ability, but also can use elite samples to improve its stability, and is more suitable for application in the WPT-MEC system.

[0059] The above experimental results show that in the face of a large number of users and a huge amount of tasks, the EDDPG algorithm can provide optimal task offloading and resource allocation solutions for multiple users in real time, significantly improving the computing efficiency and energy utilization rate of the system.

[0060] The above is only the preferred implementation manner of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent.

Claims

1. A task offloading and resource allocation method for a hybrid energy WPT-MEC system, characterized in that Assume that the hybrid energy WPT-MEC system includes a base station and N wireless devices, and the system operates according to a timeline . In each time slot, the wireless device i has a randomly arriving task , where represent the data volume, the number of CPU cycles required to process 1 bit of task, and the maximum tolerable delay for task execution, respectively; The method includes the following steps: Step 1: Establish a system model; By establishing an energy harvesting model, a communication model, and an energy consumption model, comprehensively describe the dynamic behavior of the system; Step 2: Formulate an optimization problem; By analyzing the renewable energy storage and consumption, and the grid energy replenishment mechanism, with the goal of minimizing the long-term average grid energy consumption, jointly optimize the task offloading ratio, the bandwidth ratio allocated to wireless devices i and the local computing frequency, and set constraints to transform the actual scenario requirements into a mathematical optimization problem; Step 3: Transform the optimization problem into a Markov decision process; By defining a state space that includes the task offloading volume, channel gain, and battery level , an action space that includes task offloading decisions, channel allocation ratios, and local computing frequencies , and a reward function set to a target negative value , and performing min-max normalization on the data before training to transform the optimization problem into a form suitable for solving by deep reinforcement learning algorithms; Step 4: Implement an action space reduction strategy; By fixing the task offloading decision, the resource allocation sub-problem is decomposed into two sub-constraints, and the convex optimization method is used to solve the optimal resource allocation scheme, reducing the original high-dimensional action space to an action space with only the task offloading decision N dimensional action space; Step 5: Perform data normalization processing; Use the min-max normalization method to process the state observation values and map them to a unified scale range; Step 6: Construct an improved deep reinforcement learning network; Combine the deep deterministic policy gradient algorithm with the cross-entropy method to form an enhanced deep deterministic policy gradient algorithm. The deep deterministic policy gradient algorithm provides a stable policy optimization foundation, and the cross-entropy method improves the algorithm's exploration ability and convergence speed by generating diverse samples and selecting elite samples; Step 7: Generate and output the optimal policy; Use the trained deep reinforcement learning network to dynamically generate the best task offloading and resource allocation scheme according to the real-time system state.

2. The task offloading and resource allocation method for the hybrid energy WPT-MEC system according to claim 1, characterized in that, The energy harvesting model is used to capture the dynamic changes of green energy; assuming that the RF transmission power is , the received power is expressed as follows: ; Among them, is the radio frequency received power; and are the radio frequency signal transmitting and receiving antenna gains respectively; is the radio frequency signal wavelength; d is the distance between the wireless device and the base station; Renewable energy collected by the base station obeys a normal distribution with a mean of and a variance of At time slot t The renewable energy is randomly sampled from the normal distribution.

3. The task offloading and resource allocation method for the hybrid energy WPT-MEC system according to claim 1, wherein The communication model uses the Rayleigh channel fading model to characterize the communication process between the wireless device and the base station. The channel gain is expressed as follows: ; wherein, is the communication channel gain; is the communication antenna gain; is the carrier frequency; is the wireless device i the distance between the base station; α is the path loss exponent; is an independent random channel fading factor, following an exponential distribution; The transmission rate of the wireless device is expressed as: ; Among them, is the task transmission rate of the wireless device i ; is the bandwidth ratio allocated to the wireless device i ; is the communication bandwidth; is the transmission power; is the noise power.

4. The task offloading and resource allocation method for the hybrid energy WPT-MEC system according to claim 1, wherein The energy consumption model is used to quantify the energy consumption of the wireless device and the base station. The base station energy consumption includes the energy consumption of the MEC server for assisting calculations and the energy consumption of replenishing the power of the wireless device through the WPT technology; Assume that the task offloading ratio of the wireless device i is , and the energy consumption of the MEC server is expressed as follows: ; Among them, is the energy consumption of the MEC server; is the energy consumption required for the MEC server to run one CPU; N is the number of wireless devices; i is the index of the wireless device; is the data volume; is the number of CPU cycles required to process 1 bit of task; Each time slot t The base station replenishes time slots for wireless devices t -1 consumption, the WPT energy consumption of the base station is expressed as follows: ; Among them, is the WPT energy consumption of the base station; is the loss factor of radio frequency energy; is the energy consumption of the wireless device at time slot t- 1, including the local computing energy consumption and the offloading transmission energy consumption of the task.

5. The task offloading and resource allocation method for the hybrid energy WPT-MEC system according to claim 4, wherein The specific process of Step 2 is as follows: Renewable energy is stored in a battery with limited capacity. Assuming that the battery level at time slot t is , its evolution formula over time is expressed as follows: ; Among them, is the battery level of the battery in time slot t+ 1; is the WPT energy consumption of the base station; is the energy consumption of the MEC server; is the time slot t when the renewable energy is available; When the green energy is exhausted, the grid energy is used as backup energy to continue to provide continuous energy support for the entire system. The corresponding grid energy consumption is expressed as follows: ; Among them, is the power grid energy consumption; Jointly optimize the offloading ratio of tasks , the bandwidth ratio allocated to wireless devices i , and the local computing frequency to minimize the long-term average grid energy, the optimization objective is expressed as follows: ​ ; Among them, T is the length of the timeline; t is the index of the time slot; at the same time, it satisfies the maximum execution delay constraint of the task, the sum of the channel allocation ratios is less than 1, the local computing frequency does not exceed its maximum capacity value, and the battery level does not exceed the maximum battery capacity constraint condition.

6. The task offloading and resource allocation method for the hybrid energy WPT-MEC system according to claim 1, wherein The deep deterministic policy gradient algorithm includes an environment, an Actor network, a Critic network, and an experience replay buffer; The Actor network is a three-layer neural network that, according to the deterministic policy obtains the mapping relationship from state to action. The goal is to maximize the expected total discounted reward, approximated by the Q function; its parameters are updated using the gradient ascent method: ; Among them, is the gradient of the objective function with respect to the parameter ; M is the size of the experience batch; m is the index sampled from the experience buffer; is Q the gradient of the network with respect to the parameter ; is the gradient of the given state , policy with respect to the parameter The structure of the target Actor network is the same as that of the Actor network, and the policy and parameters of the target Actor network are represented as and ; The Critic network is a three-layer neural network that evaluates the quality of actions and updates network parameters , according to the Bellman equation, the current Q-value is expressed as follows: ; wherein, is the discount factor; E[ ] is the expected value of the random variable within the brackets; is the state under which the policy generates the corresponding Q value of the action; The structure of the target Critic network is the same as that of the Critic network, and its parameters are denoted as ; Predict the action of the next state using the target Actor network and calculate the target Q value using the target Critic network y , and update the network parameters: ; Among them, is the loss function; is Q the value estimation of the network for the state and the action .

7. The task offloading and resource allocation method for the hybrid energy WPT-MEC system according to claim 6, wherein The specific process of the cross-entropy method is as follows: First, add Gaussian noise to the current Actor network parameters to generate candidate samples; then, interact the candidate samples with the environment and record the corresponding rewards, and select the elite samples; finally, update the Actor network parameters using the average value of the elite sample parameters to accelerate the convergence speed of the DDPG algorithm.

Citation Information

Patent Citations

  • Calculation unloading and resource allocation method for smart grid power supply system

    CN113630734A

  • Mobile edge computing resource allocation method and system for dynamic user random access

    CN117793805A

  • Automatic driving resource arrangement and task unloading method based on model segmentation

    CN119603653A

  • Joint optimization system and method for computation offloading and resource allocation in multi-constraint-edge environment

    WO2024065903A1

Cited By

  • Unmanned aerial vehicle assisted MEC network energy efficiency optimization method based on reinforcement learning

    CN121037917A

  • A reinforcement learning-based method for optimizing the energy efficiency of unmanned aerial vehicle-assisted MEC networks.

    CN121037917B

  • DeepSeek-based server green energy perception model training scheduling method and system

    CN121187738A

  • A training and scheduling method and system for server-side green energy awareness based on DeepSeek.

    CN121187738B