An online task offloading and resource allocation joint optimization method for wireless powered edge computing

By employing Lyapunov optimization and deep reinforcement learning methods using convolutional neural networks, the task offloading and resource allocation of wirelessly powered edge computing networks are optimized, solving the problems of computing speed and energy utilization efficiency in dynamic environments and achieving more efficient computing and energy management.

CN122373055APending Publication Date: 2026-07-10DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DALIAN MARITIME UNIVERSITY
Filing Date
2026-05-13
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In multi-user wirelessly powered edge computing networks, existing technologies struggle to effectively optimize task offloading and resource allocation in dynamic environments, especially under time-varying channel conditions and random task arrival scenarios, resulting in poor computing speed and energy utilization efficiency.

Method used

A deep reinforcement learning approach combining Lyapunov optimization and convolutional neural networks is adopted. By constructing a virtual energy queue and an Actor-Critic structure, the offloading decision, the local computing frequency of the wireless device, and the energy allocation are optimized, which is transformed into a deterministic problem per time frame. The optimal offloading strategy is solved by convex optimization and Lagrange duality.

Benefits of technology

It achieves long-term optimized performance and stability in dynamic environments, improves the weighted computing rate and energy utilization efficiency of wireless devices, and reduces data queue backlog and energy shortage problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122373055A_ABST
    Figure CN122373055A_ABST
Patent Text Reader

Abstract

The application provides an online task offloading and resource allocation joint optimization method for wireless energy supply edge computing, comprising the following steps: S1, establishing a basic framework of a wireless energy supply assisted mobile edge computing model under a random task data arrival and time-varying channel scenario; S2, according to the basic framework of the wireless energy supply assisted mobile edge computing model, performing model establishment on mobile edge computing resource allocation to obtain a mathematical model; S3, using Lyapunov optimization combined with deep reinforcement learning based on a convolutional neural network and a convex optimization method to maximize the weighted sum of the computing rates of all wireless devices and jointly optimize offloading decisions, local computing frequencies of the wireless devices, offloading transmission time allocation and offloading energy. The application solves the problem of wireless energy supply edge computing online task offloading under the conditions of a time-varying channel and random task arrival, and guarantees long-term offloading benefits of the wireless devices and long-term stability of the system through the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mobile edge computing technology, and more particularly to a joint optimization method for online task offloading and resource allocation for wirelessly powered edge computing. Background Technology

[0002] With the rapid development of communication and IoT technologies, the data generated by IoT devices is growing exponentially. However, most IoT devices have limited computing power and insufficient battery capacity, making it difficult to meet the computing and long-term operation requirements of complex mobile services. Combining Mobile Edge Computing (MEC) and Wireless Power Transfer (WPT) can effectively address the problems of insufficient computing power and energy constraints, improving the computing power and battery life of wireless devices. MEC enhances the computing power of devices by offloading computationally intensive tasks to edge servers; WPT can provide a sustainable energy supply to devices wirelessly. Users can efficiently complete different tasks by offloading computations, while effectively alleviating battery power shortages by receiving radio frequency energy, ensuring long-term efficient operation of the devices.

[0003] Existing work can integrate radio frequency power transmitters into MEC servers, applying WPT to MEC to address the dual limitations of computing power and energy in network edge devices. Therefore, in multi-user wirelessly powered edge computing networks, how to balance computational offloading and resource allocation under the constraints of data processing causality and available energy causality to maximize the weighted and computational rates of system users is a problem worth considering.

[0004] Several studies in recent years have investigated joint computation offloading and resource allocation in wirelessly powered MEC systems. However, most of these studies focus on short-term optimization, assuming known channel conditions or user computational tasks, and neglect long-term performance optimization. In multi-user wirelessly powered edge networks, more realistic scenarios need to be considered, including time-varying channel conditions and random data arrival. Furthermore, due to the uncertainty of channel conditions and task arrival, there may be situations where poor channel conditions result in less energy collection and excessive data congestion. In addition, the rapid changes in edge network environments necessitate frequent re-solution of complex computation offloading problems. These issues pose significant challenges to the long-term optimization performance and stability of wirelessly powered edge computing networks in dynamic environments. Summary of the Invention

[0005] In view of this, the purpose of this invention is to propose a joint optimization method for online task offloading and resource allocation for wirelessly powered edge computing, so as to solve the problem of task offloading and resource allocation in wirelessly powered assisted mobile edge computing networks under dynamic environments.

[0006] The technical means employed in this invention are as follows:

[0007] A joint optimization method for online task offloading and resource allocation for wirelessly powered edge computing includes the following steps: S1. Establish the basic framework of a wireless power-assisted mobile edge computing model under scenarios of random task data arrival and time-varying channel; S2. Based on the basic framework of the wireless power-assisted mobile edge computing model, a model for mobile edge computing resource allocation is established to obtain a mathematical model. S3. Based on the mathematical model obtained in S2, Lyapunov optimization is used in combination with deep reinforcement learning and convex optimization methods based on convolutional neural networks to maximize the weighted sum of computing speeds of all wireless devices under the constraints of data queue and long-term available energy. The offloading decision, local computing frequency of wireless devices, offloading transmission time allocation and offloading energy are jointly optimized, and the optimal offloading decision, local computing frequency, offloading time allocation and offloading energy of the current time frame are output.

[0008] Furthermore, S1 specifically includes the following steps: S11. Construct a wireless power supply edge computing model for a single base station with multiple wireless devices; the base station is equipped with an edge computing server, a radio frequency energy transmitter, and multiple antennas; each wireless device is equipped with a single antenna and a rechargeable battery to collect and store radio frequency energy for power supply. S12. A binary offload method is adopted, in which the wireless device performs tasks through local computing or offloaded computing; downlink power transmission and uplink computing offload are performed simultaneously on the orthogonal frequency band, and uplink computing offload uses TDMA for data transmission. S13. Divide the system time into continuous time frames of equal length, and let Indicates the first A time frame; the channel remains static within a time frame but varies between different time frames; [recorded] Let be the channel gain of all wireless devices in the t-th time frame; the arrival of task data from wireless devices follows an exponential distribution. Indicates the first The first time frame The data queue length of each wireless device; Indicates the first The first time frame The remaining energy level of each wireless device; S14. The wireless device obtains wireless energy from the base station and stores the collected energy in the battery for local computing or offloading computing; the energy used by the wireless device to process tasks comes only from the battery, and the energy collected in the current time frame can only be used in the next time frame; S15. The long-term available energy constraint of the wireless device, that is, the long-term available energy of each wireless device is higher than a preset threshold, so as to avoid excessive discharge or data processing delay caused by insufficient energy collection due to poor channel conditions. S16. Under the constraints of data queue stability and long-term available energy, with the goal of maximizing the weighted sum of computing speeds of all wireless devices, construct the basic framework of a wireless power-assisted mobile edge computing model for random task data arrival and time-varying channel scenarios.

[0009] Furthermore, in S2, the objective function of the mathematical model is as follows: , The objective function is a stochastic optimization problem with multiple time frames, where, Indicates the first The weight of each wireless device Indicates the first The wireless device in the first The rate of each time frame.

[0010] Furthermore, S3 specifically includes the following steps: S31. To meet the long-term energy constraints of wireless devices, a virtual energy queue is introduced. ,in, The energy perturbation parameter transforms the long-term available energy constraint into a stability constraint for a virtual energy queue. The objective function is transformed using Lyapunov functions, Lyapunov drift, and minimizing the drift plus penalty upper bound, thus decoupling the stochastic optimization problem of S2 multi-timeframes into a deterministic problem per timeframe. S32. Deep reinforcement learning using an Actor-Critic structure, where the Actor module is composed of a convolutional neural network. The system generates channel conditions and task load, and updates the data queue and virtual energy queue; it also adjusts the channel gain of the current time frame. Data queue and virtual energy queue Input a convolutional neural network and output a relaxed unloading decision; S33. Perform noise-preserving quantization on the relaxed offloading decision to obtain a set of candidate binary offloading strategies for all wireless devices at the current time. S34. The Critic module uses a convex optimization method to evaluate each candidate binary unloading decision. Once the binary unloading decision is determined, the objective function is transformed into a convex optimization problem. The solution is obtained through closed-form solutions and the Lagrange duality method to obtain the binary unloading decision that optimizes the objective function. At the same time, the corresponding local computing frequency, unloading time allocation, and unloading energy scheme are also obtained.

[0011] Furthermore, S31 specifically includes the following steps: S311. The stochastic optimization problem with multiple time frames is decoupled using Lyapunov theory, where the Lyapunov function is defined as the sum of squares of the data queue and the virtual energy queue. The formula for the Lyapunov function is as follows:

[0012] in, Indicates the length of the data queue. Indicates the length of the virtual energy queue; S312. Subtract the Lyapunov function value of the current time frame from the Lyapunov function value of the next time frame to obtain the Lyapunov drift. The Lyapunov drift is used to weigh the choice of resource allocation strategy. By controlling the change of the function at each step, the final value of the function can be controlled. The formula for the Lyapunov drift is as follows:

[0013] S313. Based on the Lyapunov drift function obtained in S312, map the objective function to a suitable penalty function to obtain the Lyapunov drift plus penalty function, as shown in the following formula:

[0014] in, It is a non-negative control factor, which can be adjusted The size of the data queue is used to obtain a trade-off between minimizing the data queue backlog, maximizing the available energy level, and minimizing the penalty, and finally the asymptotic optimal solution is obtained.

[0015] Furthermore, S32 specifically includes the following steps: Based on historical data, channel gain, data queue length, and virtual energy queue length are used as inputs. Through the forward propagation of a convolutional neural network, a relaxed binary offloading decision is output. The convolutional neural network includes an input layer, a hidden layer, and an output layer.

[0016] Furthermore, S33 specifically includes the following steps: Based on the relaxation unloading decision output by the convolutional neural network, it is quantized into 2K+1 candidate binary unloading decisions using a noise-preserving quantization method, including K decisions with added noise.

[0017] Furthermore, S34 specifically includes the following steps: S341. Based on the given binary unloading decision, decompose the optimization problem into a local computation subproblem and an unloading subproblem; S342. For the local computation subproblem, construct a cubic convex optimization function and obtain the optimal local computation frequency by comparing the stationary points and boundary points. S343. For the unloading subproblem of S341, introduce dual variables to construct a partial Lagrangian function; S344. By leveraging the independence of variables for each wireless device, the problem is broken down into multiple parallel sub-problems of offloading devices. S345. For each wireless device, the offloading energy is expressed as a function of offloading time and offloading rate using the relationship of communication rate, which is transformed into an optimization problem only concerning offloading time and offloading rate. S346. Hierarchical optimization is adopted: the outer layer optimizes the unloading rate, and the inner layer optimizes the unloading time. The inner layer problem is a convex optimization, and the optimal solution is located at the stationary point or the boundary, thus obtaining the optimal unloading time allocation under a given unloading rate. S347. Substitute the obtained optimal unloading time allocation into the objective function to obtain a linear programming problem only concerning the unloading rate; obtain the optimal solution by determining the sign of the coefficients; if the total time constraint is not satisfied, use the bisection method to solve the optimal dual variable, and then obtain the optimal unloading time allocation and optimal unloading energy under the optimal dual variable through linear programming. S348. The system enters the next time frame, generates new time-varying channel conditions and random task data, updates the data queue and virtual energy queue according to the energy consumed by the optimal offloading decision and the amount of data processed, and uses the updated state as the input of the convolutional neural network in the next time frame. S349. Repeat S341 to S348, and update the convolutional neural network parameters using historical data until the preset number of iterations is reached.

[0018] The present invention also provides a storage medium comprising a stored program, wherein, when the program is executed, it performs any of the above-described methods for joint optimization of online task offloading and resource allocation for wireless-powered edge computing.

[0019] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes any of the above-described online task offloading and resource allocation joint optimization methods for wirelessly powered edge computing through the computer program.

[0020] Compared with the prior art, the present invention has the following advantages: This invention proposes an online task offloading and resource allocation joint optimization model for wireless power edge computing, which is more in line with the actual scenario than models that focus on short-term optimization based on known channel conditions or user computing tasks. The online task offloading algorithm proposed in this invention, which combines deep reinforcement learning based on convolutional neural networks and Lyapunov optimization, also known as the LyCNNOP algorithm, introduces a virtual energy queue to transform the energy threshold constraint into a stability constraint of the virtual energy queue. Then, the Lyapunov function is used to decouple the stochastic optimization problem into a deterministic problem per time frame, ensuring the stability of the data queue and the virtual energy queue is stable near the threshold. For the deterministic problem at each time frame, this invention adopts an Actor-Critic structure for solving the problem. The Actor module uses a convolutional neural network to solve the binary offloading decision, while the Critic module uses convex optimization (closed-form solution and Lagrange duality) to solve the offloading problem and the local computation problem. This approach utilizes historical data and mathematical formula derivation to better complete the computational offloading. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a basic framework diagram of the present invention.

[0023] Figure 2 This is a flowchart of the algorithm of the present invention.

[0024] Figure 3 A comparison chart of the algorithm's weighted sum calculation rate when the number of wireless devices is 10.

[0025] Figure 4 A comparison chart showing the average data queue length of the algorithm when the number of wireless devices is 10.

[0026] Figure 5 A comparison chart showing the average battery energy levels of the algorithm when the number of wireless devices is 10.

[0027] Figure 6 A comparison chart of the algorithm's weighted sum calculation rate when the number of wireless devices is 20.

[0028] Figure 7 The graph shows a comparison of the average data queue length of the algorithm when the number of wireless devices is 20.

[0029] Figure 8 A comparison chart showing the average battery energy levels of the algorithm when the number of wireless devices is 20.

[0030] Figure 9 A comparison chart of the algorithm's weighted sum calculation rate when the number of wireless devices is 30.

[0031] Figure 10 The graph shows a comparison of the average data queue length of the algorithm when the number of wireless devices is 30.

[0032] Figure 11 This is a comparison chart showing the average battery energy levels of the algorithm when the number of wireless devices is 30. Detailed Implementation

[0033] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] like Figure 1 and 2 As shown, this invention provides a joint optimization method for online task offloading and resource allocation for wirelessly powered edge computing, comprising the following steps: S1. Establish the basic framework of a wireless power-assisted mobile edge computing model under scenarios of random task data arrival and time-varying channel; S11. Establish a wireless power supply edge computing model with a single base station and multiple wireless devices. The base station is equipped with an edge computing server and a radio frequency power transmitter. The base station has multiple antennas, and each device has one antenna and a rechargeable battery to collect and store energy to power its operation. A binary offload method is adopted, allowing users to process tasks through local computing or offloaded computing. To avoid interference, downlink power transmission and uplink computing offload can be performed simultaneously on orthogonal frequency bands, and uplink computing offload uses TDMA for data transmission. The system time is divided into continuous time frames of equal length. Indicates the first The channel remains static within a single time frame, but may change between different time frames. Indicates the first The channel gain of all wireless devices in each time frame is set, and the arrival of task data from the wireless devices is configured to follow an exponential distribution. Indicates the first The first time frame Length of data queue for each wireless device Indicates the first The first time frame Remaining power level of each wireless device; S12. Based on time-varying channel conditions and the arrival of random task data, the wireless device obtains wireless energy from the base station. The user stores the collected energy in the battery and uses it for local computation or offloading computation. It is assumed that the energy used by the IoT device to process tasks comes solely from the battery, and the energy collected in the current time slot can only be used in the next time slot. To avoid over-discharge of the wireless device and delays in data processing due to insufficient collected energy under poor channel conditions, a long-term available energy constraint for the wireless device is introduced, meaning the long-term available energy of the wireless device must exceed a threshold. Under the conditions of ensuring the data queue and the long-term available energy constraint of the wireless device, the weighted sum computation rate of all wireless devices is maximized, thus obtaining the basic framework of the wireless-powered mobile edge computing model under random task data arrival and time-varying channel scenarios.

[0036] S2. Based on the basic framework of the wireless power-assisted mobile edge computing model, a model for mobile edge computing resource allocation is established to obtain a mathematical model. Establish the following objective function: , in, Indicates the first The weight of each wireless device Indicates the first The wireless device in the first The rate of each time frame.

[0037] S3. Based on the mathematical model obtained in S2, Lyapunov optimization is combined with deep reinforcement learning based on convolutional neural networks and convex optimization methods to maximize the weighted sum of computing rates of all wireless devices under the constraints of data queue and long-term available energy. The offloading decision, local computing frequency of wireless devices, offloading transmission time allocation, and offloading energy are jointly optimized. The objective function in S2 is a stochastic optimization problem involving multiple time frames.

[0038] S31. Introducing a virtual energy queue ,in The energy perturbation parameter can transform the long-term available energy level constraint into a stability constraint for the virtual energy queue. Furthermore, the objective function is transformed by setting the Lyapunov function, Lyapunov drift, and minimizing the drift plus penalty upper bound. A penalty factor is introduced to obtain a new objective function, thus decoupling the stochastic optimization problem of S2 multi-timeframes into a per-timeframe deterministic problem. S311. The stochastic optimization problem with multiple time frames is decoupled using Lyapunov theory, where the Lyapunov function is defined as the sum of squares of the data queue and the virtual energy queue. The formula for the Lyapunov function is as follows:

[0039] in, Indicates the length of the data queue. Indicates the length of the virtual energy queue; S312. Subtract the Lyapunov function value of the current time frame from the Lyapunov function value of the next time frame to obtain the Lyapunov drift. The Lyapunov drift is used to weigh the choice of resource allocation strategy. By controlling the change of the function at each step, the final value of the function can be controlled. The formula for the Lyapunov drift is as follows:

[0040] S313. Based on the Lyapunov drift function obtained in S312, map the objective function to a suitable penalty function to obtain the Lyapunov drift plus penalty function, as shown in the following formula:

[0041] in, It is a non-negative control factor, which can be adjusted by... The size of the data queue is used to obtain a trade-off between minimizing the data queue backlog, maximizing the available energy level, and minimizing the penalty. The final solution obtained is the asymptotic optimal solution.

[0042] S32. Deep reinforcement learning adopts an Actor-Critic structure, where the Actor module is composed of a convolutional neural network. The system generates channel conditions and task load, and updates the data queue and virtual energy queue, including channel information. Data queue Virtual energy queue The input is fed into a convolutional neural network to obtain a relaxed unloading decision; Deep reinforcement learning methods require the use of historical data, setting up input, output, and hidden layers, and passing channel conditions, data queues, and virtual energy queues through a convolutional neural network to output relaxed binary offloading decisions.

[0043] S33. The relaxed offloading decision obtained in S32 is processed by noise-preserving quantization to obtain a series of candidate binary offloading strategies for all users at the current time. Based on the relaxed unloading decision output by the convolutional neural network, it is quantized as Group decision-making, which includes Unloading decisions for groups with added noise.

[0044] S34. The Critic module of deep reinforcement learning uses convex optimization methods to evaluate each candidate binary unloading decision, specifically closed-form solutions and Lagrange duality. Resource allocation is performed based on the objective function obtained in S31 and the binary unloading decision obtained in S33. The objective function obtained in S31 is a non-convex problem. Once the binary unloading decision is determined, the problem is transformed into a convex optimization problem. At this time, the convex optimization method is used to find the binary unloading decision that makes the objective function optimal and to obtain the corresponding local computation frequency, unloading time allocation and unloading energy scheme under this decision.

[0045] S341. Based on the given binary unloading decision, decompose the problem into a local computation subproblem and an unloading subproblem; S342. For the local computation subproblem, which is a cubic convex optimization function, the optimal local computation frequency can be obtained by comparing the relationship between the stationary points and the boundary points. S343. For the unloading subproblem, first introduce dual variables to set the partial Lagrangian function; S344. Since each device item is independent, the problem can be broken down into multiple parallel sub-problems of unloading devices. S345. For the S344 problem for each wireless device, by utilizing the relationship of communication rate, the offloading energy can be expressed as a formula in terms of offloading time and offloading rate. At this point, the problem is transformed into an optimization problem only in terms of offloading time and offloading rate. S346. Perform hierarchical optimization on the problem obtained in S345. The outer layer optimizes the unloading rate, and the inner layer optimizes the unloading time. The inner layer problem is a convex optimization problem. The optimal solution is at the stationary point or the boundary. At this time, the optimal unloading time allocation under the given unloading rate can be obtained. S347. Substituting the optimal unloading time allocation obtained in S346 into the objective function of S345 yields a function relating only to the unloading rate. This is now a linear programming problem, and the optimal solution is obtained by considering the sign of the decision coefficients. However, this may not satisfy the total time constraint. Therefore, it is necessary to find the optimal dual variable that satisfies the total time constraint. The bisection method is used to solve for the optimal dual variable, and then linear programming can be used to obtain the optimal unloading time allocation and optimal unloading energy under the optimal dual variable.

[0046] S348. The system generates the time-varying channel conditions and random task data arrival for the next time frame. Based on the energy consumed under the optimal offloading decision obtained in S347, the system updates the data queue and virtual energy queue with the amount of data processed, and updates the state of the input convolutional neural network to the state of the next time frame.

[0047] S349. Repeat the above steps and update the parameters of the convolutional neural network using historical data until the final number of iterations is reached.

[0048] This paper proposes a joint optimization method for online task offloading and resource allocation based on wireless powered edge computing networks. By using the Lyapunov optimization method, combined with deep reinforcement learning and convex optimization methods based on convolutional neural networks, the optimization objective is to maximize the weighted sum computing rate of all wireless devices, while ensuring the stability of the data queue and the available energy threshold constraint.

[0049] This embodiment conducts experiments in real-world task scenarios, testing different numbers of terminal wireless devices. The comparison algorithms used in this paper employ Lyapunov combined with fully connected neural networks (LyDNN), Lyapunov combined with stochastic partial offloading (LyROM), and Lyapunov combined with local-only computation (LyAL).

[0050] like Figure 3 The image shows a comparison of the weighted sum calculation rates of various algorithms within the same time range when there are 10 wireless devices.

[0051] like Figure 4 The figure shows a comparison of the average data queue length under different algorithms within the same time range, with 10 wireless devices.

[0052] like Figure 5 The image shows a comparison of the average battery energy levels under different algorithms within the same time period, with 10 wireless devices.

[0053] like Figure 6 The image shows a comparison of the weighted sum calculation rates of various algorithms within the same time range when there are 20 wireless devices.

[0054] like Figure 7 The figure shows a comparison of the average data queue length under different algorithms within the same time range, with 20 wireless devices.

[0055] like Figure 8 The image shows a comparison of the average battery energy levels of different algorithms within the same time period for a total of 20 wireless devices.

[0056] like Figure 9 The image shows a comparison of the weighted sum calculation rates of various algorithms within the same time range when there are 30 wireless devices.

[0057] like Figure 10 The figure shows a comparison of the average data queue length under different algorithms within the same time range, with 30 wireless devices.

[0058] like Figure 11 The image shows a comparison of the average battery energy levels of different algorithms within the same time period for a total of 30 wireless devices.

[0059] As can be seen from the figure, the LyCNNOP algorithm proposed in this patent outperforms the comparison algorithms in terms of average weighted sum calculation speed, average data queue length, and average battery energy level, and the effect is more obvious with a larger number of users.

[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A joint optimization method for online task offloading and resource allocation for wirelessly powered edge computing, characterized in that, Includes the following steps: S1. Establish the basic framework of a wireless power-assisted mobile edge computing model under scenarios of random task data arrival and time-varying channel; S2. Based on the basic framework of the wireless power-assisted mobile edge computing model, a model for mobile edge computing resource allocation is established to obtain a mathematical model. S3. Based on the mathematical model obtained in S2, Lyapunov optimization is used in combination with deep reinforcement learning and convex optimization methods based on convolutional neural networks to maximize the weighted sum of computing speeds of all wireless devices under the constraints of data queue and long-term available energy. The offloading decision, local computing frequency of wireless devices, offloading transmission time allocation and offloading energy are jointly optimized, and the optimal offloading decision, local computing frequency, offloading time allocation and offloading energy of the current time frame are output.

2. The online task offloading and resource allocation joint optimization method for wirelessly powered edge computing according to claim 1, characterized in that, S1 specifically includes the following steps: S11. Construct a wireless power supply edge computing model for a single base station with multiple wireless devices; the base station is equipped with an edge computing server, a radio frequency energy transmitter, and multiple antennas; each wireless device is equipped with a single antenna and a rechargeable battery to collect and store radio frequency energy for power supply. S12. Using binary offloading method, wireless devices process tasks through local computing or offloading computing. Downlink power transmission and uplink compute offloading occur simultaneously in orthogonal frequency bands, with uplink compute offloading using TDMA for data transmission. S13. Divide the system time into continuous time frames of equal length, and let Indicates the first A time frame; the channel remains static within a time frame but varies between different time frames; [recorded] Let be the channel gain of all wireless devices in the t-th time frame; the arrival of task data from wireless devices follows an exponential distribution. Indicates the first The first time frame Data queue length for each wireless device; Indicates the first The first time frame The remaining energy level of each wireless device; S14. The wireless device obtains wireless energy from the base station and stores the collected energy in the battery for local computing or offloading computing; the energy used by the wireless device to process tasks comes only from the battery, and the energy collected in the current time frame can only be used in the next time frame; S15. The long-term available energy constraint of the wireless device, that is, the long-term available energy of each wireless device is higher than a preset threshold, so as to avoid excessive discharge or data processing delay caused by insufficient energy collection due to poor channel conditions. S16. Under the constraints of data queue stability and long-term available energy, with the goal of maximizing the weighted sum of computing speeds of all wireless devices, construct the basic framework of a wireless power-assisted mobile edge computing model for random task data arrival and time-varying channel scenarios.

3. The online task offloading and resource allocation joint optimization method for wirelessly powered edge computing according to claim 1, characterized in that, In S2, the objective function of the mathematical model is as follows: , The objective function is a stochastic optimization problem with multiple time frames, where, Indicates the first The weight of each wireless device Indicates the first The wireless device in the first The rate of each time frame.

4. The online task offloading and resource allocation joint optimization method for wirelessly powered edge computing according to claim 3, characterized in that, S3 specifically includes the following steps: S31. To meet the long-term energy constraints of wireless devices, a virtual energy queue is introduced. ,in, The energy perturbation parameter transforms the long-term available energy constraint into a stability constraint for a virtual energy queue. The objective function is transformed using Lyapunov functions, Lyapunov drift, and minimizing the drift plus penalty upper bound, thus decoupling the stochastic optimization problem of S2 multi-timeframes into a deterministic problem per timeframe. S32. Deep reinforcement learning using an Actor-Critic structure, where the Actor module is composed of a convolutional neural network. The system generates channel conditions and task load, and updates the data queue and virtual energy queue; it also adjusts the channel gain of the current time frame. Data queue and virtual energy queue Input a convolutional neural network and output a relaxed unloading decision; S33. Perform noise-preserving quantization on the relaxed offloading decision to obtain a set of candidate binary offloading strategies for all wireless devices at the current time. S34. The Critic module uses a convex optimization method to evaluate each candidate binary unloading decision. Once the binary unloading decision is determined, the objective function is transformed into a convex optimization problem. The solution is obtained through closed-form solutions and the Lagrange duality method to obtain the binary unloading decision that optimizes the objective function. At the same time, the corresponding local computing frequency, unloading time allocation, and unloading energy scheme are also obtained.

5. The online task offloading and resource allocation joint optimization method for wirelessly powered edge computing according to claim 4, characterized in that, S31 specifically includes the following steps: S311. The stochastic optimization problem with multiple time frames is decoupled using Lyapunov theory, where the Lyapunov function is defined as the sum of squares of the data queue and the virtual energy queue. The formula for the Lyapunov function is as follows: in, Indicates the length of the data queue. Indicates the length of the virtual energy queue; S312. Subtract the Lyapunov function value of the current time frame from the Lyapunov function value of the next time frame to obtain the Lyapunov drift. The Lyapunov drift is used to weigh the choice of resource allocation strategy. By controlling the change of the function at each step, the final value of the function can be controlled. The formula for the Lyapunov drift is as follows: S313. Based on the Lyapunov drift function obtained in S312, map the objective function to a suitable penalty function to obtain the Lyapunov drift plus penalty function, as shown in the following formula: in, It is a non-negative control factor, which can be adjusted The size of the data queue is used to obtain a trade-off between minimizing the data queue backlog, maximizing the available energy level, and minimizing the penalty, and finally the asymptotic optimal solution is obtained.

6. The online task offloading and resource allocation joint optimization method for wirelessly powered edge computing according to claim 4, characterized in that, S32 specifically includes the following steps: Based on historical data, channel gain, data queue length, and virtual energy queue length are used as inputs. Through the forward propagation of a convolutional neural network, a relaxed binary offloading decision is output. The convolutional neural network includes an input layer, a hidden layer, and an output layer.

7. The online task offloading and resource allocation joint optimization method for wirelessly powered edge computing according to claim 4, characterized in that, S33 specifically includes the following steps: Based on the relaxation unloading decision output by the convolutional neural network, it is quantized into 2K+1 candidate binary unloading decisions using a noise-preserving quantization method, including K decisions with added noise.

8. The online task offloading and resource allocation joint optimization method for wirelessly powered edge computing according to claim 4, characterized in that, S34 specifically includes the following steps: S341. Based on the given binary unloading decision, decompose the optimization problem into a local computation subproblem and an unloading subproblem; S342. For the local computation subproblem, construct a cubic convex optimization function and obtain the optimal local computation frequency by comparing the stationary points and boundary points. S343. For the unloading subproblem of S341, introduce dual variables to construct a partial Lagrangian function; S344. By leveraging the independence of variables for each wireless device, the problem is broken down into multiple parallel sub-problems of offloading devices. S345. For each wireless device, the offloading energy is expressed as a function of offloading time and offloading rate using the relationship of communication rate, which is transformed into an optimization problem only concerning offloading time and offloading rate. S346. Hierarchical optimization is adopted: the outer layer optimizes the unloading rate, and the inner layer optimizes the unloading time. The inner layer problem is a convex optimization, and the optimal solution is located at the stationary point or the boundary, thus obtaining the optimal unloading time allocation under a given unloading rate. S347. Substitute the obtained optimal unloading time allocation into the objective function to obtain a linear programming problem with respect to unloading rate only; obtain the optimal solution by determining the sign of the decision coefficients; If the total time constraint is not met, the bisection method is used to solve for the optimal dual variable, and then the optimal unloading time allocation and optimal unloading energy under the optimal dual variable are obtained through linear programming. S348. The system enters the next time frame, generates new time-varying channel conditions and random task data, updates the data queue and virtual energy queue according to the energy consumed by the optimal offloading decision and the amount of data processed, and uses the updated state as the input of the convolutional neural network in the next time frame. S349. Repeat S341 to S348, and update the convolutional neural network parameters using historical data until the preset number of iterations is reached.

9. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, it performs the online task offloading and resource allocation joint optimization method for wirelessly powered edge computing as described in any one of claims 1 to 8.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the online task offloading and resource allocation joint optimization method for wireless-powered edge computing as described in any one of claims 1 to 8 through the computer program.