Task offloading and wireless energy transfer method and system

By introducing UAV-assisted task offloading and wireless power transfer methods into IoT devices, and utilizing the deep reinforcement learning TD3 algorithm to optimize task offloading and power transfer, the problem of insufficient battery power in IoT devices is solved, the system stability and efficiency are improved, and maintenance costs are reduced.

CN119545380BActive Publication Date: 2025-10-17GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411373737.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-10-17
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

Existing IoT devices suffer from insufficient battery power, which prevents them from maintaining normal functions, affecting the stability and efficiency of IoT systems. In particular, they struggle to perform tasks effectively in complex natural environments and under time-varying wireless channel conditions.

Method used

By establishing a communication and computing model between UAVs and terrestrial IoT users, the deep reinforcement learning TD3 algorithm is used to optimize task offloading and wireless power transfer. A Markov decision process is constructed to dynamically adjust task offloading and power transfer strategies, thereby optimizing the energy consumption of UAVs and system throughput.

Benefits of technology

It improves the task processing capability and system stability of IoT devices in complex environments, reduces reliance on battery replacement, lowers manual maintenance costs, and enables the system to operate efficiently in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119545380B_ABST
    Figure CN119545380B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of wireless communication, and discloses a task offloading and wireless energy transmission method and system, which establishes communication and calculation models between a UAV and a ground Internet of Things user, including a time-varying channel model for communication, a WPT channel model modeled by using a free space loss model, and a calculation and transmission model of the UAV and the ground Internet of Things user; a basic framework for task offloading and WPT of the ground Internet of Things user by using the UAV in an edge computing scenario assisted by the UAV is designed, system modeling is performed according to the framework, and a mathematical model of the entire system is obtained; an optimization problem is designed, a target function is defined as a weighted sum of system total throughput and UAV energy consumption; the entire framework is modeled as a Markov process, a reasonable state-action space is designed, and a reasonable reward function is designed by using the target function; and the double-delay deep deterministic policy gradient algorithm is used to optimize the task offloading decision and the wireless energy transmission while constraining the system total throughput and the UAV energy consumption.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to but is not limited to the technical field of wireless communication, and particularly relates to a task offloading and wireless energy transmission method and system. BACKGROUND

[0002] With the rapid advancement of 5G technology, Internet of Things users have shown a trend of meshing and extensive distribution in daily life and natural environment. The popularity of Internet of Things technology has greatly changed people's way of life, significantly improving the degree of intelligence, and people can more efficiently manage daily affairs through various intelligent devices. However, due to the design constraints of size and manufacturing cost of these Internet of Things devices, the capacity of the built-in battery is limited, and the computing resources are relatively scarce. As the devices continue to generate and collect massive amounts of data, the problem of insufficient task processing capacity is increasingly prominent. These tasks often have high computational intensity and are extremely sensitive to delay characteristics, although some Internet of Things devices are equipped with energy conversion functions such as solar or wind energy modules, which can theoretically alleviate the problem of insufficient power to some extent, but due to the uncertainty of the natural environment, these devices may encounter various challenges in actual operation. For example, sandstorms or thunderstorms can significantly reduce the efficiency of the energy conversion module, making it difficult for the device to fully utilize external resources for self-charging. In addition, the gradual consumption of battery power during long-term operation, especially in extreme environments, can cause the device to rapidly lose power, seriously affecting its normal function, and when the Internet of Things device is in a power emergency and cannot be replenished in time through the energy conversion module, the need to replace the battery arises. However, due to the large number and wide distribution of Internet of Things devices, battery replacement is not only a time-consuming and labor-intensive task, but also puts a huge workload on the maintenance team, increasing operating costs. Especially in remote or hard-to-reach areas, the difficulty of replacing the battery is further increased, greatly affecting the stability and sustainability of the Internet of Things system.

[0003] Unmanned aerial vehicles (UAVs) are attracting attention due to their flexibility, easy deployment, and other characteristics, and can serve as an aerial edge server to provide offloading and a series of services for ground users. Since the UAV flies at a high altitude, the probability of establishing a line-of-sight (LOS) communication link between the UAV and the user is high, and the UAV is relatively flexible in the air and can provide better service quality in some extreme disaster (flood, earthquake, volcanic eruption) environments. Nan Zhao et al. studied the offloading service provided by the UAV in the air for the ground Internet of Things (IoT) user, and optimized the trajectory, power, and other optimization objectives of the UAV to optimize the optimization objectives. Although these techniques can alleviate the shortage of computing resources of the ground IoT device to some extent, the battery power problem is always a thorny challenge. Although the ground IoT device can effectively complete complex task processing with the support of external computing resources, the limited battery capacity of the device restricts the continuous operation of the device. With the increase of the task amount and the extension of the running time, the gradual consumption of the battery power becomes a key factor affecting the performance and reliability of the device. Once the battery power is insufficient, the normal function of the device may not be maintained, thereby affecting the stability and efficiency of the entire IoT system.

[0004] In view of the above analysis, the existing technical problems to be solved in the prior art are that once the battery power of the existing device is insufficient, the normal function of the device may not be maintained, thereby affecting the stability and efficiency of the entire IoT system. SUMMARY

[0005] In view of the problems in the prior art, the present application provides a task offloading and wireless energy transmission method and system to solve the problem that the ground IoT device cannot effectively and stably perform task processing and calculation due to its insufficient computing resources and limited battery capacity in the case of complex natural environment and time-varying wireless channel. The present application maintains the stability of the system and keeps the system utility relatively optimal while improving the information processing capability of the ground IoT user in complex situations.

[0006] The present application is implemented as follows. A task offloading and wireless energy transmission method comprises the following steps:

[0007] S1, a communication and calculation model between the UAV and the ground IoT user is established according to the variable and complex electromagnetic interference and other situations in the real environment, including a time-varying channel model for communication, a WPT channel model modeled by a free space loss model, and a calculation and transmission model of the UAV and the ground IoT user;

[0008] S2, based on the channel and calculation model proposed in S1, a basic framework of using UAV to offload tasks and WPT for ground IoT users in the edge computing scenario assisted by UAV is designed, and the system is modeled according to the framework to obtain the mathematical model of the whole system. Under this framework, at the beginning of each time slot, the UAV will provide reasonable task offloading and WPT decision for the ground IoT users;

[0009] S3, based on the basic framework of using UAV to offload tasks and WPT for ground IoT users in the edge computing scenario assisted by UAV proposed in S2, the energy consumption of UAV and ground IoT devices and the total system throughput are calculated, and the CPU computing frequency of UAV and ground IoT users, transmission power, task size and channel state are jointly considered. By jointly formulating the optimization problem by using energy consumption and total system throughput, the objective function is defined as the weighted sum of total system throughput and UAV energy consumption, and the objective function is represented as follows:

[0010]

[0011] wherein D (t) represents the total energy consumption of UAV in time interval t, D all (t) represents the total system throughput, w1 and w2 represent two weight coefficients which are positive and negative respectively;

[0012] S4, the whole system framework is modeled as a Markov process, a reasonable state and action space is designed, and a reasonable reward function is designed by using the objective function;

[0013] S5, the TD3 algorithm of deep reinforcement learning is used to optimize the objective function, and the task offloading decision and wireless energy transmission are optimized while the total system throughput and UAV energy consumption are constrained.

[0014] Further, S1 specifically includes:

[0015] S11, a service scenario of one UAV carrying an edge computing server and a WPT radio frequency device and multiple ground users is established. Each ground user carries a receiving energy device with the same frequency as the UAV radio frequency device and a rechargeable battery to store the collected energy, and the UAV and the ground user each carry a single CPU and have a certain memory space to store the computing task;

[0016] S12, the continuous time running mechanism is changed into equal-length time slots, each time slot has a time interval of T, the time slot set is τ={1, 2, 3,..., i,..., T}, and considering the real scene, the randomly arriving computing tasks are subject to exponential distribution, and each task is independently and identically distributed;

[0017] In step S1, the UAV carrying the edge server and wireless radio charging device in the air is considered to provide services for the ground Internet of Things users, the UAV in the air is denoted as A, whose three-dimensional coordinates are A =(x A ,y A ,z A ), the ground Internet of Things users are denoted as t, whose three-dimensional coordinates are t =(x t ,y t ,z t ), so the distance between the UAV in the air and the ground Internet of Things users can be represented as D t , D t =||X A -X t ||.

[0018] Further, in step S1, the communication channel between the UAV and the ground Internet of Things users is established considering the variability of the real situation and the randomness of the channel, aiming to use the UAV in the air to provide services for the ground users. The UAV generally stays in the sky several kilometers high, and the probability of being blocked between the UAV and the ground Internet of Things users is small, so there will be relatively stable communication quality and low delay, and it is considered that the communication link between the UAV and the ground Internet of Things users is mainly dominated by the line of sight (LOS) channel. In order to reduce the communication interference between each Internet of Things user on the ground and the UAV in the air as much as possible, the communication between the UAV and a single Internet of Things user adopts the orthogonal frequency division multiple access (OFDMA) protocol. The channel state is static within a time slot, but dynamic between different time slots. It is assumed that the average channel gain follows the path loss model:

[0019]

[0020] wherein represents the average channel gain between the UAV and the tth Internet of Things user on the ground, A d represents the antenna gain of each Internet of Things user, f c represents the signal carrier frequency, d c represents the path loss exponent, and the channel gain h A,t (t) between each Internet of Things user and the UAV follows the Rayleigh distribution of independent and identical distribution.

[0021] Further, in step S1, the WPT channel is modeled, and it is considered that the link for WPT between the UAV and the ground Internet of Things users is dominated by the LOS (Line of sight) link, and the free space loss model is used to model the transmission link for WPT:

[0022]

[0023] wherein denotes the energy received by the ground IoT device at time t, v t denotes the wireless charging conversion efficiency of the tth ground IoT device, g t denotes the channel gain of the tth ground IoT device, denotes the WPT action of the UAV to the tth ground IoT device, denotes how much energy is to be transmitted for the ground IoT device.

[0024] Further, in step S1, the channel transmission rate R t between the UAV and the tth ground IoT device can be obtained according to the Shannon formula

[0025]

[0026] wherein B denotes the total bandwidth of the channel, φ denotes the power spectral density of the noise signal, P t denotes the signal transmission power of the tth ground IoT device to the UAV, T denotes the number of ground IoT users, and the consumption of the task result unloaded by the MEC server is ignored since the amount of data obtained after calculation by the MEC server is small; according to the channel transmission rate obtained above, the task uploading time of the tth ground IoT user at time t is denoted as , which is jointly determined by the task queue size of the local user, the offloading decision made by the reinforcement learning agent, and the computing task generated by the ground IoT user at time t, so the task uploading energy and the task uploading energy are denoted as:

[0027]

[0028] Further, in steps S2 and S3, a set of UAV-based task offloading and WPT framework is designed for the ground IoT user, specifically, the energy consumption of the UAV and the ground IoT device and the overall throughput of the system are calculated, in which process, the CPU computing frequency of the UAV and the ground IoT user, the transmission power, the task size, and the channel state are comprehensively considered, by jointly analyzing the energy consumption and the overall throughput of the system, an optimization problem is formulated, the objective function of which is defined as the weighted sum of the overall throughput of the system and the energy consumption of the UAV, and the objective function is denoted as:

[0029]

[0030] wherein denotes the total energy consumption of the UAV at time interval t, wherein represents the energy consumption of the UAV at time t, represents the energy consumption of the UAV at time t,

[0031]

[0032] where k represents the computing energy efficiency factor, which is related to the CPU chip structure of the edge server equipped on the UAV, f A represents the CPU computing frequency of the UAV, while D all (t) represents the total throughput of the system at time t, and respectively represent the local computing task amount of the ground IoT user and the UAV at time t, it is worth noting that the tasks processed by the UAV at each time slot are taken from its own memory space for storing tasks, rather than the tasks uploaded by each ground user at each time slot, and the local computing task amount of the UAV and the local IoT user needs to follow the following conditions:

[0033]

[0034] where a t (t) and a A (t) are the task queue storage size of the ground IoT device and the aerial UAV, respectively, represents the energy consumption of the ground IoT user for computing at time t, only when the UAV and the ground IoT device satisfy the above conditions and the following equation is satisfied:

[0035]

[0036] where f A represents the local CPU computing frequency of the UAV, Q t and Q A respectively represent the CPU required for the ground IoT device and the UAV to run 1 bit.

[0037] Further, in step S4, the basic framework of task offloading and WPT for ground IoT users using UAV in the edge computing scenario of UAV assistance proposed in S2 is modeled as an MDP decision process (s, a, r, p, g), where s represents the environment state including time-varying channel gain, randomly arriving task size, and task queue of ground IoT users, a represents the agent action including task offloading and WPT decision, r represents the reward brought by the agent performing one-step action, p is the state transition function, g is the discount factor of reward r, g e (0, 1), and the total utility of the system is only related to the current state and action, and is irrelevant to the actions and states at other times.

[0038] Further, in step S5, according to the MDP process constructed in S4, in order to cope with the continuous action space, the target function is optimized using the deep reinforcement learning TD3 algorithm, which optimizes the task offloading decision and wireless energy transmission while constraining the total throughput of the system and the energy consumption of the UAV. The TD3 algorithm is a deep reinforcement learning algorithm based on policy gradient, which mainly includes two Q networks, a p network and three target networks. In step S5, the main steps of the deep reinforcement learning algorithm TD3 are as follows:

[0039] (1) Initialize various parameters, UAV-assisted edge computing environment, agent network, and construct experience replay buffer buffer;

[0040] (2) Select action, the policy network p generates action and adds exploration noise, executes the action in the current state, and obtains reward r and next state s';

[0041] (3) Update the experience pool, store the current state s, action a, reward r, and next state s' in the experience buffer buffer;

[0042] (4) Sample from the experience buffer, randomly sample a small batch (s, a, r, s') from the experience buffer;

[0043] (5) Calculate the target action, use the target policy network p_target to generate the target action a' of the next state s', and add noise;

[0044] (6) Calculate the target Q value, use the target networks Q1_target and Q2_target to calculate the target Q value, Q'(s'a') = min(Q1'(s',a'), Q'2(s',a')), and use TD-target to approximate the output y'(t) of the Q network y'(t) = r + gQ'(s',a');

[0045] (7) Update the Q network using the TD algorithm, and the loss function TD-error of the two Q networks is as follows: Where y(t) = Q1(s,a) or Q2(s,a), after obtaining the TD-error, the gradient of the TD-error is back propagated, and the Adam optimization algorithm is used to update the neural network parameters of the Q network, when the Q network is updated, the policy gradient algorithm is used to update the pi network in a fixed iteration period, and the following is the derived policy gradient formula for updating the pi network:

[0046]

[0047] Where a is the action predicted by the pi network when the state is S;

[0048] (8) Use soft update to stabilize the parameters of the target network, slowly migrate the parameters of the main network to the target network, so as to avoid overestimation of the evaluation function and improve the stability of the system, and the soft update of the neural network parameters is as follows:

[0049] θ' = δθ + (1-δ)θ',

[0050] Where θ' represents the parameters of the target network, and θ represents the parameters of the main network;

[0051] (9) Repeat steps (2) to (8) to continue iteration until the algorithm converges.

[0052] Another object of the present application is to provide a task offloading and wireless energy transmission system for implementing the task offloading and wireless energy transmission method, comprising:

[0053] The communication and calculation model establishing module establishes the communication and calculation model between the UAV and the ground Internet of Things user according to the complex electromagnetic interference and other conditions in the real environment, including the time-varying channel model for communication, the WPT channel model modeled by the free space loss model, and the calculation and transmission model of the UAV and the ground Internet of Things user.

[0054] The task offloading and WPT basic framework obtaining module jointly designs the basic framework of task offloading and WPT of the ground Internet of Things user by the UAV in the edge calculation scene assisted by the UAV based on the channel and calculation model, models the system according to the framework, and obtains the mathematical model of the whole system.

[0055] The target function module is based on the basic framework of using a UAV to offload tasks and WPT for ground IoT users in a UAV-aided edge computing scenario, calculates the energy consumption of the UAV and the ground IoT device and the total system throughput, jointly considers the CPU computing frequency of the UAV and the ground IoT user, the transmission power, the task size and the channel state, etc., formulates an optimization problem by using the energy consumption and the total system throughput, defines the target function as the weighted sum of the total system throughput and the UAV energy consumption, and the target function is represented as follows:

[0056]

[0057] wherein represents the total energy consumption of the UAV at the time interval t, D all represents the total system throughput, and w1 and w2 represent two weight coefficients that are positive and negative respectively.

[0058] The reward function module models the entire system framework as a Markov process, designs a reasonable state and action space, and designs a reasonable reward function by using the target function.

[0059] The target function optimization module optimizes the target function by using the deep reinforcement learning TD3 algorithm, and optimizes the task offloading decision and wireless energy transmission while constraining the total system throughput and the UAV energy consumption.

[0060] Another object of the present application is to provide a computer device, which comprises a memory and a processor, the memory stores a computer program, and the computer program is executed by the processor to make the processor execute the steps of the task offloading and wireless energy transmission method.

[0061] Another object of the present application is to provide a computer readable storage medium, which stores a computer program, and the computer program is executed by the processor to make the processor execute the steps of the task offloading and wireless energy transmission method.

[0062] Another object of the present application is to provide an information data processing terminal, which comprises the task offloading and wireless energy transmission system.

[0063] In combination with the above technical solutions and the technical problems solved, the technical solution to be protected by the present application has the following advantages and positive effects:

[0064] Firstly, compared with the prior art, the unmanned aerial vehicle-aided task offloading and wireless energy transmission method based on deep reinforcement learning has the following advantages:

[0065] (1) By introducing a deep reinforcement learning algorithm, the system can dynamically adjust the task offloading and energy transmission strategy according to the changes in the environment, improving the adaptability in complex natural environments and time-varying wireless channels.

[0066] (2) Effectively balance the computing resources and energy consumption of ground Internet of Things devices, achieve optimal trade-off between energy consumption and system throughput through multi-objective optimization algorithm, and improve the overall utility of the system.

[0067] (3) Adopting adaptive strategy, the system can maintain stable operation under uncertain conditions, effectively dealing with the challenges brought by environmental and channel changes.

[0068] (4) Through online learning mechanism, the system can continuously update and optimize its strategy model over time and with environmental changes, thus maintaining long-term stability and efficient operation.

[0069] (5) Through reasonable energy management and wireless energy transmission, the dependence of ground Internet of Things devices on battery replacement is reduced, and the cost of manual maintenance is greatly reduced.

[0070] Secondly, the present application proposes a method for UAV-assisted task offloading and wireless energy transmission (WPT), aiming to solve the communication and computing problems between UAV and ground Internet of Things users in complex real environments. By establishing a communication and computing framework including time-varying channel model and WPT channel model, a joint optimization system for task offloading and energy transmission is constructed. In this system, UAV provides task offloading and energy replenishment for ground Internet of Things users through wireless energy transmission (WPT), and the overall throughput and energy consumption of the system are jointly optimized through mathematical model, finally aiming to maximize the system performance.

[0071] Specifically, the present application designs a joint optimization problem, taking the energy consumption of UAV and the total throughput of the system as the objective function, and optimizing it through deep reinforcement learning (TD3 algorithm). The key of task offloading decision and energy transmission is to dynamically adjust according to the channel state, task amount, CPU frequency, etc. in different time slots, so that the system can achieve higher task processing capability while consuming less energy.

[0072] By introducing Markov decision process (MDP) to model the decision-making process of the whole system, reasonable state, action and reward functions are defined to ensure efficient task offloading and energy transmission under different environmental conditions. TD3 algorithm, as the core part of reinforcement learning, approximates and updates the objective function through double Q network and policy network, so that the system maintains stable performance in dynamic environment.

[0073] The present application significantly improves the task offloading efficiency and energy transmission capacity of the UAV in a complex environment. Through the combination of the policy gradient method and the deep reinforcement learning algorithm, the optimization effect that cannot be solved by the prior art is achieved, and the energy consumption of the system is significantly reduced, and technical progress is achieved in task offloading and WPT.

[0074] Thirdly, the present application solves the multiple technical problems of UAV-assisted task offloading and wireless power transmission (WPT) in practical industrial applications. Firstly, in the traditional UAV task offloading and energy transmission system, it is often difficult to guarantee the communication quality and system stability in the face of complex environmental conditions and dynamic channel state. The present application enhances the communication stability between the UAV and the ground Internet of Things users by constructing a time-varying channel model and a WPT channel model, effectively dealing with complex electromagnetic interference and channel randomness problems, and improving the reliability of task transmission.

[0075] Secondly, the existing UAV energy consumption optimization method is relatively limited, and it is difficult to balance the system total throughput and the energy efficiency of the UAV. The present application proposes a joint optimization strategy based on the system total throughput and the energy consumption of the UAV, which adjusts and optimizes the task offloading and wireless energy transmission in real time through deep reinforcement learning (TD3 algorithm), so that the system can realize the maximum task processing capacity under the premise of low energy consumption. This improvement greatly improves the overall performance of the system.

[0076] Thirdly, the traditional task offloading scheme lacks flexibility in dealing with dynamic changes in the environment and tasks, resulting in system response lag and uneven resource allocation problems. The present application models the task offloading and wireless energy transmission framework as a Markov decision process (MDP), dynamically adjusts the resource allocation of the UAV and the ground Internet of Things users, and ensures efficient task offloading and energy transmission under different environmental conditions, significantly improving the response speed and adaptability of the system.

[0077] Finally, the present application realizes significant technical progress in the UAV-assisted edge computing scenario in industrial applications, especially in the fields of Internet of Things, big data processing and intelligent devices. Through the combination of deep reinforcement learning algorithm and optimization problem, the system can realize intelligent decision and resource optimization in complex industrial environment, reduce the energy consumption of UAV, improve the task offloading efficiency and energy transmission capacity of the system, and provide an effective technical solution for large-scale application of UAV in practical industry. BRIEF DESCRIPTION OF DRAWINGS

[0078] Figure 1 is the task offloading and wireless energy transmission method flowchart provided by the embodiment of the present application.

[0079] Figure 2It is a UAV-assisted edge computing scenario provided by the embodiment of the application.

[0080] Figure 3 It is a graph of average reward per round of training of the reinforcement learning algorithms TD3 and DDPG provided by the embodiment of the application.

[0081] Figure 4 It is a system total throughput curve provided by the embodiment of the application.

[0082] Figure 5 It is a UAV energy consumption graph provided by the embodiment of the application.

[0083] Figure 6 It is a task offloading and wireless energy transmission system structure diagram provided by the embodiment of the application. DETAILED DESCRIPTION

[0084] In order to make the purpose, technical scheme and advantages of the application clearer, the application will be further described in detail below in combination with embodiments. It should be understood that the specific embodiments described herein are only used to explain the application and not to limit the application.

[0085] The application provides a method for optimizing task offloading decisions and wireless energy transmission (WPT) of ground Internet of Things (IoT) users in a UAV-assisted edge computing scenario by using DRL. First, based on a real scenario, considering time-varying wireless channels and complex natural environments, a system model of the UAV-assisted edge computing scenario is established. Second, the time and energy consumption of the UAV and the ground IoT users for computation and offloading are calculated by using the established system model. Third, a reasonable optimization function is constructed according to the established system model and system simulation quantities. The entire system framework is modeled as a Markov decision process, and a reasonable state-action space and reward function are formulated. Finally, a deep reinforcement learning algorithm is used to optimize the established optimization function as a whole, while optimizing the task offloading decisions and wireless energy transmission by constraining the system total throughput and UAV energy consumption.

[0086] As shown in Figure 1 and Figure 2 , the application provides a task offloading and wireless energy transmission method based on deep reinforcement learning in a UAV-assisted edge computing scenario, which specifically includes the following steps:

[0087] S1, a communication and computation model between the UAV and the ground IoT users is established according to the variable and complex electromagnetic interference and other conditions in the real environment, including a time-varying channel model for communication, a WPT channel model modeled by using a free space loss model, and a computation and transmission model of the UAV and the ground IoT users, as shown in Figure 2 ;

[0088] S11, establish a service scenario by a UAV carrying an edge computing server and a WPT radio device and multiple ground users. Each ground user carries a receiving energy device and a rechargeable battery to store the collected energy, and the UAV and the ground user each carry a single CPU and have a certain memory space to store computing tasks;

[0089] S12, convert the continuous time running mechanism into equal-length time slots, each time slot has a time interval of T, the set of time slots is τ = {1, 2, 3, …, i, …, T}, and considering the real scene, the randomly arriving computing tasks are subject to an exponential distribution, and each task is independently and identically distributed;

[0090] As shown in Figure 2 , the UAV in the air is represented by A, and its three-dimensional coordinates are A = (x A , y A , z A ), the ground Internet of Things user is represented by t, and its three-dimensional coordinates are t = (x t , y t , z t ), so the distance between the UAV in the air and the Internet of Things user on the ground can be represented by D t , D t = ||X A -X t ||; considering the variability of the real situation and the randomness of the channel, the communication channel between the UAV and the ground Internet of Things user is established, aiming to use the UAV in the air to provide services to the ground user. The UAV generally stays in the sky several kilometers high, and the probability of being blocked between the UAV and the ground Internet of Things user is small, so there will be relatively stable communication quality and low delay, and it is considered that the communication link between the UAV and the ground Internet of Things user is mainly dominated by the line of sight (LOS) channel. In order to reduce the communication interference between each Internet of Things user on the ground and the UAV in the air as much as possible, the communication between the UAV and a single Internet of Things user adopts the orthogonal frequency division multiple access (OFDMA) protocol. The channel state is static within a time slot, but dynamic between different time slots. It is assumed that the average channel gain follows the path loss model:

[0091]

[0092] Where represents the average channel gain between the UAV and the tth Internet of Things user on the ground, A d represents the antenna gain of each Internet of Things user, f c represents the signal carrier frequency, d cdenotes the path loss exponent, while h A,t (t) are Rayleigh distributed independently.

[0093] Then, the WPT channel is modeled, considering that the link between the UAV and the ground IoT users for WPT is dominated by the LOS (Line of Sight) link, and the free space loss model is used to model the transmission link of WPT:

[0094]

[0095] wherein denotes the energy received by the ground IoT device at time t, v t denotes the wireless charging conversion efficiency of the t-th ground IoT device, g t denotes the channel gain of the t-th ground IoT device, denotes the WPT action of the UAV to the t-th ground IoT device, and denotes how much energy is to be transmitted for the ground IoT device.

[0096] After obtaining the distance and channel gain of the UAV and the ground IoT users, the channel transmission rate R t between the UAV and the t-th ground IoT device can be obtained according to the Shannon formula:

[0097]

[0098] wherein B denotes the total bandwidth of the channel, φ denotes the power spectral density of the noise signal, P t denotes the signal transmission power of the t-th ground IoT device to the UAV, and T denotes the number of ground IoT users. Since the amount of data obtained after calculation by the MEC server is very small, the consumption of the task result after offloading calculation by the MEC server is ignored; according to the channel transmission rate obtained above, the task upload time of the t-th ground IoT user at time t is denoted by , which is jointly determined by the task queue size of the local user, the offloading decision made by the reinforcement learning agent, and the computing task generated by the ground IoT user at time t. Therefore, the task upload time and the task upload energy are denoted as:

[0099]

[0100] S2, based on the channel and the calculation model proposed in S1, a basic framework of UAV-assisted edge computing scenario using UAV to offload tasks and WPT for ground IoT users is designed, and the system is modeled according to the framework to obtain the mathematical model of the whole system. Under this framework, at the beginning of each time slot, the UAV will provide reasonable task offloading and WPT decision for the ground IoT users;

[0101] S3, based on the basic framework of UAV-assisted edge computing scenario using UAV to offload tasks and WPT for ground IoT users proposed in S2, the energy consumption of UAV and ground IoT devices and the total throughput of the system are calculated, and the CPU computing frequency, transmission power, task size and channel state of UAV and ground IoT users are jointly considered. By jointly formulating the optimization problem using energy consumption and total system throughput, the objective function is defined as the weighted sum of total system throughput and UAV energy consumption.

[0102] A set of UAV-based task offloading and WPT framework is designed. Specifically, the energy consumption of UAV and ground IoT devices and the overall throughput of the system are calculated. In this process, the CPU computing frequency, transmission power, task size and channel state of UAV and ground IoT users are comprehensively considered. By jointly analyzing energy consumption and total system throughput, an optimization problem is formulated, and the objective function is defined as the weighted sum of total system throughput and UAV energy consumption. The objective function is expressed as follows:

[0103]

[0104] Wherein represents the total energy consumption of UAV in time interval t, Wherein represents the energy consumption of UAV at time t, represents:

[0105]

[0106] Wherein k represents the computing energy efficiency factor, which is related to the CPU chip structure of the edge server equipped with UAV, A represents the CPU computing frequency of UAV, and D all (t) represents the total throughput of the system at time t, and respectively represent the local computing task amount of ground IoT users and UAV at time t, it is worth noting that the task processed by UAV in each time slot is taken from its own memory space for storing tasks, rather than the task uploaded by each ground user at each time slot, and the local computing task amount of UAV and local IoT users needs to follow the following conditions:

[0107]

[0108] wherein α t (t) and α A (t) are the task queue storage size of ground IoT devices and aerial UAV respectively, represents the energy consumed by ground IoT users for computing at time t, only when UAV and ground IoT devices meet the above conditions and the following equation is satisfied:

[0109]

[0110] wherein f A represents the CPU computing frequency of UAV locally, Θ t and Θ A respectively represent the CPU required for ground IoT devices and UAV to run 1 bit;

[0111] S4, model the entire system framework as a Markov process, design reasonable state and action space, and design a reasonable reward function by using the objective function;

[0112] Model the UAV-assisted task offloading and WPT framework as an MDP decision process (s, a, r, p, γ), wherein s represents the environment state including time-varying channel gain, randomly arrived task size and task queue of ground IoT users, a represents the agent action including task offloading and WPT decision, r represents the reward brought by the agent executing one step of action, p is the state transition function, γ is the discount factor of reward r, γ ∈ (0, 1), and the total utility of the system is only related to the current state and action, and is irrelevant to the action and state at other times.

[0113] S5, use the deep reinforcement learning TD3 algorithm to optimize the objective function, and optimize the task offloading decision and wireless energy transmission while constraining the total throughput of the system and the energy consumption of UAV;

[0114] According to the MDP process constructed in S4, in order to cope with the continuous action space, the target function is optimized by using the deep reinforcement learning TD3 algorithm, which optimizes the task offloading decision and wireless energy transmission while constraining the total throughput of the system and the energy consumption of the UAV. The TD3 algorithm is a deep reinforcement learning algorithm based on policy gradient, which mainly includes two Q networks, a pi network and three target networks. In step S5, the main steps of the deep reinforcement learning algorithm TD3 are as follows:

[0115] (1) Initialize various parameters, UAV-assisted edge computing environment, agent network, and construct experience replay buffer buffer;

[0116] (2) Select action, the policy network pi generates action and adds exploration noise, executes the action in the current state, and obtains reward r and next state s';

[0117] (3) Update the experience pool, store the current state s, action a, reward r, and next state s' in the experience buffer buffer;

[0118] (4) Sample from the experience buffer, randomly sample a small batch (s, a, r, s') from the experience buffer;

[0119] (5) Calculate the target action, use the target policy network pi_target to generate the target action a' of the next state s', and add noise;

[0120] (6) Calculate the target Q value, use the target networks Q1_target and Q2_target to calculate the target Q value Q'(s'a')=min(Q1'(s',a'),Q'2(s',a')), and use TD-target to approximate the output of the Q network y'(t)=r+gammaQ'(s',a');

[0121] (7) Update the Q network using the TD algorithm, the loss function TD-error of the two Q networks is as follows: Where y(t)=Q1(s,a)or Q2(s,a), after obtaining the TD-error, the gradient of the TD-error is back propagated, and the Adam optimization algorithm is used to update the neural network parameters of the Q network. When the Q network is updated, the policy gradient algorithm is used to update the pi network in a fixed iteration period. The following is the derived policy gradient formula for updating the pi network:

[0122]

[0123] Where a is the action predicted by the pi network in state S;

[0124] (8) Use soft update to stabilize the parameters of the target network, slowly migrate the parameters of the main network to the target network, so as to avoid overestimation of the evaluation function and improve the stability of the system. The soft update of the neural network parameters is as follows:

[0125] θ' = δθ + (1-δ)θ',

[0126] Where θ' represents the parameters of the target network, and θ represents the parameters of the main network;

[0127] (9) Repeat steps (2) to (8) for continuous iteration until the algorithm converges.

[0128] As shown in Figure 3 , it represents the whole learning process of TD3 and DDPG algorithm. The actor and critic networks of Agent in both algorithms use sigmoid activation function. It can be seen that in the initial stage of model training, the average reward of Agent of both algorithms gradually increases. This is because Agent constantly improves its neural network in the interaction process with the environment, so that the action output of the pi network is consistent with the direction of increasing reward. After comparison, the overall convergence reward of the TD3 algorithm is greater than that of the DDPG algorithm, and a higher average reward can be obtained after the model converges. This is because the TD3 algorithm reduces the overestimation bias of the environment Q value by introducing double Q learning, thereby improving the stability and performance of the training.

[0129] As shown in Figure 4 , it represents the system total throughput curve obtained by interacting with the test environment for 5000 steps using the trained TD3 algorithm and DDPG algorithm agent. The number of users is 10, and the task arrival rate is 2.4 Mbps. As can be seen, when the task arrival rate is 2.4 Mbps, both algorithms can maintain system stability, keeping the task queue in a reasonable state. However, it is worth noting that when the time step is between 0-1000 steps, the system throughput caused by the TD3 algorithm is greater than that of the other two comparison algorithms, resulting in a smaller task storage space pressure in the system using the TD3 algorithm agent than the DDPG algorithm.

[0130] As shown in Figure 5 , it represents the UAV energy consumption graph obtained by interacting with the test environment for 5000 steps using the trained TD3 algorithm and DDPG algorithm agent. As can be seen, the average energy consumption of the system using the TD3 algorithm is slightly less than that of the DDPG algorithm, which indicates that the TD3 algorithm can more effectively balance system throughput and energy consumption when dealing with task offloading and wireless energy transmission problems, and achieve better energy consumption management in complex environments. In addition, due to its double Critic network structure, the TD3 algorithm can effectively reduce the deviation of Q value estimation, thereby further improving energy efficiency.

[0131] As Figure 6 shown, the task offloading and wireless energy transmission system provided by the embodiment of the application comprises:

[0132] The communication and calculation model establishing module establishes a communication and calculation model between the UAV and the ground Internet of Things user according to the variable and complex electromagnetic interference and the like in the real environment, including a time-varying channel model for communication, a WPT channel model modeled by using a free space loss model, and a calculation and transmission model of the UAV and the ground Internet of Things user.

[0133] The task offloading and WPT basic framework obtaining module jointly designs a basic framework for task offloading and WPT of the ground Internet of Things user by using the UAV in the UAV-assisted edge computing scenario based on the channel and calculation model, and obtains a mathematical model of the entire system by modeling the system according to the framework. Under the framework, the UAV will provide reasonable task offloading and WPT decisions for the ground Internet of Things user at the start of each time slot.

[0134] The objective function module calculates the energy consumption of the UAV and the ground Internet of Things device and the total system throughput based on the basic framework for task offloading and WPT of the ground Internet of Things user by using the UAV in the UAV-assisted edge computing scenario, jointly considers the CPU calculation frequency of the UAV and the ground Internet of Things user, the transmission power, the task size, and the channel state, and defines the objective function as a weighted sum of the total system throughput and the UAV energy consumption by jointly formulating an optimization problem by using the energy consumption and the total system throughput, and the objective function is expressed as follows:

[0135]

[0136] wherein D(t) represents the total energy consumption of the UAV at the time interval t, D all (t) represents the total system throughput, and w1 and w2 represent two weight coefficients that are positive and negative, respectively.

[0137] The reward function module models the entire system framework into a Markov process, designs a reasonable state and action space, and designs a reasonable reward function by using the objective function.

[0138] The objective function optimization module optimizes the objective function by using a deep reinforcement learning TD3 algorithm, and optimizes the task offloading decisions and the wireless energy transmission while constraining the total system throughput and the UAV energy consumption.

[0139] The application embodiment of the present application provides a computer device, which comprises a memory and a processor, the memory stores a computer program, and the computer program is executed by the processor to make the processor execute the steps of the task offloading and wireless energy transmission method.

[0140] The application embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by the processor to make the processor execute the steps of the task offloading and wireless energy transmission method.

[0141] The application embodiment of the present application provides an information data processing terminal, which comprises a task offloading and wireless energy transmission system.

[0142] The following are two embodiments of the present application in industrial application:

[0143] 1. UAV task offloading and wireless energy transmission application in the field of intelligent agriculture

[0144] In intelligent agriculture, unmanned aerial vehicles (UAVs) are commonly used for large-scale farmland monitoring, pest detection, and soil composition analysis. However, in the face of vast farmland, the battery power of ground sensors and monitoring devices is easily depleted, and the amount of task data is large, making computation and transmission efficiency a bottleneck. Through the method of the present application, UAVs can not only offload tasks to ground sensors and process big data in real time, but also charge ground devices through wireless energy transmission (WPT) to extend their working life. This application significantly improves the continuous working ability of the agricultural monitoring system, reduces the cost of frequent battery replacement, and improves the efficiency of farmland monitoring and data processing.

[0145] 2. Internet of Things device maintenance and energy transmission application in smart cities

[0146] In the Internet of Things (IoT) application of smart cities, the operation of environmental sensors, traffic monitoring devices, and public facilities in the city depends on power supply and network connection. However, due to the wide distribution and scattered geographical location of devices, energy supply and communication problems often become bottlenecks. The present application introduces a UAV-assisted task offloading and wireless energy transmission system, which can transmit energy to IoT devices distributed in different areas of the city during cruising, while reducing local computing pressure through task offloading and uploading large amounts of data to the cloud or edge server for processing. This application improves the maintenance efficiency of smart city devices, reduces maintenance costs, and ensures the continuity and real-time performance of urban environmental data.

[0147] It should be noted that embodiments of the present application can be realized by hardware, software, or a combination of software and hardware. The hardware portion can be realized by a special logic; the software portion can be stored in a memory and executed by a proper instruction execution system, such as a microprocessor or a specially designed hardware. A person of ordinary skill in the art can understand that the above-mentioned apparatus and method can be realized by computer executable instructions and / or included in processor control codes, such as a carrier medium, such as a magnetic disk, CD or DVD-ROM, a programmable memory, such as a read-only memory (firmware), or a data carrier, such as an optical or electronic signal carrier. The apparatus of the present application and its modules can be realized by a hardware circuit, such as a very large scale integrated circuit or a gate array, a semiconductor, such as a logic chip, a transistor, or a programmable hardware device, such as a field programmable gate array, a programmable logic device, or the like, by software executed by various types of processors, or by a combination of the above-mentioned hardware circuit and software, such as firmware.

[0148] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any modification, equivalent replacement, and improvement within the technical range disclosed by the present application, and within the spirit and principle of the present application, should be included in the protection scope of the present application.

Claims

1. A task offloading and wireless energy transmission method, characterized in that: The steps include: S1. Based on the changing environmental conditions and complex electromagnetic interference in real-world environments, a communication and computation model between UAVs and terrestrial IoT users was established, including a time-varying channel model for communication, a WPT channel model using a free-space loss model, and a computation and transmission model between UAVs and terrestrial IoT users. S2. Based on the channel and computing models proposed in S1, a basic framework for using UAVs to offload tasks and perform WPT for terrestrial IoT users in a UAV-assisted edge computing scenario is jointly designed. Based on this framework, system modeling is performed to obtain a mathematical model of the entire system. Within this framework, the UAV at the start of each time slot will provide reasonable task offloading and WPT decisions for terrestrial IoT users. S3. Based on the basic framework of using UAV to offload tasks and perform WPT for terrestrial IoT users in the UAV-assisted edge computing scenario proposed in S2, the energy consumption of UAV and terrestrial IoT devices and the total system throughput are calculated. Taking into account the CPU computing frequency, transmission power, task size, and channel status of UAV and terrestrial IoT users, the energy consumption and total system throughput are jointly formulated into an optimization problem. The objective function is defined as the weighted sum of the total system throughput and the UAV energy consumption. The objective function is expressed as follows: in represents the total energy consumption of the UAV in time interval t, D all (t) represents the total system throughput, w1 and w2 represent two weight coefficients, one positive and one negative; S4. Model the entire system framework as a Markov process, design a reasonable state and action space, and use the objective function to design a reasonable reward function; S5. Use the deep reinforcement learning TD3 algorithm to optimize the objective function, while constraining the total system throughput and UAV energy consumption to optimize task offloading decisions and wireless energy transmission.

2. The task offloading and wireless energy transmission method according to claim 1, wherein: S1 specifically includes: S11. Establish a service scenario consisting of a UAV equipped with an edge computing server and a WPT radio frequency device and multiple ground users, where each ground user is equipped with an energy receiving device that uses the same frequency as the UAV radio frequency device and a rechargeable battery to store the collected energy. The UAV and ground users are both equipped with a single CPU and a certain amount of memory space to store computing tasks; S12. Convert the continuous time operation mechanism into time slots of equal length, with the time interval of each time slot being Τ, and the time slot set being τ = {1, 2, 3, ..., i, ..., T}. Considering the real-world scenario, the randomly arriving computing tasks are made to obey an exponential distribution, and each task is independent and identically distributed. In S1, it is considered that the UAV in the air equipped with edge servers and wireless radio frequency charging devices provides services to the IoT users on the ground. The UAV in the air is represented by A and its three-dimensional coordinates are € A =(x A, y A, z A ), the ground IoT user is represented by t and its three-dimensional coordinates are € t =(x t, y t, z t ), so the distance between the UAV in the air and the IoT user on the ground is D t Indicates that D t =||X A -X t ||.

3. The task offloading and wireless energy transmission method according to claim 1, wherein: In step S1, a communication channel between the UAV and the ground IoT user is established, taking into account the variability of the real situation and the randomness of the channel. The purpose is to use the UAV in the air to provide services to the ground users. The UAV generally stays at an altitude of several kilometers, and the probability of being blocked by the ground IoT user is relatively low, so there will be relatively stable communication quality and low latency. It is therefore believed that the communication link between the UAV and the ground IoT user is mainly dominated by the line-of-sight (LOS) channel. In order to minimize the interference between each IoT user on the ground and the UAV in the air, the communication between the UAV and a single IoT user adopts the orthogonal frequency division multiple access (OFDMA) protocol. The channel state is static within a time slot, but dynamic within different time slots. It is assumed that the average channel gain follows the path loss model: in represents the average channel gain between the UAV and the t-th IoT user on the ground, A d represents the antenna gain of each IoT user, f c Indicates the signal carrier frequency, d c represents the path loss index, and the channel gain h between each IoT user and UAV A,t (t) obeys independent and identically distributed Rayleigh distribution.

4. The task offloading and wireless energy transmission method according to claim 1, wherein: In step S1, the WPT channel is modeled. It is assumed that the link between the UAV and the terrestrial IoT user for WPT is dominated by the LOS (Line of Sight) link. The free space loss model is used to model the WPT transmission link: in represents the energy received by the ground IoT device at time t, v t ∈(0,1) represents the wireless charging conversion efficiency of the tth ground IoT device, g t It represents the channel gain of the t-th ground IoT device. It represents the WPT action of the UAV to the tth terrestrial IoT device and indicates how much energy to transmit to the terrestrial IoT device.

5. The task offloading and wireless energy transmission method according to claim 1, wherein: In step S1, the channel transmission rate R between the UAV and the tth ground IoT device is obtained according to the Shannon formula t : Where B represents the total channel bandwidth, φ represents the power spectral density of the noise signal, and P t represents the signal transmission power of the t-th ground IoT device to the UAV, T represents the number of ground IoT users. Since the amount of data obtained after the calculation by the MEC server is very small, the consumption of the task results after the MEC server offloads the calculation is ignored. According to the channel transmission rate obtained above, represents the amount of tasks uploaded by the tth ground IoT user at time t. This amount of tasks is determined by the size of the local user's task queue, the offloading decision made by the reinforcement learning agent, and the computing tasks generated by the ground IoT user at time t. Therefore, the task upload time T t up (t) and task upload energy Expressed as:

6. The task offloading and wireless energy transmission method according to claim 1, wherein: In steps S2 and S3, a UAV-based task offloading and WPT framework was designed for terrestrial IoT users. Specifically, the energy consumption of the UAV and terrestrial IoT devices, as well as the overall system throughput, were calculated. During this process, the CPU computing frequency, transmission power, task size, and channel state of the UAV and terrestrial IoT users were comprehensively considered. By jointly analyzing energy consumption and total system throughput, an optimization problem was formulated. The objective function was defined as the weighted sum of total system throughput and UAV energy consumption. The objective function is expressed as follows: in represents the total energy consumption of the UAV in time interval t, in Indicates that the UAV calculates energy consumption at time t, Expressed as: Where k represents the computational energy efficiency factor, which is related to the CPU chip structure of the edge server equipped by the UAV, and f A Indicates the CPU calculation frequency of UAV, and D all (t) represents the total throughput of the system at time t, and They represent the local computing task load of the ground IoT user and the UAV at time t, respectively. It is worth noting that the tasks processed by the UAV in each time slot are taken from its own memory space used to store tasks, rather than tasks uploaded by each ground user in each time slot. In addition, the local computing task load of the UAV and the local IoT user must comply with the following conditions: where α t (t) and α A (t) are the task queue storage sizes of ground IoT devices and aerial UAVs, represents the energy consumed by the ground IoT user for computing at time t, Only when UAV and ground IoT devices meet the above conditions and Only then the following equation is satisfied: where f A Indicates the CPU calculation frequency of the UAV local, Θ t and Θ A They represent the number of CPUs required for ground IoT devices and UAVs to run 1 bit respectively.

7. The task offloading and wireless energy transmission method according to claim 1, wherein: In step S4, according to the basic framework of using UAV to perform task offloading and WPT for terrestrial IoT users in the UAV-assisted edge computing scenario proposed in S2, the UAV-assisted task offloading and WPT framework is modeled as an MDP decision process (s, a, r, p, γ), where s represents the environmental state including the time-varying channel gain, the size of randomly arriving tasks, and the task queue of terrestrial IoT users, a represents the agent action including task offloading and WPT decision, r represents the reward brought by the agent performing one step of action, p is the state transition function, γ is the discount factor of the reward r, γ∈(0,1), and the total utility of the system is only related to the current state and action, but not to the actions and states at other times.

8. The task offloading and wireless energy transmission method according to claim 1, wherein: In step S5, based on the MDP process constructed in S4, in order to cope with the continuous action space, the deep reinforcement learning TD3 algorithm is used to optimize the objective function. While constraining the total system throughput and UAV energy consumption, the task offloading decision and wireless energy transmission are optimized. The TD3 algorithm is a deep reinforcement learning algorithm based on policy gradients. It mainly includes two Q networks, one π network and three target networks. In step S5, the main steps of the deep reinforcement learning algorithm TD3 are as follows: (1) Initialize various parameters, UAV-assisted edge computing environment, agent network, and build the experience replay pool buffer; (2) Select an action. The policy network π generates an action and adds exploration noise. The action is executed in the current state to obtain the reward r and the next state s'. (3) Update the experience pool and store the current state s, action a, reward r, and next state s' into the experience buffer; (4) Sampling from the experience buffer: randomly sample a small batch (s, a, r, s') from the experience buffer; (5) Calculate the target action, use the target policy network π_target to generate the target action a' for the next state s', and add noise; (6) Calculate the target Q value. Use the target network Q1_target and Q2_target to calculate the target Q value, Q'(s'a') = min(Q'1(s',a'),Q'2(s',a')). Use TD-target to approximate the output of the Q network y'(t) = r + γQ'(s',a'); (7) Using the TD algorithm to update the Q network, the loss function TD-error of the two Q networks is expressed as follows: Where y(t) = Q1(s,a) or Q2(s,a). After obtaining the TD-error, the gradient of the TD-error is back-propagated, and the Adam optimization algorithm is used to update the neural network parameters of the Q network. When the Q network is updated, the policy gradient algorithm is used to update the π network in a fixed iteration cycle. The following is the derived policy gradient formula for updating the π network: Where a is the action predicted by the π network when the state is S; (8) Use soft update to stabilize the parameters of the target network and slowly migrate the parameters of the main network to the target network, thereby avoiding overestimation of the valuation function and improving the stability of the system. The soft update of the neural network parameters is as follows: θ'=δθ+(1-δ)θ', Where θ' represents the parameters of the target network, and θ represents the parameters of the main network; (9) Repeat steps (2) to (8) and continue iterating until the algorithm converges.

9. A task offloading and wireless energy transmission system that implements the task offloading and wireless energy transmission method according to any one of claims 1 to 8, characterized in that: include: The communication and computing model building module establishes a communication and computing model between the UAV and terrestrial IoT users based on the changing environmental conditions and complex electromagnetic interference in real-world environments. This includes a time-varying channel model for communication, a WPT channel model using a free-space loss model, and a computing and transmission model for the UAV and terrestrial IoT users. The module develops a basic framework for task offloading and WPT. Based on channel and computational models, it jointly designs a basic framework for using UAVs to perform task offloading and WPT for terrestrial IoT users in UAV-assisted edge computing scenarios. Based on this framework, system modeling is performed to obtain a mathematical model of the entire system. Within this framework, the UAV at the start of each time slot will provide reasonable task offloading and WPT decisions for terrestrial IoT users. The objective function module, based on the basic framework of using UAVs to offload tasks and perform WPT for terrestrial IoT users in UAV-assisted edge computing scenarios, calculates the energy consumption of UAVs and terrestrial IoT devices and the total system throughput. Taking into account the CPU computing frequency, transmission power, task size, and channel status of UAVs and terrestrial IoT users, the objective function is defined as the weighted sum of the total system throughput and the UAV energy consumption by jointly formulating the optimization problem using energy consumption and total system throughput. The objective function is expressed as follows: in represents the total energy consumption of the UAV in time interval t, D all (t) represents the total system throughput, w1 and w2 represent two weight coefficients, one positive and one negative; The reward function module models the entire system framework as a Markov process, designs a reasonable state and action space, and uses the objective function to design a reasonable reward function; The objective function optimization module uses the deep reinforcement learning TD3 algorithm to optimize the objective function, while constraining the total system throughput and UAV energy consumption to optimize task offloading decisions and wireless energy transmission.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the task offloading and wireless energy transmission method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • UAV (Unmanned Aerial Vehicle) energy consumption optimization method in air-ground cooperative mobile edge computing system

    CN117939537A

  • Unmanned plane assisted unmanned ship task unloading method based on deep reinforcement learning

    CN118574156A