A dynamic task offloading method of a smart agricultural unmanned aerial vehicle
By employing a dynamic task offloading method that combines edge computing drones and cloud collaboration, along with a dynamic dual-delay deep deterministic policy gradient algorithm, the user scheduling strategy and flight parameters of drones are adjusted in real time. This solves the problem of balancing latency and energy consumption in smart agriculture and improves the adaptability and convergence of the model.
Patent Information
- Application Number
- CN202510215618.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2045-02-26
AI Technical Summary
In smart agriculture scenarios, existing technologies struggle to balance latency and energy consumption, and reinforcement learning methods cannot dynamically adjust to adapt to changes in system state, leading to a decline in model performance.
A dynamic task offloading method is adopted, which combines edge computing drones and cloud collaboration with a dynamic dual-latency deep deterministic strategy gradient algorithm and adaptive exploration noise to adjust user scheduling strategy, offloading ratio, flight speed and launch power in real time to optimize latency and energy consumption.
It improves the model's adaptability and convergence, reduces system costs, solves the problem of model performance degradation caused by environmental changes, and achieves a dynamic balance between latency and energy consumption.
Smart Images

Figure CN120075899B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of smart agriculture and edge computing technology, and particularly relates to a dynamic task unloading method for smart agriculture drones. Background Technology
[0002] Edge computing is a new computing architecture that moves communication, computing, and storage services from the network core to the edge of the wireless network. Among them, drone-assisted edge computing is an emerging concept that has appeared in recent years with the development of drone technology and edge computing technology. It combines the flexibility of drones with the efficient data processing capabilities of edge computing, and can adapt to many application scenarios that traditional centralized computing systems cannot meet.
[0003] In smart agriculture scenarios, traditional agricultural drones can interact with the agricultural environment, but they struggle with computationally demanding tasks, such as pest image analysis. Therefore, deploying edge computing models on top of traditional agricultural drones is an effective solution. However, current technologies typically only offload tasks to the edge, neglecting the auxiliary role of the cloud. Thus, deploying a cloud-edge collaborative model and implementing appropriate task offloading are pressing technical challenges in smart agriculture scenarios.
[0004] Timely analysis of collected image information and improved endurance of edge computing models are crucial for smart agriculture, thus considering latency and energy consumption is essential. However, there is a trade-off between decision latency and energy consumption. Local execution of computationally intensive tasks struggles to meet latency requirements; while transmitting data to the edge or cloud for processing can save local computing resources, the data transmission process increases latency and energy consumption. Therefore, achieving a balance between these two factors is key to designing an efficient system.
[0005] Furthermore, existing technologies typically only consider one of the metrics, latency or energy consumption, and neglect the impact of real-time changes in system state on energy consumption and latency preferences. Therefore, for smart agriculture scenarios, how to reduce latency and energy consumption, as well as analyze the impact of system state changes on energy consumption and latency preferences, are urgent problems to be solved.
[0006] Because the model under consideration comprises a three-layer structure consisting of a data collection drone, an edge computing drone, and the cloud, the resulting optimization problem is quite complex and falls under the category of Mixed-Integer Nonlinear Programming (MINLP). The state and action spaces are both high-dimensional and continuous. Traditional optimization methods in existing technologies often encounter problems of excessive computational resource requirements and long computation times when dealing with large-scale, high-dimensional MINLP problems. Reinforcement learning, on the other hand, has several significant advantages. However, current reinforcement learning applications often use static parameters, which cannot be adjusted according to system changes, potentially leading to performance degradation when the environment changes. Therefore, there is an urgent need to research a reinforcement learning method that dynamically adjusts parameters. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this application proposes a dynamic task unloading method for smart agriculture drones, which fully considers the real-time changes in system status and helps to cope with complex environmental changes in smart agriculture.
[0008] Firstly, this application proposes a dynamic task unloading method for smart agricultural drones, including:
[0009] The computational task involves using multiple data drones to collect images of pests and sending the status information of the data drones in the current time slot to the edge computing drone.
[0010] A pre-trained task offloading model deployed in an edge computing drone is used. The state information of the data drone in the current time slot and the standardized weighting coefficients are input into the trained task offloading model to obtain the dynamically optimal user scheduling strategy, the offloading ratio of the edge computing drone, the offloading ratio of the data drone, the flight speed of the edge computing drone, the flight angle of the edge computing drone, and the transmission power of the edge computing drone. According to the dynamically optimal user scheduling strategy, a communication link is established with one of the multiple data drones, and the offloading ratio of the data drone is sent to the data drone with the established communication link.
[0011] The edge computing drone adjusts its position based on its flight speed and flight angle.
[0012] The data drone that has established a communication link will unload the computational task of the pest image from the data drone and transfer it to the edge computing drone that has been repositioned.
[0013] After the edge computing drone has been repositioned, it offloads the computational task of the pest image to the cloud server based on the offload ratio of the edge computing drone.
[0014] The standardized weighting coefficients are obtained through the following process:
[0015] Based on the remaining total workload of the data drone and the remaining total system time under the current state, the initial delay coefficient under the current state is calculated.
[0016] The initial energy consumption coefficient for the current state is calculated based on the remaining available power of the edge computing drone and the total power of the edge computing drone in the current state.
[0017] The initial time delay coefficient and initial energy consumption coefficient are processed to obtain the standardized weighted coefficients under the current state.
[0018] The standardized weighted coefficients under the current state include: the standardized time delay coefficient and the standardized energy consumption coefficient, calculated as follows:
[0019]
[0020] Where α is the standardized time delay coefficient, β is the standardized energy consumption coefficient, α1 is the normalized time delay coefficient, and β1 is the normalized energy consumption coefficient.
[0021] The pre-trained task unloading model is pre-trained on a cloud server, and the training process includes:
[0022] Based on the communication model, computation model, energy harvesting model, latency model, and energy consumption model, a task offloading model is constructed.
[0023] Based on the task offloading model, a Markov decision process is constructed to obtain a task offloading model based on reinforcement learning.
[0024] A dynamic dual-delay deep deterministic policy gradient algorithm is used to solve the task offloading model based on reinforcement learning and obtain the optimal dynamic offloading policy.
[0025] The trained reinforcement learning-based task offloading model is deployed to the edge computing drone.
[0026] The task offloading model, constructed based on the communication model, computation model, energy harvesting model, latency model, and energy consumption model, is calculated as follows:
[0027]
[0028] C3:q(i)∈{(x(i),y(i)),H2|x(i)∈[0,L1],y(i)∈[0,L2]}
[0029] C4:p k (i)∈{(x k (i),y k (i)),H1|xk (i)∈[0,L1],y k (i)∈[0,L2]}
[0030]
[0031] C7: v(i)∈[0,v max ],φ(i)∈[0,2π]
[0032]
[0033] C9:P min ≤P up,c (i)≤P max
[0034] Where, τ k (i) represents the user scheduling variable for the i-th time slot in the communication model, where I is the total number of time slots, K is the total number of data drones, and T is the total number of data drones. total,k (i) represents the total latency of processing all computational tasks in the i-th time slot in the latency model, E total,k (i) represents the total energy consumption for processing all computational tasks in the energy consumption model, R. k,e (i) represents the offloading ratio of computational tasks from the data drone to the edge computing drone in the computational model, R. e,c (i) represents the offloading ratio of computing tasks from the edge computing drone to the cloud server in the computational model, where α is the standardized latency coefficient, β is the standardized energy consumption coefficient, μ1 is the first dimensionless factor, μ2 is the second dimensionless factor, and E b P represents the total power consumption of the edge computing drone. min P is the minimum transmit power of the edge computing drone. max Let be the maximum transmit power of the edge computing drone, q(i) be the position coordinate of the edge computing drone in the i-th time slot, x(i) be the x-axis value of the position coordinate of the edge computing drone in the i-th time slot, y(i) be the y-axis value of the position coordinate of the edge computing drone in the i-th time slot, and p be the maximum transmit power of the edge computing drone. k (i) represents the position coordinates of the k-th data UAV in the i-th time slot, x k (i) represents the x-axis value of the position coordinates of the k-th data UAV in the i-th time slot, and y k (i) represents the y-axis value of the position coordinates of the k-th data UAV in the i-th time slot, E hover2 (i) represents the energy consumption of the edge computing drone hovering in the i-th time slot, E fly (i) represents the energy consumption of the UAV during flight at the edge of the i-th time slot in the energy consumption model, E tr,E-UAV,e,c(i) represents the energy consumption of the edge computing drone transmitting computing tasks to the cloud server in the i-th time slot of the energy consumption model, E com,E-UAV (i) represents the computational energy consumption of the edge computing drone in the i-th time slot of the energy consumption model, φ(i) represents the flight angle of the edge computing drone in the i-th time slot, T represents the length of the computational task unloading cycle, and P represents the computational energy consumption of the edge computing drone in the i-th time slot. up,c (i) represents the transmit power of the edge computing drone in the i-th time slot, and L1 and L2 represent the flight range of the edge computing drone and the data drone within a rectangular area with side lengths L1 and L2, respectively. k (i) represents the amount of data collected by the data-collecting UAV in the i-th time slot, D represents the amount of data that the system needs to process throughout the entire cycle, and E represents the amount of data collected by the data-collecting UAV in the i-th time slot. charge (i) represents the energy collected in the energy harvesting model of the i-th time slot, v(i) represents the flight speed of the edge computing drone, C1 indicates that within a time slot, the edge computing drone can only establish a communication link with one data drone, C2 indicates that both offloading methods adopt a partial offloading strategy, and the two offloading methods include: the data drone offloading the computing task to the edge computing drone, and the edge computing drone offloading the computing task to the cloud server, C3 represents the flight range of the edge computing drone, C4 represents the flight range of the data drone, C5 indicates that the energy consumption of the edge computing drone must be less than the sum of the available electricity and the collected solar energy, C6 indicates that the transmission and computing delay in each time slot must be less than the length of the time slot, C7 represents the range of the flight angle and flight speed of the edge computing drone, C8 indicates that the data drone, the edge computing drone, and the cloud server must complete the specified tasks, and C9 indicates that the transmission power of the edge computing drone must meet the specified range.
[0035] Based on the task offloading model, a Markov decision process is constructed to obtain a reinforcement learning-based task offloading model, including:
[0036] The remaining available power of the edge computing drone, the location coordinates of the edge computing drone, the location coordinates of the data drone, the remaining total workload of the data drone, and the workload collected by the data collection drone that has established a communication link with the edge computing drone in the current time slot are used as the state space of the Markov decision process.
[0037] The optimization variables are used as the action space of the Markov decision process. The optimization variables include: user scheduling strategy, the offloading ratio of data drones to edge computing drones, the offloading ratio of edge computing drones to cloud servers, the flight speed of edge computing drones, the flight angle of edge computing drones, and the transmission power of edge computing drones.
[0038] The probability that the system will transition to the next state after taking an action in the current state is taken as the state transition probability of the Markov decision process.
[0039] The immediate reward obtained after taking an action in the current state, i.e., the negative value of the optimization objective in the current time slot, is used as the reward function of the Markov decision process.
[0040] A Markov decision process is constructed using the state space, action space, state transition probabilities, and reward function to obtain a task offloading model based on reinforcement learning.
[0041] The method employs a dynamic dual-delay deep deterministic policy gradient algorithm to solve the task offloading model based on reinforcement learning, including:
[0042] Exploration action noise is added using an adaptive attenuation mechanism. The calculation formula for the exploration action noise is as follows:
[0043] σ t =max(σ min ,σ t-1 ×δ)
[0044] Where, σ t Let σ be the variance of the exploration action noise at the current time t. min To explore the minimum value of the variance of motion noise, σ t-1 Let δ be the variance of the exploration action noise at the previous time t-1, and δ be the attenuation coefficient.
[0045] The dynamic dual-delay deep deterministic policy gradient algorithm includes: an Actor network and a Critic network;
[0046] The method of using a dynamic dual-delay deep deterministic policy gradient algorithm to solve the reinforcement learning-based task offloading model also includes:
[0047] The average error of the two Critic networks is exponentially smoothed to obtain an error metric.
[0048] If the error metric is greater than or equal to the maximum preset threshold, the delay update step of the Actor network is increased by one step.
[0049] If the error metric is less than or equal to the minimum preset threshold, the delay update step of the Actor network is reduced by one step.
[0050] If the error metric is greater than the minimum preset threshold and less than the maximum preset threshold, the delay update step of the Actor network is not changed.
[0051] Secondly, this application proposes an electronic device, including: one or more processors, and a memory for storing instructions, which, when executed by the one or more processors, cause the one or more processors to perform the dynamic task unloading method of a smart agricultural drone.
[0052] Thirdly, this application proposes a computer-readable storage medium storing executable instructions that, when executed, cause a processor to perform the aforementioned dynamic task unloading method for a smart agricultural drone.
[0053] Beneficial effects:
[0054] This application proposes a dynamic task unloading method for smart agricultural drones. Based on the system's state information in the current time slot, the standardized weighting coefficients are changed in real time to dynamically adjust the user scheduling strategy, unloading ratio, flight speed, flight angle, and transmission power. This solves the problem of model performance degradation caused by environmental changes and static parameters, improves the model's adaptability and convergence, and reduces system costs. Attached Figure Description
[0055] Figure 1 A flowchart of a dynamic task unloading method for a smart agricultural drone according to an embodiment of this application;
[0056] Figure 2 A schematic diagram of dynamic task unloading of a smart agricultural drone according to an embodiment of this application;
[0057] Figure 3 The convergence performance of the dynamic TD3 (D-TD3) algorithm in this application embodiment under different learning rates;
[0058] Figure 4 The convergence performance of the dynamic TD3 (D-TD3) algorithm in this application embodiment under different discount factors;
[0059] Figure 5 A comparison chart of the convergence performance and average system cost of the dynamic TD3 (D-TD3) algorithm in this application embodiment with other solution algorithms;
[0060] Figure 6 A comparison chart of the average system cost of the dynamic TD3 (D-TD3) algorithm and other solution algorithms under different bandwidth conditions in this application embodiment;
[0061] Figure 7 A schematic diagram illustrating the variation of the dynamic weighting coefficients with time slots in an embodiment of this application;
[0062] Figure 8 A comparison chart of the average system cost of the cloud-edge collaboration scheme in this application and other schemes under different task sizes;
[0063] Figure 9 A comparison chart of the average system cost of some unloading strategies in this application embodiment and other unloading strategies under different task sizes;
[0064] Figure 10 A comparison chart of the average system cost of some unloading strategies and other unloading strategies under different bandwidth conditions in this application embodiment. Detailed Implementation
[0065] The specific embodiments of this application will be further described in detail below with reference to the accompanying drawings and examples. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0066] Existing reinforcement learning applications often use static parameters, which cannot be adjusted according to system changes, potentially leading to performance degradation when the environment changes. To address this issue, this application proposes a dynamic task offloading method for smart agricultural drones. Based on traditional agricultural drones, it proposes a cloud-edge collaborative computing model for smart agriculture scenarios and comprehensively considers multiple optimization variables, effectively solving the challenge of limited computing power in traditional agricultural drones.
[0067] This application's method comprehensively considers both latency and energy consumption in smart agriculture, proposing a joint optimization problem that minimizes the weighted sum of latency and energy consumption. It also proposes dynamic weight adjustment to fully account for real-time changes in system state, helping to cope with complex environmental changes in smart agriculture. For the joint optimization problem, a dynamic dual-latency deep deterministic policy gradient algorithm (D-TD3) is proposed, combining adaptive exploration noise and adaptive policy update to improve convergence performance.
[0068] Example 1:
[0069] This embodiment proposes a dynamic task unloading method for smart agricultural drones, such as... Figure 1 , Figure 2 As shown, it includes:
[0070] Step S1: The computational task of collecting pest images using multiple data drones, and sending the status information of the data drones in the current time slot to the edge computing drone;
[0071] The status information of the data drone in the current time slot includes: the location information of the data drone (i.e., p(i) is the location coordinate of the data drone in the i-th time slot), and the remaining total workload of the data drone (i.e., D). remain The remaining total workload of the data drone in the current state), and the workload of the data drone collecting pest images in the current time slot (i.e., D). k (i) represents the amount of data collected by the data collection drone in the i-th time slot.
[0072] In this embodiment, K data drones (hereinafter referred to as D-UAVs) are used. k This represents the computational task of the k-th data drone collecting images of pests. K is a D-UAV. k Flying on a plane at a lower altitude of H1, the set of which is represented as k∈{1,2,…,K}, the position coordinates of the k-th data drone are p. k (i)=[x k (i),y k (i),H1] T Due to the traditional D-UAV k The primary tasks are monitoring and collecting pest image data. Due to limited computing power, edge computing drones (hereinafter referred to as E-UAVs) need to be deployed, along with connections to remote cloud servers. The E-UAV flies at a relatively high altitude of H2. The edge computing drone flies at a higher altitude than the data drone, denoted as E = {e}. The position coordinates of the edge computing drone are q(i) = [x(i), y(i), H2]. T The cloud server is deployed at a relatively far distance M = (m1, m2, 0).
[0073] Step S2: Using the pre-trained task offloading model deployed in the edge computing drone, the state information of the data drone in the current time slot and the standardized weighting coefficients are input into the trained task offloading model to obtain the dynamically optimal user scheduling strategy, the offloading ratio of the edge computing drone, the offloading ratio of the data drone, the flight speed of the edge computing drone, the flight angle of the edge computing drone, and the transmission power of the edge computing drone. Based on the dynamically optimal user scheduling strategy, a communication link is established with one of the multiple data drones, and the offloading ratio of the data drone is sent to the data drone with the established communication link.
[0074] In this embodiment, various parameters of the pre-trained task offloading model are stored in the edge computing drone. The data drone sends the remaining total task volume of the data drone, the location information of the data drone in the current state, and the amount of task volume collected by the data drone in the current time slot to the edge drone in real time. The edge drone inputs the system state information and weighting coefficients in the current time slot into the pre-trained task offloading model to obtain the dynamically optimal user scheduling strategy, the offloading ratio of the edge computing drone, the offloading ratio of the data drone, the flight speed of the edge computing drone, the flight angle of the edge computing drone, and the transmission power of the edge computing drone. A communication link is established according to the dynamically optimal user scheduling strategy to send the offloading ratio of the data drone to the data drone.
[0075] This embodiment employs a dynamic update strategy for latency and energy consumption coefficients, meaning they change in real time during task processing. Existing technologies use models with fixed weight coefficients, which lack flexibility and cannot adapt to changes in different environments and tasks. In some cases, the system may face varying latency or energy consumption requirements. If the weight coefficients are fixed, the system cannot adjust to these changes. For example, as the workload or network load changes, the system may need to adjust priorities. Fixed weights lead to inefficiency when facing new challenges and an inability to respond promptly to environmental changes.
[0076] The standardized weighting coefficients are obtained through the following process:
[0077] Step S2.1: Calculate the initial delay coefficient in the current state based on the remaining total task load of the data drone and the remaining total system time.
[0078] In this embodiment, the latency coefficient represents the weight of latency in the optimization problem, i.e., the system's latency requirements. Its value should change in real time with changes in the system state throughout the task processing, rather than remaining fixed. The system's emphasis on latency is affected by the remaining workload and remaining time. Therefore, the initial latency coefficient is defined as:
[0079]
[0080] Where α0 is the initial time delay coefficient, D remain t represents the remaining total workload of the data drone in the current state. remain This represents the total remaining time of the system, i.e., the total remaining time of the entire system.
[0081] Step S2.2: Based on the remaining available power of the edge computing drone and its total power in the current state, calculate the initial energy consumption coefficient for the current state. The calculation formula is as follows:
[0082]
[0083] Where β0 is the initial energy consumption coefficient in the current state, E remain This represents the remaining available power of the edge computing drone in its current state. Lower available power indicates higher overall system energy consumption, highlighting that reducing energy consumption is a pressing issue. b This represents the total power consumption of the edge computing drone.
[0084] Step S2.3: Process the initial time delay coefficient and the initial energy consumption coefficient to obtain the standardized weighted coefficients under the current state;
[0085] In this embodiment, the two unnormalized factors α0 and β0 are then normalized to their maximum and minimum values to obtain the normalized time delay coefficient and the normalized energy consumption coefficients α1 and β1. To ensure that the sum of the two coefficients is 1, further standardization is performed.
[0086] The standardized weighted coefficients in the current state include: a standardized time delay coefficient and a standardized energy consumption coefficient. In this embodiment, to ensure that the energy consumption coefficient and the time delay coefficient add up to 1, the final standardized weighted coefficient is defined as follows:
[0087]
[0088]
[0089] Where α is the standardized latency coefficient, β is the standardized energy consumption coefficient, α1 is the normalized latency coefficient, and β1 is the normalized energy consumption coefficient. This ensures that the preference for latency and energy consumption dynamically adjusts with changes in system state during task processing, thus fully considering the impact of system state.
[0090] In the implementation of UAV-assisted edge computing, during the initial stage of task processing, the UAV has relatively sufficient battery power and a large remaining workload, requiring faster processing. Therefore, the latency coefficient is relatively high. Towards the end of the task processing, the UAV's battery is relatively scarce, and the remaining workload is smaller. The system prioritizes conserving power to maintain the operation of the UAV-assisted edge computing system, resulting in a higher energy consumption coefficient. Thus, the overall performance shows a gradual decrease in latency and a gradual increase in energy consumption. This dynamic offloading model, by dynamically adjusting weight coefficients, allows the system to adapt more flexibly to the current state, avoiding over-optimization of any single objective and thereby improving the overall performance of the UAV-assisted edge computing system.
[0091] Step S3: The edge computing drone adjusts its position based on its flight speed and flight angle;
[0092] Step S4: The data drone that has established a communication link unloads the computational task of the pest image from the data drone and transfers it to the edge computing drone that has been repositioned.
[0093] Step S5: After the edge computing drone has been repositioned, it offloads the computation task of the pest image to the cloud server using the transmission power of the edge computing drone, based on the offload ratio of the edge computing drone.
[0094] In this embodiment, the task unloading model needs to be pre-trained on a cloud server. The training process includes:
[0095] Step S100: Construct a task offloading model based on the communication model, computation model, energy harvesting model, latency model, and energy consumption model;
[0096] In this embodiment, the process of establishing the communication model, computation model, energy harvesting model, delay model, and energy consumption model is described in detail first.
[0097] (1) Communication model:
[0098] The entire communication and mission offloading cycle has a length of T, which is uniformly divided into I time slots, with the time slot set i∈{1,2,3,4...i}. Each time slot includes a flight duration t. fly And hovering duration t hover At the beginning of each time slot, the E-UAV flies to a new position and hovers until the time slot ends. This embodiment assumes the D-UAV... k It moves randomly within a designated area at low speed. During each time slot, when establishing communication with the E-UAV, it remains hovering. Furthermore, it is stipulated that within a single time slot, the E-UAV can only communicate with one D-UAV. k When establishing communication, the user scheduling constraints are defined as follows:
[0099] τ k (i)∈{0,1}
[0100]
[0101] Where, τ k (i) represents the user scheduling strategy in the i-th time slot of the communication model, i.e., how to select the data drones to establish a communication link.
[0102] After establishing a communication link, D-UAV k The collected pest image information can be unloaded, and the E-UAV can continue to unload the unloaded tasks to the cloud server. In each time slot, the updated location of the E-UAV is:
[0103] q(i+1) = [x(i+1), y(i+1), H2]T
[0104] The flight angles of the E-UAV all satisfy φ(i)∈[0,2π], and the flight speeds satisfy v(i)∈[0,v]. max Therefore, we have:
[0105] x(i+1)=x(i)+t fly v(i)cosφ(i)
[0106] y(i+1)=y(i)+t f1y (i)sinφ(i)
[0107] Among them, v max This is the maximum flight speed.
[0108] Communication models include D-UAV k For communication with E-UAV and communication between E-UAV and the cloud, since there are almost no buildings obstructing the smart agriculture scenario, this embodiment assumes that only the LOS link is considered for both channels, ignoring the impact of NLOS. The channel gains of the two communication links are as follows:
[0109]
[0110] Where η0 is the channel gain per unit distance, g k (i) is D-UAV k The Euclidean distance between the E-UAV and the UAV, g e (i) represents the Euclidean distance between the E-UAV and the cloud server. Therefore, Shannon's formula yields two wireless transmission rates:
[0111]
[0112] Where, r k,e (i)D-UAV k Link between E-UAV, r e,c (i) represents the wireless transmission rate of the link between the E-UAV and the cloud server, B k,e For D-UAV k Link between B and E-UAV e,c P is the bandwidth of the link between the E-UAV and the cloud server. up,e (i) is D-UAV k Links between E-UAV and P up,c (i) represents the transmission power of the link between the E-UAV and the cloud server, σ 2 This represents the noise power of the two links.
[0113] (2) Calculation model:
[0114] The computational model refers to D-UAV k After collecting computing tasks, how to offload the tasks results in two offload ratios: the offload ratio R of E-UAV. k,E (i) The proportion of E-UAV offloading to the cloud R E (i) Both types of uninstallation are partial uninstallations, thus satisfying the following:
[0115] 0≤R k,e (i)≤1
[0116] 0≤R e,c (i)≤1
[0117] Among them, R k,e (i) represents the offloading ratio of computational tasks from the data drone to the edge computing drone in the computational model, R. e,c (i) represents the ratio of edge computing drones offloading computing tasks to cloud servers in the computing model.
[0118] (3) Energy harvesting model:
[0119] In future drone-assisted edge computing (MEC), energy harvesting technology will be more widely used in drones due to battery capacity limitations. The E-UAV in this invention is equipped with energy harvesting capabilities, enabling it to continuously collect solar energy packets for energy acquisition. The arrival of these energy packets follows a Poisson process, meaning the number of energy packets arriving e(i) in each time slot follows a Poisson distribution. Assuming the energy carried by each energy packet is fixed, the energy E collected in each time slot... charge (i) can be described as a scaling transformation of the Poisson distribution. The collected energy packets will be stored in the batteries of the E-UAV to maintain flight hovering, mission computation, and transmission.
[0120] (4) Delay Model:
[0121] The time delay model refers to D-UAV k The collected computational tasks can be computed locally, on the E-UAV, or in the cloud, resulting in local computation latency, transmission latency to the E-UAV, computation latency within the E-UAV, and transmission latency from the E-UAV to the cloud. Since cloud servers have greater computing power, this embodiment ignores the computational latency of cloud servers. The local D-UAV... k The computation delay is:
[0122]
[0123] Among them, t D-UAV,k (i) represents the computation delay of the k-th data UAV in the i-th time slot, D k(i) represents the i-th time slot D-UAV k The collected computational task size, u represents the number of CPU cycles required to process each unit bit of task, f D-UAV D-UAV k The calculated frequency of D-UAV. k The transmission delay to E-UAV is expressed as:
[0124]
[0125] Among them, t tr,k,e (i) represents the computation latency in the i-th time slot of the k-th data drone transmitting the computation task to the i-th time slot of the edge computing drone, r k,e (i) represents the transmission rate derived from Shannon's formula. The calculated delay of E-UAV is expressed as:
[0126]
[0127] Among them, t E-UAV (i) represents the computation delay in the i-th time slot of the edge computing drone, f E-UAV The calculation formula for E-UAV transmission to the cloud, representing the calculation frequency of E-UAV, is as follows:
[0128]
[0129] Among them, t tr,e,c (i) represents the computation latency in the i-th time slot when the edge computing drone transmits the computing task to the cloud server, r e,c (i) represents the transmission rate derived from Shannon's formula.
[0130] Therefore, the total delay is expressed as:
[0131] T total,k (i)=max{t D-UAV,k (i),t tr,k,e (i)+t E-UAV,c (i),t tr,k,e (i)+t tr,e,c (i)}
[0132] Among them, T total,k (i) represents the total delay in the i-th time slot in the delay model.
[0133] (5) Energy consumption model:
[0134] Energy consumption model refers to the propulsion energy consumption of UAVs and D-UAVs. kThe energy consumption includes local computation, transmission energy consumption to the E-UAV, computational energy consumption of the E-UAV, and transmission energy consumption from the E-UAV to the cloud server. Since the cloud server has strong computing power, its computational energy consumption is ignored in this embodiment. Throughout the entire cycle, the D-UAV... k Both the D-UAV and E-UAV fly or hover in the air, significantly impacting energy consumption. Therefore, this embodiment also considers the propulsion energy consumption of the UAV. Since the D-UAV moves randomly at low speed within a defined area and remains hovered while establishing communication with the E-UAV, the propulsion energy consumption is simplified to the energy consumption during the entire time slot of hovering, expressed as:
[0135] E hover1 (i)=KP hover1 (i)(t fly +t hover )
[0136] Among them, E hover1 (i) represents the energy consumption of the data UAV during hovering in the i-th time slot, K represents the total number of D-UAVs, and P hover1 (i) represents the hovering power of the data UAV in the i-th time slot, which is a fixed value. The propulsion energy consumption of the E-UAV is divided into two parts, and its hovering energy consumption is expressed as:
[0137] E hover2 (i)=P hover2 (i)t hover
[0138] Among them, E hover2 (i) represents the energy consumption of the UAV during hovering within the i-th time slot, calculated at the edge. hover2 (i) is the hovering power of the UAV calculated at the edge of the i-th time slot, which is a fixed value. The energy consumption of the E-UAV during flight is defined as:
[0139]
[0140] Among them, E fly (i) represents the energy consumption of the edge computing drone during the i-th time slot, and W2 represents the mass of the edge computing drone.
[0141] Transmission energy consumption is related to transmission power and transmission time, where the transmission power P transmitted to the cloud is the most significant factor. up,c (i) is an optimization variable in this optimization problem, D-UAV k The energy consumption for transmission to E-UAV and the energy consumption for transmission from E-UAV to the cloud are expressed as follows:
[0142] E tr,D-UAV,k,e (i)=P up,e (i)t tr,k,e (i)
[0143] E tr,E-UAV,e,c (i)=P up,c (i)t tr,e,c (i)
[0144] Among them, E tr,D-UAV,k,e (i) represents the transmission energy consumption of the k-th data drone transmitting computing tasks to the i-th time slot of the edge computing drone, E tr,E-UAV,e,c (i) represents the transmission energy consumption of the edge computer transmitting computing tasks to the cloud server in the i-th time slot, P. up,e (i) represents the D-UAV in the i-th time slot. k Link between E-UAV, P up,c (i) represents the transmission power of the link between the E-UAV and the cloud server.
[0145] Computational energy consumption is related to computational power and computation time, where computational power is related to computational frequency, and is defined as follows:
[0146]
[0147] Where, p com,D-UAV p is the computing power of the data drone. com,E-UAV Let f be the computing power of the edge computing drone, ε be the impact factor of chip structure on CPU processing, and f be the computing power of the edge computing drone. D-UAV f is the computing frequency of the data drone. E-UAV Let be the computation frequency of the edge computing drone. Therefore, the computational energy consumption of the two drones can be expressed as:
[0148]
[0149] Among them, E com,D-UAV (i) represents the computational energy consumption of the data UAV in the i-th time slot, E com,E-UAV (i) represents the computational energy consumption of the edge drone in the i-th time slot, t D-UAV,k (i) is D-UAV k The time required to transfer computing tasks to E-UAV, t E-UAV (i) is the time required for E-UAV to transfer computing tasks to the cloud.
[0150] In summary, the total energy consumption of the entire system is:
[0151] E total,k (i)=E hover1 (i)+E hover2 (i)+E fly (i)+E tr,D-UAV,k,e (i)+E tr,E-UAV,e,c (i)
[0152] +Ecom,D-UAV (i)+E com,E-UAV (i)
[0153] Among them, E total,k (i) represents the total energy consumption of the i-th time slot, E hover1 (i) represents the propulsion energy consumption of the data UAV in the i-th time slot, E hover2 (i) calculates the energy consumption of the UAV when hovering at the edge of the i-th time slot, E fly (i) represents the energy consumption of the edge computing drone during flight in the i-th time slot, E tr,D-UAV,k,e (i) represents the transmission energy consumption of the k-th data drone transmitting computing tasks to the i-th time slot of the edge computing drone, E tr,E-UAV,e,c (i) represents the energy consumption for the edge computer to transmit computing tasks to the cloud server in the i-th time slot, E com,D-UAV (i) represents the computational energy consumption of the data UAV in the i-th time slot, E com,E-UAV (i) represents the computational energy consumption of the i-th time slot of the edge drone.
[0154] The task offloading model, constructed based on the communication model, computation model, energy harvesting model, latency model, and energy consumption model, is calculated as follows:
[0155]
[0156] C3:q(i)∈{(x(i),y(i)),H2|x(i)∈[0,L1],y(i)∈[0,L2]}
[0157] C4:p k (i)∈{(x k (i),y k (i)),H1|x k (i)∈[0,L1],y k (i)∈[0,L2]}
[0158]
[0159] C7: v(i)∈[0,v max ],φ(i)∈[0,2π]
[0160]
[0161] C9:P min ≤P up,c (i)≤P max
[0162] Where, τ k (i) represents the user scheduling variable for the i-th time slot in the communication model, where I is the total number of time slots, K is the total number of data drones, and T is the total number of data drones.total,k (i) represents the total latency of processing all computational tasks in the i-th time slot in the latency model, E total,k (i) represents the total energy consumption for processing all computational tasks in the energy consumption model, R. k,e (i) represents the offloading ratio of computational tasks from the data drone to the edge computing drone in the computational model, R. e,c (i) represents the offloading ratio of computing tasks from the edge computing drone to the cloud server in the computational model, where α is the standardized latency coefficient, β is the standardized energy consumption coefficient, μ1 is the first dimensionless factor, and μ2 is the second dimensionless factor to eliminate the influence caused by different dimensions. E b P represents the total power consumption of the edge computing drone. min P is the minimum transmit power of the edge computing drone. max Let be the maximum transmit power of the edge computing drone, q(i) be the position coordinate of the edge computing drone in the i-th time slot, x(i) be the x-axis value of the position coordinate of the edge computing drone in the i-th time slot, y(i) be the y-axis value of the position coordinate of the edge computing drone in the i-th time slot, and p be the maximum transmit power of the edge computing drone. k (i) represents the position coordinates of the k-th data UAV in the i-th time slot, x k (i) represents the x-axis value of the position coordinates of the k-th data UAV in the i-th time slot, and y k (i) represents the y-axis value of the position coordinates of the k-th data UAV in the i-th time slot, E hover2 (i) represents the energy consumption of the drone hovering at the edge within the i-th time slot, E fly (i) represents the energy consumption of the UAV during flight at the edge of the i-th time slot in the energy consumption model, E tr,E-UAV,e,c (i) represents the energy consumption of the edge computing drone transmitting computing tasks to the cloud server in the i-th time slot of the energy consumption model, E com,E-UAV (i) represents the computational energy consumption of the edge computing drone in the i-th time slot of the energy consumption model, φ(i) represents the flight angle of the edge computing drone in the i-th time slot, T represents the length of the computational task unloading cycle, and P represents the computational energy consumption of the edge computing drone in the i-th time slot. up,c (i) represents the transmit power of the edge computing drone in the i-th time slot, and L1 and L2 represent the flight range of the edge computing drone and the data drone within a rectangular area with side lengths L1 and L2, respectively. k (i) represents the amount of data collected by the data collection drone in the i-th time slot. D E represents the amount of tasks the system needs to process throughout the entire cycle. charge(i) represents the energy collected in the energy harvesting model of the i-th time slot, v(i) represents the flight speed of the edge computing drone, C1 indicates that within a time slot, the edge computing drone can only establish a communication link with one data drone, C2 indicates that both offloading methods adopt a partial offloading strategy, and the two offloading methods include: the data drone offloading the computing task to the edge computing drone, and the edge computing drone offloading the computing task to the cloud server, C3 represents the flight range of the edge computing drone, C4 represents the flight range of the data drone, C5 indicates that the energy consumption of the edge computing drone must be less than the sum of the available electricity and the collected solar energy, C6 indicates that the transmission and computing delay in each time slot must be less than the length of the time slot, C7 represents the range of the flight angle and flight speed of the edge computing drone, C8 indicates that the data drone, the edge computing drone, and the cloud server must complete the specified tasks, and C9 indicates that the transmission power of the edge computing drone must meet the specified range.
[0163] Step S101: Based on the task offloading model, construct a Markov decision process to obtain a reinforcement learning-based task offloading model, including:
[0164] Step S101.1: Use the remaining available power of the edge computing drone, the location coordinates of the edge computing drone, the location coordinates of the data drone, the remaining total task of the data drone, and the amount of task collected by the data collection drone that has established a communication link with the edge computing drone in the current time slot as the state space of the Markov decision process.
[0165] Step S101.2: Use the optimization variables as the action space of the Markov decision process. The optimization variables include: user scheduling strategy, i.e., which data collection drone to establish a communication link with in the current time slot, the offloading ratio of the data drone to the edge computing drone, the offloading ratio of the edge computing drone to the cloud server, the flight speed of the edge computing drone, the flight angle of the edge computing drone, and the transmission power of the edge computing drone.
[0166] Step S101.3: The probability that the system will transition to the next state after taking an action in the current state is taken as the state transition probability of the Markov decision process;
[0167] Step S101.4: The immediate reward obtained after taking an action in the current state, i.e., the negative value of the optimization objective in the current time slot, is used as the reward function of the Markov decision process;
[0168] Step S101.5: Construct a Markov decision process using the state space, action space, state transition probability, and reward function to obtain a task offloading model based on reinforcement learning.
[0169] In this embodiment, based on the task offloading model optimization problem, a joint optimization problem of minimizing the weighted sum of energy consumption and latency is proposed. The optimization variables are both discrete and continuous, and the objective function and constraints are nonlinear, belonging to a mixed-integer nonlinear programming problem. The following three aspects illustrate why deep reinforcement learning outperforms traditional optimization methods in this optimization problem: 1) Dynamic environment and high-dimensional space: In UAV-assisted edge computing systems, the location of user equipment, task requirements, and network status may change dynamically. Traditional optimization methods struggle to adapt to such dynamic environments because they typically assume that problem parameters are fixed or static. Reinforcement learning can learn policies in such dynamic environments, continuously interacting with the environment to update policies, gradually adapting to changes, and thus finding the optimal scheduling and resource allocation scheme. 2) Complex nonlinear optimization problems: The weighted sum of latency and energy consumption involves nonlinear relationships (such as transmission rate, channel conditions, dynamic changes in computing resources, etc.). Traditional linear programming and dynamic programming often struggle to effectively model and solve such nonlinear problems. Reinforcement learning, especially deep reinforcement learning, can handle high-dimensional nonlinear problems, approximating complex functions through neural networks without explicitly building optimization models. 3) Continuous Control Requirements: In UAV environments, control variables (such as the UAV's flight path, task allocation ratio, and transmission power) are continuous. Traditional discretization methods often result in insufficient accuracy or excessive computational load. Reinforcement learning directly processes the continuous action space, learning precise control strategies that can adapt to the needs of task scheduling and energy optimization.
[0170] This embodiment constructs an MDP (Markov Decision Process) based on the above optimization problem to describe the current joint optimization problem. The MDP mainly includes a state space, action space, state transition probabilities, and a reward function, where S represents the state space, which in this embodiment is represented as: the remaining battery power of the E-UAV, the location information of the two UAVs, the remaining total task size, and the D-UAV performing task offloading. k The task size is defined as A, which represents the action space and is specifically defined in this embodiment as the six optimization variables mentioned above. State transition describes the probability that the system will transition from state s to the next state s' after performing action a.
[0171] The reward function defines the reward value obtained after performing action a in state s. The optimization problem in this embodiment is to minimize the weighted sum of latency and energy consumption of the entire system after completing the total computation task. The reward is defined as the negative value of the weighted sum of latency and energy consumption. Therefore, maximizing the reward is equivalent to minimizing the weighted sum of latency and energy consumption. The reward within a single time slot is defined as follows:
[0172]
[0173] Where, r tThe reward for selecting an action in the current state within the t-th time slot, s t Let a be the state information within the t-th time slot. t The action selected within the t-th time slot.
[0174] Based on the real-time changes in the state of the smart agriculture system during the task calculation process, dynamic weights are implemented to construct a dynamic reward function.
[0175] Step S102: Use the dynamic dual-delay deep deterministic policy gradient algorithm to solve the task offloading model based on reinforcement learning and obtain the optimal dynamic offloading policy;
[0176] Step S103: Deploy the trained reinforcement learning-based task offloading model to the edge computing drone.
[0177] In this embodiment, based on the TD3 algorithm (Twin Delayed DeepDeterministic policy gradient algorithm), a dynamic policy delay update mechanism and an adaptive exploration noise mechanism are introduced to form the D-TD3 (Dynamic-TD3) algorithm and solve the above optimization problem.
[0178] Traditional Deep Q-Network (DQN) algorithms cannot effectively handle problems with continuous action spaces. This is because finding the maximum Q-value in a continuous action space is extremely difficult. DQN is suitable for discrete action spaces, but performs poorly for high-dimensional and continuous problems. The DDPG algorithm, however, can effectively solve problems with continuous action spaces. Based on an Actor-Critic structure, the Actor network outputs specific actions, while the Critic network evaluates the value of those actions. By avoiding directly maximizing the Q-value, DDPG is better suited for continuous action spaces.
[0179] The TD3 algorithm is an improvement upon DDPG, which suffers from overestimation bias due to the approximation error of the Q-value function. Overestimation can cause policy failure during training, leading to local optima. To address this issue, the TD3 algorithm incorporates the following improvements:
[0180] 1) Dual Q-Network Reduces Overestimation Bias: DDPG uses a single Critic network to estimate state-action values (Q-values). This design can lead to overestimation bias, caused by the accumulation of errors in the Q-value function approximation, potentially causing the policy to get stuck in local optima. TD3 introduces a dual Q-network mechanism, using two independent Critic networks to estimate Q-values. When updating the Critic network, TD3 selects the smaller of the two Q-values as the target value, significantly reducing overestimation bias. The formula for calculating the target value is as follows:
[0181]
[0182] Where y is the target Q value, r is the immediate reward obtained at the current moment, and γ is the discount factor, which represents the weight of future rewards relative to current rewards and takes a value of 0 to 1. To evaluate the network's Q-function, which is used to estimate the action π to be taken in state s′. φ′ The long-term return value obtained by (s′).
[0183] 2) Delayed Update Strategy Improves Stability: In DDPG, the Actor and Critic networks are updated simultaneously. This frequent updating can lead to drastic policy fluctuations, especially in high-dimensional continuous action spaces, potentially causing instability during training. TD3 employs a delayed update strategy. The Critic network updates every time, while the Actor network updates only after the Critic network has been updated multiple times. This strategy reduces the instability caused by the update frequency. Delayed updates reduce policy oscillations, improve the system's performance stability in dynamic environments, and make the policy more likely to converge.
[0184] 3) Reduced Target Policy Noise to Minimize Policy Fluctuations: In DDPG, the target policy directly depends on the estimates from the Critic network. This direct dependence can lead to policy fluctuations during training, especially in noisy environments. TD3 introduces noise interference into the target policy. When the Critic network is updated, TD3 adds small noise perturbations to the target action and then uses this noisy action to update the Q-value.
[0185] This noise makes policy updates smoother, avoiding drastic changes in actions and thus improving policy stability. Compared to DDPG, TD3 reduces drastic policy oscillations during training.
[0186] In this embodiment, further improvements have been made to the TD3 network architecture described above, as follows:
[0187] 1) Adaptive decay mechanism for exploration noise: A relatively large amount of exploration noise is added at the beginning of training, and then gradually decays. The decay formula is expressed as:
[0188] σ t =max(σ min ,σ t-1 ×δ)
[0189] Where, σ t Let σ be the variance of the exploration action noise at the current time t. min To explore the minimum value of the variance of motion noise, σ t-1 Let be the noise variance of the exploration action at the previous time t-1, and δ be the attenuation coefficient with a value of 0.995. The initial noise variance is 0.9, which gradually decreases to 0.01 and remains constant.
[0190] During actual training, the noise level begins to decay from 0.9. When the noise variance is greater than 0.1, the policy leans more towards exploration, allowing for the exploration of more state-action spaces and increasing the likelihood of finding the global optimum. When episode 429 is reached, the noise variance decays to 0.1, beginning to balance exploration and exploitation. When episode 817 is reached, the noise variance decays to 0.01 and remains constant, indicating a shift towards exploitation, reducing unnecessary exploration and making the policy more stable.
[0191] Adaptive exploration noise dynamically adjusts its magnitude based on the training state of the reinforcement learning model. This promotes thorough exploration in the early stages while ensuring accurate decision-making in later stages. In this way, the system can flexibly adapt to different environments and task requirements, improving training efficiency and model stability while avoiding overexploration and overexploitation. This accelerates convergence to the optimal policy and ultimately enhances performance.
[0192] 2) Critic error-driven adaptive strategy update:
[0193] The dynamic dual-delay deep deterministic policy gradient algorithm includes: an Actor network and a Critic network;
[0194] The method of using a dynamic dual-delay deep deterministic strategy gradient algorithm to solve the transformed task unloading model also includes:
[0195] Step a: Perform exponential smoothing on the average error of the two Critic networks to obtain the error metric;
[0196] Step b: When the error metric is greater than or equal to the maximum preset threshold, the delay update step of the Actor network is increased by one step;
[0197] Step c: If the error metric is less than or equal to the minimum preset threshold, then reduce the delay update step of the Actor network by one step;
[0198] Step d: If the error metric is greater than the minimum preset threshold and less than the maximum preset threshold, then the delay update step of the Actor network is not changed.
[0199] The average error of the two Critic networks is exponentially smoothed to obtain an error metric.
[0200] In this embodiment, delayed policy updates are one of the improvements of the TD3 algorithm based on the DDPG algorithm. However, updating the policy too quickly or too slowly can lead to premature convergence to a local optimum. Specifically, if the Actor network updates frequently before the Critic network converges, the policy may be optimized under inaccurate Q-value estimates, thus getting stuck in a suboptimal solution. If the Actor network updates too slowly, it may miss the best opportunity to optimize the policy, resulting in an inefficient training process or even stagnation on poor policies.
[0201] In other inventions of UAV-assisted edge computing, the TD3 algorithm often employs a fixed update strategy, with a delay of 2 or 3 steps. Fixed-frequency updates cannot adaptively adjust the update pace based on the current performance of the Critic network. When the error is large, frequent Actor updates may lead to suboptimal Actor action selection, affecting the convergence of the entire training process. Furthermore, with a fixed update frequency, the system cannot adjust the update strategy based on environmental conditions or Critic errors, potentially causing the model to fail to adapt to new dynamic changes when environmental conditions change, thus impacting the system's adaptability.
[0202] In this embodiment, the update frequency of the Actor network is dynamically adjusted according to the error magnitude of the Critic network. The specific update strategy is as follows: the average error of the two Critic networks is calculated and exponentially smoothed to avoid instability caused by error fluctuations or environmental uncertainties, thereby obtaining a stable error metric L. loss If L loss Larger, i.e., L loss >threshold max (i.e., the maximum preset threshold), delay update steps +1, that is, increase the delay update steps and decrease the update frequency, threshold max The value is 0.1. If L loss Smaller, i.e., L loss <threshold min(i.e., minimum preset threshold), delay update steps -1, decrease delay update steps, increase policy update frequency, threshold min The value is 0.01.
[0203] This allows for increased update frequency when the error is small, accelerating convergence; while reducing update frequency when the error is large, avoiding drastic policy fluctuations and instability; thus ensuring policy updates are performed under accurate Q-value network estimation, avoiding getting trapped in local optima, and improving policy quality.
[0204] The method employs a dynamic dual-delay deep deterministic strategy gradient algorithm to solve the transformed task unloading model, including:
[0205] Exploration action noise is added using an adaptive attenuation mechanism. The calculation formula for the exploration action noise is as follows:
[0206] σ t =max(σ min ,σ t-1 ×δ)
[0207] Where, σ t Let σ be the variance of the exploration action noise at the current time t. min To explore the minimum value of the variance of motion noise, σ t-1 Let δ be the variance of the exploration action noise at the previous time t-1, and δ be the attenuation coefficient.
[0208] To verify the effectiveness of the method in this embodiment, the following simulation experiments are provided, and the simulation parameters are shown in Table 1 below:
[0209] Table 1 Simulation Parameters
[0210] parameter numerical values D-UAV quantity 4 Inter-drone link bandwidth 1MHz Drone cloud link bandwidth 1MHZ Drone flight time per time slot 3s Number of time slots 40 Service period 320s E-UAV calculation frequency 1.2GHz D-UAV frequency calculation 0.4GHz Drone battery capacity 240KJ maximum speed of drone 15m / s Site length and width 100m D-UAV height 10m E-UAV height 110m Number of CPU cycles required to process one bit 1000 cycles / s noise power -100dBm Total computational workload 100Mbits D-UAV transmission power 0.1w Cloud location (8700m, 8700m, 0)
[0211] like Figure 3As shown, the convergence performance is best when the learning rate is 0.0001 and 0.0005. When the learning rate of the Actor or Critic is too large, the gradient update amplitude is large, and the model may oscillate back and forth during gradient descent, or even diverge in extreme cases. This makes it difficult for the network parameters to converge to the optimal value, and it cannot effectively learn a suitable policy. An excessively large learning rate also makes it impossible for the Critic network in the D-TD3 algorithm to stably learn the value function of the environment, causing the Actor to be unable to make optimal decisions based on reliable value estimates, thus increasing the instability of the policy. When the learning rate is too small, the update step size of the network is very small, and the parameter adjustment amplitude is small, resulting in a very slow learning speed of the algorithm, requiring more training time to converge to a better policy. Especially in high-dimensional state space environments, a small learning rate will greatly prolong the training time. An excessively small learning rate may cause the model to get stuck in local optima, because the step size is too small, making it difficult to overcome the possible local optimum traps, resulting in insufficient performance of the algorithm in the end.
[0212] like Figure 4 As shown, a higher discount factor will exhibit greater volatility in the early stages because the strategy needs to consider long-term effects, and all historical experience is evaluated during updates. A lower discount factor may result in faster convergence because the model places less weight on future rewards, thus favoring short-term rewards and leading to faster updates; however, the converged return level may be lower because the model lacks long-term strategy planning capabilities. The selection of the discount factor generally follows the principle of choosing a larger discount factor while ensuring good convergence performance; for this purpose, a discount factor of 0.9 is chosen.
[0213] like Figure 5 As shown in the comparison chart of the time latency and energy consumption weighted sum training results of different algorithms under the same environment, it is clear that D-TD3 has a lower system cost. DDPG performs poorly, mainly due to its tendency to overestimate during the estimation process. Although DDPG introduces a target network, its Q-value is often overestimated, affecting the stability of learning. TD3 effectively reduces this bias by taking the smaller value in the dual Q-network update. Furthermore, DDPG is prone to policy overfitting when generating actions based on the target policy, while TD3 adds noise when generating target actions, thus ensuring the smoothness and robustness of policy learning. In addition, TD3 introduces delayed policy updates, enabling the critic network to provide more stable gradient information, avoiding fluctuations during policy updates, and ultimately improving convergence performance. In comparison, these improvements make TD3 more efficient and stable than DDPG in tasks handling continuous action spaces. The D-TD3 proposed in this embodiment incorporates an adaptive policy update mechanism to prevent the agent's policy from converging prematurely and avoiding getting trapped in local optima.
[0214] like Figure 6The diagram shows a comparison of the system costs of the three algorithms under different bandwidth sizes. Clearly, when the task size is the same, the D-TD3 algorithm proposed in this embodiment consistently maintains the lowest system cost. Furthermore, a larger bandwidth results in a higher transmission rate and shorter transmission time, thus reducing latency and energy consumption.
[0215] like Figure 7 As shown in the figure, the dynamic weighting coefficients involved in this embodiment are used because the system state changes in real time during different time slots in the task processing. For example, the remaining task load, remaining time, and remaining power are all closely related to the emphasis on latency and energy consumption. Therefore, the relevant energy consumption latency coefficients are defined according to the real-time changing physical quantities. The design of dynamic weighting coefficients fully adapts to the changes in system state. In the early stage of the calculation, the power is sufficient and the amount of task to be completed is large, so the focus is on latency. In the later stage of the calculation, the power is critical, so the focus is on energy consumption to improve the service duration of E-UAV.
[0216] like Figure 8 As shown, the cloud-edge collaborative model involved in this embodiment, compared with the cloudless model where everything is offloaded to the cloud, clearly demonstrates a larger workload and higher system cost under different task sizes. The cloudless model relies entirely on traditional drones and E-UAVs, resulting in slower task processing speeds and increased energy consumption due to the high power consumption required to maintain E-UAV computation. Offloading everything to the cloud means uploading all computational tasks to a remote cloud server for processing. When the cloud server is far away or network bandwidth is limited, data transmission latency and energy consumption increase significantly. Therefore, the cloud-edge collaborative model has the lowest system cost.
[0217] like Figure 9 , Figure 10 The diagram shows a comparison of system costs between the partial offloading strategy and other offloading strategies used in this embodiment under different computing task sizes and bandwidths. Clearly, the larger the computing task, the higher the system cost; conversely, the higher the bandwidth, the lower the system cost. The partial offloading strategy refers to D-UAV. k The collected computational tasks can be computed locally or offloaded to the E-UAV or the cloud, with both offload ratios ranging from 0 to 1. Binary offload refers to either an offload ratio of 0 or 1, meaning either no offload or full offload. This lacks flexibility in task allocation, hinders better decision-making, and is unsuitable for dynamically changing environments, potentially leading to poor system performance in terms of latency and energy consumption. Random offload refers to an offload ratio ranging from 0 to 1, lacking flexibility and optimization capabilities. It cannot dynamically adjust the offload strategy based on task requirements, network status, and computing resources, resulting in poor performance in terms of latency and energy consumption. Therefore, partial offload is more efficient and suitable for task offloading in smart agriculture drones.
[0218] This embodiment proposes a dynamic task offloading method for smart agricultural drones. A task offloading model is constructed based on communication, computation, energy harvesting, latency, and energy consumption models. A dynamically optimal task offloading strategy is obtained based on a pre-trained model. During the optimization of the task offloading model, a dynamic, dual-latency deep deterministic policy gradient algorithm is employed to solve the model. This embodiment comprehensively considers multiple optimization variables, effectively addressing the challenge of limited computing power in traditional agricultural drones and resolving the performance degradation issue caused by environmental changes, thus improving the model's adaptability. Furthermore, the combination of adaptive exploration noise and adaptive policy updates enhances the convergence performance of the task offloading model.
[0219] Example 2:
[0220] This embodiment proposes an electronic device, including: one or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the dynamic task unloading method of a smart agricultural drone.
[0221] The electronic device may be a mobile phone, computer, or tablet computer, etc., and includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements a dynamic task offloading method for a smart agricultural drone as described in the embodiments. It is understood that the electronic device may also include an input / output (I / O) interface and communication components.
[0222] The processor is used to execute all or part of the steps in the dynamic task unloading method for a smart agricultural drone as described in the above embodiments. The memory is used to store various types of data, which may include, for example, instructions for any application or method in the electronic device, as well as application-related data.
[0223] The processor can be implemented as an Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), controller, microcontroller, microprocessor, or other electronic components, and is used to execute the dynamic task unloading method for a smart agricultural drone described in the above embodiments.
[0224] Example 3:
[0225] This embodiment proposes a computer-readable storage medium that stores executable instructions. When these instructions are executed, if they are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0226] The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the dynamic task unloading method for a smart agricultural drone described in various embodiments of this application.
[0227] The aforementioned storage media include: flash memory, hard disk, multimedia card, card-type memory (e.g., SD (Secure Digital Memory Card) or DX (Memory Data Register, MDR) memory, random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, server, APP (Application) application store, and other media capable of storing program verification codes. These media store computer programs, and when executed by a processor, they can implement the various steps of the aforementioned dynamic task unloading method for a smart agricultural drone.
[0228] The various embodiments in this application are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0229] The scope of protection of this application is not limited to the embodiments described above. Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from the scope and spirit of this disclosure. If such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, then the intent of this disclosure also includes such modifications and variations.
Claims
1. A dynamic task unloading method for a smart agricultural drone, characterized in that, include: The computational task involves using multiple data drones to collect images of pests and sending the status information of the data drones in the current time slot to the edge computing drone. A pre-trained task offloading model deployed in an edge computing drone is used. The state information of the data drone in the current time slot and the standardized weighting coefficients are input into the trained task offloading model to obtain the dynamically optimal user scheduling strategy, the offloading ratio of the edge computing drone, the offloading ratio of the data drone, the flight speed of the edge computing drone, the flight angle of the edge computing drone, and the transmission power of the edge computing drone. According to the dynamically optimal user scheduling strategy, a communication link is established with one of the multiple data drones, and the offloading ratio of the data drone is sent to the data drone with the established communication link. The edge computing drone adjusts its position based on its flight speed and flight angle. The data drone that has established a communication link will unload the computational task of the pest image from the data drone and transfer it to the edge computing drone that has been repositioned. After the edge computing drone has been repositioned, it offloads the computational task of the pest image to the cloud server based on the offload ratio of the edge computing drone.
2. The dynamic task unloading method for a smart agricultural drone according to claim 1, characterized in that, The standardized weighting coefficients are obtained through the following process: Based on the remaining total workload of the data drone and the remaining total system time under the current state, the initial delay coefficient under the current state is calculated. The initial energy consumption coefficient for the current state is calculated based on the remaining available power of the edge computing drone and the total power of the edge computing drone in the current state. The initial time delay coefficient and initial energy consumption coefficient are processed to obtain the standardized weighted coefficients under the current state.
3. The dynamic task unloading method for a smart agricultural drone according to claim 2, characterized in that, The standardized weighted coefficients under the current state include: the standardized time delay coefficient and the standardized energy consumption coefficient, calculated as follows: ; ; in, To standardize the time delay factor, To standardize the energy consumption coefficient, The normalized time delay coefficient, This is the normalized energy consumption coefficient.
4. The dynamic task unloading method for a smart agricultural drone according to claim 1, characterized in that, The pre-trained task unloading model is pre-trained on a cloud server, and the training process includes: Based on the communication model, computation model, energy harvesting model, latency model, and energy consumption model, a task offloading model is constructed. Based on the task offloading model, a Markov decision process is constructed to obtain a task offloading model based on reinforcement learning. A dynamic dual-delay deep deterministic policy gradient algorithm is used to solve the task offloading model based on reinforcement learning and obtain the optimal dynamic offloading policy. The trained reinforcement learning-based task offloading model is deployed to the edge computing drone.
5. The dynamic task unloading method for a smart agricultural drone according to claim 4, characterized in that, The task offloading model, constructed based on the communication model, computation model, energy harvesting model, latency model, and energy consumption model, is calculated as follows: ; in, Let I be the user scheduling variable for the i-th time slot in the communication model, where I is the total number of time slots and K is the total number of data drones. This represents the total latency for processing all computational tasks in the i-th time slot in the latency model. This represents the total energy consumption for processing all computational tasks in the energy consumption model. This refers to the offloading ratio of computational tasks from data drones to edge computing drones in the computational model. The computational model uses edge computing drones to offload computational tasks to cloud servers. To standardize the time delay factor, To standardize the energy consumption coefficient, As the first dimensionless factor, As the second dimensionless factor, The total battery power of the edge computing drone. This represents the minimum transmit power for edge computing drones. This represents the maximum transmit power of the edge computing drone. Let i be the position coordinates of the edge computing drone in the i-th time slot. Let x be the x-axis value of the position coordinates of the edge computing drone in the i-th time slot. Let y be the position coordinate of the edge computing drone in the i-th time slot. Let be the position coordinates of the k-th data UAV in the i-th time slot. Let x be the x-axis value of the position coordinates of the k-th data UAV in the i-th time slot. Let y be the position coordinate of the k-th data UAV in the i-th time slot. The energy consumption of the drone hovering at the edge is calculated in the i-th time slot. To calculate the energy consumption of the UAV during flight in the i-th time slot of the energy consumption model, This represents the energy consumption of the edge computing drone transmitting computing tasks to the cloud server in the i-th time slot of the energy consumption model. The energy consumption of the edge computing drone in the i-th time slot is given by the energy consumption model. The flight angle of the UAV is calculated at the edge of the i-th time slot, and T is the length of the task unloading cycle. Calculate the UAV's transmit power at the edge of the i-th time slot. , This indicates that the flight range of edge computing drones and data drones is represented by a side length of [missing information]. , Within the rectangular area, Let D be the amount of data collected by the UAV in the i-th time slot, and D be the amount of data the system needs to process throughout the entire cycle. The energy collected in the i-th time slot energy harvesting model, To calculate the flight speed of edge computing drones, This means that within a time slot, an edge computing drone can only establish a communication link with one data drone. This indicates that both offloading methods employ a partial offloading strategy. The two methods include: the data drone offloading computing tasks to an edge computing drone, and the edge computing drone offloading computing tasks to a cloud server. This indicates the flight range of the edge computing drone. Indicates the flight range of the data drone. This indicates that the energy consumption of edge computing drones is less than the sum of available electricity and collected solar energy. This means that the transmission and computation delay in each time slot must be less than the length of the time slot. This indicates the range of flight angles and speeds for edge computing drones. This indicates that data drones, edge computing drones, and cloud servers are to complete designated tasks. This indicates that the transmission power of the edge computing drone must meet the specified range. H1 is the flight altitude of the data drone, and H2 is the flight altitude of the edge computing drone.
6. The dynamic task unloading method for a smart agricultural drone according to claim 4, characterized in that, The step of constructing a Markov decision process based on the task offloading model to obtain a reinforcement learning-based task offloading model includes: The remaining available power of the edge computing drone, the location coordinates of the edge computing drone, the location coordinates of the data drone, the remaining total workload of the data drone, and the workload collected by the data collection drone that has established a communication link with the edge computing drone in the current time slot are used as the state space of the Markov decision process. The optimization variables are used as the action space of the Markov decision process. The optimization variables include: user scheduling strategy, the offloading ratio of data drones to edge computing drones, the offloading ratio of edge computing drones to cloud servers, the flight speed of edge computing drones, the flight angle of edge computing drones, and the transmission power of edge computing drones. The probability that the system will transition to the next state after taking an action in the current state is taken as the state transition probability of the Markov decision process. The immediate reward obtained after taking an action in the current state, i.e., the negative value of the optimization objective in the current time slot, is used as the reward function of the Markov decision process. A Markov decision process is constructed using the state space, action space, state transition probabilities, and reward function to obtain a task offloading model based on reinforcement learning.
7. The dynamic task unloading method for a smart agricultural drone according to claim 4, characterized in that, The method employs a dynamic dual-delay deep deterministic policy gradient algorithm to solve the task offloading model based on reinforcement learning, including: Exploration action noise is added using an adaptive attenuation mechanism. The calculation formula for the exploration action noise is as follows: ; in, Let be the variance of the exploration action noise at the current time t. To explore the minimum value of the motion noise variance, Let be the variance of the exploration action noise at the previous time t-1. This is the attenuation coefficient.
8. The dynamic task unloading method for a smart agricultural drone according to claim 4, characterized in that, The dynamic dual-delay deep deterministic policy gradient algorithm includes: an Actor network and a Critic network; The method of using a dynamic dual-delay deep deterministic policy gradient algorithm to solve the reinforcement learning-based task offloading model also includes: The average error of the two Critic networks is exponentially smoothed to obtain an error metric. If the error metric is greater than or equal to the maximum preset threshold, the delay update step of the Actor network is increased by one step. If the error metric is less than or equal to the minimum preset threshold, the delay update step of the Actor network is reduced by one step. If the error metric is greater than the minimum preset threshold and less than the maximum preset threshold, the delay update step of the Actor network is not changed.
9. An electronic device, characterized in that, include: One or more processors, and a memory for storing instructions that, when executed by the one or more processors, cause the one or more processors to perform a dynamic task unloading method for a smart agricultural drone as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed, cause the processor to perform a dynamic task unloading method for a smart agricultural drone as described in any one of claims 1 to 8.