Dynamic task unloading method for intelligent agricultural unmanned aerial vehicle

By dynamically adjusting the task offloading strategy in smart agricultural drones, combined with real-time optimization of edge computing drones, the problems of delay and energy consumption balance in smart agricultural scenarios are solved, and the adaptability of the model and the overall performance of the system are improved.

CN120075899AActive Publication Date: 2025-05-30NORTHEASTERN UNIV AT QINHUANGDAO

Patent Information

Application Number
CN202510215618.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In the smart agriculture scenario, it is difficult for the existing technology to balance time delay and energy consumption, and ignores the impact of real-time changes in system state on the energy consumption time delay preference, resulting in a decline in model performance when environmental changes.

Method used

A dynamic mission offloading method for smart agricultural drones is proposed. The calculation task of collecting pest images through multiple data drones, and the pre-trained mission offloading model in the edge computing drone is used to dynamically adjust user scheduling strategies, offload ratios, flight speeds, flight angles and transmission power, and optimize according to real-time changes in system status.

Benefits of technology

Dynamic task offloading that balances latency and energy consumption is achieved, improving the adaptability and convergence of the model, reducing system costs, and avoiding model performance degradation due to environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075899A_ABST
    Figure CN120075899A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic task unloading method for a smart agriculture unmanned aerial vehicle, and belongs to the technical field of smart agriculture and edge computing. Establishing a communication link with one data unmanned aerial vehicle in the plurality of data unmanned aerial vehicles according to the dynamic optimal user scheduling strategy, and sending the unloading ratio of the data unmanned aerial vehicle to the data unmanned aerial vehicle with the established communication link; the data unmanned aerial vehicle which has established the communication link unloads the calculation task of the pest image to the edge calculation unmanned aerial vehicle of which the position has been adjusted according to the unloading ratio of the data unmanned aerial vehicle; the edge calculation unmanned aerial vehicle with the adjusted position unloads the calculation task of the pest image to the cloud server at the transmitting power of the edge calculation unmanned aerial vehicle according to the unloading ratio of the edge calculation unmanned aerial vehicle. According to the method, the problem that the model performance is reduced due to environmental changes and static parameters is solved, the adaptability and convergence of the model are improved, and the system cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of smart agriculture and edge computing, and particularly relates to a dynamic task offloading method for a smart agriculture unmanned aerial vehicle. Background Technique

[0002] Edge computing is a new computing architecture that sinks services such as communication, computing, and storage from the network core to the wireless network edge. Among them, unmanned aerial vehicle (UAV)-assisted edge computing is an emerging concept that has emerged in recent years with the development of UAV technology and edge computing technology. It combines the flexibility of UAVs and the high-efficiency data processing capabilities of edge computing, and can adapt to many application scenarios that cannot be satisfied by traditional centralized computing systems.

[0003] In the smart agriculture scenario, traditional agricultural UAVs can interact with the agricultural environment, but it is difficult to complete tasks with high computing requirements, such as pest image analysis. Therefore, deploying an edge computing model relying on traditional agricultural UAVs is an effective solution. However, in current technologies, tasks are usually only offloaded to the edge side, ignoring the auxiliary role of the cloud. Therefore, in the smart agriculture scenario, deploying a cloud-edge collaborative model and performing reasonable task offloading are technical problems that need to be solved urgently.

[0004] Timely analyzing the collected image information and improving the endurance of the edge computing model are the keys to smart agriculture. Therefore, it is necessary to consider latency and energy consumption. However, there is a certain trade-off relationship between decision latency and energy consumption. Local execution of high-computation tasks is difficult to meet the latency requirements; while transmitting data to the edge or the cloud for processing can save local computing resources, but the data transmission process will increase latency energy consumption. Therefore, how to balance the two is the key to designing an efficient system.

[0005] In addition, existing technologies usually only consider one of the indicators of latency or energy consumption, and ignore the impact of the real-time change of the system state on the energy consumption and latency preferences. Therefore, in the smart agriculture scenario, how to reduce latency and energy consumption, and analyze the impact of system state changes on energy consumption and latency preferences are all problems that need to be solved urgently.

[0006] Since the considered model is a three-layer structure composed of data collection drones, edge computing drones, and the cloud, the constructed optimization problem is relatively complex and belongs to the Mixed-Integer Nonlinear Programming (MINLP) problem. The state-action space is high-dimensional and continuous. Traditional optimization methods in the existing technology often encounter problems such as excessive computational resource requirements and long computational time when dealing with large-scale and high-dimensional MINLP. Reinforcement learning has several significant advantages. However, the application of reinforcement learning in the existing technology often uses static parameters and cannot be adjusted according to system changes. Therefore, it may lead to a decline in model performance when the environment changes. Therefore, it is urgent to study a reinforcement learning method with dynamically adjustable parameters. Summary of the Invention

[0007] Aiming at the deficiencies of the existing technology, this application proposes a dynamic task offloading method for smart agriculture drones, which fully considers the real-time changes of the system state and helps to cope with the complex environmental changes in smart agriculture.

[0008] In the first aspect, this application proposes a dynamic task offloading method for smart agriculture drones, including:

[0009] Adopt multiple data drones to collect the computing tasks of pest images and send the state information of the data drones in the current time slot to the edge computing drones;

[0010] Use the pre-trained task offloading model deployed in the edge computing drones to input the state information of the data drones and the normalized weighting coefficients in the current time slot into the trained task offloading model to obtain the dynamically optimal user scheduling strategy, the offloading ratio of the edge computing drones, the offloading ratio of the data drones, the flight speed of the edge computing drones, the flight angle of the edge computing drones, and the transmission power of the edge computing drones. And establish a communication link with one of the multiple data drones according to the dynamically optimal user scheduling strategy, and send the data drone offloading ratio to the data drone with the established communication link;

[0011] The edge computing drone adjusts its position according to the flight speed and flight angle of the edge computing drone;

[0012] The data drone with the established communication link offloads the computing task of the pest image to the edge computing drone after the position has been adjusted according to the data drone offloading ratio;

[0013] The edge computing drone after the position has been adjusted offloads the computing task of the pest image to the cloud server at the transmission power of the edge computing drone according to the offloading ratio of the edge computing drone.

[0014] The process of obtaining the standardized weighting coefficient includes:

[0015] Calculate the initial delay coefficient in the current state based on the remaining total task volume of the data UAV and the remaining total time of the system;

[0016] Calculate the initial energy consumption coefficient in the current state based on the remaining available power of the edge computing UAV and the total power of the edge computing UAV;

[0017] Process the initial delay coefficient and the initial energy consumption coefficient to obtain the standardized weighting coefficient in the current state.

[0018] The standardized weighting coefficient in the current state includes: the standardized delay coefficient and the standardized energy consumption coefficient, and the calculation formula is as follows:

[0019]

[0020] Among them, α is the standardized delay coefficient, β is the standardized energy consumption coefficient, α 1 is the delay coefficient after normalization, and β 1 is the energy consumption coefficient after normalization.

[0021] The pre-trained task offloading model is pre-trained in the cloud server, and the training process includes:

[0022] Construct a task offloading model according to the communication model, calculation model, energy harvesting model, delay model, and energy consumption model;

[0023] Construct a Markov decision process according to the task offloading model to obtain a task offloading model based on reinforcement learning;

[0024] Use the dynamic double delay deep deterministic policy gradient algorithm to solve the task offloading model based on reinforcement learning to obtain the optimal dynamic offloading strategy;

[0025] Deploy the trained task offloading model based on reinforcement learning to the edge computing UAV.

[0026] The calculation formula for constructing the task offloading model according to the communication model, calculation model, energy harvesting model, delay model, and energy consumption model is as follows:

[0027]

[0028] C 3 :q(i)∈{(x(i),y(i)),H 2 |x(i)∈[0,L 1 ,y(i)∈[0,L 2}

[0029] C 4 :p k (i) ∈ {(x k (i), y k (i)), H 1 |x k (i) ∈ [0, L 1 , y k (i) ∈ [0, L 2}

[0030]

[0031] C 7 : v(i) ∈ [0, v max , φ(i) ∈ [0, 2π]

[0032]

[0033] C 9 : P min ≤ P up,c (i) ≤ P max

[0034] Among them, τ k (i) is the user scheduling variable of the i-th time slot in the communication model, I is the total number of time slots, K is the total number of data drones, T total,k (i) is the total delay for processing all computing tasks in the i-th time slot in the delay model, E total,k (i) is the total energy consumption for processing all computing tasks in the energy consumption model, R k,e (i) is the offloading ratio of the data drone to offload the computing task to the edge computing drone in the computing model, R e,c (i) is the offloading ratio of the edge computing drone to offload the computing task to the cloud server in the computing model, α is the normalized delay coefficient, β is the normalized energy consumption coefficient, μ 1 is the first dimensionless factor, μ 2 is the second dimensionless factor, E b is the total power of the edge computing drone, P min is the minimum value of the transmission power of the edge computing drone, P max is the maximum value of the transmission power of the edge computing drone, q(i) is the position coordinate of the edge computing drone in the i-th time slot, x(i) is the x-axis value in the position coordinate of the edge computing drone in the i-th time slot, y(i) is the y-axis value in the position coordinate of the edge computing drone in the i-th time slot, p k (i) is the position coordinate of the k-th data drone in the i-th time slot, x k (i) is the x-axis value in the position coordinate of the k-th data drone in the i-th time slot, yk (i) is the y-axis value of the position coordinates of the k-th data drone in the i-th time slot, E hover2 (i) is the energy consumption when the edge computing drone hovers in the i-th time slot, E fly (i) is the energy consumption when the edge computing drone flies in the i-th time slot in the energy consumption model, E tr,E-UAV,e,c (i) is the transmission energy consumption when the edge computing drone transmits the computing task to the cloud server in the i-th time slot of the energy consumption model, E com,E-UAV (i) is the computing energy consumption of the edge computing drone in the i-th time slot of the energy consumption model, φ(i) is the flight angle of the edge computing drone within the i-th time slot, T is the length of the computing task offloading period, P up,c (i) is the transmission power of the edge computing drone within the i-th time slot, L 1 and L 2 indicate that the flight ranges of the edge computing drone and the data drone are within a rectangular area with side lengths of L 1 and L 2 respectively, D k (i) is the amount of tasks collected by the data collection drone within the i-th time slot, D is the amount of tasks that the system needs to process during the entire period, E charge (i) is the energy collected in the energy collection model within the i-th time slot, v(i) is the flight speed of the edge computing drone, C 1 indicates that within one time slot, the edge computing drone can only establish a communication link with one data drone, C 2 indicates that both offloading methods adopt the partial offloading strategy. The two offloading methods include: the data drone offloads the computing task to the edge computing drone, and the edge computing drone offloads the computing task to the cloud server, C 3 indicates the flight range of the edge computing drone, C 4 indicates the flight range of the data drone, C 5 indicates that the energy consumption of the edge computing drone should be less than the sum of the available power and the energy of the collected solar energy, C 6 indicates that the transmission and computing delay within each time slot should be less than the length of the time slot, C 7 indicates the range of the flight angle and flight speed of the edge computing drone, C 8 indicates that the data drone, the edge computing drone, and the cloud server should complete the specified tasks, C 9 indicates that the transmission power of the edge computing drone should conform to the specified range.

[0035] According to the task offloading model, a Markov decision process is constructed to obtain a task offloading model based on reinforcement learning, including:

[0036] Take the remaining available power of the edge computing UAV, the position coordinates of the edge computing UAV, the position coordinates of the data UAV, the remaining total task volume of the data UAV, and the amount of tasks collected by the data collection UAV that has established a communication link with the edge computing UAV in the current time slot as the state space of the Markov decision process;

[0037] Take the optimization variables as the action space of the Markov decision process. The optimization variables include: user scheduling strategy, offloading ratio of the data UAV to offload computing tasks to the edge computing UAV, offloading ratio of the edge computing UAV to offload computing tasks to the cloud server, flight speed of the edge computing UAV, flight angle of the edge computing UAV, and transmission power of the edge computing UAV;

[0038] Take the probability that the system transfers to the next state after taking a certain action in the current state as the state transition probability of the Markov decision process;

[0039] Take the immediate reward obtained after taking a certain action in the current state, that is, the negative value of the optimization objective in the current time slot, as the reward function of the Markov decision process;

[0040] Construct a Markov decision process with the state space, action space, state transition probability, and reward function to obtain a task offloading model based on reinforcement learning.

[0041] The method of using the dynamic double-delayed deep deterministic policy gradient algorithm to solve the task offloading model based on reinforcement learning includes:

[0042] Add exploration action noise through an adaptive decay mechanism. The exploration action noise is calculated as follows:

[0043] σ t =max(σ min ,σ t-1 ×δ)

[0044] where σ t is the exploration action noise variance at the current time t, σ min is the minimum value of the exploration action noise variance, σ t-1 is the exploration action noise variance at the previous time t - 1, and δ is the decay coefficient.

[0045] The dynamic double-delayed deep deterministic policy gradient algorithm includes: an Actor network and a Critic network;

[0046] The method of using the dynamic double-delayed deep deterministic policy gradient algorithm to solve the task offloading model based on reinforcement learning further includes:

[0047] Perform exponential smoothing on the average error of the two Critic networks to obtain an error metric;

[0048] When the error metric is greater than or equal to the maximum preset threshold, the delay update step of the Actor network is increased by one step;

[0049] When the error metric is less than or equal to the minimum preset threshold, the delay update step of the Actor network is decreased by one step;

[0050] When the error metric is greater than the minimum preset threshold and less than the maximum preset threshold, the delay update step of the Actor network remains unchanged.

[0051] In a second aspect, the present application provides an electronic device, including: one or more processors, and a memory for storing instructions, which when executed by the one or more processors, cause the one or more processors to execute the dynamic task offloading method of a smart agricultural drone as described above.

[0052] In a third aspect, the present application provides a computer-readable storage medium storing executable instructions, which when executed cause a processor to execute the dynamic task offloading method of a smart agricultural drone as described above.

[0053] Advantageous effects:

[0054] The present application provides a dynamic task offloading method for a smart agricultural drone. According to the state information of the system in the current time slot, the normalized weighting coefficient is changed in real time, and the user scheduling strategy, offloading ratio, flight speed, flight angle, and transmission power are dynamically adjusted, solving the problem of model performance degradation caused by environmental changes and static parameters, improving the adaptability and convergence of the model, and reducing the system cost. Description of the drawings

[0055] Figure 1 Flowchart of a dynamic task offloading method for a smart agricultural drone according to an embodiment of the present application;

[0056] Figure 2 Schematic diagram of a dynamic task offloading for a smart agricultural drone according to an embodiment of the present application;

[0057] Figure 3 Convergence performance of the dynamic TD3 (D-TD3) algorithm according to an embodiment of the present application at different learning rates;

[0058] Figure 4 Convergence performance of the dynamic TD3 (D-TD3) algorithm according to an embodiment of the present application at different discount factors;

[0059] Figure 5Comparison chart of the convergence performance and average system cost between the dynamic TD3 (D-TD3) algorithm of the embodiments of the present application and other solution algorithms;

[0060] Figure 6 Comparison chart of the average system cost between the dynamic TD3 (D-TD3) algorithm of the embodiments of the present application and other solution algorithms under different bandwidth conditions;

[0061] Figure 7 Schematic diagram of the dynamic weighting coefficient changing with time slots in the embodiments of the present application;

[0062] Figure 8 Comparison chart of the average system cost between the cloud-edge collaboration scheme of the embodiments of the present application and other schemes under different task sizes;

[0063] Figure 9 Comparison chart of the average system cost between the partial offloading strategy of the embodiments of the present application and other offloading strategies under different task sizes;

[0064] Figure 10 Comparison chart of the average system cost between the partial offloading strategy of the embodiments of the present application and other offloading strategies under different bandwidth sizes. Detailed implementation manners

[0065] The following further describes in detail the specific implementation manners of the present application in conjunction with the accompanying drawings and embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention fall within the protection scope of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0066] In the prior art, the application of reinforcement learning often uses static parameters and cannot be adjusted according to system changes. Therefore, it may lead to a decline in model performance when the environment changes. To solve this problem, the present application proposes a dynamic task offloading method for intelligent agricultural drones. Based on traditional agricultural drones, a cloud-edge collaboration computing model in the intelligent agricultural scenario is proposed, and multiple optimization variables are comprehensively considered, effectively solving the challenge of limited computing power of traditional agricultural drones.

[0067] The method of the present application comprehensively considers two indicators of time delay and energy consumption in intelligent agriculture, proposes a joint optimization problem of minimizing the weighted sum of time delay and energy consumption, and proposes dynamic weight adjustment, fully considering the real-time changes of the system state, which helps to cope with the complex environmental changes in intelligent agriculture. For the joint optimization problem, a dynamic double-delay deep deterministic policy gradient algorithm (D-TD3) is proposed, which combines adaptive exploration noise and adaptive policy update to improve the convergence performance.

[0068] Example 1:

[0069] This example proposes a dynamic task offloading method for intelligent agricultural drones, as Figure 1 , Figure 2 shown, including:

[0070] Step S1: Use multiple data drones to collect the computing tasks of pest images and send the status information of the data drones in the current time slot to the edge computing drones;

[0071] The status information of the data drones in the current time slot includes: the position information of the data drones (i.e., p(i) is the position coordinate of the data drone in the i-th time slot), the remaining total task volume of the data drones (i.e., D remain is the remaining total task volume of the data drone in the current state), and the task volume of the pest images collected by the data drones in the current time slot (i.e., D k (i) is the task volume size collected by the data collection drone in the i-th time slot).

[0072] In this example, K data drones (hereinafter, the k-th data drone is represented by D-UAV k ) are used to collect the computing tasks of pest images. The K D-UAVs k fly on a plane at a lower altitude of H 1 . Their set is represented as k ∈ {1, 2,..., K}, and the position coordinate of the k-th data drone is p k (i) = [x k (i), y k (i), H 1 T . Since the main tasks of traditional D-UAVs k are monitoring and pest image data collection, and their computing power is poor, it is necessary to deploy edge computing drones (hereinafter, the edge computing drone is represented by E-UAV) and link to the remote cloud server. Among them, the E-UAV flies on a plane at a higher altitude of H 2 . The flight altitude of the edge computing drone is greater than that of the data drone, which is represented as E = {e}. The position coordinate of the edge computing drone is q(i) = [x(i), y(i), H 2 T , and the cloud server is deployed at a relatively long distance M = (m 1 , m 2 , 0).

[0073] ​​Step S2: Use the pre-trained task offloading model deployed in the edge computing UAV to input the status information of the data UAV and the normalized weighting coefficient in the current time slot into the trained task offloading model, and obtain the dynamic optimal user scheduling strategy, the offloading ratio of the edge computing UAV, the offloading ratio of the data UAV, the flight speed of the edge computing UAV, the flight angle of the edge computing UAV, and the transmission power of the edge computing UAV. Then, establish a communication link with one of the multiple data UAVs according to the dynamic optimal user scheduling strategy, and send the data UAV offloading ratio to the data UAV with the established communication link;

[0074] In this embodiment, various parameters of the pre-trained task offloading model are saved in the edge computing UAV. The data UAVs send the remaining total task volume of the data UAVs in the current state, the location information of the data UAVs in the current state, and the task volume collected by the data UAVs in the current time slot to the edge UAVs in real time. The edge UAVs input the system status information and the weighting coefficient in the current time slot into the pre-trained task offloading model, and obtain the dynamic optimal user scheduling strategy, the offloading ratio of the edge computing UAV, the offloading ratio of the data UAV, the flight speed of the edge computing UAV, the flight angle of the edge computing UAV, and the transmission power of the edge computing UAV. Then, establish a communication link according to the dynamic optimal user scheduling strategy and send the data UAV offloading ratio to the data UAVs;

[0075] In this embodiment, a dynamic update strategy is adopted for the delay coefficient and the energy consumption coefficient, that is, they change in real time during the task processing. In the prior art, a model with fixed weight coefficients is used, which lacks flexibility and cannot adapt to changes in different environments and tasks. In some cases, the system may face different delay or energy consumption requirements. If the weight coefficients are fixed, the system cannot adjust according to these changes in requirements. For example, as the task volume changes or the network load changes, the system may need to adjust the priority. Fixed weights will cause the system to be inefficient when facing new challenges and unable to respond to environmental changes in a timely manner.

[0076] The process of obtaining the normalized weighting coefficient includes:

[0077] Step S2.1: Calculate the initial delay coefficient in the current state according to the remaining total task volume of the data UAVs in the current state and the remaining total time of the system;

[0078] In this embodiment, the delay coefficient represents the weight of the delay in the optimization problem, that is, the system's requirement for delay. It should change in real time with the change of the system state during the whole task processing, rather than being fixed. The system's emphasis on delay is affected by the remaining task volume and the remaining time. Therefore, the initial delay coefficient is defined as:

[0079]

[0080] Among them, α 0 is the initial time delay coefficient, D remain is the remaining total task volume of the data UAV in the current state, t remain is the remaining total time of the system, that is, the total remaining time of the entire system.

[0081] Step S2.2: According to the remaining available power of the edge computing UAV and the total power of the edge computing UAV in the current state, calculate the initial energy consumption coefficient in the current state. The calculation formula is as follows:

[0082]

[0083] Among them, β 0 is the initial energy consumption coefficient in the current state, E remain is the remaining available power of the edge computing UAV in the current state. The lower the available power, the greater the energy consumption of the entire system, indicating that reducing energy consumption is an urgent problem to be solved. E b is the total power of the edge computing UAV.

[0084] Step S2.3: Process the initial time delay coefficient and the initial energy consumption coefficient to obtain the standardized weighted coefficient in the current state;

[0085] In this embodiment, then perform min-max normalization on the two unnormalized factors α 0 , β 0 to obtain the normalized time delay coefficient and the normalized energy consumption coefficient α 1 , β 1 . In order to ensure that the sum of the two coefficients is 1, further standardization is performed.

[0086] The standardized weighted coefficient in the current state includes: the standardized time delay coefficient and the standardized energy consumption coefficient. In this embodiment, to ensure that the energy consumption coefficient and the time delay coefficient add up to 1, the final standardized weighted coefficient is defined as:

[0087]

[0088]

[0089] Among them, α is the standardized time delay coefficient, β is the standardized energy consumption coefficient, α 1 is the normalized time delay coefficient, β 1 is the normalized energy consumption coefficient. Thus, it is ensured that during the task processing, the preferences for time delay and energy consumption are dynamically adjusted as the system state changes, thereby fully considering the influence of the system state.

[0090] In the specific implementation process of UAV-assisted edge computing, at the initial stage of task processing, the UAV has relatively sufficient power and a large amount of remaining tasks, so it is necessary to accelerate the processing of computing tasks, and thus the delay coefficient is relatively large. At the end of task processing, the UAV's power is relatively scarce and the remaining tasks are few. The system tends to save power resources to maintain the service operation of the UAV-assisted edge computing system, so the energy consumption coefficient is relatively large. Therefore, overall, it shows a situation where the delay coefficient gradually decreases and the energy consumption coefficient gradually increases. In this dynamic offloading model, by dynamically adjusting the weight coefficient, the system can more flexibly adapt to the current state, avoid over-optimizing a certain goal, and thus improve the overall performance of the UAV-assisted edge computing system.

[0091] Step S3: The edge computing UAV adjusts its position according to the flight speed and flight angle of the edge computing UAV.

[0092] Step S4: The data UAV that has established a communication link offloads the computing task of the pest image to the edge computing UAV whose position has been adjusted according to the offloading ratio of the data UAV.

[0093] Step S5: The edge computing UAV whose position has been adjusted offloads the computing task of the pest image to the cloud server at the transmission power of the edge computing UAV according to the offloading ratio of the edge computing UAV.

[0094] In this embodiment, the task offloading model needs to be pre-trained in the cloud server, and the training process includes:

[0095] Step S100: Construct a task offloading model according to the communication model, computing model, energy harvesting model, delay model, and energy consumption model.

[0096] In this embodiment, the establishment processes of the communication model, computing model, energy harvesting model, delay model, and energy consumption model are first described in detail.

[0097] (1) Communication model:

[0098] The length of the entire communication and task offloading cycle is T, which is evenly divided into I time slots, and the time slot set is i ∈ {1, 2, 3, 4... I}. Each time slot includes a flight duration t fly and a hovering duration t hover , and at the beginning of each time slot, the E-UAV flies to a new position and hovers until the end of the time slot. In this embodiment, it is assumed that the D-UAV k moves randomly at a low speed within a specified area. In each time slot, when establishing communication with the E-UAV, it remains in a hovering state, and it is stipulated that in one time slot, the E-UAV can only establish communication with one D-UAV k , and its user scheduling constraint is defined as:

[0099] τ k (i) ∈ {0, 1}

[0100]

[0101] where τ k (i) is the user scheduling strategy in the i-th time slot of the communication model, that is, how to select the data UAVs for establishing communication links.

[0102] After establishing the communication link, the D-UAV k can offload the collected pest image information. After receiving the offloaded task, the E-UAV can also continue to offload it to the cloud server. In each time slot, the updated position of the E-UAV is:

[0103] q(i + 1) = [x(i + 1), y(i + 1), H 2 T

[0104] where the flight angles of the E-UAV all satisfy φ(i) ∈ [0, 2π], and the flight speed satisfies v(i) ∈ [0, v max , so there is:

[0105] x(i + 1) = x(i) + t fly v(i)cosφ(i)

[0106] y(i + 1) = y(i) + t f1y v(i)sinφ(i)

[0107] where v max is the maximum flight speed.

[0108] The communication model includes the communication between the D-UAV k and the E-UAV, and the communication between the E-UAV and the cloud. Since there are almost no building obstructions in the smart agriculture scenario, this embodiment assumes that only LOS links are considered for the two channels, and the influence of NLOS is ignored. The channel gains of the two communication links are respectively:

[0109]

[0110] where η 0 is the channel gain per unit distance, g k (i) is the Euclidean distance between the D-UAV k and the E-UAV, and g e (i) is the Euclidean distance between the E-UAV and the cloud server. Thus, the Shannon formula yields two wireless transmission rates:

[0111] ​

[0112] Among them, r k,e (i) The link between D-UAV k and E-UAV, r e,c (i) is the wireless transmission rate of the link between E-UAV and the cloud server, B k,e is the link between D-UAV k and E-UAV, B e,c is the bandwidth of the link between E-UAV and the cloud server, P up,e (i) is the link between D-UAV k and E-UAV, P up,c (i) is the transmission power of the link between E-UAV and the cloud server, σ 2 is the noise power of the two links.

[0113] (2) Calculation model:

[0114] The calculation model refers to how D-UAV k performs task offloading after collecting computing tasks, thus generating two offloading ratios: the offloading ratio R k,E (i) of E-UAV, the offloading ratio R E (i) of E-UAV to the cloud. Both offloadings are partial offloadings, thus satisfying:

[0115] 0 ≤ R k,e (i) ≤ 1

[0116] 0 ≤ R e,c (i) ≤ 1

[0117] Among them, R k,e (i) is the offloading ratio of the data UAV to the edge computing UAV in the calculation model, and R e,c (i) is the offloading ratio of the edge computing UAV to the cloud server in the calculation model.

[0118] (3) Energy harvesting model:

[0119] In future UAV-assisted mobile edge computing (MEC), due to the limitation of battery capacity, energy harvesting technology will be more widely applied in UAVs. The E-UAV in the present invention is equipped with an energy harvesting function and can obtain energy by continuously collecting solar energy packets. The arrival of the energy packets follows a Poisson process, that is, the number of energy packets arriving in each time slot e(i) satisfies the Poisson distribution. Assuming that the energy carried by each energy packet is fixed, the energy E charge (i) collected in each time slot can be described as a scale transformation of the Poisson distribution. The collected energy packets will be stored in the battery of the E-UAV for maintaining flight hover, computing and transmission of tasks.

[0120] (4) Delay model:

[0121] The delay model means that for the D-UAV k The collected computing tasks can be computed locally, on the E-UAV, or in the cloud, thus resulting in local computing delay, transmission delay to the E-UAV, delay in computing on the E-UAV, and delay in transmitting from the E-UAV to the cloud. Since the cloud server has strong computing power, the computing delay of the cloud server is ignored in this embodiment. Among them, the local D-UAV k The computing delay is:

[0122]

[0123] Among them, t D-UAV,k (i) is the computing delay of the k-th data UAV in the i-th time slot, D k (i) represents the size of the computing tasks collected by the D-UAV in the i-th time slot, u represents the number of CPU cycles required to process each unit-bit task, and f k represents the computing frequency of the D-UAV D-UAV The transmission delay from the D-UAV to the E-UAV is expressed as: k The transmission delay from the D-UAV to the E-UAV is expressed as: k The transmission delay from the D-UAV to the E-UAV is expressed as:

[0124]

[0125] Among them, t tr,k,e (i) is the computing delay of the k-th data UAV transmitting the computing tasks to the edge computing UAV in the i-th time slot, and r k,e (i) is the transmission rate derived from the Shannon formula above. The computing delay of the E-UAV is expressed as:

[0126]

[0127] Among them, t E-UAV (i) is the computing delay of the edge computing UAV in the i-th time slot, and f E-UAV is the computing frequency of the E-UAV. The calculation formula for transmitting from the E-UAV to the cloud is expressed as:

[0128]

[0129] Among them, t tr,e,c (i) is the computing delay of the edge computing UAV transmitting the computing tasks to the cloud server in the i-th time slot, and r e,c (i) is the transmission rate derived from the Shannon formula above.

[0130] Therefore, the total delay is expressed as:

[0131] T total,k (i) = max{t D-UAV,k (i), t tr,k,e (i) + t E-UAV,c (i), t tr,k,e (i) + t tr,e,c}

[0132] Among them, T total,k (i) is the total delay in the i-th time slot in the delay model.

[0133] (5) Energy consumption model:

[0134] The energy consumption model refers to the propulsion energy consumption of the UAV, the energy consumption generated by local computing of the D-UAV k , the transmission energy consumption transmitted to the E-UAV, the computing energy consumption of the E-UAV, and the transmission energy consumption transmitted from the E-UAV to the cloud server. Since the cloud server has strong computing power, the computing energy consumption of the cloud server is ignored in this embodiment. During the entire cycle, both the D-UAV k and the E-UAV are flying or hovering in the air, and it has a great impact on energy consumption. Therefore, the propulsion energy consumption of the UAV is also considered in this embodiment. Because the D-UAV moves randomly at a low speed within a specified area and hovers when establishing communication with the E-UAV, the propulsion energy consumption is simplified to the energy consumption of hovering throughout the time slot, expressed as:

[0135] E hover1 (i) = KP hover1 (i)(t fly + t hover )

[0136] Among them, E hover1 (i) is the hovering energy consumption of the data UAV in the i-th time slot, K represents the total number of D-UAVs, and P hover1 (i) is the hovering power of the data UAV in the i-th time slot, which is a fixed value. The propulsion energy consumption of the E-UAV is divided into two parts, and its hovering energy consumption is expressed as:

[0137] E hover2 (i) = P hover2 (i)t hover

[0138] Among them, E hover2 (i) is the hovering energy consumption of the edge computing UAV in the i-th time slot, and P hover2 (i) is the hovering power of the edge computing UAV in the i-th time slot, which is a fixed value. The energy consumption during flight of the E-UAV is defined as:

[0139]

[0140] Among them, E fly (i) is the energy consumption during the flight of the edge computing UAV in the i-th time slot, and W 2 is the mass of the edge computing UAV.

[0141] The transmission energy consumption is related to the transmission power and transmission time. Among them, the transmission power P up,c (i) is an optimization variable of this optimization problem, and D-UAV k The transmission energy consumption from D-UAV to E-UAV and the transmission energy consumption from E-UAV to the cloud are respectively expressed as:

[0142] E tr,D-UAV,k,e (i) = P up,e (i)t tr,k,e (i)

[0143] E tr,E-UAV,e,c (i) = P up,c (i)t tr,e,c (i)

[0144] Among them, E tr,D-UAV,k,e (i) is the transmission energy consumption when the k-th data UAV transmits the computing task to the edge computing UAV in the i-th time slot, and E tr,E-UAV,e,c (i) is the transmission energy consumption when the edge computer transmits the computing task to the cloud server in the i-th time slot, and P up,e (i) is the link between D-UAV k and E-UAV in the i-th time slot, and P up,c (i) is the transmission power of the link between E-UAV and the cloud server.

[0145] The computing energy consumption is related to the computing power and computing time. Among them, the computing power is related to the computing frequency, and is respectively defined as:

[0146]

[0147] Among them, p com,D-UAV is the computing power of the data UAV, p com,E-UAV is the computing power of the edge computing UAV, ε is the influence factor of the chip structure on the cpu processing, and f D-UAV is the computing frequency of the data UAV, and f E-UAV is the computing frequency of the edge computing UAV. Therefore, the computing energy consumption of the two UAVs is expressed as:

[0148]

[0149] Among them, E com,D-UAV (i) is the computing energy consumption of the data UAV in the i-th time slot, and E com,E-UAV(i) is the computing energy consumption of the edge UAV in the i-th time slot, t D-UAV,k (i) is the D-UAV k The time required to transmit the computing task to the E-UAV, t E-UAV (i) is the time required for the E-UAV to transmit the computing task to the cloud.

[0150] In summary, the total energy consumption of the entire system is:

[0151] E total,k (i) = E hover1 (i) + E hover2 (i) + E fly (i) + E tr,D-UAV,k,e (i) + E tr,E-UAV,e,c (i)

[0152] + E com,D-UAV (i) + E com,E-UAV (i)

[0153] Among them, E total,k (i) is the total energy consumption in the i-th time slot, E hover1 (i) is the propulsion energy consumption of the data UAV in the i-th time slot, E hover2 (i) is the energy consumption when the edge computing UAV hovers in the i-th time slot, E fly (i) is the energy consumption when the edge computing UAV flies in the i-th time slot, E tr,D-UAV,k,e (i) is the transmission energy consumption when the k-th data UAV transmits the computing task to the edge computing UAV in the i-th time slot, E tr,E-UAV,e,c (i) is the transmission energy consumption when the edge computer transmits the computing task to the cloud server in the i-th time slot, E com,D-UAV (i) is the computing energy consumption of the data UAV in the i-th time slot, E com,E-UAV (i) is the computing energy consumption of the edge UAV in the i-th time slot.

[0154] The construction of the task offloading model according to the communication model, computing model, energy harvesting model, delay model, and energy consumption model is as follows:

[0155]

[0156] C 3 : q(i) ∈ {(x(i), y(i)), H 2 | x(i) ∈ [0, L 1 , y(i) ∈ [0, L 2}

[0157] C 4 : p k (i) ∈ {(x k (i), yk (i)), H 1 |x k (i) ∈ [0, L 1 , y k (i) ∈ [0, L 2}

[0158]

[0159] C 7 : v(i) ∈ [0, v max , φ(i) ∈ [0, 2π]

[0160]

[0161] C 9 : P min ≤ P up,c (i) ≤ P max

[0162] Among them, τ k (i) is the user scheduling variable of the i-th time slot in the communication model, I is the total number of time slots, K is the total number of data drones, T total,k (i) is the total delay for processing all computing tasks in the i-th time slot in the delay model, E total,k (i) is the total energy consumption for processing all computing tasks in the energy consumption model, R k,e (i) is the offloading ratio of the data drone to offload the computing task to the edge computing drone in the computing model, R e,c (i) is the offloading ratio of the edge computing drone to offload the computing task to the cloud server in the computing model, α is the normalized delay coefficient, β is the normalized energy consumption coefficient, μ 1 is the first dimensionless factor, μ 2 is the second dimensionless factor, eliminating the influence caused by different dimensions, E b is the total power of the edge computing drone, P min is the minimum value of the transmission power of the edge computing drone, P max is the maximum value of the transmission power of the edge computing drone, q(i) is the position coordinate of the edge computing drone in the i-th time slot, x(i) is the x-axis value in the position coordinate of the edge computing drone in the i-th time slot, y(i) is the y-axis value in the position coordinate of the edge computing drone in the i-th time slot, p k (i) is the position coordinate of the k-th data drone in the i-th time slot, x k (i) is the x-axis value of the position coordinate of the k-th data drone in the i-th time slot, y k (i) is the y-axis value of the position coordinate of the k-th data drone in the i-th time slot, E hover2(i) is the energy consumption when the edge - computing UAV hovers in the i - th time slot, E fly (i) is the energy consumption when the edge - computing UAV flies in the i - th time slot in the energy - consumption model, E tr,E-UAV,e,c (i) is the transmission energy consumption when the edge - computing UAV transmits the computing task to the cloud server in the i - th time slot of the energy - consumption model, E com,E-UAV (i) is the computing energy consumption of the edge - computing UAV in the i - th time slot of the energy - consumption model, φ(i) is the flight angle of the edge - computing UAV in the i - th time slot, T is the length of the computing - task offloading period, P up,c (i) is the transmission power of the edge - computing UAV in the i - th time slot, L 1 、L 2 represents that the flight ranges of the edge - computing UAV and the data UAV are within a rectangular area with side lengths of L 1 、L 2 ,D k (i) is the amount of tasks collected by the data - collection UAV in the i - th time slot, D is the amount of tasks that the system needs to process during the entire period, E charge (i) is the energy collected in the energy - harvesting model in the i - th time slot, v(i) is the flight speed of the edge - computing UAV, C 1 represents that within one time slot, the edge - computing UAV can only establish a communication link with one data UAV, C 2 represents that both offloading methods adopt the partial - offloading strategy. The two offloading methods include: the data UAV offloads the computing task to the edge - computing UAV, and the edge - computing UAV offloads the computing task to the cloud server, C 3 represents the flight range of the edge - computing UAV, C 4 represents the flight range of the data UAV, C 5 represents that the energy consumption of the edge - computing UAV should be less than the sum of the available power and the energy of the collected solar energy, C 6 represents that the transmission and computing delay in each time slot should be less than the length of the time slot, C 7 represents the range of the flight angle and flight speed of the edge - computing UAV, C 8 represents that the data UAV, the edge - computing UAV, and the cloud server should complete the specified tasks, C 9 represents that the transmission power of the edge - computing UAV should meet the specified range.

[0163] Step S101: According to the task - offloading model, construct a Markov decision process to obtain a task - offloading model based on reinforcement learning, including:

[0164] Step S101.1: Take the remaining available power of the edge computing UAV, the position coordinates of the edge computing UAV, the position coordinates of the data UAV, the remaining total task volume of the data UAV, and the task volume collected by the data collection UAV that has established a communication link with the edge computing UAV in the current time slot as the state space of the Markov decision process;

[0165] Step S101.2: Take the optimization variables as the action space of the Markov decision process. The optimization variables include: the user scheduling policy, i.e., which data collection UAV to establish a communication link with in the current time slot, the offloading ratio of the data UAV to offload the computing task to the edge computing UAV, the offloading ratio of the edge computing UAV to offload the computing task to the cloud server, the flight speed of the edge computing UAV, the flight angle of the edge computing UAV, and the transmission power of the edge computing UAV;

[0166] Step S101.3: Take the probability that the system transfers to the next state after taking a certain action in the current state as the state transition probability of the Markov decision process;

[0167] Step S101.4: Take the immediate reward obtained after taking a certain action in the current state, i.e., the negative value of the optimization objective in the current time slot, as the reward function of the Markov decision process;

[0168] Step S101.5: Construct a Markov decision process with the state space, action space, state transition probability, and reward function to obtain a task offloading model based on reinforcement learning.

[0169] In this embodiment, based on the optimization problem of the task offloading model, this embodiment proposes a joint optimization problem with the minimum weighted sum of energy consumption and delay. Among them, the optimization variables are both discrete and continuous, and the objective function and constraints are all non-linear, belonging to a mixed-integer non-linear programming problem. The following will be described from three aspects. Under the current optimization problem, deep reinforcement learning is superior to traditional optimization methods: 1) Dynamic environment and high-dimensional space: In the UAV-assisted edge computing system, the positions, task requirements, and network states of user devices may change dynamically. Traditional optimization methods are difficult to adapt to such a dynamic environment because they usually assume that the problem parameters are fixed or static. Reinforcement learning can learn strategies in such a dynamic environment, update the strategies by continuously interacting with the environment, gradually adapt to the changes, and thus find the optimal scheduling and resource allocation scheme. 2) Complex non-linear optimization problems: The weighted sum of delay and energy consumption involves non-linear relationships (such as dynamic changes in transmission rate, channel conditions, computing resources, etc.). Traditional linear programming and dynamic programming are often difficult to effectively model and solve when facing such non-linear problems. Reinforcement learning, especially deep reinforcement learning, can handle high-dimensional non-linear problems, approximate complex functions through neural networks, and do not require explicit construction of an optimization model. 3) Continuous control requirements: In the UAV environment, control variables (such as the flight path of the UAV, task allocation ratio, transmission power, etc.) are continuous. Traditional discretization methods often lead to insufficient accuracy or excessive computational complexity. Reinforcement learning directly processes the continuous action space, obtains an accurate control strategy through learning, and can adapt to the requirements of task scheduling and energy optimization.

[0170] This embodiment constructs an MDP (Markov decision process) based on the above optimization problem to describe the current joint optimization problem. The MDP mainly includes a state space, an action space, a state transition probability, a reward function, etc. Among them, S represents the state space, which is expressed in this embodiment as: the remaining battery power of the E-UAV, the position information of the two UAVs, the remaining total task size, and the task size of the D-UAV for task offloading. k A represents the action space, which is specifically defined as the above six optimization variables in this embodiment. The state transition describes the probability that the system transfers from state s to the next state s' after executing action a.

[0171] The reward function defines the reward value obtained after executing action a in state s. The optimization problem of this embodiment is to minimize the weighted sum of delay and energy consumption after the entire system completes the total computing task. The reward is defined as the negative value after weighting the delay and energy consumption. Therefore, maximizing the reward is equivalent to minimizing the weighted sum of delay and energy consumption. The reward within a single time slot is defined as:

[0172]

[0173] where r tThe reward for selecting a certain action in the current state at the t-th time slot, s t The state information at the t-th time slot, a t The action selected at the t-th time slot.

[0174] According to the real-time changes in the state of the smart agriculture system during the task calculation process, a dynamic weight is implemented, thereby constructing a dynamic reward function.

[0175] Step S102: Use the dynamic double-delayed deep deterministic policy gradient algorithm to solve the task offloading model based on reinforcement learning, and obtain the optimal dynamic offloading policy;

[0176] Step S103: Deploy the trained task offloading model based on reinforcement learning to the edge computing drone.

[0177] In this embodiment, based on the TD3 algorithm (Twin Delayed Deep Deterministic policy gradient algorithm), a dynamic policy delay update mechanism and an adaptive exploration noise mechanism are introduced to form the D-TD3 (Dynamic-TD3) algorithm and solve the above optimization problem.

[0178] The traditional Deep Q-Network (DQN) algorithm cannot effectively handle problems with continuous action spaces. This is because in a continuous action space, it is very difficult to find the maximum Q value. DQN is suitable for discrete action spaces and performs poorly for high-dimensional and continuous problems. The DDPG algorithm can effectively solve the continuous action space problem and is based on the Actor-Critic structure, where the Actor network outputs specific actions, and the Critic network evaluates the value of that action. DDPG is more suitable for continuous action spaces by avoiding directly maximizing the Q value.

[0179] The TD3 algorithm is an improvement based on DDPG. The DDPG algorithm has the problem of overestimation bias. This is caused by the approximation error of the Q-value function. Overestimation will cause the algorithm to have a policy failure during training and fall into a local optimal solution. To address this problem, the TD3 algorithm has made the following improvements:

[0180] 1) Double Q-network reduces overestimation bias: DDPG uses a single Critic network to estimate the state-action value (Q-value). This design may lead to overestimation bias, which is caused by the error accumulation in the Q-value function approximation and may trap the policy in a local optimum. TD3 introduces a double Q-network mechanism, using two independent Critic networks to estimate the Q-value. When updating the Critic network, TD3 selects the smaller of the two Q-values as the target value, which can significantly reduce the overestimation bias. The formula for calculating the target value is as follows:

[0181]

[0182] where y is the target Q-value, r is the immediate reward obtained at the current time, γ is the discount factor, representing the weight of future rewards relative to the current reward, and its value ranges from 0 to 1. is the evaluation network Q function, used to estimate the long-term return value obtained by taking the action π φ′ (s′).

[0183] 2) Delayed update strategy improves stability: In DDPG, the Actor network and the Critic network are updated simultaneously. This frequent update can lead to drastic fluctuations in the policy, especially in high-dimensional continuous action spaces, which may cause the training process to be unstable. TD3 adopts a delayed update strategy. The Critic network is updated each time, while the Actor network is updated only after the Critic network has been updated multiple times. This strategy can reduce the instability caused by the update frequency. The delayed update reduces policy oscillation and improves the performance stability of the system in a dynamic environment, making the policy easier to converge.

[0184] 3) Target policy noise reduces policy fluctuations: In DDPG, the target policy directly depends on the estimated value of the Critic network. This direct dependence may cause the policy to fluctuate during the training process, especially in an environment with large noise. TD3 introduces noise interference into the target policy. When updating the Critic network, TD3 adds small noise perturbations to the target action, and then uses this action with added noise for Q-value update.

[0185] This noise makes the policy smoother during update, avoiding drastic changes in actions, thereby improving the stability of the policy. Compared with DDPG, TD3 reduces drastic policy oscillations during the training process.

[0186] In this embodiment, further improvements are made under the above TD3 network architecture, and the improvements are as follows:

[0187] 1) Explore the adaptive attenuation mechanism of noise: Through the adaptive attenuation mechanism, at the initial stage of training, a relatively large exploration action noise is added, and then it gradually attenuates. The attenuation formula is expressed as:

[0188] σ t =max(σ min ,σ t-1 ×δ)

[0189] where, σ t is the exploration action noise variance at the current time t, σ min is the minimum value of the exploration action noise variance, σ t-1 is the exploration action noise variance at the previous time t - 1, and δ is the attenuation coefficient with a value of 0.995. The initial noise variance is 0.9, which gradually attenuates to 0.01 and remains unchanged.

[0190] In the actual training process, the noise starts to attenuate from 0.9. When the noise variance is greater than 0.1, the policy is more inclined to exploration, which can explore more state-action spaces and increase the possibility of finding the global optimal solution. When the episode reaches 429, the noise variance attenuates to 0.1, and exploration and exploitation start to be balanced. When the episode reaches 817, the noise variance attenuates to 0.01 and starts to remain unchanged, and the policy turns to exploitation, reducing unnecessary exploration and making the policy more stable.

[0191] The adaptive exploration noise can dynamically adjust the size of the noise according to the training state of the reinforcement learning model, which can not only promote sufficient exploration in the initial stage but also ensure accurate decision-making in the later stage. In this way, the system can flexibly adapt to different environmental and task requirements, avoid over-exploration and over-exploitation, improve the training efficiency and the stability of the model, thus accelerating the convergence to the optimal policy and improving the final performance.

[0192] 2) Adaptive policy update driven by Critic error:

[0193] The dynamic double-delayed deep deterministic policy gradient algorithm includes: an Actor network and a Critic network;

[0194] When using the dynamic double-delayed deep deterministic policy gradient algorithm to solve the transformed task offloading model, it also includes:

[0195] Step a: Perform exponential smoothing on the average error of the two Critic networks to obtain an error metric;

[0196] Step b: When the error metric is greater than or equal to the maximum preset threshold, increase the delayed update step of the Actor network by one step;

[0197] Step c: When the error metric is less than or equal to the minimum preset threshold, reduce the delay update step of the Actor network by one step;

[0198] Step d: When the error metric is greater than the minimum preset threshold and less than the maximum preset threshold, do not change the delay update step of the Actor network.

[0199] Perform exponential smoothing on the average error of the two Critic networks to obtain the error metric;

[0200] In this embodiment, the delay policy update is one of the improvement points of the TD3 algorithm based on the DDPG algorithm. However, if the update frequency of the policy is too fast or too slow, it may lead to premature convergence to a local optimal solution. Specifically, if the Actor network is updated frequently before the Critic network converges, it may cause the policy to be optimized under inaccurate Q-value estimates, thus falling into a suboptimal solution. If the Actor network is updated too slowly, it may miss the best opportunity to optimize the policy, resulting in an inefficient training process and even stagnating at a poor policy.

[0201] In other inventions of drone-assisted edge computing, the TD3 algorithm often adopts a fixed policy update, and the delay update step often adopts 2 or 3. Fixed-frequency updates cannot adaptively adjust the update pace according to the current performance of the Critic network. When the error is large, if frequent Actor updates are still performed, it may lead to sub-optimal action selection of the Actor, thus affecting the convergence of the entire training process. If the fixed policy update frequency is used, the system cannot adjust the update policy according to the environmental state or the error of the Critic, which may cause the model to be unable to adapt to new dynamic changes when the environmental state changes, thus affecting the adaptability of the system.

[0202] In this embodiment, the update frequency of the Actor network is dynamically adjusted according to the error size of the Critic. The specific policy update process is as follows: Calculate the average error of the two Critic networks and perform exponential smoothing to avoid instability caused by error fluctuations or environmental uncertainties, so as to obtain a stable error metric L loss , if L loss is larger, that is, L loss > threshold max (i.e., the maximum preset threshold), the delay update step +1, that is, increase the delay update step and decrease the update frequency, and the value of threshold max is 0.1. If L loss is smaller, that is, L loss < threshold min(i.e., the minimum preset threshold), delay the update step by -1, decrease the delay update step, increase the policy update frequency, threshold min The value is 0.01.

[0203] In this way, the update frequency can be increased when the error is small to accelerate convergence; while the update frequency can be decreased when the error is large to avoid drastic fluctuations in the policy and instability; thus ensuring policy updates under accurate Q-value network estimation, avoiding falling into local optimal solutions, and improving the policy quality.

[0204] The dynamic double-delay deep deterministic policy gradient algorithm is used to solve the transformed task offloading model, including:

[0205] Add exploration action noise through an adaptive attenuation mechanism. The calculation formula for the exploration action noise is as follows:

[0206] σ t = max(σ min , σ t-1 × δ)

[0207] where, σ t is the exploration action noise variance at the current time t, σ min is the minimum value of the exploration action noise variance, σ t-1 is the exploration action noise variance at the previous time t-1, and δ is the attenuation coefficient.

[0208] To verify the effectiveness of the method in this embodiment, the following simulation experiments are provided. The simulation parameters are shown in Table 1 as follows:

[0209] Table 1 Simulation Parameters

[0210] Parameter Value Number of D-UAVs 4 Bandwidth of Inter-UAV Link 1 MHz Bandwidth of UAV-Cloud Link 1 MHz Flight Time of UAV per Time Slot 3s Number of Time Slots 40 Service Period 320s Computing Frequency of E-UAV 1.2 GHz Computing Frequency of D-UAV 0.4 GHz Battery Capacity of UAV 240 KJ Maximum Speed of UAV 15 m / s Length and Width of Site 100m Height of D-UAV 10m Height of E-UAV 110m Number of CPU Cycles Required for Unit bit Processing 1000 cycles / s Noise Power -100 dBm Total Computing Task Volume 100 Mbits Transmission Power of D-UAV 0.1w Cloud Location (8700 m, 8700 m, 0)

[0211] As Figure 3As shown, the convergence performance is best when the learning rate is 0.0001 and 0.0005. When the learning rate of the Actor or Critic is too large, the gradient update amplitude is large, and the model may oscillate back and forth during the gradient descent process, and even diverge in extreme cases. This will cause the network parameters to be difficult to converge to the optimal value and unable to effectively learn a suitable strategy. A too large learning rate will also make the Critic network in the D-TD3 algorithm unable to stably learn the value function of the environment, resulting in the Actor being unable to make optimal decisions based on reliable value estimates, thereby increasing the instability of the strategy. When the learning rate is too small, the update step size of the network is very small, and the parameter adjustment amplitude is small, resulting in a very slow learning speed of the algorithm, and more training time is required to converge to a better strategy. Especially in an environment with a high-dimensional state space, a small learning rate will greatly extend the training time. A too small learning rate may cause the model to fall into a local optimum because the step size is too small to cross the possible local optimum trap, resulting in insufficient final performance of the algorithm.

[0212] As Figure 4 shown, a higher discount factor will have larger fluctuations in the early stage because the strategy needs to consider long-term effects and all historical experiences will be evaluated during the update. A lower discount factor may have a faster convergence process because the model assigns less weight to future rewards. Therefore, the preference for short-term rewards makes the update faster, but the converged reward level may be lower because the model lacks long-term strategy planning ability. The selection of the discount factor generally follows the principle of taking a larger discount factor under the condition of ensuring good convergence performance. Therefore, the discount factor is selected as 0.9.

[0213] As Figure 5 shown, the comparison chart of the weighted sum of delay and energy consumption training results of different algorithms in the same environment. Obviously, the system cost of D-TD3 is smaller. The effect of DDPG is poor, which is mainly attributed to the fact that DDPG is prone to overestimation bias during the estimation process. Although DDPG introduces a target network, its Q value is still often overestimated, which affects the stability of learning. TD3 effectively reduces this bias by taking the smaller value in the double Q-network update. At the same time, DDPG is prone to policy overfitting when generating actions with the target policy, while TD3 adds noise when generating target actions, thus ensuring the smoothness and robustness of policy learning. In addition, TD3 introduces delayed policy updates, enabling the critic network to provide more stable gradient information, avoiding fluctuations during policy updates, and ultimately improving the convergence performance. In contrast, these improvements make TD3 more efficient and stable than DDPG in handling tasks with continuous action spaces. The proposed D-TD3 in this embodiment adds an adaptive policy update mechanism to prevent the agent's policy from converging prematurely and avoid falling into local optimal solutions.

[0214] As Figure 6As shown, it represents the comparison of the system costs of three algorithms under different bandwidth sizes. Obviously, when the task sizes are the same, the D-TD3 algorithm proposed in this embodiment always maintains the lowest system cost. In addition, the larger the bandwidth, the higher the transmission rate, the shorter the transmission time, and thus the lower the delay and energy consumption.

[0215] As Figure 7 shown, for the dynamic weighting factor involved in this embodiment, during the task processing, the system state changes in real time at different time slots. For example, the remaining task volume, remaining time, and remaining power are all closely related to the emphasis on delay and energy consumption. Therefore, relevant energy consumption and delay coefficients are defined according to the physically changing quantities in real time, and the design of dynamic weighting factors is adopted to fully adapt to the changes in the system state. At the beginning of the calculation, the power is sufficient and the amount of tasks to be completed is large, so the emphasis is on delay. In the later stage of the calculation, the power is urgent, so the emphasis is on energy consumption, improving the service duration of the E-UAV.

[0216] As Figure 8 shown, the comparison of the cloud-edge collaboration model involved in this embodiment with the case of no cloud and all offloading to the cloud under different task sizes. Obviously, when the amount of tasks to be completed is large, the resulting system cost is large. The mode without cloud means relying entirely on traditional UAVs and E-UAVs, resulting in a slow task processing speed and also consuming a large amount of power to maintain the E-UAV calculation process, leading to an increase in energy consumption. All offloading to the cloud means uploading all computing tasks to a remote cloud server for processing. When the cloud server is far away or the network bandwidth is limited, the data transmission delay and energy consumption increase significantly. Therefore, the system cost of cloud-edge collaboration is the smallest.

[0217] As Figure 9 , Figure 10 shown, it is a comparison chart of the system costs of the partial offloading strategy involved in this embodiment and other offloading strategies under different computing task sizes and different bandwidths. Obviously, the larger the computing task, the greater the system cost, and the larger the bandwidth, the smaller the system cost. The partial offloading strategy means that the computing tasks collected by the D-UAV k can be locally computed or offloaded to the E-UAV or the cloud, and the range of the two offloading ratios is 0-1. Binary offloading means that the values of the two offloading ratios are 0 or 1, that is, either no offloading or all offloading, which cannot perform flexible task allocation, cannot seek better decisions, is not suitable for the dynamically changing environment, and may lead to poor performance in terms of delay and energy consumption of the system. Random offloading means that the values of the two offloading ratios are random numbers between 0 and 1, lacking flexibility and optimization ability, and cannot dynamically adjust the offloading strategy according to task requirements, network status, and computing resources, resulting in poor performance in terms of delay and energy consumption. Therefore, partial offloading is more efficient and suitable for task offloading under intelligent agricultural UAVs.

[0218] This embodiment proposes a dynamic task offloading method for a smart agricultural drone. A task offloading model is constructed based on a communication model, a computing model, an energy harvesting model, a latency model, and an energy consumption model, and a dynamic optimal task offloading strategy is obtained according to the pre-trained model. In the process of optimizing the task offloading model, a dynamic double-delay deep deterministic policy gradient algorithm is adopted to solve the task offloading model. This embodiment comprehensively considers multiple optimization variables, effectively solves the challenge of limited computing power of traditional agricultural drones, solves the problem of model performance degradation due to environmental changes, and improves the adaptability of the model. At the same time, by combining adaptive exploration noise and adaptive policy update, the convergence performance of the task offloading model is improved.

[0219] Embodiment 2:

[0220] This embodiment proposes an electronic device, including: one or more processors, and a memory. The memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the dynamic task offloading method of a smart agricultural drone described above.

[0221] The electronic device can be a mobile phone, a computer, a tablet computer, etc., including a memory and a processor. A computer program is stored on the memory, and when the computer program is executed by the processor, it implements the dynamic task offloading method of a smart agricultural drone as described in the embodiment. It can be understood that the electronic device can also include an input / output (I / O) interface and a communication component.

[0222] Among them, the processor is used to execute all or part of the steps in the dynamic task offloading method of a smart agricultural drone as described in the above embodiment. The memory is used to store various types of data, which can include, for example, instructions of any application program or method in the electronic device, as well as data related to the application program.

[0223] The processor can be implemented by an application specific integrated circuit (ASIC), a digital signal processor (DSP), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components, and is used to execute the dynamic task offloading method of a smart agricultural drone as described in the above embodiment.

[0224] Embodiment 3:

[0225] This embodiment provides a computer-readable storage medium storing executable instructions, which, when implemented in the form of software functional units and sold or used as independent products, can be stored in a computer-readable storage medium.

[0226] This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of a dynamic task offloading method for an intelligent agricultural drone described in various embodiments of this application.

[0227] The aforementioned storage medium includes: flash memory, hard disk, multimedia card, card-type memory (such as SD (Secure Digital Memory Card) or DX (abbreviation for Memory Data Register, MDR), memory data register, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, server, APP (abbreviation for Application, application software) application mall, and various other media that can store program verification codes. A computer program is stored thereon, and when the computer program is executed by a processor, it can implement each step of the aforementioned dynamic task offloading method for an intelligent agricultural drone.

[0228] The various embodiments in this application are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.

[0229] The protection scope of this application is not limited to the above embodiments. Obviously, those skilled in the art can make various changes and deformations to this disclosure without departing from the scope and spirit of this disclosure. If these changes and deformations fall within the scope of the claims of this disclosure and their equivalent technologies, the intention of this disclosure also includes these changes and deformations.

Claims

1. A dynamic task offloading method for a smart agricultural drone, characterized in that: include: The computing task of using multiple data drones to collect pest images and sending the status information of the data drones in the current time slot to the edge computing drones; Using the pre-trained task offloading model deployed in the edge computing drone, the state information of the data drone in the current time slot and the standardized weighted coefficient are input into the trained task offloading model to obtain the dynamically optimal user scheduling strategy, the offloading ratio of the edge computing drone, the offloading ratio of the data drone, the flight speed of the edge computing drone, the flight angle of the edge computing drone, and the transmission power of the edge computing drone. According to the dynamically optimal user scheduling strategy, a communication link is established with one of the multiple data drones, and the data drone offloading ratio is sent to the data drone with the established communication link. The edge computing drone adjusts its position according to the flight speed and flight angle of the edge computing drone; The data drone that has established the communication link unloads the computation task of the pest image to the edge computing drone that has adjusted its position according to the data drone offloading ratio; The edge computing drone, which has adjusted its position, offloads the computational task of the pest image to the cloud server at the transmission power of the edge computing drone according to the offloading ratio of the edge computing drone.

2. The method for dynamic task offloading of a smart agricultural drone according to claim 1, characterized in that: The process of obtaining the standardized weighted coefficient includes: According to the remaining total task volume of the data drone and the remaining total time of the system in the current state, the initial delay coefficient in the current state is calculated; According to the remaining available power of the edge computing drone in the current state and the total power of the edge computing drone, the initial energy consumption coefficient in the current state is calculated; The initial delay coefficient and the initial energy consumption coefficient are processed to obtain a standardized weighted coefficient in the current state.

3. The method for dynamic task offloading of a smart agricultural drone according to claim 2, characterized in that: The standardized weighted coefficients in the current state include: a standardized delay coefficient and a standardized energy consumption coefficient, which are calculated as follows: Among them, α is the standardized delay coefficient, β is the standardized energy consumption coefficient, α1 is the normalized delay coefficient, and β1 is the normalized energy consumption coefficient.

4. The method for dynamic task offloading of a smart agricultural drone according to claim 1, characterized in that: The pre-trained task offloading model is pre-trained in the cloud server, and the training process includes: Construct task offloading model based on communication model, computation model, energy collection model, latency model, and energy consumption model; According to the task offloading model, a Markov decision process is constructed to obtain a task offloading model based on reinforcement learning; A dynamic double-delay deep deterministic policy gradient algorithm is used to solve the task offloading model based on reinforcement learning and obtain the optimal dynamic offloading strategy; The trained reinforcement learning-based task offloading model is deployed to the edge computing drone.

5. The method for dynamic task offloading of a smart agricultural drone according to claim 4, characterized in that: The task offloading model constructed according to the communication model, calculation model, energy collection model, delay model, and energy consumption model is calculated as follows: C1: C2: C3:q(i)∈{(x(i),y(i)),H2|x(i)∈[0,L1],y(i)∈[0,L2]} C4:p k (i)∈{(x k (i),and k (i)),H1|x k (i)∈[0,L1],and k (i)∈[0,L2]} C5:, C6: C7:v(i)∈[0,v max ],φ(i)∈[0,2π] C8: C9:P min ≤P up,c (i)≤P max Among them, τ k (i) is the user scheduling variable for the i-th time slot in the communication model, I is the total number of time slots, K is the total number of data drones, and T total,k (i) is the total delay of processing all computing tasks in the i-th time slot in the delay model, E total,k (i) is the total energy consumption of processing all computing tasks in the energy consumption model, R k,e (i) is the offloading ratio of the data drone to the edge computing drone in the computing model, R e,c (i) is the offloading ratio of edge computing drones to cloud servers in the computing model, α is the standardized delay coefficient, β is the standardized energy consumption coefficient, μ1 is the first dimensionless factor, μ2 is the second dimensionless factor, and E b is the total power of the edge computing drone, P min is the minimum value of the edge computing drone’s transmission power, P max is the maximum value of the transmission power of the edge computing drone, q(i) is the position coordinate of the edge computing drone in the i-th time slot, x(i) is the x-axis value of the position coordinate of the edge computing drone in the i-th time slot, y(i) is the y-axis value of the position coordinate of the edge computing drone in the i-th time slot, p k (i) is the position coordinate of the kth data drone in the i-th time slot, x k (i) is the x-axis value of the position coordinate of the kth data drone in the i-th time slot, y k (i) is the y-axis value of the position coordinate of the kth data drone in the i-th time slot, E hover2 (i) is the energy consumption of the edge computing drone when hovering in the i-th time slot, E fly (i) is the energy consumption of the edge computing drone during flight in the i-th time slot in the energy consumption model, E tr,E-UAV,e,c (i) is the transmission energy consumption of the edge computing drone transmitting the computing task to the cloud server in the i-th time slot of the energy consumption model, E com,E-UAV (i) is the computational energy consumption of the edge computing drone in the i-th time slot of the energy consumption model, φ(i) is the flight angle of the edge computing drone in the i-th time slot, T is the length of the computational task offloading cycle, P up,c (i) is the transmission power of the edge computing drone in the i-th time slot, L1 and L2 indicate that the flight range of the edge computing drone and the data drone is within the rectangular area with side lengths L1 and L2, and D k (i) is the amount of tasks collected by the data collection drone in the i-th time slot, D is the amount of tasks that the system needs to process in the entire cycle, and E charge (i) is the energy collected in the energy collection model in the i-th time slot, v(i) is the flight speed of the edge computing drone, C1 means that within a time slot, the edge computing drone can only establish a communication link with one data drone, C2 means that both offloading methods adopt a partial offloading strategy, and the two offloading methods include: the data drone offloads the computing task to the edge computing drone, and the edge computing drone offloads the computing task to the cloud server, C3 means the flight range of the edge computing drone, C4 means the flight range of the data drone, C5 means the energy consumption of the edge computing drone must be less than the sum of the available power and the collected solar energy, C6 means the transmission and calculation delay in each time slot must be less than the length of the time slot, C7 means the range of the flight angle and flight speed of the edge computing drone, C8 means the data drone, edge computing drone and cloud server must complete the specified tasks, and C9 means the transmission power of the edge computing drone must be within the specified range.

6. The method for dynamic task offloading of a smart agricultural drone according to claim 4, characterized in that: According to the task offloading model, a Markov decision process is constructed to obtain a task offloading model based on reinforcement learning, including: The remaining available power of the edge computing drone, the location coordinates of the edge computing drone, the location coordinates of the data drone, the remaining total task volume of the data drone, and the task volume collected by the data collection drone that has established a communication link with the edge computing drone in the current time slot are used as the state space of the Markov decision process; The optimization variables are used as the action space of the Markov decision process, and the optimization variables include: user scheduling strategy, offloading ratio of the data drone to the edge computing drone, offloading ratio of the edge computing drone to the cloud server, flight speed of the edge computing drone, flight angle of the edge computing drone, and transmission power of the edge computing drone; The probability that the system will transfer to the next state after taking a certain action in the current state is taken as the state transition probability of the Markov decision process; The timely reward obtained after taking an action in the current state, that is, the negative value of the optimization target in the current time slot, is used as the reward function of the Markov decision process; A Markov decision process is constructed using the state space, action space, state transition probability and reward function to obtain a task offloading model based on reinforcement learning.

7. The method for dynamic task offloading of a smart agricultural drone according to claim 4, characterized in that: The dynamic double-delay deep deterministic policy gradient algorithm is used to solve the task offloading model based on reinforcement learning, including: The exploration action noise is added through the adaptive attenuation mechanism. The exploration action noise is calculated as follows: s t =max(σ min ,s t-1 ×d) Among them, σ t is the exploration action noise variance at the current time t, σ min To explore the minimum value of the action noise variance, σ t-1 is the exploration action noise variance at the previous time t-1, and δ is the attenuation coefficient.

8. The method for dynamic task offloading of a smart agricultural drone according to claim 4, characterized in that: The dynamic double-delay deep deterministic policy gradient algorithm includes: an Actor network and a Critic network; The method adopts a dynamic double-delay deep deterministic policy gradient algorithm to solve a task offloading model based on reinforcement learning, and further includes: Exponentially smooth the average error of the two critic networks to obtain the error metric; When the error metric is greater than or equal to a maximum preset threshold, the number of delayed update steps of the Actor network is increased by one step; When the error metric is less than or equal to a minimum preset threshold, the number of delayed update steps of the Actor network is reduced by one step; When the error metric is greater than a minimum preset threshold and less than a maximum preset threshold, the number of delayed update steps of the Actor network is not changed.

9. An electronic device, characterized in that: include: One or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the dynamic task offloading method of a smart agricultural drone as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: It stores executable instructions, which, when executed, enable the processor to execute the dynamic task unloading method for a smart agricultural drone as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Unmanned aerial vehicle assisted MEC system joint task scheduling and motion trail optimization method

    CN116257335A

  • Air-ground integrated Internet of Things joint resource allocation method based on deep double-Q network and federated learning

    CN116600316A

  • Mobile edge computing auxiliary unloading method based on unmanned aerial vehicle assistance

    CN117596571A

  • Unmanned plane assisted unmanned ship task unloading method based on deep reinforcement learning

    CN118574156A

  • Task unloading method and device based on space-air-ground integrated network

    CN119211875A

Cited By

  • Vehicle information energy power distribution method and device based on depth deterministic strategy gradient

    CN120980660A