Multi-uav air-ground collaborative task offloading optimization method for smart agriculture
By optimizing the task offloading method through multi-UAV air-ground collaboration, the problem of information acquisition and decision-making response lag of UAVs in agricultural scenarios is solved, realizing efficient and reliable monitoring of crop growth status and pest and disease control, and improving the system's task completion rate and energy efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUILIN UNIVERSITY OF TECHNOLOGY
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-23
AI Technical Summary
Existing technologies lack systematic optimization mechanisms for agricultural scenarios in terms of drone deployment, task offloading decisions, and resource coordination. This results in significant lags in information acquisition and decision response across large areas of farmland, making it difficult to achieve real-time and reliable monitoring of crop growth status and pest and disease control.
A task offloading optimization method for multi-UAV air-ground collaboration is constructed. By improving the particle swarm optimization algorithm and the dual-delay deep deterministic policy gradient algorithm, the joint processing of tasks at the UAV end or the ground terminal is dynamically determined. Combined with the adaptive inertial weight and dynamic learning factor adjustment mechanism, the access position and offloading decision of UAVs are optimized to generate resource allocation scheme.
It improves the data response speed in smart agriculture, ensures that tasks are completed within an effective time, reduces drone energy consumption, enhances the continuous supply capacity of drones, and achieves efficient response and stable execution of agricultural tasks.
Smart Images

Figure CN122264412A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV)-assisted edge computing, and particularly to a task offloading optimization method for multi-UAV air-ground collaboration in smart agriculture. Background Art
[0002] In modern agricultural production, crop cultivation management is gradually shifting from experience-driven to data-driven and intelligent decision-making. With the advancement of high-standard farmland construction and large-scale planting, higher requirements for timeliness and computational intelligence are put forward for crop growth status monitoring, precise pest control, and refined regulation of irrigation and fertilization. The traditional method relying on fixed monitoring equipment or manual inspection is difficult to achieve accurate perception in a large area of farmland. Especially in the case of sudden pests and diseases, extreme weather or abnormal operations, there is an obvious lag in information acquisition and decision-making response. UAVs, with their advantages of flexibility and mobility, have become an important means for agricultural remote sensing, plant protection and environmental perception. However, they face problems such as limited computing resources, communication and energy consumption constraints in practical applications. Relying solely on local processing or remote servers is difficult to balance real-time performance and reliability.
[0003] Under this background, the collaborative construction of a UAV-assisted edge computing system between UAVs and farm ground terminal devices has become an important technical direction to support smart agriculture. In the system, ground terminal devices often act as information initiators or data transmitters by mobile devices, and usually have relatively low computing power. Traditional edge computing often calculates tasks locally or offloads them to edge servers, ignoring its efficient processing method of linkage. Forming a hierarchical collaborative computing architecture with UAVs and fixed edge computing power can dynamically determine the joint processing and execution of tasks on terminal devices or UAVs according to the characteristics of operation tasks, UAV status and communication conditions, so as to ensure the timeliness of tasks and the ability of continuous service.
[0004] In summary, the existing solutions still lack a systematic optimization mechanism for UAV deployment, task offloading decision-making and resource collaboration in agricultural scenarios. Summary of the Invention
[0005] The purpose of the present invention is to provide a task offloading optimization method for multi-UAV air-ground collaboration in smart agriculture, aiming to propose a collaborative computing method between UAVs and service terminal devices in each area of the farm for typical agricultural applications such as crop cultivation and pest control, improve the response speed of data in smart agriculture, ensure that tasks are completed in time within their effective time, reduce the energy consumption of UAVs, and improve the continuous supply ability of UAVs.
[0006] To achieve the above object, the present invention provides a task offloading optimization method for multi-UAV air-ground collaboration in smart agriculture, including the following steps:
[0007] Step 1: In the agricultural production area scenario, construct a system containing a set of ground terminals. With drones The system model is used to establish a task model, a communication service model, a computation model, and a motion model of the UAV, thereby obtaining the optimization objective;
[0008] Step 2: Construct a deployment fitness function based on the crop information distribution density function, and use an improved particle swarm optimization algorithm that includes adaptive inertia weights and dynamic learning factor adjustment mechanisms to optimize the initial access position of the UAV;
[0009] Step 3: Based on the optimal initial deployment location, a dual-delay deep deterministic strategy gradient algorithm based on preference prediction and synchronous unloading is used to learn the optimal collaborative unloading decision.
[0010] Step 4: Determine the uninstallation target and uninstallation method, allocate resources, and generate the uninstallation strategy and resource allocation plan for the entire network time.
[0011] Optionally, the execution process of step 1 includes the following steps:
[0012] Step 1.1: Incorporate crop information and ground terminal equipment information into the modeling process to construct a system model;
[0013] Step 1.2: Construct the motion model of the UAV and, based on the time delay constraints of the mission types generated in different time slots, construct the mission model;
[0014] Step 1.3: Construct communication service models for different situations based on communication conditions such as obstacles or terrain;
[0015] Step 1.4: Construct computing models based on different computing service carriers;
[0016] Step 1.5: Based on the established models, clarify the optimization objective, namely, to minimize the total energy consumption of the system while ensuring that all tasks are completed within the specified time delay.
[0017] Optionally, the improved particle swarm optimization algorithm described in step 2 aims to minimize the communication distance between the agricultural terminal equipment and the nearest drone, and introduces drone safety distance constraints during the optimization process to avoid flight conflicts between drones in low-altitude agricultural operations.
[0018] Optionally, the execution process of step 2 includes the following steps:
[0019] Step 2.1: Represent the candidate access positions of the UAV in the airspace as particle position vectors, and introduce crop information distribution as an important constraint and guiding factor for deployment optimization;
[0020] Step 2.2: By mapping the density of crop information, the intensity of data collection demand, and the distribution of ground equipment within the farmland area to the particle fitness evaluation index, the fitness value for each round is calculated; the inertia weight is dynamically increased or decreased based on the change in fitness value, thereby avoiding getting trapped in local optima or improving local fine search.
[0021] Step 2.3: Calculate the center position of the current particle swarm, and calculate the swarm diversity index based on the distance of all particles to the swarm center. Dynamically adjust the individual learning factor and social learning factor to accelerate the approach to the optimal solution.
[0022] Step 2.4: After updating the inertia weights and learning factors, the algorithm calculates the new velocity of each particle in the current iteration according to the velocity update formula, and truncates the components that exceed the maximum velocity constraint; then, it calculates the new position of the particle based on the updated velocity, and after updating the position, it compares the fitness of the particle with the current global optimum and updates it accordingly; the algorithm continues to iterate until the maximum number of iterations is reached, or the global optimum changes less than a set threshold in multiple consecutive iterations.
[0023] Optionally, in step 3, the task unloading decision is modeled as a multi-agent partially observable Markov decision process, in which agricultural terminal equipment and drones are regarded as agents with independent decision-making capabilities. The decision process is solved using a dual-delay deep deterministic policy gradient algorithm based on preference prediction and synchronous unloading. Global state information is introduced in the centralized training phase, and each agent makes decisions based only on local observations in the distributed execution phase.
[0024] Optionally, in step 3, a pre-processed preference network is introduced as a pre-network for deep reinforcement learning, and a suitable UAV is selected for synchronous unloading. While the UAV is processing other tasks, the ground terminal device can first calculate an appropriate amount of data for local unloading to avoid the local device from waiting indefinitely and wasting local resources. Then, another part of the tasks is unloaded to the UAV cache module for waiting. After the UAV completes the calculation of the previous task, the second task is unloaded until the system time ends, thus obtaining the resource allocation and unloading strategy of the entire network.
[0025] Optionally, during the execution of step 4, the agent outputs continuous unloading decision actions based on its current observation state. The decision actions include at least the task unloading ratio and the corresponding unloading target indication. The unloading ratio is used to determine the threshold to distinguish between complete unloading and partial unloading. When the unloading ratio is 0, the task is completely executed on the local device. When the unloading ratio is 1, the task is completely unloaded to the target drone. In other cases, partial unloading is performed.
[0026] Subsequently, based on the unloading target and unloading method, the system jointly allocates computing and communication resources, including allocating corresponding transmission power, communication bandwidth, and computing frequency for the unloading task. It also verifies and adjusts the resource allocation results based on the UAV's current load status and remaining processing time. Through continuous updating and storage of unloading decisions and resource allocation results within each time slot, a dynamic unloading strategy sequence and resource allocation scheme for the entire network runtime are ultimately formed to guide the collaborative execution of multiple UAVs and ground equipment.
[0027] This invention provides a task offloading optimization method for multi-UAV air-ground collaboration in smart agriculture. First, addressing the uneven distribution of crop equipment and limited computing power of ground terminals in agricultural production areas, a system model is constructed. With the optimization objectives of improving task completion rate and reducing system energy consumption, an improved particle swarm optimization algorithm is introduced to initialize the deployment of UAV access locations. Through adaptive inertia weights and dynamic learning factor adjustment mechanisms, rapid convergence and efficient search for optimal UAV coverage locations in farmland scenarios are achieved. Next, the agricultural task offloading process is modeled as a multi-agent partially observable Markov decision process. An improved dual-delay deep deterministic policy gradient algorithm is used to learn the collaborative offloading strategy between multiple UAVs and ground terminals, and an attention mechanism is introduced during the decision-making process to characterize different differential contributions. Simultaneously, an offloading prior guidance mechanism based on preference prediction is designed to pre-determine the priority distribution of each UAV as an offloading target. Combined with synchronous offloading constraints, this guides ground terminals to perform partial local computation and cache remaining tasks during busy UAV periods to avoid resource idleness and task blocking. Finally, based on the offloading targets and offloading methods, multiple resources are jointly allocated to generate a dynamic offloading and resource allocation strategy covering the entire system runtime. The above methods enable efficient response and stable execution of heterogeneous agricultural tasks in complex environments, improving the task computing rate and overall operational efficiency of the system in smart agriculture scenarios. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a flowchart illustrating a task offloading optimization method for multi-UAV air-to-ground collaboration in smart agriculture, as proposed in this invention.
[0030] Figure 2 This is a schematic diagram of the IPSO process of the method of the present invention.
[0031] Figure 3 This is the SOPM3 flowchart of the method of the present invention.
[0032] Figure 4 This is a schematic diagram comparing the total system energy consumption of the SOPM3 algorithm with other comparison algorithms under different task data sizes in this embodiment of the invention.
[0033] Figure 5 This is a schematic diagram comparing the task completion rates of the SOPM3 algorithm with other comparative algorithms under different task data sizes in an embodiment of the present invention. Detailed Implementation
[0034] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0035] The following explains the English abbreviations used in this invention:
[0036] IPSO: An improved particle swarm optimization algorithm;
[0037] SOPM3: A dual-delay deep deterministic policy gradient algorithm based on preference prediction and synchronous unloading.
[0038] Please see Figure 1 This invention provides a task offloading optimization method for multi-UAV air-to-ground collaboration in smart agriculture, comprising the following steps:
[0039] Step 1: In the agricultural production area scenario, construct the system model, the UAV motion model, the task model, the communication service model, and the computation model, and determine the optimization objective;
[0040] Step 2: Based on the distribution of ground crop information, an improved particle swarm optimization algorithm is used to initialize the deployment of UAVs and obtain the optimal access location;
[0041] Step 3: Based on the optimal initial deployment location, a dual-delay deep deterministic strategy gradient algorithm based on preference prediction and synchronous unloading is used to learn the optimal collaborative unloading decision.
[0042] Step 4: Determine the uninstallation target and uninstallation method, allocate resources, and generate the uninstallation strategy and resource allocation plan for the entire network time.
[0043] The following provides further explanation in conjunction with the specific implementation steps:
[0044] 1. Step 1 involves model building and optimization objective determination. Specifically, the entire system consists of m UAVs and k agricultural ground terminals. The agricultural ground terminals can generate tasks based on crop or environmental conditions. These tasks typically have different latency constraints and must be completed within a certain timeframe. When the computing power of the ground terminal equipment is insufficient to complete the task within the specified time or a better-efficiency method is available, the task will be offloaded to edge computing or computed locally. The task model expression is:
[0045] .in Indicates the task size. This indicates the specified completion time for the task. This indicates the number of cycles required to compute each bit of data. The entire system uses TDMA for multiple access to the wireless channel, and the total system duration is defined as... Divided into Each time slot is of equal length. Indicates a specific index and The duration of each time slot is .
[0046] Specifically, the propulsion power and energy consumption model of the UAV is expressed as follows:
[0047]
[0048]
[0049] 2. Based on the varying terrain of agricultural sites in different regions, a flexible communication channel model is adopted. Specifically, the UAV can receive tasks not only from ground equipment but also from other UAVs, thus possessing both ground-to-air and air-to-air communication modes. The UAV is equipped with multiple antennas, enabling simultaneous reception... Multiple tasks are performed, but only one task can be computed at a time. The transmission rate is shown by Shannon's formula (3).
[0050]
[0051] in , The reference distance is The channel gain at that time. As shown in Equation (4), for non-line-of-sight channels with a large number of obstacles, the widely used block fading channel is used in this invention to describe it.
[0052]
[0053] Where F is the CDF function , It is the SNR difference between the actual modulation scheme and the theoretical Gaussian signaling. It is the path loss index for large-scale fading. It is the Marcum-Q function. As shown in formula (5), the logistic regression model in statistics is used to describe the Los channel probability, and the parameters a and b can be adjusted to cope with agricultural scenarios with different terrains.
[0054]
[0055] The weighted air-to-ground transmission rate is expressed as in This represents the probability of a non-line-of-sight channel.
[0056] 3. Based on different offloading methods, computational energy consumption is divided into computational energy consumption of a single local device, computational energy consumption of a single UAV, and computational energy consumption of both local device and UAV simultaneously. The offloading ratio for partial offloading is represented as λt∈[0,1]. Transmission energy consumption is divided into energy consumption for transmission from the local device to the UAV and energy consumption for mutual forwarding between UAVs.
[0057] The expression for transmission energy consumption is: The expression for calculating energy consumption is: ,in This indicates the computing frequency of the drone or local device. This is expressed as the time of transmission or the time of computation.
[0058] 4. Thus, the energy consumption of the entire system is as shown in formula (6).
[0059] The overall goal of the system is to minimize the time required for all tasks to be completed within their allotted time. If the time from when a task is generated to when it ends is within the specified time, the task is considered complete; otherwise, it is considered incomplete.
[0060] Please see Figure 2 In step 2, based on ground crop information or the distribution of terminal devices, an improved particle swarm optimization algorithm is used to optimize the deployment of UAVs, aiming to minimize the access distance between the UAVs and all ground devices. Specifically, inertial weights and dynamic complementary factors are added to the basic particle swarm optimization algorithm, making it more flexible and avoiding getting trapped in local optima.
[0061] The inertia weight update formula is shown in equation (7).
[0062]
[0063] in The formula for the custom inertial step size adjustment and the group particle center is shown in equation (8).
[0064]
[0065] The complementary dynamic factor adjustment formulas are shown in formulas (9) and (10).
[0066]
[0067]
[0068] in This represents the maximum possible diversity of the group. .when The larger the value, the more dispersed the particle distribution, indicating an initial stage of exploration.
[0069] The execution process of steps 3 and 4 is as follows: Figure 3As shown, this process is executed after optimizing the deployment location of the UAV based on the IPSO algorithm. A dual-delay deep deterministic policy gradient algorithm based on preference prediction and synchronous offloading is used to complete the offloading decision and resource allocation. Based on this, the preference prediction network of the local terminal agent outputs a prior probability distribution about the target UAV for offloading based on the encoded state, used to characterize the relative offloading priority of different UAVs in the current time slot. Simultaneously, the actor network generates continuous actions based on the same state. Subsequently, the system combines the actions of all terminal devices and UAVs into a joint action and executes it synchronously in the environment. Based on the joint action, the environment calculates the transmission latency, computation latency, and corresponding energy consumption of the task, and determines whether the task enters the UAV buffer, triggers forwarding, or performs partial local computation based on the synchronous offloading mechanism. After executing the action, the environment returns the system-level reward value and the system state for the next time slot. After obtaining the interaction results, the interaction results are placed in a buffer pool. A batch of historical experiences is further randomly sampled from the experience pool for network updates. For each sampled data, the target actor network first generates the target action for the next time moment, and applies pruning noise to the target action to achieve policy smoothing. Subsequently, the target Q-value is calculated using the target critic network, and the smaller value between the two critic outputs is taken as the temporal difference objective. Then, the parameters of the two critic networks are updated by minimizing the mean squared error between the current critic output and the target Q-value. After the critic network is updated, the actor network is not updated immediately. Instead, according to a delayed update mechanism, the actor network is updated only based on the deterministic policy gradient when a preset update interval condition is met. The above-described time-slot-based interaction, storage, and update process continues within one round until the system time ends or the task termination condition is met. After training, the optimal unloading decision and resource allocation scheme are obtained to maximize the overall system benefit. This multi-agent reinforcement learning framework includes... Each agent continuously updates its strategy to approximate the optimal decision through constant interaction with the dynamic edge computing environment. For any agent... Internally, it maintains the following three core network structures: 1) Action network 1) Responsible for generating decision-making actions. 2) Dual-commentator network and 3) Preference prediction network Based on the current state, the prior preference information for task unloading is output to guide the exploration direction in the high-dimensional hybrid action space. The specific update process can be represented as follows:
[0070] The sample batch size in the experience pool is denoted as B. The loss function of the agent's dual-critic network is shown in Equation (11).
[0071]
[0072] in Let represent the action of each agent. The action network is updated using the deterministic policy gradient formula, as shown in Equation (12).
[0073]
[0074] The target network soft update is shown in equation (13).
[0075]
[0076]
[0077]
[0078] in To update the coefficients, Represents an actor network. Indicates the critic network. This indicates a preference for certain networks.
[0079] Furthermore, by comparing with several relevant baseline algorithms, the present invention further demonstrates its beneficial effects:
[0080] The relevant baseline algorithms are described below:
[0081] 1) Multi-Agent Deep Deterministic Policy Gradient (MADDPG): Employs a centralized Critic and a single Q network.
[0082] 2) Cooperation without long-term optimization (CNL): This method uses a greedy approximation algorithm based on mathematics to solve the problem without considering long-term benefits.
[0083] 3) Computational Capacity Optimal Policy (CCOP): All tasks are offloaded to the drone, while considering long-term benefits.
[0084] 4) Local only (LO): The task is executed entirely on the agricultural terminal.
[0085] The relevant experimental parameters are shown in Table 1:
[0086] Table 1 Experimental parameters
[0087]
[0088] See Table 2 below and Figure 4 Figure 5 Under different task data generated, the method proposed in this invention improves both the total system energy consumption and task completion rate compared to the baseline method.
[0089] Table 2 Comparison of energy consumption of the method of this invention with other baseline methods under different task data sizes.
[0090]
[0091] Figure 4 and Figure 5 Tasks are randomly generated according to a Poisson distribution. If all generated tasks are completed within the specified tolerance time after the entire time slot has elapsed, then the task is considered complete, denoted as... Figure 5 The success rate of [previous method] is [not specified], while the method proposed in this invention maintains a 100% task completion rate under the set parameters. During this period, the sum of the UAV's flight energy consumption, the system's computational energy consumption, and transmission energy consumption is expressed as [formula missing]. Figure 4 In terms of total energy consumption, the method proposed in this invention reduces energy consumption by an average of 1.8% compared to other best-performing methods under different task data sizes. The results demonstrate the effectiveness of the proposed method.
[0092] In summary, this invention addresses the challenges of smart agricultural production scenarios, specifically the highly heterogeneous task types and latency requirements of collaborative operations involving multiple ground terminals and drones within farmland areas. It proposes a joint offloading optimization method involving multiple drones and ground terminal equipment. In an agricultural wireless edge computing environment, this method comprehensively considers factors such as crop information distribution and limited terminal device computing power. It quickly and efficiently completes drone access location selection, task offloading decisions, and the generation of time-sequential computation and communication resource allocation schemes, thereby improving the efficiency of agricultural task completion within limited latency constraints and enhancing the overall system computing power. Through multi-drone collaborative deployment and joint decision-making mechanisms, this invention effectively alleviates performance bottlenecks caused by distance differences and uneven load distribution, avoiding common coverage blind spots and service capability fluctuations in agricultural scenarios. Furthermore, compared to existing general multi-drone offloading methods, this invention further integrates the latency-sensitive characteristics of agricultural tasks in applications such as crop monitoring, pest and disease control, and environmental perception. It differentiates the timeliness requirements of different tasks through modeling and decision-making, effectively reducing the probability of task timeouts and deadline breaches, and improving the overall computational completion rate of various agricultural tasks.
[0093] The above description discloses only one preferred embodiment of the present invention, and should not be construed as limiting the scope of the present invention. Those skilled in the art will understand that all or part of the processes of the above embodiments can be implemented, and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.
Claims
1. A task offloading optimization method for multi-UAV air-to-ground collaboration in smart agriculture, characterized in that, Includes the following steps: Step 1: In the agricultural production area scenario, construct a system containing a set of ground terminals. With drones The system model is used to establish a task model, a communication service model, a computation model, and a motion model of the UAV, thereby obtaining the optimization objective; Step 2: Construct a deployment fitness function based on the crop information distribution density function, and use an improved particle swarm optimization algorithm that includes adaptive inertia weights and dynamic learning factor adjustment mechanisms to optimize the initial access position of the UAV; Step 3: Based on the optimal initial deployment location, a dual-delay deep deterministic strategy gradient algorithm based on preference prediction and synchronous unloading is used to learn the optimal collaborative unloading decision. Step 4: Determine the uninstallation target and uninstallation method, allocate resources, and generate the uninstallation strategy and resource allocation plan for the entire network time.
2. The task offloading optimization method for multi-UAV air-to-ground collaboration in smart agriculture as described in claim 1, characterized in that, The execution process of step 1 includes the following steps: Step 1.1: Incorporate crop information and ground terminal equipment information into the modeling process to construct a system model; Step 1.2: Construct the motion model of the UAV and, based on the time delay constraints of the mission types generated in different time slots, construct the mission model; Step 1.3: Construct communication service models for different situations based on communication conditions such as obstacles or terrain; Step 1.4: Construct computing models based on different computing service carriers; Step 1.5: Based on the established models, clarify the optimization objective, namely, to minimize the total energy consumption of the system while ensuring that all tasks are completed within the specified time delay.
3. The task offloading optimization method for multi-UAV air-to-ground collaboration in smart agriculture as described in claim 2, characterized in that, The improved particle swarm optimization algorithm described in step 2 aims to minimize the communication distance between the agricultural terminal equipment and the nearest drone, and introduces a safe distance constraint for drones during the optimization process to avoid flight conflicts between drones in low-altitude agricultural operations.
4. The task offloading optimization method for multi-UAV air-to-ground collaboration in smart agriculture as described in claim 3, characterized in that, The execution process of step 2 includes the following steps: Step 2.1: Represent the candidate access positions of the UAV in the airspace as particle position vectors, and introduce crop information distribution as an important constraint and guiding factor for deployment optimization; Step 2.2: By mapping the density of crop information, the intensity of data collection demand, and the distribution of ground equipment within the farmland area to the particle fitness evaluation index, the fitness value for each round is calculated; the inertia weight is dynamically increased or decreased based on the change in fitness value, thereby avoiding getting trapped in local optima or improving local fine search. Step 2.3: Calculate the center position of the current particle swarm, and calculate the swarm diversity index based on the distance of all particles to the swarm center. Dynamically adjust the individual learning factor and social learning factor to accelerate the approach to the optimal solution. Step 2.4: After updating the inertia weights and learning factors, the algorithm calculates the new velocity of each particle in the current iteration according to the velocity update formula, and truncates the components that exceed the maximum velocity constraint. The new position of the particle is then calculated based on the updated velocity. After the position update is completed, the particle's fitness is compared with the current global optimum and updated accordingly. The algorithm continues to iterate. The process continues until the maximum number of iterations is reached, or the global optimal solution changes less than a set threshold in multiple consecutive iterations.
5. The task offloading optimization method for multi-UAV air-to-ground collaboration in smart agriculture as described in claim 4, characterized in that, In step 3, the task unloading decision is modeled as a multi-agent partially observable Markov decision process, in which agricultural terminal equipment and drones are regarded as agents with independent decision-making capabilities. The decision process is solved using a dual-delay deep deterministic policy gradient algorithm based on preference prediction and synchronous unloading. Global state information is introduced in the centralized training phase, and each agent makes decisions based only on local observations in the distributed execution phase.
6. The task offloading optimization method for multi-UAV air-to-ground collaboration in smart agriculture as described in claim 5, characterized in that, The dual-delay deep deterministic strategy gradient algorithm based on preference prediction and synchronous unloading introduces an attention mechanism to characterize the collaborative relationship between different drones and the farm edge computing center. Based on the unloading decision, a synchronous unloading mechanism is introduced according to the status of the terminal equipment and drone equipment to maximize resource utilization while reducing energy consumption and improving unloading efficiency.
7. The task offloading optimization method for multi-UAV air-to-ground collaboration in smart agriculture as described in claim 6, characterized in that, In step 3, a pre-processed preference network is introduced as a pre-network for deep reinforcement learning, and a suitable UAV is selected for synchronous unloading. While the UAV is processing other tasks, the ground terminal device can first calculate an appropriate amount of data for local unloading to avoid the local device from waiting indefinitely and wasting local resources. Then, another part of the tasks is unloaded to the UAV cache module for waiting. After the UAV completes the calculation of the previous task, the second task is unloaded until the system time ends, thus obtaining the resource allocation and unloading strategy of the entire network.
8. The task offloading optimization method for multi-UAV air-to-ground collaboration in smart agriculture as described in claim 7, characterized in that, During the execution of step 4, the agent outputs continuous unloading decision actions based on its current observation state. The decision actions include at least the task unloading ratio and the corresponding unloading target indication. The unloading ratio is used to distinguish between complete unloading and partial unloading. When the unloading ratio is 0, the task is completely executed on the local device. When the unloading ratio is 1, the task is completely unloaded to the target drone. In other cases, partial unloading is performed.