Unmanned aerial vehicle plant protection path optimization system based on reinforcement learning
By constructing a digital farmland environment model and a reinforcement learning strategy decision-making module, and combining real-time environmental perception adjustment and path execution feedback optimization, the adaptability and efficiency issues of UAV plant protection path planning in complex farmland environments were solved, achieving highly efficient path optimization results.
Patent Information
- Application Number
- CN202511679265.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-02-10
AI Technical Summary
Existing drone-based plant protection path planning technology has poor adaptability in complex and ever-changing farmland environments, lacks real-time dynamic adjustment capabilities, struggles to achieve the optimal balance between control effectiveness, energy efficiency, and operation time, suffers from low computational efficiency, insufficient environmental modeling, and a lack of learning ability.
A digital farmland environment model is constructed, employing a reinforcement learning strategy decision-making module and a real-time environment perception and adjustment module, combined with a near-end strategy optimization algorithm, and forming a closed-loop optimization system through a path execution feedback optimization module to achieve intelligent path planning and real-time optimization.
It improves the accuracy and adaptability of path planning, saves 25% to 35% of flight time, reduces energy consumption by 20%, and increases the utilization rate of liquid medicine by 30%, achieving the optimal balance between prevention and control effect, energy efficiency and operation time.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent agriculture technology, specifically to a reinforcement learning-based drone plant protection path optimization system, which is applied to intelligent path planning and real-time optimization of agricultural plant protection drones. Background Technology
[0002] With the rapid development of precision agriculture, drone-based plant protection technology has become an important means of pest and disease control in modern agriculture. By carrying spraying equipment, plant protection drones can achieve efficient operations on large areas of farmland, offering advantages over traditional ground-based pesticide application, such as higher operational efficiency, higher pesticide utilization, and less damage to crops. Path planning, as the core technology of plant protection drone operations, directly affects the control effect, energy consumption, and operation time.
[0003] Existing drone-based plant protection path planning technologies are mainly based on traditional algorithms. For example, Chinese patent application CN120450179A discloses a drone-based pest monitoring method based on intelligent path planning. This method divides the monitoring area into multiple sub-regions, calculates a disaster index by detecting meteorological data, lighting conditions, and historical pest infestation data, and classifies the area into key and secondary areas based on the disaster index. Then, it uses the A algorithm for path planning. This method adds a priority weight parameter to the A algorithm, setting a lower path cost for key areas, thus allowing the drone to prioritize flying to high-risk areas.
[0004] However, the aforementioned existing technologies have the following shortcomings: First, while the A* algorithm, as a traditional heuristic search algorithm, can find the shortest path, it relies on a pre-defined heuristic function and cost function, making it poorly adaptable to complex and ever-changing farmland environments. It cannot dynamically adjust decision-making strategies based on real-time environmental changes, making it difficult to achieve an optimal balance between control effectiveness, energy efficiency, and operation time. Second, existing methods have relatively simple environmental modeling, mainly based on static meteorological data and historical statistical information. They lack comprehensive modeling of dynamic factors such as farmland topography, crop distribution, and real-time spread of pests and diseases, resulting in insufficient accuracy and adaptability in path planning. Third, existing methods employ fixed regional divisions and path planning strategies. Although they can update monitoring routes during flight, they are essentially rule-based replanning, lacking the ability to learn and optimize from historical experience, and unable to quickly adapt and continuously improve in different farmland scenarios. Fourth, traditional algorithms suffer from computational bottlenecks in large-scale, complex terrain control areas, making real-time global optimization difficult and affecting the overall efficiency of UAV operations.
[0005] Therefore, there is a need for a drone-based plant protection path planning technology that can adaptively learn the optimal flight strategy, comprehensively consider multi-dimensional environmental factors, and achieve dynamic real-time optimization, in order to overcome the above-mentioned shortcomings of existing technologies. Summary of the Invention
[0006] To address the problems of existing technologies, this invention provides a reinforcement learning-based drone-based agricultural drone path optimization system, aiming to solve the technical problems of poor adaptability, insufficient environmental modeling, lack of learning ability, and low computational efficiency of traditional path planning algorithms. This invention constructs a digital farmland environment model, trains a path decision network using a proximal policy optimization algorithm, and combines hierarchical Markov decision process modeling and environmental perception enhancement mechanisms to achieve intelligent planning and real-time optimization of agricultural drone paths, achieving an optimal balance between control effectiveness, energy efficiency, and operation time.
[0007] To achieve the above-mentioned objectives, the present invention provides the following technical solution: A reinforcement learning-based drone-based agricultural path optimization system includes: The digital farmland environment modeling module constructs a digital farmland environment model that includes information on topography, crops, and pests and diseases. It represents the control area as a grid structure, with each grid cell containing elevation, crop type, pest and disease density, and environmental parameters, and outputs an environmental state vector. The reinforcement learning strategy decision-making module is connected to the digital farmland environment modeling module. It constructs a path decision network based on the near-end strategy optimization algorithm, receives the environmental state vector, and outputs flight action decisions through the strategy network. The flight action decisions include flight direction, flight speed, and spraying control. The flight action benefits are evaluated according to the reward function, which comprehensively considers the prevention and control coverage, energy consumption, and operation time. The real-time environmental perception and adjustment module is connected to the reinforcement learning strategy decision module. It dynamically adjusts flight action decisions based on environmental change information monitored in real time by airborne sensors. The environmental change information includes wind direction changes, remaining drug solution, and newly discovered disease areas. The environmental change information is fused with the output of the strategy network to generate adjusted flight commands. The path execution feedback optimization module is connected to the real-time environment perception adjustment module and the reinforcement learning strategy decision module. It collects actual operation effect data after executing flight commands. The actual operation effect data includes actual prevention and control coverage, energy consumption and operation time. It calculates the actual reward value based on the actual operation effect data and feeds the actual reward value back to the reinforcement learning strategy decision module to update the parameters of the policy network.
[0008] The beneficial effects of this invention are as follows: First, a high-fidelity digital farmland environment model is constructed through the digital farmland environment modeling module, which integrates terrain data, crop distribution and pest and disease monitoring information to provide a comprehensive and accurate representation of the environmental state for reinforcement learning algorithms. This overcomes the shortcomings of simple environmental modeling in existing technologies and improves the accuracy of path planning.
[0009] Second, a reinforcement learning strategy decision-making module based on a near-end strategy optimization algorithm is adopted. Through continuous interaction with the environment, it autonomously learns the optimal flight strategy. Compared with traditional heuristic methods such as the A* algorithm, it has stronger adaptability and optimization capabilities and can automatically find the optimal balance between prevention and control effect, energy efficiency and operation time.
[0010] Third, the real-time environmental perception and adjustment module enables real-time response to dynamic environmental factors such as wind direction changes, remaining pesticide solution, and newly discovered disease areas. It can dynamically adjust the flight plan according to the actual situation, and has stronger adaptability and robustness compared to the fixed path scheme.
[0011] Fourth, the path execution feedback optimization module collects actual operation effect data and feeds it back to the strategy decision module, forming a complete closed-loop optimization mechanism. This enables the system to continuously learn and improve from actual operation experience, thereby achieving continuous performance improvement.
[0012] Fifth, the four modules form a deeply coupled collaborative system. The output of the digital farmland environment modeling module directly drives the decision-making process of the reinforcement learning strategy decision-making module. The output of the strategy decision-making module guides the adjustment strategy of the real-time environment perception adjustment module. The adjusted flight command is evaluated by the path execution feedback optimization module and then influences the parameter update of the strategy decision-making module. This realizes mutual promotion and superposition effect among the modules, and the overall system performance shows a non-linear growth characteristic of 1+1>2.
[0013] In validation across multiple crop types, compared to traditional grid path planning, this invention saves 25% to 35% of flight time, reduces energy consumption by 20%, and increases pesticide utilization by 30%, providing an efficient and reliable path optimization solution for intelligent plant protection. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the overall structure of the system of the present invention; Figure 2 This is a schematic diagram of the network structure of the reinforcement learning strategy decision-making module of the present invention; Figure 3 This is a schematic diagram of the path optimization process of the present invention. Detailed Implementation
[0015] Please refer to the attached document. Figures 1-3 The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0016] Reference Figure 1 This invention provides a reinforcement learning-based UAV agricultural path optimization system, comprising a digital farmland environment modeling module 1, a reinforcement learning strategy decision-making module 2, a real-time environmental perception and adjustment module 3, and a path execution feedback optimization module 4. These four modules form a deeply coupled collaborative system, achieving closed-loop optimization through forward data transmission and backward feedback mechanisms.
[0017] The Digital Farmland Environment Modeling Module 1 is used to construct a high-fidelity digital farmland environment model, providing an accurate representation of the environmental state for reinforcement learning algorithms. This module integrates multi-source data, including topographic data, crop distribution information, and pest and disease monitoring information, representing the control area as a multi-level grid structure.
[0018] Specifically, the digital farmland environment modeling module 1 first acquires the geographic information of the control area. Topographic data is obtained through a high-precision digital elevation model with a resolution of 1m × 1m, containing the altitude, slope, and aspect information of each grid cell. Crop distribution information is acquired through multispectral remote sensing imagery, enabling the identification of crop type, growth stage, and growth status. Pest and disease monitoring information comes from historical monitoring records and real-time monitoring equipment, including the frequency, severity, and spread trends of pests and diseases in each area.
[0019] Digital farmland environment modeling module 1 divides the prevention and control area into The grid structure has each grid cell measuring 5m x 5m. Each grid cell... It contains multiple state attributes, represented in vector form: , in, The average elevation of the grid cells. Encode crop types, This is the pest and disease density index. For wind direction angle, For wind speed, For temperature, Humidity. Crop type codes are represented by integers, such as 1 for rice, 2 for wheat, and 3 for corn. The pest and disease density index is calculated by combining historical monitoring data and real-time detection results, with a value range of 0 to 1, where 0 indicates no pests or diseases and 1 indicates extremely severe pests and diseases.
[0020] Digital farmland environment modeling module 1 further constructs environmental state vectors. This information is used as input to the reinforcement learning policy decision module 2. The environment state vector contains the current location information of the UAV, the remaining drug level, the battery level, and the state information of the surrounding grid cells. , in, , , For drones at all times The three-dimensional coordinate position, This represents the percentage of remaining medication. This represents the remaining battery charge percentage. This is a local grid state matrix centered on the current location of the UAV, including the surrounding... The state information of each grid cell, in this embodiment The value is 7, which means extracting the local environment state of 7×7.
[0021] The digital farmland environment modeling module 1 uses a sliding window mechanism to dynamically update the environmental state. As the drone moves, the local mesh state matrix... Real-time updates ensure that policy decisions are always based on the latest environmental information. Preferably, the environmental state vector is updated at a frequency of 10Hz, that is, once every 0.1s, which ensures real-time performance while avoiding excessive computational burden.
[0022] To improve the accuracy of the environmental model, the digital farmland environmental modeling module 1 introduces a dynamic environmental modeling mechanism. The pest and disease density index is dynamically updated based on a diffusion model that considers the impact of meteorological factors such as wind direction, wind speed, and temperature on pest and disease spread. Specifically, the pest and disease density index is updated dynamically at a time step. The subsequent update uses the diffusion equation: , in, This represents the natural spread coefficient of pests and diseases. Let Laplace operator represent the space diffusion term. Meteorological influencing factors, related to wind direction Wind speed ,temperature Related, For grid cells The neighborhood set. Meteorological influence factors are obtained by fitting historical data, preferably in the following form: , in, , , , The fitting coefficients are determined based on different crops and pest types. In this example, for rice planthoppers, the fitting coefficients are 0.02, 0.15, 0.008, and 0.003, respectively.
[0023] The environmental state vector output by the digital farmland environment modeling module 1 serves as the input to the reinforcement learning policy decision-making module 2, driving the subsequent path decision-making process. By constructing a high-fidelity digital farmland environment model, this invention overcomes the shortcomings of existing technologies in terms of simple environmental modeling, providing accurate and reliable state information for path optimization.
[0024] The reinforcement learning policy decision-making module 2 is the core module of this invention. It constructs a path decision network based on a proximal policy optimization algorithm and learns the optimal flight strategy through continuous interaction with the environment. This module receives the environmental state vector output by the digital farmland environment modeling module 1 and outputs the UAV's flight action decisions through the policy network.
[0025] Reference Figure 2 The path decision network adopts an Actor-Critic architecture, consisting of a policy network (Actor) and a value network (Critic). The policy network is responsible for outputting the probability distribution of actions based on the current state, while the value network is responsible for evaluating the value of the state. The two networks share information by sharing a low-level feature extraction layer, thereby improving learning efficiency.
[0026] The input to the policy network is the environment state vector. The output is the action probability distribution. The network structure employs a fully connected neural network, comprising a feature extraction layer, a policy output layer, and a value output layer. The feature extraction layer contains three hidden layers with 512, 256, and 128 neurons respectively, using the ReLU activation function. The policy output layer uses the Softmax activation function and outputs the probability distribution of the discrete action space. The value output layer outputs a scalar value representing the value estimate of the current state.
[0027] Flight maneuver decisions include flight direction, flight speed, and spray control commands. Flight direction is discretized into eight directions: east, southeast, south, southwest, west, northwest, north, and northeast, coded from 0 to 7. Flight speed is discretized into three levels: low speed (3 m / s), medium speed (5 m / s), and high speed (7 m / s), coded from 0 to 2. Spray control commands are binary decisions: 0 indicates spraying off, and 1 indicates spraying on. Therefore, the action space dimension is... The policy network outputs a 48-dimensional probability vector.
[0028] The reinforcement learning policy decision-making module 2 employs the proximal policy optimization algorithm to train the policy network. The proximal policy optimization algorithm is a policy gradient method that improves sample utilization efficiency while ensuring training stability by limiting the magnitude of policy updates. The core idea of this algorithm is to prevent performance degradation caused by excessively large policy updates by truncating the importance sampling ratio.
[0029] The optimization objective of the policy network is to maximize the expected cumulative reward. At each time step... Drones performing actions Afterwards, the environment returns an immediate reward. And transition to the new state Accumulated rewards are defined as the sum of discounts on future rewards: , in, This is a discount factor, ranging from 0 to 1. In this embodiment, it is set to 0.99, indicating the importance attached to future rewards.
[0030] The near-end policy optimization algorithm updates policy parameters by maximizing the following objective function. : , in, The probability ratio between the old and new strategies. , For the new strategy in the state Select action The probability, For the old strategy in state Select action The probability, Let be the dominance function, representing the state . Next action Advantages relative to the average level For the truncation function, Limited to Within the interval, To truncate the parameter, the value is set to 0.2 in this embodiment.
[0031] Advantage function Value network estimation: , in, For value network state The value estimate, These are the parameters of the value network. The value network is trained by minimizing the temporal difference error: , The reward function is a key design element of the reinforcement learning policy decision-making module 2, directly affecting the quality of the learned policy. This invention designs a composite reward function that comprehensively considers prevention and control coverage, energy consumption, and operation time: , in, Incentives for comprehensive prevention and control coverage As an energy efficiency reward, Rewards for time efficiency As a penalty item, , , , These are weighting coefficients, with values of 0.5, 0.2, 0.2, and 0.1 respectively, which can be adjusted according to actual application requirements.
[0032] Prevention and control coverage reward Calculated based on the area and density of pests and diseases covered by the drone spraying operation. If the drone is at a certain time... If you are spraying pesticides and located in an area infested with pests and diseases, you will receive a positive reward: , in, To cover the reward coefficient, a value of 10 is used. This represents the pest and disease density index of the current grid cell. The effective coverage area for a single spray is determined based on the nozzle parameters; in this embodiment, it is 25m. 2 .
[0033] Energy efficiency rewards We encourage drones to adopt energy-saving flight methods. Energy consumption mainly includes flight energy consumption and spraying energy consumption. Flight energy consumption is related to flight speed, while spraying energy consumption is related to spraying flow rate. , in, For a moment Flight speed, In spraying mode, the value is 0 or 1. This is the flight energy consumption coefficient, with a value of 0.01. The spraying energy consumption coefficient is set to 0.5. Flight speed and energy consumption have a square relationship, consistent with the physical laws of air resistance.
[0034] Time efficiency reward Encourage drones to complete tasks quickly. Impose a fixed time cost penalty at each time step to incentivize drones to optimize their paths and reduce task time. , in, This is the time cost coefficient, with a value of 0.1. The time step is 0.1s in this embodiment.
[0035] Penalty items Used to constrain unsafe drone behavior, including boundary crossing penalties, collision penalties, and duplicate coverage penalties: , in, The number of times the drone sprays the area again after it has already been sprayed is the number of times it has been used to cover the area.
[0036] Reinforcement learning strategy decision-making module 2 employs an experience replay mechanism to improve sample utilization efficiency. During the UAV's mission execution, the state transitions at each time step... Data is stored in an experience pool. During training, batches of data are randomly sampled from the experience pool for gradient updates, breaking the temporal correlation between samples and improving training stability. The experience pool capacity is set to 100,000 experience records, and the batch size is set to 256.
[0037] Module 2 of the reinforcement learning strategy decision-making module also employs a target network mechanism to stabilize the training process. The target value calculation of the value network uses a separate target network. The parameters of the target network are copied and updated from the value network at regular intervals; in this embodiment, this is set to be updated every 1000 steps. This delayed update mechanism reduces the fluctuation of value estimation and improves training stability.
[0038] To accelerate the training process, reinforcement learning policy decision-making module 2 employs a parallel simulation environment. The system runs multiple independent simulation environment instances simultaneously, each executing a different policy sampling trajectory, and the experience from all instances is stored uniformly in an experience pool. In this embodiment, eight parallel environments are set up, increasing the training data acquisition speed by eight times and significantly shortening the training time.
[0039] The flight action decisions output by reinforcement learning policy decision module 2 are transmitted to real-time environment perception and adjustment module 3 to generate actual flight control commands. Through deep reinforcement learning based on a proximal policy optimization algorithm, this invention achieves the ability to adaptively learn the optimal flight strategy, exhibiting stronger environmental adaptability and optimization capabilities compared to the traditional A* algorithm.
[0040] The real-time environment perception and adjustment module 3 dynamically adjusts the flight action decisions output by the reinforcement learning strategy decision-making module 2 based on real-time environmental change information monitored by the UAV's onboard sensors, generating the final flight commands. This module is key to the real-time adaptability of this invention, enabling it to cope with dynamic changes in the farmland environment.
[0041] The real-time environmental perception and adjustment module 3 receives three types of environmental change information: wind direction changes, remaining pesticide solution levels, and newly discovered diseased areas. This information is collected in real time by sensors carried by the UAV, including wind speed and direction sensors, pesticide solution flow sensors, and multispectral cameras.
[0042] Wind direction changes significantly impact drone flight and pesticide spraying effectiveness. When a significant change in wind direction or speed is detected, the real-time environmental perception and adjustment module 3 adjusts the flight direction and spraying timing. Specifically, the module maintains a historical wind direction queue of length 10, storing wind direction and speed data for the most recent 10 time steps. Wind direction stability is determined by calculating the wind direction variance. , in, For the first From a historical perspective The average wind direction angle, For wind direction variance. When At that time, the wind direction was determined to be unstable, among which The wind direction stability threshold is set at 25 degrees. 2 .
[0043] In cases of unstable wind direction, the real-time environmental perception adjustment module 3 adopts a conservative strategy, reducing flight speed and pausing spraying until the wind direction stabilizes. The adjustment rules are as follows: , in, For the adjusted flight maneuvers, To maintain the current flight direction, This is the low speed setting; 0 indicates spraying is off. The raw flight maneuvers output by the strategy decision-making module.
[0044] Monitoring the remaining liquid pesticide solution is a crucial step in ensuring operational continuity. The real-time environmental sensing and adjustment module 3 continuously monitors the percentage of remaining liquid pesticide solution. ,when At that time, among them The minimum medicine threshold is set at 15%. The module automatically plans the return path, guiding the drone to the nearest resupply point for medicine replenishment. The return path is a straight line to ensure reaching the resupply point with the shortest distance.
[0045] Newly discovered disease areas are detected in real time using an airborne multispectral camera. The camera collects multispectral reflectance data of the crop canopy, and abnormal areas are identified using the normalized difference vegetation index (NDI) and the spectral characteristics of pests and diseases. When a new disease area is detected, the real-time environmental perception and adjustment module 3 feeds back the location information of the area and the estimated pest and disease density to the digital farmland environment modeling module 1 to update the environmental model. Simultaneously, the module assesses the urgency of the new disease area; if the pest and disease density exceeds an emergency threshold... If the value is 0.8, the flight path will be adjusted to prioritize the area for prevention and control. , in, This indicates the generation of coordinates for traveling to the new diseased area. Flight maneuver sequence, This represents the estimated pest and disease density in newly affected areas.
[0046] The real-time environmental perception and adjustment module 3 employs a fusion strategy to integrate environmental perception information with policy decision outputs. This fusion strategy uses a priority mechanism, processing adjustment requests in the order of safety, urgency, and optimization. Safety adjustments, including obstacle avoidance and boundary constraints, have the highest priority. Urgent adjustments, including pesticide replenishment and emergency disease area response, have the second highest priority. Optimization adjustments, including wind direction adaptation and path fine-tuning, have the lowest priority.
[0047] The adjusted flight commands output by the real-time environment perception and adjustment module 3 are sent to the UAV flight control system for execution, and simultaneously transmitted to the path execution feedback optimization module 4 for effect evaluation. Through the real-time environment perception and adjustment mechanism, this invention achieves rapid response to dynamic environmental changes, significantly improving the system's adaptability and robustness.
[0048] The path execution feedback optimization module 4 collects the actual operation effect data after the UAV executes flight commands, calculates the actual reward value, and feeds the reward value back to the reinforcement learning strategy decision module 2 to update the strategy network parameters, forming a complete closed-loop optimization mechanism.
[0049] The path execution feedback optimization module 4 collects three types of actual operation effect data: actual control coverage rate, actual energy consumption, and actual operation time. The actual control coverage rate is calculated by comparing the pest and disease density distribution before and after the operation. The initial pest and disease density distribution before the operation is obtained through historical monitoring data and real-time detection. After the operation, the final distribution of pest and disease density was obtained through delayed monitoring. The control coverage rate is defined as the proportion of the area where pest and disease density has decreased to the total area affected by pests and diseases. , in, The threshold for determining pest and disease density is set at 0.3. The area of a single grid cell is 25m² in this embodiment. 2 , This is an indicator function that takes the value 1 when the condition is met, and 0 otherwise.
[0050] Actual energy consumption is calculated using changes in battery power recorded by the drone's battery management system and the amount of medicine consumed recorded by the medicine flow sensor. Total energy consumption is expressed as a weighted sum of electrical energy consumption and medicine consumption: , in, The unit for electrical energy consumption during flight is watt-hour (Wh). This represents the total amount of medicine consumed, expressed in liters (L). The conversion factor for the cost of the liquid medicine is set at 50Wh / L, which converts the consumption of the liquid medicine into equivalent energy cost.
[0051] Actual working time The entire operation is recorded directly via a timer, from the moment the drone takes off until it returns to the resupply point after completing its pest control mission. The operation time includes flight time, spraying time, and resupply time.
[0052] The path execution feedback optimization module 4 calculates the actual reward value based on the collected actual operation performance data. The actual reward function uses the same evaluation criteria as the reward function defined in the reinforcement learning strategy decision module 2, but is calculated based on real measurement data. , Among them, the weighting coefficient , , Consistent with the training phase, the values were 0.5, 0.2, and 0.2 respectively. To ensure comparability of values across different dimensions, all data were normalized. Prevention and control coverage is now expressed as a percentage from 0 to 1; energy consumption is normalized to energy consumption per unit area; and operating time is normalized to operating time per unit area. , in, The total area to be controlled is expressed in square meters.
[0053] The path execution feedback optimization module 4 feeds back the calculated actual reward value to the reinforcement learning strategy decision-making module 2. The feedback mechanism uses an incremental update method, incorporating practical experience... High-value samples are stored in the experience pool, where This represents the initial state at the start of the task. For a complete sequence of actions, This represents the final state at the end of the task. High-value samples have a higher sampling probability during training, prompting the policy network to focus on learning from real-world task experience.
[0054] To quantify the performance improvement of the strategy, the path execution feedback optimization module 4 maintains historical performance records, recording the prevention and control coverage, energy consumption, and operation time for each task. The learning effect of the strategy network is evaluated by comparing the performance metrics of multiple consecutive tasks. If the average performance of the most recent 10 tasks improves by more than 5% compared to the previous 10 tasks, the strategy optimization is deemed effective, and the current learning strategy continues. If the performance improvement is less than 5% or declines, the learning rate is adjusted or the exploration ratio is increased to avoid the strategy getting trapped in local optima.
[0055] The path execution feedback optimization module 4 also supports an online learning mode. During the actual deployment phase, the system continuously collects job data and updates the policy network, enabling continuous learning and improvement from practical experience. Online learning employs mini-batch gradient updates, with the learning rate set to 1 / 10 of that used in the offline training phase, ensuring that the policy gradually adapts to the new environment while maintaining stability.
[0056] To ensure the security of online learning, the path execution feedback optimization module 4 adopts a conservative update strategy. After each parameter update, the performance of the new strategy is verified in a simulation environment. The new strategy is only officially deployed if its performance is superior to the old one. If the performance of the new strategy deteriorates, it reverts to the old strategy parameters to avoid job failures due to strategy degradation.
[0057] Through the closed-loop feedback mechanism established by the path execution feedback optimization module 4, this invention achieves continuous learning and optimization from actual operational experience, enabling the system performance to continuously improve with the increase in the number of operations. This mechanism is a significant advantage of this invention compared to existing technologies, overcoming the limitation of traditional algorithms that cannot learn and improve from experience.
[0058] The four modules of this invention form a deeply coupled collaborative system, achieving closed-loop optimization through forward data transmission and backward feedback mechanisms. The coupling relationships and synergistic effects among the modules are as follows: First, the output environment state vector of the digital farmland environment modeling module 1 is directly used as the input of the reinforcement learning policy decision-making module 2, forming a deep parameter-level coupling between the two. The accuracy of the environment state directly affects the quality of policy decision-making. A high-fidelity environment model provides a reliable state representation for policy learning, enabling the policy network to learn a more accurate state-action mapping relationship. This coupling relationship achieves synergistic effects between environmental perception and decision-making reasoning. Compared with independent environment modeling and path planning, the coordinated system improves the quality of path optimization by about 30%.
[0059] Second, the flight action decisions output by the reinforcement learning policy decision-making module 2 are transmitted to the real-time environment perception and adjustment module 3. The latter adjusts the actions based on real-time environmental changes and generates the final flight command. A deep logical coupling is formed between the two: policy decision-making provides the basic action plan, while environment perception and adjustment provide safety and adaptability guarantees. This coupling achieves complementary advantages between global optimization and local adjustment. The policy network focuses on global path planning to maximize long-term returns, while the adjustment module handles short-term emergencies and safety constraints. Their collaborative work enables the system to possess both global optimality and local flexibility.
[0060] Third, the process of the real-time environmental perception adjustment module 3 executing flight commands is monitored throughout by the path execution feedback optimization module 4, which collects actual operational performance data and calculates actual reward values. A deep state-level coupling is formed between the two; the output state of the execution module directly determines the evaluation result of the feedback module, and the quality of feedback depends on the accurate measurement of the execution effect. This coupling relationship ensures the authenticity and timeliness of the feedback information, providing reliable learning signals for strategy optimization.
[0061] Fourth, the path execution feedback optimization module 4 feeds back the actual reward value to the reinforcement learning strategy decision-making module 2, forming a complete closed-loop feedback mechanism. This is a key link in the backward feedback, realizing the reverse influence from the execution result to the decision-making strategy. Through this feedback path, successful experiences and lessons learned in actual operations are integrated into the parameter updates of the policy network, enabling continuous policy optimization. This closed-loop feedback mechanism realizes a complete optimization cycle of forward transmission → performance evaluation → backward feedback → parameter adjustment, which is the core mechanism for achieving continuous learning capability in this invention.
[0062] Fifth, the information on newly identified diseased areas detected by the real-time environmental perception and adjustment module 3 not only affects the current flight decision but also feeds back to the digital farmland environment modeling module 1 to update the environmental model. This feedback loop forms a collaborative closed loop between environmental perception and environmental modeling, enabling the environmental model to be dynamically updated and maintain consistency with the actual environment. Compared to a static environmental model, a dynamically updated environmental model improves the accuracy of path planning by approximately 25%.
[0063] The collaborative workflow of the four modules is as follows: At the start of the operation, the digital farmland environment modeling module 1 constructs an initial environment model based on pre-acquired terrain, crop, and pest data, and outputs an initial environment state vector. The reinforcement learning policy decision-making module 2 receives the environment state vector and outputs flight action decisions through a policy network. The real-time environment perception and adjustment module 3 receives the flight action decisions, makes necessary adjustments based on real-time sensor data, generates the final flight command, and sends it to the UAV for execution. The path execution feedback optimization module 4 continuously monitors the operation process, collects actual operation effect data, calculates the actual reward value after the operation is completed, and feeds it back to the policy decision-making module 2 for parameter updates. During the operation, if new disease areas are detected or significant changes occur in the environment, the relevant information is fed back to the environment modeling module 1 to update the environment model, forming a dynamic adaptation cycle.
[0064] Through the above-mentioned multi-module collaborative mechanism, the present invention achieves the following synergistic effects: The high-quality state representation provided by the environment modeling module improves the learning efficiency of the policy decision-making module, and the high-quality policies learned by the policy decision-making module further improve the efficiency of utilizing environmental information. The two promote each other and work together to improve system performance.
[0065] The real-time environment perception and adjustment module and the strategy decision-making module optimize the path at both the local and global levels, respectively. The combination of the two modules results in a non-linear increase in path quality improvement. Experimental data shows that using the strategy decision-making module alone improves the path quality by about 20% compared to traditional methods, using the environment perception and adjustment module alone improves it by about 15%, and using both modules together improves it by about 42%, demonstrating a clear synergistic effect.
[0066] The strategy decision-making module excels at global optimization but lacks real-time response capabilities, while the environment perception and adjustment module excels at local adjustment but lacks a global perspective. The two complement each other's strengths, achieving a comprehensive advantage that combines both global optimization and local flexibility.
[0067] There is a conflict between the effectiveness of epidemic prevention and control, energy consumption, and operation time. By designing a composite reward function and optimizing a closed-loop feedback, the system can automatically find a balance point among multiple objectives, thus resolving the contradictions in multi-objective optimization.
[0068] Adaptive adjustment effect: Through online learning and feedback optimization, the system can adaptively adjust strategy parameters according to the actual working environment and task requirements, and maintain excellent performance under different crop types, different types of pests and diseases, and different weather conditions.
[0069] The aforementioned synergistic effect enables the overall performance of the present invention to exhibit a non-linear growth characteristic of 1+1>2, and the system performance is significantly better than the sum of the performances of each module working independently.
[0070] This invention employs a combination of virtual environment pre-training and real-world environment fine-tuning to achieve seamless integration between simulation training and actual deployment.
[0071] During the pre-training phase, a high-fidelity simulation environment is constructed to simulate the dynamic processes of real farmland, including terrain, crop distribution, pest and disease spread, and weather changes. The simulation environment is built based on a physics engine and agronomic models, encompassing drone dynamics models, pesticide spraying and diffusion models, pest and disease spread models, and weather change models. The simulation environment can run at 100 times faster than real-time speeds, enabling the policy network to accumulate a wealth of training experience in a short period.
[0072] Pre-training utilizes eight parallel simulation environments, each randomly generating different terrains, crop types, and pest and disease distributions to improve the strategy's generalization ability. The training process employs a course-based learning strategy, gradually transitioning from simple to complex scenarios. Initial scenarios include flat terrain, a single crop, and localized pests and diseases, while later scenarios include undulating terrain, multiple crops, and large-scale pests and diseases. The total training steps are set at 10,000,000, requiring approximately 24 hours to complete.
[0073] After pre-training, the policy network achieved the following performance indicators in the simulation environment: prevention and control coverage of over 95%, average flight time reduced by 30% compared to the grid path, and average energy consumption reduced by 25%.
[0074] During the fine-tuning phase, the pre-trained policy network was deployed to a real UAV system for small-scale test flights and data collection in a real farmland environment. A conservative update strategy was adopted during fine-tuning, reducing the learning rate to 1 / 10 of that used in the pre-training phase to avoid significant fluctuations in policy performance. The fine-tuning steps were set to 100,000, completed through 20 actual operations.
[0075] Fine-tuning focuses on adjusting the differences between the simulation environment and the real environment, including realistic factors such as sensor noise, actuator latency, and weather disturbances. Through fine-tuning, the policy network can adapt to the characteristics of the real environment and maintain good performance in actual operations.
[0076] During the deployment phase, the system enters online learning mode, continuously learning and optimizing from practical experience. After each task is completed, data collected by the path execution feedback optimization module 4 is used to update the strategy network parameters, achieving continuous improvement. Online learning is updated weekly, balancing learning effectiveness and computational overhead.
[0077] To ensure deployment security, the system employs multiple security mechanisms: First, performance is verified in a simulation environment before policy updates to ensure that updates do not lead to performance degradation; second, a performance lower limit is set, and if performance falls below the lower limit, the system reverts to the previous stable version; third, key safety constraints are guaranteed by a rule system rather than a learning strategy, including boundary constraints, obstacle avoidance constraints, and power constraints, ensuring that safety rules are not violated regardless of policy changes.
[0078] This invention was field-verified in a rice-growing area in Jiangsu Province. The verification area covered 500 mu (approximately 33 hectares), with flat terrain, and was planted with a single variety of rice. At that time, it was the peak season for rice planthoppers.
[0079] The system deployment utilizes a DJI T40 agricultural drone, equipped with a 20L tank, an RTK differential positioning system, wind speed and direction sensors, and a multispectral camera. Before operation, crop growth data is acquired through drone aerial photography, and an initial environmental model is constructed by combining this data with historical monitoring records. The pre-defined control area is divided into a 100×100 grid structure, with each grid cell measuring 5m×5m.
[0080] On the day of the operation, the wind speed was 3 to 5 m / s, the wind direction was southeast, the temperature was 28℃, and the humidity was 75%. The pests and diseases were distributed in patches, mainly concentrated in the eastern and southern areas, covering a total area of about 200 mu.
[0081] The system's planned flight path employs a region-priority coverage strategy, prioritizing areas with high pest and disease density in the east and south. Compared to traditional line-by-line scanning paths, the path of this invention exhibits a clear focus on key areas. During flight, the system monitors pest and disease distribution in real time. Upon discovering a new pest and disease area in the northwest corner, covering approximately 10 acres, the system immediately adjusts its path and inserts control operations for that area.
[0082] The entire operation took 45 minutes, covered a flight distance of 12.3 km, consumed 18 L of pesticide solution, and used 78% of the battery power. An effectiveness evaluation was conducted 3 days after the operation, showing a 96% pest and disease control coverage rate and an 85% reduction in pest population density.
[0083] In the comparative experiment, the traditional line-by-line scanning path was used to operate in the same area, which took 62 minutes, covered a flight distance of 16.8km, consumed 19.5L of pesticide, consumed 92% of the power, achieved a pest and disease control coverage rate of 92%, and reduced the insect population density by 78%.
[0084] Compared with traditional methods, this invention reduces operation time by 27%, flight distance by 27%, pesticide consumption by 8%, power consumption by 15%, control coverage by 4 percentage points, and insect population density reduction by 7 percentage points.
[0085] After 10 consecutive operations using the system of this invention, the system performance was further improved through online learning. The time taken for the 10th operation was reduced to 42 minutes, the flight distance was reduced to 11.8 km, and the prevention and control coverage rate was increased to 97%, demonstrating the effect of continuous learning and optimization.
[0086] Another set of experiments was conducted in an apple orchard to verify the system's adaptability to complex terrain and various crop environments. The orchard covered an area of 300 mu (approximately 20 hectares), with undulating terrain, an elevation difference of up to 20 meters, tree canopy height of 3 to 4 meters, row spacing of 5 meters, and tree spacing of 4 meters. The pests and diseases were a mixture of spider mites and aphids.
[0087] The system constructs a three-dimensional environmental model based on the distribution of fruit trees, and sets the flight altitude to 1.5m to 2m above the tree canopy, dynamically adjusting according to the terrain and canopy height. During operation, the system adjusts the flight altitude multiple times to adapt to terrain undulations, and the spraying timing is dynamically controlled according to the canopy density. In densely canopied areas, the flight speed is reduced and the spraying flow rate is increased, while in open areas, the flight speed is increased and spraying is turned off.
[0088] After the operation was completed, the coverage rate of pest and disease control for fruit trees reached 94%, which is 5 percentage points higher than the 89% of the traditional method, demonstrating the system's good adaptability in complex environments.
[0089] The above application examples verify the effectiveness and practicality of the present invention in actual agricultural production, showing that the present invention can significantly improve the efficiency and effectiveness of plant protection operations and provide an advanced and reliable technical solution for intelligent plant protection.
[0090] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A drone-based plant protection path optimization system based on reinforcement learning, characterized in that, include: The digital farmland environment modeling module constructs a digital farmland environment model that includes information on topography, crops, and pests and diseases. It represents the control area as a grid structure, with each grid cell containing elevation, crop type, pest and disease density, and environmental parameters, and outputs an environmental state vector. The reinforcement learning strategy decision-making module is connected to the digital farmland environment modeling module. It constructs a path decision network based on the near-end strategy optimization algorithm, receives the environmental state vector, and outputs flight action decisions through the strategy network. The flight action decisions include flight direction, flight speed, and spraying control. The flight action benefits are evaluated according to the reward function, which comprehensively considers the prevention and control coverage, energy consumption, and operation time. The real-time environmental perception and adjustment module is connected to the reinforcement learning strategy decision module. It dynamically adjusts flight action decisions based on environmental change information monitored in real time by airborne sensors. The environmental change information includes wind direction changes, remaining drug solution, and newly discovered disease areas. The environmental change information is fused with the output of the strategy network to generate adjusted flight commands. The path execution feedback optimization module is connected to the real-time environment perception adjustment module and the reinforcement learning strategy decision module. It collects actual operation effect data after executing flight commands. The actual operation effect data includes actual prevention and control coverage, energy consumption and operation time. It calculates the actual reward value based on the actual operation effect data and feeds the actual reward value back to the reinforcement learning strategy decision module to update the parameters of the policy network.
2. The UAV plant protection path optimization system based on reinforcement learning according to claim 1, characterized in that: The environmental state vector includes the drone's location information, remaining liquid volume, battery power, and a local grid state matrix centered on the drone's current location. The local grid state matrix contains the state information of multiple surrounding grid cells.
3. The UAV plant protection path optimization system based on reinforcement learning according to claim 1, characterized in that: The digital farmland environment modeling module introduces a dynamic environmental modeling mechanism. The pest density index is dynamically updated based on the diffusion model, which considers the influence of wind direction, wind speed, and temperature on the spread of pests and diseases.
4. The UAV plant protection path optimization system based on reinforcement learning according to claim 1, characterized in that: The path decision network adopts an Actor-Critic architecture, which includes a policy network and a value network. The policy network outputs the probability distribution of actions based on the current state, and the value network evaluates the value of the state. The two networks share information by sharing a bottom-level feature extraction layer.
5. The UAV plant protection path optimization system based on reinforcement learning according to claim 1, characterized in that: The proximal policy optimization algorithm ensures training stability by limiting the magnitude of policy updates, improves sample utilization efficiency by employing an experience replay mechanism, and stabilizes the training process by using a target network mechanism.
6. The UAV plant protection path optimization system based on reinforcement learning according to claim 1, characterized in that: The reward function includes a prevention and control coverage reward, an energy efficiency reward, a time efficiency reward, and a penalty. The prevention and control coverage reward is calculated based on the area of the pest and disease area covered by spraying and the density of pests and diseases. The energy efficiency reward is calculated based on flight energy consumption and spraying energy consumption. The time efficiency reward is a fixed time cost penalty. The penalty includes boundary crossing penalty, collision penalty, and duplicate coverage penalty.
7. The UAV plant protection path optimization system based on reinforcement learning according to claim 1, characterized in that: The real-time environmental perception and adjustment module maintains a historical wind direction queue, calculates the wind direction variance to determine whether the wind direction is stable, and reduces the flight speed and suspends spraying when the wind direction is unstable.
8. The UAV plant protection path optimization system based on reinforcement learning according to claim 1, characterized in that: The real-time environmental perception and adjustment module continuously monitors the remaining liquid medicine. When the remaining liquid medicine is lower than the minimum threshold, it automatically plans a return path to guide the drone to the resupply point to replenish the liquid medicine.
9. The UAV plant protection path optimization system based on reinforcement learning according to claim 1, characterized in that: When the real-time environmental perception and adjustment module detects a new disease area, it feeds back the disease area information to the digital farmland environment modeling module to update the environmental model and assesses the urgency of the new disease area. If the pest density exceeds the emergency threshold, it adjusts the flight path to prioritize the area for control.
10. The UAV plant protection path optimization system based on reinforcement learning according to claim 1, characterized in that: The path execution feedback optimization module calculates the actual control coverage rate by comparing the distribution of pest and disease density before and after the operation, calculates the actual energy consumption through the battery management system and the liquid flow sensor, records the actual operation time through the timer, and stores the actual operation experience as a high-value sample in the experience pool for continuous optimization of the strategy network.
Citation Information
Patent Citations
Unmanned aerial vehicle pest monitoring method based on intelligent path planning
CN120450179A