Method, device and equipment for planning coverage path of pesticide spraying unmanned aerial vehicle
By building a farmland geometric feature model and using deep reinforcement learning technology to optimize drone path planning, the problems of insufficient coverage and poor adaptability in irregular farmland areas were solved, achieving efficient pesticide spraying effects.
Patent Information
- Application Number
- CN202510744451.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-26
AI Technical Summary
Existing drone path planning methods have problems with low planning efficiency, insufficient coverage, and poor adaptability when facing irregular farmland areas. They are difficult to adapt to complex terrain, resulting in waste of resources and increased mechanical energy consumption.
By constructing a geometric feature model of the farmland area, combining deep reinforcement learning technology, designing a dynamic reward function and a priority experience replay mechanism, building a deep Q network model, and optimizing the UAV flight path, we can achieve adaptability and coverage improvement in complex farmland environments.
It improves the operating efficiency of drones in complex farmland environments, reduces omissions and repeated spraying during operations, and improves the efficiency and effectiveness of pesticide spraying.
Smart Images

Figure CN120704351A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of agricultural intelligent technology, and specifically relates to a coverage path planning method, device and equipment for a pesticide spraying drone, which is suitable for regularizing the coverage path of drones in scenarios such as pesticide spraying and crop monitoring. Background Art
[0002] In traditional agricultural pesticide spraying, manual and ground-based mechanical spraying are the primary methods. However, manual spraying is labor-intensive and inefficient, especially in complex terrain or large farmlands, where uniformity and timeliness of spraying are difficult to ensure. While mechanical spraying is more efficient than manual spraying, it is limited by the shape of farmland boundaries and is prone to problems such as crop crushing and redundant turning paths, resulting in pesticide waste or unsatisfactory results. With advances in science and technology, intelligent agriculture is gradually developing. Drones, with their flexibility, efficient operation capabilities, and precise low-altitude control, can quickly cover large areas of farmland. This not only reduces manpower input in pesticide spraying operations, but also improves operational efficiency and reduces the waste of pesticide resources. This provides a new solution for farmland pesticide spraying.
[0003] However, existing drone path planning methods have problems such as low planning efficiency, insufficient coverage, and poor adaptability when faced with irregular farmland areas. For example, when faced with irregular fields along rivers, traditional methods have difficulty adapting to complex terrain, resulting in waste of resources and increased mechanical energy consumption. In reality, there are a large number of irregularly shaped fields, but most existing research focuses on fields with regular shapes such as rectangles or regular polygons, which is difficult to meet actual needs. In addition, when existing plant protection drones operate in irregular areas, they often have problems such as inaccurate route planning, large areas of omissions, and repeated spraying, resulting in pesticide waste and environmental pollution. Therefore, a coverage path planning method is needed to solve this problem, enhance the adaptability of drones in complex farmland environments, and improve the operating efficiency of drones. Summary of the Invention
[0004] The purpose of the present invention is to provide a coverage path planning method, device and equipment for a pesticide spraying drone. By analyzing the geometric characteristics of the farmland area to construct an environmental model, and combining deep reinforcement learning technology to plan and adjust the path, the limitations of traditional methods in complex environments are overcome, so that the drone can automatically adjust the flight path according to the geometric characteristics of the farmland, optimize the coverage rate and path length, and adapt to the complex and changeable farmland environment, thereby improving the efficiency and effectiveness of pesticide spraying.
[0005] To achieve the above objectives, the present invention provides a coverage path planning method for a pesticide spraying drone, which works by establishing a deep Q-network reinforcement learning model that combines regional geometric feature analysis with a heuristic optimal angle calculation method. Specifically, the method includes: constructing an irregular farmland environment model containing geometric features such as centroid, area, and rotation angle; designing a dynamic reward function that integrates multi-dimensional objectives such as coverage improvement and exploration diversity; combining a priority experience replay mechanism and constructing a deep Q-network model using a structure containing a main network and a target network; using the model to train the drone's coverage path planning, and simultaneously using a linear soft update method to achieve smooth updating of network parameters and realize refined control of the drone's flight angle; performing path planning based on the trained deep Q-network model to generate an optimal path that adapts to the farmland boundary shape and reduces coverage blind spots, thereby solving the problems of drone operation efficiency and environmental adaptability in complex farmland environments.
[0006] The present invention provides a coverage path planning method for a pesticide spraying UAV, comprising:
[0007] Preferably, a complex farmland environment model is constructed, specifically:
[0008] The actual irregular farmland area is approximately converted into a corresponding ordered set of vertices and arc segment data information, and the area is rasterized. The grid size is the drone's spraying radius, which is used for subsequent coverage calculations. The geometric features of the area are then calculated, and the horizontal and vertical coordinates of all vertex coordinates are averaged to obtain the coordinates of the area's centroid. The corresponding edges are then connected with the centroid as the origin, splitting the polygon into multiple triangles. The area of each triangle is calculated and added together to obtain the area of the polygon. The total area of the area is calculated based on the arc segment data and added to the area of the arc segment. The convexity of the polygon is then determined to determine the overall shape characteristics of the farmland.
[0009] The region is also rotated to simulate the drone's coverage of farmland under different maneuvers. A rotation matrix is constructed, and each vertex coordinate is rotated around the center of mass. The resulting coordinates form the rotated region. The rotation angle is then normalized and mapped to the range [0, 1]. This serves as a key state variable to complete the construction of the irregular farmland environment model.
[0010] Preferably, the area subject angle is calculated using a heuristic method based on environmental characteristics, specifically:
[0011] According to the constructed environment model, based on the heuristic method, the optimal angle θ is obtained by fusing the features of the polygon main part, arc part and sawtooth part. opt , of the form:
[0012] θ opt =α·θ pca +β·θ h +γ·θ l +δ·θ arc
[0013] where θ pca is the principal component analysis angle, obtained by calculating the covariance matrix of the vertex coordinates and performing eigenvalue decomposition; θ h is the high-frequency component angle, that is, for the sawtooth processing, the corresponding angle is found through Fourier transform; θ l is the angle corresponding to the longest side; θ arc is the angle of the arc segment, which is calculated by determining the center and radius of the arc segment; the sum of the weight coefficients is 1. According to the angle characteristics and regional feature information, the weights of the corresponding parts are reasonably set and summed to obtain the heuristic optimal angle.
[0014] Preferably, the action space and reward function of the drone in the farmland environment are designed as follows:
[0015] According to the environmental information and regional geometric characteristics, the discrete angle change value is set as a set of optional actions, including [-1°, -0.3°, -0.1°, 0°, 0.1°, 0.3°, 1°], corresponding to seven actions {a1, a2, L, a7}, representing the angle that the drone can adjust each time.
[0016] According to the current state and environment information, a hierarchical design and dynamic weighting method is adopted, with maximizing coverage as the core goal, and a reward function is designed in the form of:
[0017] R=R cov +R exp +R div +R glo
[0018] Among them, R cov It is a reward for improving coverage. After each drone performs an action, the ratio of the newly covered area to the uncovered area is calculated and a corresponding reward is given. The specific form is as follows:
[0019] R cov =Δc*ω1
[0020] Among them, Δc is the change in coverage, and ω1 is the corresponding weight.
[0021] Among them, R exp It is the exploration reward. The reward part is designed based on the state visit frequency and the angle change amplitude. The specific form is as follows:
[0022]
[0023] Among them, f(s) is the access frequency of the current state s, that is, the number of visits to the current angle. If the state has not been recorded, the default access frequency is 1. The lower the state access number, the larger the value of this part, and the greater the reward; a i is the currently selected action, which indicates the angle change. The greater the angle change, the higher the reward for that part. t is the number of steps in the current round. As the number of steps in the current round increases, the decay will gradually slow down.
[0024] Among them, R div This is a diversity reward. Combining the change in action angles with the coverage results, positive rewards are given to action combinations that improve coverage, penalties are imposed on repeated ineffective actions that fail to improve coverage, and additional rewards are given to actions that break the historical maximum coverage. In addition, there are rewards for angle diversity. The specific forms are as follows:
[0025]
[0026] Among them, I c It is an indicator function, which is 1 if the current coverage is greater than the historical maximum coverage, otherwise it is 0; D a This is the angle diversity component, which measures the angle difference between the current state and the previous states. If the number of historical visits to a state is greater than 1, the average angle of all historical states is calculated, and the absolute difference between the current state and the mean is used to represent the diversity; otherwise, it is 0. The greater the difference between the current state and the mean, the more novel the current exploration direction is, and the higher the diversity reward value.
[0027] Among them, R glo The reward is guided by the global maximum coverage rate. The reward part is designed based on the relationship between the current coverage rate and the historical best coverage rate. If the current coverage rate exceeds the historical best, a larger positive reward will be given; if the current coverage rate is lower than the best, a penalty will be given. The coverage rates corresponding to the last three actions are compared, and if there is an upward trend, an additional reward will be given. The specific form is as follows:
[0028] R glo =R g1 +R g2
[0029]
[0030] Among them, R g1 Is the basic reward, c t is the current coverage, c b is the best coverage rate in history, and rewards are given based on the relationship between the current coverage rate and the best coverage rate in history; R g2It is an additional reward. When the current coverage rate is greater than the coverage rates corresponding to the previous three actions, a certain reward will be given.
[0031] Preferably, according to the set action space and reward function, a deep Q network reinforcement learning model based on the priority experience replay mechanism is constructed, specifically:
[0032] In a reinforcement learning system for pesticide spraying drone coverage path planning, flight angle is a key feature of the state space, reflecting the drone's flight direction. This feature is normalized and mapped to the interval [0, 1] as state information. Geometric features of the farmland area are also converted into a state representation for the reinforcement learning model, including the current area's center of mass and area, to provide a reference for subsequent decision-making.
[0033] According to the flight characteristics of the UAV in the farmland environment, the action is set as the angle adjustment value, including large adjustment in units of 1, small fine adjustment in units of 0.3 and 0.1, and maintaining the original direction.
[0034] Based on the action space and reward function, a deep reinforcement learning network model is constructed. The main network adopts a three-layer fully connected neural network. The input layer dimension corresponds to the state space dimension, receives state information of normalized rotation angle and regional features, and the output layer dimension corresponds to the action space dimension. It outputs the Q value corresponding to each action to guide decision-making. The target network has the same structure as the main network and regularly synchronizes the main network parameters θ to provide a stable target Q value for training.
[0035] The priority experience replay mechanism is introduced. When the UAV interacts with the environment, the five-tuple (s t ,a t ,r(s t ,a t ),s t+1 d) The data will be stored in the experience replay pool and initially assigned a priority. Later in the training process, the priority of each experience sample will be dynamically adjusted based on the error between the predicted Q value and the target Q value. The larger the error, the higher the priority. During training, samples are sampled from the replay pool based on the priority, with high-priority samples being learned first. This improves data utilization efficiency and accelerates model convergence.
[0036] Preferably, the model is used to train the coverage path planning of the UAV, and a linear soft update method is used to achieve smooth update of network parameters, specifically:
[0037] Initialize the environment, set the polygonal vertices of the farmland boundary, and determine the farmland area. At the same time, configure the parameter weights and obtain the approximate optimal angle based on the heuristic method, which will be used as the initial angle for subsequent drone training.
[0038] Initialize the drone's configuration, including the network weights, the experience replay pool capacity, and the training data batch size. Use the Adam adaptive learning rate optimizer. Also define the total number of training runs and the initial learning rate.
[0039] Start training according to the configuration. At the beginning of each training round, the environment is reset and the drone is set to the initial state. Then, an action is selected based on the current state based on the ε-greedy strategy. The drone executes the selected action, interacts with the environment, updates the current state, and returns a reward value. At the same time, it determines whether the current round of training is completed based on the termination condition. The data generated by each interaction is packaged into a five-tuple (s t ,a t ,r(s t ,a t ),s t+1 ,d) and stored in the experience replay pool. When the data in the experience pool exceeds the capacity, the data that came in first will be deleted to allow the later experience to come in.
[0040] When the number of experiences in the replay pool exceeds the set training batch size, a group of experiences will be sampled from the replay pool according to priority for training learning, as follows:
[0041] When experience is stored in the experience pool, it is assigned an initial priority, which is set to the maximum priority in the current replay pool. When the number of experiences in the experience pool exceeds the training batch size, a sampling probability is calculated by normalizing the priority of each experience to the power of α. Based on the sampling probability, a batch of experiences is sampled from the experience pool for training.
[0042] During the training process, each time a set of experiences is sampled, the predicted Q value and the target Q value are calculated by the forward propagation method, and then the error δ of each experience is calculated. i , and get the priority of the i-th experience by adding a small positive number to the absolute value of the error, and then update the calculated priority to the experience replay pool to complete the dynamic adjustment of the experience priority. Based on the sampling experience, the importance sampling weight ω is introduced i To correct the deviation, the loss function is obtained by multiplying the sampling weight of each experience by the mean square error and taking the average. Then backpropagation is performed to calculate the gradient of the loss function with respect to the network parameters, and the optimizer is used to update the main network parameters θ to minimize the loss function.
[0043] During training, a soft update is performed every certain rounds. Based on the parameter τ, the parameters calculated by weighted averaging are copied to the target network to update the target network parameters. The parameter τ is set to a smaller value, making the update process slower.
[0044] According to the training process, if the coverage of the current path exceeds the historical best performance, the best performance record is updated and the current network parameters are saved; if the training reaches the maximum number of steps or the coverage reaches a certain threshold, the current round of training is completed.
[0045] The above process is repeated in training iterations until the number of iterations is greater than the total number of iterations.
[0046] Preferably, a trained deep Q network model can be obtained through the training process. The next action is output according to the current state of the drone. The drone performs the action and obtains the state at the next moment. The process is repeated to perform path planning until the drone coverage path in the farmland area is obtained. Specifically,
[0047] Initialize the farmland environment information and the initial flight angle of the drone to obtain the initial state information s0; based on the current state information, the trained model will predict the optimal action a t The drone will execute the action and change its current state. The environment will then provide feedback with the corresponding reward value and new state information. The model will then predict the optimal action based on this new state information. This process repeats until the optimal flight angle for optimal coverage is found, resulting in the drone's coverage path within the farmland area and completing the path planning.
[0048] To solve the above technical problems, the present invention also provides a coverage path planning device for a pesticide spraying drone, comprising:
[0049] The acquisition module receives characteristic information of irregular farmland areas, generates a list of boundary vertices and arc segment parameters, simulates and displays the shape of the area, and establishes an environmental model; it also receives the drone's spraying radius and reinforcement learning-related parameters to provide a basis for path planning.
[0050] The calculation module calculates the geometric features of the polygonal area, including the center of mass and area, and calculates the heuristic optimal angle based on the extracted geometric features and weight parameters to provide initial path planning guidance for the drone. It dynamically calculates the reward value and priority during the reinforcement learning training process to complete model training and path planning reasoning.
[0051] The output module is responsible for outputting the path planning results generated by the calculation module to the UAV control system, including the UAV's flight path, as well as performance indicators such as coverage and optimal angle, to evaluate the effectiveness of optimizing the current path planning strategy.
[0052] To solve the above technical problems, the present invention also provides a coverage path planning device for a pesticide spraying UAV, comprising:
[0053] Memory is used to store computer programs and data information.
[0054] When the processor executes the program, the method for drone coverage path planning is implemented.
[0055] Beneficial effects: Compared with existing technical methods, the present invention discloses a coverage path planning method for a pesticide spraying drone, which is mainly aimed at the coverage path planning problem in an irregularly shaped farmland environment. The farmland area is represented as polygonal vertices and rasterized to simulate the pesticide spraying range of the drone. Based on the geometric characteristics of the farmland area, a multi-feature fusion heuristic method is used to obtain the approximately optimal flight angle, which is used as the initial angle for path planning training. The action space in the environment is set, and a reward mechanism is designed by comprehensively considering factors such as coverage and exploration. A deep reinforcement learning network model is constructed, and a priority experience replay mechanism is introduced to dynamically adjust the sampling probability according to the priority of the experience, so that the network prioritizes learning important experience. A soft update method is used to achieve smooth update of the target network parameters to avoid the problem of unstable training caused by drastic fluctuations in the target value. According to the reinforcement learning model, the global optimal angle corresponding to the maximum coverage rate is found, that is, the optimal flight angle of the drone, to obtain the coverage path of the area. In the drone's farmland coverage task, the above invention improves the operation efficiency and effectively reduces the problems of large area omissions and repeated spraying during operation, providing a new method for the coverage path planning problem of pesticide spraying in irregular farmland areas.
[0056] The present invention also provides a coverage path planning device and equipment for a pesticide spraying UAV, which has the same beneficial effects as the above method. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 A schematic diagram of a flow chart of a coverage path planning method for a pesticide spraying UAV provided by the present invention;
[0058] Figure 2 Schematic diagram of the farmland environment model provided by the present invention;
[0059] Figure 3 Schematic diagram of the model framework provided by the present invention;
[0060] Figure 4 The result of the coverage path planning of the pesticide spraying drone provided by the present invention;
[0061] Figure 5 This is a structural schematic diagram of a coverage path planning device for a pesticide spraying UAV provided by the present invention.
[0062] Figure 6 This is a structural schematic diagram of a coverage path planning device for a pesticide spraying UAV provided by the present invention. DETAILED DESCRIPTION
[0063] To better illustrate the purpose and advantages of the present invention, the following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0064] Example:
[0065] Please refer to Figure 1 , Figure 1 The present invention provides a flow chart of a method for planning a coverage path for a pesticide spraying drone, specifically including:
[0066] Step 1: Construct a complex farmland environment model and calculate the regional main body angle using a heuristic method based on environmental characteristics.
[0067] Specifically, a set of vertex coordinates and arc segment data information is used to represent the actual irregular farmland area, forming the basic framework of the farmland, and the area is rasterized. The grid size, i.e., the spraying radius of the drone, is set to 1, as shown in the following example: Figure 2 As shown. Based on the geometric features of the area, the horizontal and vertical coordinates of all vertex coordinates are averaged to obtain the coordinates of the center of mass. The corresponding edges are connected with the center of mass as the origin, and the polygon is divided into multiple triangles. The area of each triangle is calculated based on the side length and the distance from the center of mass to each edge, and then accumulated. The area of the arc portion is added to obtain the total area of the farmland area. A rotation matrix is constructed, and the coordinates of each vertex of the polygon are rotated with the center of mass as the rotation center to obtain the rotated vertex coordinates. The rotation angle of the area is then mapped to the interval [0,1] through normalization and used as an important state variable to complete the construction of the farmland environment model.
[0068] Specifically, according to the boundary characteristics of the farmland area, the corresponding weight coefficients of the polygon main part, arc part and sawtooth part are set. The specific form is as follows:
[0069] θ opt =0.3·θ pca +0.3·θ h +0.3·θ l +0.1·θ arc
[0070] Based on the above calculation, the optimal flight angle θ in the region can be obtained opt .
[0071] Step 2: Design the action space and reward function of the drone in the farmland environment.
[0072] Specifically, based on the environmental information and geometric characteristics, the discrete angle change value is set as a set of optional actions, including [-1°, -0.3°, -0.1°, 0°, 0.1°, 0.3°, 1°], corresponding to seven actions {a1, a2, L, a7}, representing the angle that the drone can adjust each time.
[0073] Specifically, based on the current state and environmental information, a hierarchical design and dynamic weighting method are adopted, with maximizing coverage as the core goal. At the same time, factors such as exploring unknown areas, encouraging path diversity, and promoting the overall planning direction are integrated to design a reward function. The specific form is as follows:
[0074] R=R cov +R exp +R div +R glo
[0075] Each time the drone performs an action, the ratio of the newly covered area to the uncovered area is calculated and a corresponding reward is given in the form of:
[0076] R cov =Δc*8000
[0077] The reward is based on the frequency of state visits and the magnitude of angle changes, in the form of:
[0078]
[0079] Combining the action angle change and coverage results, the corresponding reward is in the form of:
[0080]
[0081] The reward for comparing coverage trend with the historical best coverage is in the form of:
[0082] R glo =R g1 +R g2
[0083]
[0084] Step 3: Based on the set action space and reward function, construct a deep Q network reinforcement learning model based on the priority experience replay mechanism.
[0085] Specifically, in the reinforcement learning architecture for pesticide spraying drone coverage path planning, the normalized flight angle serves as the key feature in the state space, reflecting the drone's flight direction. The geometric characteristics of irregular farmland, including the current region's centroid and area, are also used as state information.
[0086] Specifically, according to the flight characteristics of the UAV in the farmland environment, the action is set as the angle adjustment value, including large adjustments in units of 1, small adjustments in units of 0.3 and 0.1, and maintaining the original direction.
[0087] Specifically, set the experience pool capacity to 10 4 , the priority experience replay parameters α is 0.6, β is 0.4, and a deep reinforcement learning network model based on priority experience replay is constructed. The main network adopts a three-layer fully connected neural network. The input layer dimension corresponds to the state space dimension, receives the state information of normalized angle and regional features, and the output layer dimension is consistent with the action space dimension; the target network has the same structure as the main network, and the main network parameters θ are synchronized regularly.
[0088] Step 4: Please refer to Figure 3 , Figure 3 This is a schematic diagram of the model framework provided by the present invention. The model is used to train the coverage path planning of UAVs, and a linear soft update method is used to achieve smooth updates of network parameters.
[0089] Specifically, the environment is initialized, the polygonal vertices of the farmland boundary are set, and the farmland area is determined. At the same time, the parameter weights are configured, and the approximate optimal angle is obtained based on the heuristic method, which is used as the initial angle for subsequent drone training.
[0090] Specifically, initialize the configuration of the drone, including the network weight parameters θ and θ t , the batch size of the experience replay pool and training data, and the Adam adaptive learning rate optimizer is selected. At the same time, the capacity of the experience replay pool is defined as 10 4 , total training times 10 3 , determine the initial value of learning rate α=10 -3 .
[0091] Specifically, training is started based on the configuration. At the beginning of each training round, the environment is reset and the drone is set to the initial state. Then, the ε-greedy strategy is used to select an action based on the current state. The drone executes the selected action, interacts with the environment, updates the current state, and returns a reward value. At the same time, it determines whether the current round of training is completed based on the termination condition. The data generated by each interaction is packaged into a five-tuple (s t ,a t ,r(s t ,a t ),s t+1 ,d) and store it in the experience replay pool. When the data in the experience pool exceeds the capacity, the data that came in first will be deleted to allow the later experience to come in.
[0092] When the number of experiences in the replay pool is greater than the training batch size, a set of experiences will be sampled from the replay pool according to priority for training learning, as follows:
[0093] When experience is stored in the experience pool, it is assigned an initial priority, which is set to the maximum priority in the current playback pool. When the number of experiences in the experience pool exceeds the training batch size, a sampling probability is calculated by normalizing the priority of each experience to the power of 0.6. Based on the sampling probability, a batch of experiences is sampled from the experience pool for training.
[0094] During the training process, each time a set of experiences is sampled, the predicted Q value and the target Q value are calculated by the forward propagation method, and then the error δ of each experience is calculated. i , and obtain the priority of the i-th experience by adding a small positive number to the absolute value of the error, and then update the calculated priority to the experience replay pool to complete the dynamic adjustment of the experience priority.
[0095] According to the sampling experience, based on the priority experience playback parameter, the importance sampling weight ω is introduced i To correct the deviation, the loss function is obtained by multiplying the sampling weight of each experience by the mean square error and taking the average. Then backpropagation is performed to calculate the gradient of the loss function with respect to the network parameters, and the optimizer is used to update the main network parameters θ to minimize the loss function.
[0096] During training, a soft update is performed every certain rounds. Based on the parameter τ, the parameters calculated by weighted averaging are copied to the target network to update the target network parameters. The parameter τ is set to 0.005 to make the update process slower.
[0097] According to the training process, if the coverage of the current path exceeds the historical best performance, the best performance record is updated and the current network parameters are saved. If the training reaches the maximum number of steps or the coverage reaches the set threshold, the current round of training is completed.
[0098] The process is repeated in training iterations until the number of iterations is greater than the total number of iterations.
[0099] Step 5: Through the training process, a trained deep Q network model can be obtained. The next action is output according to the current state of the drone. The drone executes the action and obtains the state at the next moment. The process is repeated to perform path planning until the drone coverage path in the farmland area is obtained.
[0100] Specifically, the environmental information of the farmland and the initial flight angle of the drone are initialized to obtain the initial state information s0; based on the current state information, the trained model will predict the optimal action a tThe drone will then execute the action and change its current state. The environment will then provide feedback with the corresponding reward value and new state information. The model will then predict the optimal action based on this new state information. This process repeats until the optimal flight angle for optimal coverage is found, resulting in the drone's coverage path for the farmland area.
[0101] To better illustrate the embodiments of the present invention, please refer to Figure 4 , Figure 4 This is a plot of the coverage path planning results for a pesticide spraying drone provided by the present invention. As can be seen from the plot, the path effectively covers irregular boundaries, improving overall coverage. Furthermore, the path planning is completed in a short time with a high reward, resulting in an optimal coverage path suitable for irregular farmland. In summary, practice demonstrates that this method can achieve coverage path planning for pesticide spraying drones in irregular farmland, achieving the intended purpose of the invention.
[0102] To better illustrate the embodiments of the present invention, please refer to Figure 5 , Figure 5 A schematic structural diagram of a coverage path planning device for a pesticide spraying drone provided by the present invention includes:
[0103] The acquisition module receives characteristic information of irregular farmland areas, generates a list of boundary vertices and arc segment parameters, simulates and displays the shape of the area, and establishes an environmental model; it also receives the drone's spraying radius and reinforcement learning-related parameters to provide a basis for path planning.
[0104] The calculation module calculates the geometric features of the polygonal area, including the center of mass and area, and calculates the heuristic optimal angle based on the extracted geometric features and weight parameters to provide initial path planning guidance for the drone. It dynamically calculates the reward value and priority during the reinforcement learning training process to complete model training and path planning reasoning.
[0105] The output module is responsible for outputting the path planning results generated by the calculation module to the UAV control system, including the UAV's flight path, as well as performance indicators such as coverage and optimal angle, to evaluate the effectiveness of optimizing the current path planning strategy.
[0106] For an introduction to the coverage path planning device for the pesticide spraying drone provided in an embodiment of the present invention, please refer to the coverage path planning method embodiment for the pesticide spraying drone in the aforementioned embodiment, and the embodiment of the present invention will not be repeated here.
[0107] Please refer to Figure 6 , Figure 6 The present invention provides a coverage path planning device for a pesticide spraying drone, comprising:
[0108] Memory is used to store computer programs and data information.
[0109] When the processor executes the program, the method for drone coverage path planning is implemented.
[0110] For an introduction to the coverage path planning device for the pesticide spraying drone provided in an embodiment of the present invention, please refer to the coverage path planning method embodiment for the pesticide spraying drone in the aforementioned embodiment, and the embodiment of the present invention will not be repeated here.
[0111] It should be noted that, in this specification, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0112] The above description is merely a preferred embodiment of the present invention and is intended only to explain the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments or to substitute equivalents for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A coverage path planning method for a pesticide spraying drone, characterized in that: The following steps are involved: (1) Construct a farmland environment model and calculate the regional subject angle using a heuristic method based on environmental characteristics; (2) Design the action space and reward function of the UAV in the farmland environment; (3) According to the set action space and reward function, a deep Q network reinforcement learning model based on the priority experience replay mechanism is constructed under the environment model; (4) Using the reinforcement learning model to train UAV path planning, a linear soft update method is used to achieve smooth updates of network parameters; (5) Using the trained deep Q network model, the next action is output according to the current state of the drone. The drone executes the action and obtains the state at the next moment. The above process is repeated for path planning until the drone coverage path in the farmland area is obtained.
2. The coverage path planning method for a pesticide spraying UAV according to claim 1, characterized in that: The environment model constructed in step 1 specifically includes: The boundaries of irregular farmland areas are represented using vertex coordinates, and geometric analysis of polygons is performed. The polygon's centroid is calculated to determine its geometric center. The polygon's area is calculated to assess the area's size and coverage. Rotation operations are used to adjust the area's orientation during path optimization. By analyzing the geometric features of the polygonal area, regional characteristics and dynamic information are extracted, serving as the basis for heuristic optimal angle calculations and providing a reference for subsequent path planning.
3. The coverage path planning method for a pesticide spraying UAV according to claim 1, characterized in that: The specific form of calculating the optimal angle using the heuristic method in step 1 is: i opt =a·θ pca +β·θ h +γ·θ l +δ·θ arc where θ pca is the principal component analysis angle, obtained by calculating the covariance matrix of the vertex coordinates and performing eigenvalue decomposition; θ h It is the angle of high-frequency components, that is, for the sawtooth processing, the corresponding angle is found through Fourier transform; θ l is the angle corresponding to the longest side; θ arc is the angle of the arc segment, which is calculated by determining the center and radius of the arc segment. The optimal angle θ is obtained by multimodal feature fusion calculation. opt According to these angle features and regional feature information, the weights are reasonably set and summed to obtain the heuristic angle, which provides guidance for the initial path planning of the UAV.
4. The coverage path planning method for a pesticide spraying UAV according to claim 1, characterized in that: The action space and reward function of the drone in the complex farmland environment in step 2 are specifically: According to the environmental information and geometric characteristics, the discrete angle change values are set as optional actions, including [-1°, -0.3°, -0.1°, 0°, 0.1°, 0.3°, 1°], a total of seven actions, representing the angles that the drone can adjust each time. According to the state space, the reward function is set by comprehensively considering multiple factors such as coverage, exploration, and global maximum coverage guidance. The specific form is: R=R cov +R exp +R div +R glo Among them, R cov is the coverage change reward, in the form of: R cov =Δc*ω1 Among them, Δc is the change in coverage, and ω1 is the corresponding weight. R exp It is an exploration reward based on the frequency of state visits, in the form of: Among them, f(s) is the access frequency of the current state s. If the state has not been recorded, the default access frequency is 1; a i is the currently selected action; t is the number of steps in the current round. R div is a diversity reward in the form of: Among them I c is an indicator function, which is 1 if the current coverage is greater than the historical maximum coverage, otherwise it is 0; D a This is the angle diversity part. If the number of historical access states is greater than 1, the average angle of all historical states is calculated and expressed as the absolute value of the difference between the current angle and the mean. Otherwise, it is 0. R glo is the global maximum coverage guidance reward, which is of the form: R glo =R g1 +R g2 where R g1 Is the basic reward, c t is the current coverage, c b is the best coverage rate in history, and rewards are given based on the relationship between the current coverage rate and the best coverage rate in history; R g2 It is an extra reward. If the current coverage rate is greater than the coverage rate corresponding to the previous three actions, a certain reward will be given. The reward function is used to provide comprehensive feedback incentives for the drone's behavioral decisions.
5. The coverage path planning method for a pesticide spraying UAV according to claim 1, characterized in that: The implementation method of building the reinforcement learning model in step 3 includes: The deep Q-network employs a dual-network architecture consisting of a main network and a target network. This network consists of three fully connected neural networks. The input layer receives state information, namely the normalized rotation angle, and the output layer outputs the Q value corresponding to each action. The main network is responsible for outputting action Q values in real time based on the current state to guide the agent's decision-making. The target network periodically copies parameters from the main network to provide relatively stable target Q values for training. At the same time, a priority experience replay mechanism is used to store and sample experience. The experience replay pool will dynamically adjust its priority based on the TD error of the experience, and perform weighted sampling based on the priority during sampling. Experience with high priority has a higher probability of being sampled, thereby improving learning efficiency and model convergence speed.
6. The coverage path planning method for a pesticide spraying UAV according to claim 1, characterized in that: In step 4, the reinforcement learning model is used to carry out path planning training and parameter update, specifically: The drone interacts with the environment, continuously acquires experience and stores it in the priority experience replay pool. In each interaction, an action is selected based on the current state. After the action is executed, the environment gives a reward and the next state. The interaction process generates data information in the form of a five-tuple (s t ,a t ,r(s t ,a t ),s t+1 ,d),This experience is stored in the experience replay pool for planning training. At the same time, a soft update method is used. After each training session, the parameters of the target network are gradually updated by a factor of τ = 0.005, achieving a smooth parameter transition, allowing the parameters of the target network to slowly converge with those of the main network. This mechanism ensures the stability of network parameters, allowing the drone to effectively complete coverage missions in complex environments.
7. The coverage path planning method for a pesticide spraying UAV according to any one of claims 1 to 6, characterized in that: In step 5, the trained model is used to perform path planning, specifically: During initialization, the angles obtained by the heuristic method are used as state information. At each planning step, the model predicts the optimal action based on the current state information. After the drone executes the action, the environment provides a reward signal and the next state information based on its new position and coverage. Termination is determined based on the completion of the farmland coverage or the number of steps, and the state is updated. By repeating these steps, the drone will gradually complete the farmland coverage path planning and output the planned path with the maximum coverage rate.
8. A coverage path planning device for a pesticide spraying drone, characterized in that: Used to perform the method according to any one of claims 1 to 7, comprising: The acquisition module receives characteristic information of irregular farmland areas, generates a list of boundary vertices and arc segment parameters, simulates and displays the shape of the area, and establishes an environmental model; it also receives the drone's spraying radius and reinforcement learning-related parameters to provide a basis for path planning. The calculation module calculates the geometric features of the polygonal area, including the center of mass and area, and calculates the heuristic optimal angle based on the extracted geometric features and weight parameters to provide initial path planning guidance for the drone. It dynamically calculates the reward value and priority during the reinforcement learning training process to complete model training and path planning reasoning. The output module is responsible for outputting the path planning results generated by the calculation module to the UAV control system, including the UAV's flight path, as well as performance indicators such as coverage and optimal angle, to evaluate the effectiveness of optimizing the current path planning strategy.
9. A coverage path planning device for a pesticide spraying drone, characterized in that: Including processor and memory: Memory is used to store computer programs and data information. When the processor executes the program, the steps of the drone coverage path planning method described in any one of claims 1 to 7 are implemented.