An orchard multi-plant-protection vehicle path planning and dynamic scheduling method

By using a multi-vehicle path planning and dynamic scheduling method, combined with drone aerial photography data and deep reinforcement learning, the problems of unscientific path planning and poor coordination in hilly orchards were solved, achieving efficient, uniform and continuous pest and disease control in orchard operations.

CN119791081BActive Publication Date: 2025-10-17CHINA AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411866718.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-10-17
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Existing plant protection vehicles have problems with unscientific path planning, poor coordination, and weak obstacle identification and local path planning capabilities in hilly orchards, resulting in low operating efficiency, high energy consumption, uneven pest and disease control, and affecting orchard ecology and yield.

Method used

A multi-vehicle path planning and dynamic scheduling method is adopted, and drone aerial photography data is used to generate an operation base map. Combined with deep reinforcement learning and sensor systems, accurate division of orchard operation areas and path optimization are achieved. LiDAR and other sensors are equipped for obstacle identification, and real-time collaborative operations between vehicles are achieved through a dynamic scheduling module.

Benefits of technology

It has achieved accurate and efficient path planning for plant protection operations in hilly orchards, improved operation efficiency and resource utilization, reduced energy consumption and operation interruptions, and ensured the continuity and uniformity of orchard pest and disease control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119791081B_ABST
    Figure CN119791081B_ABST
Patent Text Reader

Abstract

The application discloses a kind of orchard multi-plant protection vehicle path planning and dynamic scheduling method, comprising: S1, through job grid bottom drawing generation and job point drawing module, unmanned aerial vehicle aerial photography data are generated job bottom drawing;S2, through pest monitoring module, real-time monitoring of pest, and the pest point coordinates monitored are transmitted to dispatch center;S3, call job grid bottom drawing generation and job point drawing module to draw job point;S4, path planning module is path planning prediction for N plant protection vehicles;S5, N plant protection vehicles receive job path information, and set out to execute job;S6, in the process of executing job, provide obstacle avoidance scheme through vehicle-mounted local path planning module;S7, through dynamic scheduling module, new task is distributed to the plant protection vehicle that has completed job, when fault information appears, unfinished task is distributed to nearby plant protection vehicle, until all tasks are completed.The application can accurately and efficiently complete orchard multi-plant protection vehicle path planning and dynamic scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent agriculture, and relates to a method for path planning and dynamic scheduling of orchard plant protection vehicles. BACKGROUND

[0002] In the field of orchard plant protection operation, traditional plant protection methods are mainly relied on. In recent years, with the development of science and technology, plant protection vehicle technology has also begun to be applied to hilly large-scale orchard operation. However, in actual application, both the traditional plant protection method and the existing plant protection vehicle technology have many problems, which are specifically as follows:

[0003] (1) The traditional plant protection method relies on manual backpack sprayers or simple mechanical spraying equipment. In hilly orchards, this method has serious limitations. In hilly areas, the terrain is undulating and complex, and when manually operating, the operator needs to walk in different terrains such as steep slopes and depressions, which is extremely labor-intensive and consumes a lot of physical strength, resulting in extremely low operation efficiency. At the same time, since it is difficult for manual operation to accurately determine the application position and dosage, it is difficult to ensure uniform and accurate application, and it is easy to cause missed spraying areas, so that pests and diseases breed and spread in these areas, affecting the growth of fruit trees and the yield of fruits; while in the heavy spraying area, not only the pesticide is wasted, the cost is increased, but also the excessive use of pesticides may cause soil compaction, water pollution and other environmental problems, destroy the ecological balance, and pose a threat to the sustainable development of the orchard.

[0004] (2) The existing plant protection vehicle technology

[0005] (a) Lack of scientific path planning: Most plant protection vehicles lack effective multi-start balanced traversal path planning methods. When facing complex terrains, various fruit tree distributions and variable pest and disease conditions in hilly orchards, it is difficult for them to intelligently consider the complex terrains, fruit tree distributions and pest and disease conditions in hilly orchards to plan intelligent paths, resulting in unreasonable operation paths, and the plant protection vehicles may drive in a roundabout way in the orchard, increasing unnecessary driving distance, increasing energy consumption, prolonging operation time and significantly reducing operation efficiency. For example, in hilly areas, vehicles may frequently go up and down slopes, consume a lot of energy and easily wear out vehicle parts. In addition, some plant protection vehicles do not fully consider the slope factor in hilly terrains when planning paths. When driving in steep slope areas, the vehicle needs to overcome a large gravitational potential energy, the energy consumption increases significantly, and vehicle faults such as engine overheating and brake failure are easily caused, which seriously affects the continuity of operation. Once the vehicle breaks down, the operation will be stalled during maintenance, which will miss the best opportunity for pest and disease control, allowing pests and diseases to spread and harm the orchard.

[0006] (b) Poor coordination between plant protection vehicles: Existing plant protection vehicle technologies often ignore the dynamic coordination between multiple plant protection vehicles. In large-scale hilly orchards, single plant protection vehicle operation cannot meet the rapid and efficient plant protection needs. Multi-plant protection vehicle coordination can significantly improve the operation efficiency. However, the existing plant protection vehicle technology lacks effective orchard communication system and dynamic scheduling mechanism, and the plant protection vehicles cannot share real-time location information and operation status, which often leads to some plant protection vehicles completing the task in advance and being idle, while some areas have serious plant diseases and insect pests, resulting in excessive workload of plant protection vehicles. When plant protection vehicles break down or operation progress is blocked, the operation plan cannot be adjusted in time, the overall operation efficiency is low, the advantages of multi-plant protection vehicle coordination cannot be fully realized, and efficient plant protection of the orchard cannot be achieved.

[0007] (c) Weak obstacle recognition and local path planning capability: In terms of obstacle recognition and local path planning, the sensor system of many plant protection vehicles is imperfect and cannot accurately recognize various obstacles such as people, stones, and gullies. Some plant protection vehicles equipped with some sensors lack effective data fusion and analysis capabilities, resulting in low accuracy and reliability of obstacle monitoring. When encountering obstacles, plant protection vehicles cannot quickly plan a reasonable local detour path, and often have to stop operation and wait for manual intervention, which easily causes operation interruption, seriously affects operation progress, misses the best disease and pest control opportunity, increases the risk of disease and pest spread, and further adversely affects the yield and quality of the orchard. SUMMARY

[0008] In view of the problems existing in the prior art, the purpose of the present application is to provide a method for planning and dynamically scheduling multiple plant protection vehicles in an orchard.

[0009] In order to achieve the above purpose, the present application adopts the following technical solutions:

[0010] A method for planning and dynamically scheduling multiple plant protection vehicles in an orchard, comprising the following steps:

[0011] S1, generating an operation bottom map from unmanned aerial vehicle aerial photography data through an operation grid bottom map generation and operation point drawing module;

[0012] S2, real-time monitoring of pests by a pest monitoring module, and transmitting the monitored pest point coordinates to a dispatch center;

[0013] S3, the dispatch center calls the operation grid bottom map generation and operation point drawing module to draw operation points according to the operation point information, and then allocates n operation areas to the existing n plant protection vehicles in the warehouse according to the drawn operation point map;

[0014] S4, the path planning module makes path planning prediction for N plant protection vehicles, returns the obtained path planning prediction result set to the dispatch center, and distributes the planned work path to the existing N plant protection vehicles;

[0015] S5, N plant protection vehicles receive work path information from the dispatch center, and start from the warehouse to perform work;

[0016] S6, during the execution of the plant protection vehicle, the position information of each plant protection vehicle is tracked in real time and transmitted back to the dispatch center, and when an obstacle is encountered, an obstacle avoidance scheme is provided through the local path planning module of the plant protection vehicle to avoid the obstacle;

[0017] S7, when a plant protection vehicle completes its task in the work area, it sends the completion information to the dispatch center, which calls the dynamic scheduling module to assign new tasks to the plant protection vehicles that have completed the work, and when a plant protection vehicle has a fault information, the dispatch center immediately notifies other nearest idle plant protection vehicles, and according to the position and task situation of the fault plant protection vehicle, the unfinished task is assigned to the nearby plant protection vehicle until all tasks are completed.

[0018] Preferably, in step S1, the work grid bottom map generation and work point drawing module generates the work bottom map from the unmanned aerial vehicle aerial photography data, comprising the following steps:

[0019] S11, using the monitoring equipment carried by the unmanned aerial vehicle, scanning and photographing the orchard to obtain unmanned aerial vehicle aerial photography data;

[0020] S12, import the obtained unmanned aerial vehicle aerial photography data into GIS software to generate digital elevation model (DEM) and orthographic image of the orchard as work bottom map.

[0021] Preferably, in step S2, the pest monitoring module monitors the pests in real time, and sends the coordinates of the monitored pest points back to the dispatch center, comprising the following steps:

[0022] The insect traps and sensors arranged in the orchard monitor the pest situation in real time, and when a pest point (or tree body) is monitored, the coordinates of the pest point are sent back to the dispatch center.

[0023] Preferably, in step S3, the work point drawing is performed by calling the work grid bottom map generation and work point drawing module, comprising the following steps:

[0024] S31, using GIS software to mark the work points on the digital elevation model (DEM) and orthographic image of the orchard, the work points being the points of the insect traps or sensors that have captured insects;

[0025] S32. Create a digital elevation model (DEM), an orthophoto map, and operation points as separate layers, each with independent properties and display settings; wherein the DEM layer is set to a three-dimensional terrain display mode to highlight topographic features; the orthophoto map layer is set to an image display mode to clearly display image information such as fruit tree distribution; and the operation point layer is set to an editable and annotated point feature layer to facilitate drawing, modifying, and annotating the operation points;

[0026] S33. After the setting is completed, a three-dimensional operation diagram is generated.

[0027] Preferably, in step S4, the path planning module performs path planning for N plant protection vehicles, including the following steps:

[0028] S41. Analyze the density and spatial layout of the operation points in the operation area. Take the minimum total distance of multiple plant protection vehicles and the minimum distance difference as the optimization goal. Suppose the number of plant protection vehicles is N, and the path length of each plant protection vehicle is , then the total distance is , the objective function is: and ,in represents the standard deviation of the path length of each plant protection vehicle;

[0029] S42. Path planning is performed using a long short-term memory network (LSTM) in deep reinforcement learning. Specifically, a state space is constructed based on the collected state information. The state space of the plant protection vehicle is defined as the location information of the current operation point and hilly terrain characteristics (slope and aspect data). The movement direction of the plant protection vehicle is defined as the action space. The reward function is set to the objective function and energy consumption in step S41. The objectives are to minimize the total distance traveled by multiple plant protection vehicles and the difference in distance traveled, while minimizing the total energy consumption of the plant protection vehicles. Task completion is a positive reward, and encountering an obstacle is a negative reward. Both are linearly added to the reward function with fixed values.

[0030] S43. Perform model training using the three-dimensional operation diagram stored in the dispatching center;

[0031] S44, using the trained model to generate path predictions for n plant protection vehicles;

[0032] S45. Return the predicted path result set of the plant protection vehicle to the dispatch center.

[0033] Preferably, in step S43, the model training using the three-dimensional operation diagram stored in the dispatching center includes the following steps:

[0034] S431. The plant protection vehicle begins interacting in the environment simulated by the three-dimensional operation map according to an initial random strategy. A value network is designed: it consists of a target network and a prediction network, both of which adopt the LSTM network structure. The target network has the same structure as the prediction network, but is updated less frequently to stabilize the training process.

[0035] S432. For each time step t, use the epsilon strategy for action selection, randomly select an action with a probability of epsilon, use the target network with a probability of 1-epsilon, and select the action with the largest Q value among all actions in that state. Epsilon is initialized to 1 and gradually decreases over time.

[0036] S433: Execute the selected action Afterwards, get rewarded , observe the next state , the experience tuple The playback buffer is set to store 1000 experience tuples. When the buffer is not full, the new experience tuple is stored in it. When the buffer is full, the new experience tuple replaces the oldest tuple.

[0037] S434. Randomly extract a batch of experience tuples from the experience replay buffer. Suppose the batch size is , for each experience tuple in the batch , calculate the target Q value according to the target network, the calculation formula is:

[0038]

[0039] in, is the target Q value, It's a reward. is the discount factor, set , is the target network (with parameters ) for the next state Take action Q value estimation;

[0040] S435, the extracted state sequence ( Each state in the prediction network is input into the forward propagation and the Q value estimation for each action is output. , where θ is the parameter of the prediction network;

[0041] S436. Calculate the loss function based on the target Q value and the predicted Q value, using the mean square error (MSE) as the loss function: ;

[0042] S437、based on the calculated loss function, the gradient of the loss function to the prediction network parameter θ is calculated by the back propagation algorithm, and the network parameter is updated according to the gradient using the Adam optimizer, which adjusts the learning rate according to the first moment estimate and the second moment estimate of the gradient, and updates the parameter, so that the loss function gradually decreases;

[0043] S438、every training a certain number of time steps, copy a part of the parameters of the prediction network to the target network to update the parameters of the target network, so that the calculation of the target Q value is more stable, and the training process is avoided to be too volatile;

[0044] S439, repeat steps S432-S438 until the predetermined training round is reached or the network performance on the validation set no longer improves.

[0045] Preferably, in step S6, the on-board local path planning module on the orchard plant protection vehicle provides an obstacle avoidance scheme, including the following steps:

[0046] S61, using the sensor system composed of laser radar, camera, ultrasonic sensor and other sensors equipped on the orchard plant protection vehicle, real-time collection of environmental information around the vehicle. Process and analyze the data collected by the sensor to determine the position and range of the obstacle in the orchard coordinate system;

[0047] S62, use RRT* algorithm to plan local path, input obstacle position, shape, range information and hill slope, slope direction, terrain undulation information from the dispatch center, output obstacle avoidance path;

[0048] S63, the plant protection vehicle avoids obstacles according to the newly planned path.

[0049] Preferably, in step S62, the RRT* algorithm is used to plan the local path, and the obstacle position, shape, range information and hill slope, slope direction, terrain undulation information from the dispatch center are input, and the obstacle avoidance path is output, including the following steps:

[0050] S621, set the current position as the starting point start and the reachable obstacle avoidance point beside the obstacle as the target point goal, create an RRT tree T containing only start, and define the reachable obstacle avoidance point as a point that can be reached and can complete effective obstacle avoidance when the plant protection vehicle reaches the point, define the cost function as the energy consumption of the plant protection vehicle, wherein the energy consumption represents a function about the hill slope, slope direction, terrain, path length, and the environment space is all reachable grid points except the obstacle;

[0051] S622, random sampling: randomly sample a point x_rand in the environment space;

[0052] S623, find the nearest node x near from the tree T, and extend a certain distance in the direction from x near to get a new node x new;

[0053] S624, detect whether a collision occurs between the path from x near to x new and the obstacle according to the position, shape and range information of the obstacle, if a collision occurs, return to step S622 to continue the next iteration;

[0054] S625, find the nearest node x min from x new in the tree T, and calculate the cost cost(x min, x new) from x min to x new;

[0055] S4626, for the nodes x near neighbors in the tree T within a certain range from x new, calculate the cost cost(x near, x new) of connecting to x new through x near, and select the connection mode with the minimum cost;

[0056] S627, add x new to the tree T, and update the connection relationship and cost information between nodes;

[0057] S628, optimize part of the nodes in the tree T with the target point goal as the target, improve the quality of the path by adjusting the connection relationship and cost information, repeat steps S622-S628 until a path is found or the maximum iteration number is reached;

[0058] S629, if the newly added node x new is close to the target point goal, a feasible path is found, and the algorithm ends.

[0059] Advantages:

[0060] Compared with the prior art, the orchard plant protection vehicle path planning and dynamic scheduling method has the following outstanding advantages:

[0061] (1) Precise and efficient path planning: the method combining multi-start balanced traversal path planning and deep reinforcement learning is adopted. In the path planning process, the topographic features of hilly orchards are fully considered, and more reasonable driving routes can be selected by obtaining topographic data (such as slope and slope direction).

[0062] (2) Intelligent cooperative job scheduling: By establishing an effective communication system and a dynamic scheduling module in the orchard, real-time communication and cooperative work between multiple plant protection vehicles are realized. The plant protection vehicles can share location information and work status in real time, and the scheduling center can reasonably allocate tasks according to the actual situation. When a plant protection vehicle completes a task or fails, the task allocation and work plan can be quickly adjusted to ensure the continuity and efficiency of the work, fully utilize the advantages of cooperative work of multiple plant protection vehicles, improve the efficiency of plant protection work in the entire orchard, and realize the optimal allocation of resources.

[0063] (3) Strong obstacle response capability: Equipped with advanced sensor systems and perfect obstacle identification and local path planning modules, multiple sensors (such as laser radar, camera, ultrasonic sensor, etc.) are used in combination to improve the accuracy and reliability of obstacle monitoring. Deep learning algorithms are used to analyze sensor data to continuously optimize the obstacle identification model, which can quickly and accurately identify various obstacles. Once obstacles are found, the optimal detour path can be quickly planned to effectively enhance the adaptability of the plant protection vehicle in complex hilly orchard environment, reduce the risk of work interruption, ensure the smooth progress of work, and improve the reliability of orchard plant protection work. BRIEF DESCRIPTION OF DRAWINGS

[0064] Figure 1 It is a schematic diagram of the architecture of a multi-plant protection vehicle path planning and dynamic scheduling system in an orchard.

[0065] Figure 2 It is a schematic diagram of the process of a multi-plant protection vehicle path planning and dynamic scheduling method in an orchard.

[0066] Figure 3 It is a schematic diagram of the process of path planning performed by the path planning module.

[0067] Figure 4 It is a schematic diagram of the execution process of the dynamic scheduling module. DETAILED DESCRIPTION

[0068] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0069] Referring to Figure 1 , the present application provides a path planning and dynamic scheduling system for plant protection vehicles in an orchard, comprising: a pest monitoring module, a work grid base map generation and work point drawing module, a path planning module, an orchard plant protection vehicle, a dynamic scheduling module, and a scheduling center. Wherein,

[0070] The job grid base map generation and job point drawing module is used for generating a job base map from the aerial photography data of the unmanned aerial vehicle, and importing the generated job base map into GIS software to generate a digital elevation model (DEM) and an orthographic image map of the orchard.

[0071] The pest monitoring module is used for monitoring pests in real time through the monitoring equipment arranged in the orchard, and sending the specific position (point coordinates) of the monitored pests back to the dispatch center.

[0072] The path planning module is used for path planning and prediction for n plant protection vehicles, and returns the path prediction results of the n plant protection vehicles to the dispatch center.

[0073] The dynamic scheduling module is used for: (1) receiving information that some plant protection vehicles complete the planning tasks assigned by the dispatch center and job information layout, selecting, according to the distribution of the uncompleted job points, the job points closest to the completed plant protection vehicles in reverse order from the uncompleted paths for the idle plant protection vehicles, and assigning the job points to the plant protection vehicles. The job information layout is updated, the newly assigned job points are marked on the map, and the vehicle task information and the updated job information layout information are returned to the dispatch center; (2) receiving fault information of a plant protection vehicle in a certain area from the dispatch center. The dispatch center immediately informs other nearest idle plant protection vehicles, and assigns the uncompleted tasks of the fault plant protection vehicle to the nearby plant protection vehicles according to the position and task situation of the fault plant protection vehicle, and finally returns the scheduling information to the dispatch center.

[0074] The plant protection vehicle loads a local path planning module, which is used for collecting environmental information around the plant protection vehicle during driving of the plant protection vehicle, and planning a local path according to the collected information, so that the plant protection vehicle avoids obstacles according to the newly planned path, i.e. an obstacle avoidance path.

[0075] Embodiment 1

[0076] Referring to Figures 2-4 A method for planning and dynamically scheduling paths of multiple plant protection vehicles in an orchard includes the following steps:

[0077] S1, generating a job base map from aerial photography data of the unmanned aerial vehicle through a job grid base map generation and job point drawing module; the specific execution process includes the following steps:

[0078] S11, using the laser radar and camera carried by the unmanned aerial vehicle and other devices to scan and photograph the orchard, and obtaining aerial photography data of the unmanned aerial vehicle;

[0079] S12, importing the obtained aerial photography data of the unmanned aerial vehicle into GIS software to generate a digital elevation model (DEM) and an orthographic image map of the orchard;

[0080] S2, real-time monitoring of pests by the pest monitoring module, and transmitting the monitored pest point coordinates to the dispatch center; the specific execution process includes the following steps: real-time monitoring of pest conditions by the insect traps and sensors arranged in the orchard, and when a pest point (or tree body) is monitored, sending the pest point coordinates (position information) back to the dispatch center.

[0081] S3, the dispatch center calls the job grid base map generation and job point drawing module according to the job point information to draw the job point, and then allocates n job areas to the existing n plant protection vehicles in the warehouse according to the drawn job point map; the specific execution process includes the following steps:

[0082] S31, using GIS software to mark the to-be-worked points on the digital elevation model (DEM) and orthographic image map of the orchard, the to-be-worked points being the points of the traps or sensors that have captured insects;

[0083] S32, creating the digital elevation model (DEM), orthographic image map and job point as different layers respectively, each layer having independent attributes and display settings, and generating a three-dimensional job map after the settings are completed; wherein the DEM layer is set to a three-dimensional terrain display mode to highlight the topographic features; the orthographic image layer is set to an image display mode to clearly display image information such as the distribution of fruit trees; and the job point layer is set to a point element layer that can be edited and labeled to facilitate the drawing, modification and labeling of the job points.

[0084] S33, generating a three-dimensional job map after the settings are completed, and saving the generated three-dimensional job map to the dispatch center.

[0085] S4, the path planning module makes path planning prediction for N plant protection vehicles, and returns the obtained path planning prediction result set to the dispatch center, and then distributes the planned job path to the existing N plant protection vehicles; the specific execution process includes the following steps:

[0086] S41, analyzing the density and spatial layout of the job points in the job area, taking the minimum total path length and the minimum path distance difference of the multiple plant protection vehicles as the optimization target, setting the number of plant protection vehicles as N, and the path length of each plant protection vehicle as , the total path length is , and the objective function is: and , wherein represents the standard deviation of the path length of each plant protection vehicle;

[0087] S42, using the long short-term memory network (LSTM) in deep reinforcement learning for path planning;

[0088] Specifically, a state space is constructed based on the collected state information. The state space of the plant protection vehicle is defined as the location information of the current operation point and the hilly terrain characteristics (slope and aspect data). The movement direction of the plant protection vehicle is defined as the action space. The reward function is set as the objective function and energy consumption in step S300. The goal is to minimize the total distance traveled by multiple plant protection vehicles and the difference in distance traveled, while minimizing the total energy consumption of the plant protection vehicle operations. At the same time, task completion is regarded as a positive reward, and encountering an obstacle is regarded as a negative reward. Both are linearly added to the reward function with fixed values.

[0089] S43. Use the three-dimensional operation diagram stored in the dispatch center to perform model training. The specific execution process includes the following steps:

[0090] S431: The plant protection vehicle begins interacting in the environment simulated by the 3D operation map according to an initial random strategy. A value network is designed: it consists of a target network and a prediction network. Both networks use the LSTM network structure. The target network has the same structure as the prediction network, but is updated less frequently to stabilize the training process.

[0091] S432. For each time step t, use the epsilon strategy for action selection, randomly select an action with a probability of epsilon, use the target network with a probability of 1-epsilon, and select the action with the largest Q value among all actions in that state. Epsilon is initialized to 1 and gradually decreases over time.

[0092] S433: Execute the selected action Afterwards, get rewarded , observe the next state , the experience tuple The playback buffer is set to store 1000 experience tuples. When the buffer is not full, the new experience tuple is stored in it. When the buffer is full, the new experience tuple replaces the oldest tuple.

[0093] S434. Randomly extract a batch of experience tuples from the experience replay buffer. Suppose the batch size is , for each experience tuple in the batch , calculate the target Q value according to the target network, the calculation formula is:

[0094]

[0095] in, is the target Q value, It's a reward. is the discount factor, set , is the target network (with parameters ) to the next state take an action Q-value estimation

[0096] S435, input the extracted state sequence (S1, S2, …, Sn) to the prediction network for forward propagation, output the Q-value estimation of each action where θ is the parameter of the prediction network

[0097] S436, calculate the loss function according to the target Q-value and the predicted Q-value, use mean square error (MSE) as the loss function:

[0098] S437, based on the calculated loss function, calculate the gradient of the loss function to the prediction network parameter θ by the back propagation algorithm, use the Adam optimizer to update the network parameter according to the gradient, the Adam optimizer will adjust the learning rate according to the first moment estimation and the second moment estimation of the gradient, and update the parameter, so that the loss function gradually decreases

[0099] S438, copy a part of the parameters of the prediction network to the target network every certain number of time steps to update the parameters of the target network, so that the calculation of the target Q-value is more stable, and the training process is avoided. Too much fluctuation

[0100] S439, repeat steps S432-S438 until a predetermined number of training rounds is reached or the network performance on the validation set no longer improves.

[0101] S44, use the trained model to generate n plant protection vehicle path predictions

[0102] S45, return the plant protection vehicle path prediction set to the scheduling center.

[0103] The intelligent path planning technology of multi-factor fusion is adopted in the application, and the multi-start balanced traversal is combined with deep reinforcement learning: a multi-start balanced traversal path planning method is adopted, combined with the long short-term memory network (LSTM) in deep reinforcement learning, and various factors in the work area are fully considered for path planning. When constructing the state space, the work points of the orchard are taken as the basis, and the current work point position, hilly terrain characteristics (such as slope and slope direction data) are comprehensively considered to define the moving direction action space of the plant protection vehicle. On the optimization problem, the minimum total distance of multiple plant protection vehicles and the minimum distance difference are taken as the optimization objectives, and the energy consumption of the plant protection vehicle is also considered, and the reward function is set to include the above objective functions and energy consumption factors.

[0104] S5, The plant protection vehicle receives the work path information from the scheduling center and starts from the warehouse to perform the work ​​

[0105] S6, during the execution of the plant protection vehicle, when encountering an obstacle, an obstacle avoidance scheme is provided by a local path planning module carried by the plant protection vehicle to enable the plant protection vehicle to avoid the obstacle; the specific execution process includes the following steps:

[0106] S61, using a sensing device such as a laser radar, a camera, an ultrasonic sensor, etc. equipped on the orchard plant protection vehicle, environmental information around the vehicle is collected in real time, and the environmental information data collected by the sensing device is processed and analyzed to determine the position and range of the obstacle in the orchard coordinate system.

[0107] S62, using an RRT* algorithm to plan a local path, inputting the position, shape, and range information of the obstacle, and the slope, slope direction, and terrain undulation information of the hill from the dispatch center, and outputting a newly planned path, i.e., an obstacle avoidance path; the specific execution process of using the RRT* algorithm to plan the local path includes the following steps:

[0108] S621, setting the current position as a starting point start and the reachable obstacle avoidance point beside the obstacle as a target point goal, creating an RRT tree T containing only start, and defining the reachable obstacle avoidance point as a point that can be reached and can effectively avoid the obstacle when the plant protection vehicle reaches the point, and defining a cost function as the energy consumption of the plant protection vehicle, wherein the energy consumption represents a function related to the hill slope, slope direction, terrain, and path length, and the environmental space is all reachable grid points except the obstacle;

[0109] S622, randomly sampling a point x_rand in the environmental space;

[0110] S623, finding the nearest node x_near in the tree T, and extending a certain distance in the direction of x_near to obtain a new node x_new;

[0111] S624, detecting whether a collision occurs between the path x_near and x_new and the obstacle according to the position, shape, and range information of the obstacle, and if a collision occurs, returning to step S622 for the next iteration;

[0112] S625, finding the nearest node x_min to x_new in the tree T, and calculating the cost cost(x_min, x_new) from x_min to x_new;

[0113] S626, for the nodes x_near_neighbors in the tree T within a certain range from x_new, calculating the cost cost(x_near, x_new) of connecting x_near to x_new, and selecting the connection mode with the minimum cost;

[0114] S627, add x_new to the tree T, and update the connection relationship and cost information between nodes;

[0115] S628, optimize part of the nodes in the tree T with the target point goal as the target, and improve the quality of the path by adjusting the connection relationship and cost information;

[0116] S629, if the newly added node x_new is close to the target point goal, a feasible path is found, and the algorithm ends.

[0117] S63, the plant protection vehicle avoids obstacles according to the obstacle avoidance path.

[0118] The application carries out local path planning for obstacle avoidance: through fusion processing of various sensor data, various obstacles are identified, including people, stones, gullies, trees and other different types of obstacles, deep learning algorithm is used to analyze the sensor data, obstacle identification model is continuously optimized, identification speed and accuracy are improved, when obstacles are detected, the RRT algorithm is used to quickly plan the optimal bypass path combined with the slope and slope direction of the hill.

[0119] S7, when a certain plant protection vehicle completes its task in the working area, the completion information is sent to the scheduling center, the dynamic scheduling module is called, and the plant protection vehicle that has completed the task is allocated a new task, when a certain plant protection vehicle has a fault information, the scheduling center immediately notifies other nearest idle plant protection vehicles, and according to the position and task situation of the fault plant protection vehicle, the uncompleted task of the fault plant protection vehicle is allocated to the nearby plant protection vehicle, until all tasks are completed. The dynamic scheduling and cooperative working mechanism, the specific execution process includes the following two aspects:

[0120] (1) receiving the information that a certain plant protection vehicle completes the planning task allocated by the scheduling center and the working information layout, according to the distribution of uncompleted working points, selecting the working points close to the completed plant protection vehicle in reverse order from the uncompleted path for the idle plant protection vehicle, and allocating the working points to the plant protection vehicle; and updating the working information layout, marking the newly allocated working points on the map, returning the vehicle task information and the updated working information layout information to the scheduling center;

[0121] (2) receiving the fault information of a certain plant protection vehicle in a certain area from the scheduling center, the scheduling center immediately notifies other nearest idle plant protection vehicles, and according to the position and task situation of the fault plant protection vehicle, the uncompleted task of the fault plant protection vehicle is allocated to the nearby idle vehicle to replace the task, and the working area division between the working bodies is dynamically adjusted according to the actual situation, and finally the scheduling information is returned to the scheduling center.

[0122] The above merely describes preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art, according to the technical solution and the improvement concept of the present application, makes equivalent replacement or change within the technical range disclosed by the present application, which should be covered in the protection scope of the present application.

Claims

1. A method for path planning and dynamic scheduling of multiple plant protection vehicles in an orchard, characterized by: The following steps are involved: S1, generate an operation base map from the drone aerial photography data through the operation grid base map generation and operation point drawing module; S2. Real-time monitoring of pests is performed through the pest monitoring module, and the coordinates of the monitored pest points are transmitted to the dispatching center; S3. The dispatch center calls the operation grid base map generation and operation point drawing module based on the operation point information to draw the operation points, and then allocates n operation areas to the n existing plant protection vehicles in the warehouse according to the drawn operation point map; S4. The path planning module makes path planning predictions for N plant protection vehicles, returns the obtained path planning prediction results to the dispatch center, and then distributes the planned operation paths to the existing N plant protection vehicles; The path planning module performs path planning for N plant protection vehicles, including the following steps: S41. Analyze the density and spatial layout of the operation points in the operation area. Take the minimum total distance of multiple plant protection vehicles and the minimum distance difference as the optimization goal. Suppose the number of plant protection vehicles is N, and the path length of each plant protection vehicle is , then the total distance is , the objective function is: and ,in represents the standard deviation of the path length of each plant protection vehicle; S42. Path planning is performed using a long short-term memory network in deep reinforcement learning. Specifically, a state space is constructed based on the collected state information. The state space of the plant protection vehicle is defined as the location information of the current operation point and the hilly terrain characteristics. The movement direction of the plant protection vehicle is defined as the action space. The reward function is set to the objective function and energy consumption in step S41. The goal is to minimize the total distance traveled by multiple plant protection vehicles and the difference in distance traveled, while minimizing the total energy consumption of the plant protection vehicle operations. Task completion is a positive reward, and encountering an obstacle is a negative reward. Both are linearly added to the reward function with fixed values. S43, using the three-dimensional operation diagram stored in the dispatching center to perform model training; S44, using the trained model to generate path predictions for n plant protection vehicles; S45. Returning the predicted path result set of the plant protection vehicle to the dispatch center; S5. N plant protection vehicles receive the operation path information from the dispatch center and start from the warehouse to perform the operation; S6. When the plant protection vehicle encounters an obstacle during operation, an obstacle avoidance plan is provided by the onboard local path planning module of the plant protection vehicle, so that the plant protection vehicle avoids the obstacle; the steps include: S61. Using the sensing equipment equipped on the orchard plant protection vehicle, collect environmental information around the vehicle in real time, and process and analyze the environmental information data collected by the sensing equipment to determine the position and range of the obstacle in the orchard coordinate system; S62. Use the RRT* algorithm to plan the local path, input the obstacle location, shape, range information and the hill slope, aspect, and terrain information from the dispatch center, and output the obstacle avoidance path; including the steps of: S621. Set the current position as the starting point start and the reachable obstacle avoidance point next to the obstacle as the target point goal, and create an RRT tree T containing only start, and define the reachable obstacle avoidance point as a point that is reachable and can effectively avoid obstacles when the plant protection vehicle reaches the point, and define the cost function as the energy consumption of the plant protection vehicle, wherein the energy consumption represents a function of the hill slope, aspect, terrain, and path length, and the environmental space is all reachable grid points except the obstacles; S63, the plant protection vehicle avoids obstacles according to the obstacle avoidance path; S7. When a plant protection vehicle completes the task in its own operation area, it sends the completion information to the dispatch center. The dispatch center calls the dynamic scheduling module to assign new tasks to the plant protection vehicles that have completed the operation. When a plant protection vehicle has a fault message, the dispatch center immediately notifies other plant protection vehicles that will be idle soon, and assigns its unfinished tasks to nearby plant protection vehicles based on the location and task status of the faulty plant protection vehicle until all tasks are completed.

2. A method for path planning and dynamic scheduling of multiple plant protection vehicles in an orchard according to claim 1, characterized in that: In step S1, the operation grid base map generation and operation point drawing module generates an operation base map from the drone aerial photography data, including the following steps: S11. Use the monitoring equipment carried by the drone to scan and photograph the orchard to obtain drone aerial photography data; S12. Import the acquired drone aerial photography data into GIS software to generate a digital elevation model and orthophoto map of the orchard as the base map for the operation.

3. A method for path planning and dynamic scheduling of multiple plant protection vehicles in an orchard according to claim 2, characterized in that: In step S2, the pest monitoring module is used to monitor pests in real time, and the coordinates of the monitored pest points are sent back to the dispatching center, including the following steps: the pest situation is monitored in real time by insect traps and sensors arranged in the orchard, and when a pest point is detected, the coordinates of the pest point are sent back to the dispatching center.

4. A method for path planning and dynamic scheduling of multiple plant protection vehicles in an orchard according to claim 3, characterized in that: In step S3, the calling of the operation grid base map generation and operation point drawing module to draw the operation points includes the following steps: S31. Using GIS software, mark the points to be operated on the digital elevation model and orthophoto map of the orchard, where the points to be operated are the locations of traps or sensors that have captured insects; S32. Create the digital elevation model, orthophoto map, and operation points as different layers, each with independent properties and display settings; wherein the digital elevation model layer is set to a three-dimensional terrain display mode to highlight the terrain features; the orthophoto map layer is set to an image display mode to clearly display the fruit tree distribution image information; and the operation point layer is set to an editable and annotated point feature layer to facilitate the drawing, modification, and annotation of the operation points; S33. After the setting is completed, a three-dimensional operation diagram is generated, and the generated three-dimensional operation diagram is saved to the dispatching center.

5. The method for path planning and dynamic scheduling of multiple plant protection vehicles in an orchard according to claim 4, characterized in that: In step S43, the model training is performed using the three-dimensional operation diagram stored in the dispatching center, including the following steps: S431: The plant protection vehicle begins interacting in the environment simulated by the 3D operation map according to the initial random strategy. A value network is designed: it consists of a target network and a prediction network, both of which use the LSTM network structure. S432. For each time step t, use the epsilon strategy for action selection, randomly select an action with a probability of epsilon, use the target network with a probability of 1-epsilon, and select the action with the largest Q value among all actions in that state. Epsilon is initialized to 1 and gradually decreases over time. S433: Execute the selected action Afterwards, get rewarded , observe the next state , the experience tuple The playback buffer is set to store 1000 experience tuples. When the buffer is not full, the new experience tuple is stored in it. When the buffer is full, the new experience tuple replaces the oldest tuple. S434. Randomly extract a batch of experience tuples from the experience replay buffer. Suppose the batch size is , for each experience tuple in the batch , calculate the target Q value according to the target network, the calculation formula is: in, is the target Q value, It's a reward. is the discount factor, set , is the target network's next state Take action Q value estimation; S435, input the extracted state sequence into the prediction network for forward propagation, and output the Q value estimate for each action , where θ is the parameter of the prediction network; S436. Calculate the loss function based on the target Q value and the predicted Q value, using the mean square error as the loss function: ; S437. Based on the calculated loss function, the gradient of the loss function with respect to the prediction network parameter θ is calculated using the backpropagation algorithm. The network parameters are updated according to the gradient using the Adam optimizer. The Adam optimizer adjusts the learning rate according to the first-order moment estimate and the second-order moment estimate of the gradient and updates the parameters so that the loss function gradually decreases. S438. After each training time step, a portion of the parameters of the prediction network is copied to the target network to update the parameters of the target network. This can make the calculation of the target Q value more stable and avoid excessive fluctuations in the training process. S439. Repeat steps S432-S438 until the predetermined number of training rounds is reached or the network performance on the validation set no longer improves.

6. The method for path planning and dynamic scheduling of multiple plant protection vehicles in an orchard according to claim 5, characterized in that: In step S62, the RRT* algorithm is used to plan the local path, input the obstacle location, shape, range information and the slope, aspect, and terrain information of the hill from the dispatch center, and output the obstacle avoidance path, which also includes the following steps: S622, random sampling: randomly sample a point x_rand in the environment space; S623. Find the nearest node x_near from the tree T, and use x_near as the starting point to extend a certain distance in the direction to obtain a new node x_new. S624. Check whether the path between x_near and x_new collides with an obstacle based on the obstacle's position, shape, and range information. If a collision occurs, return to step S622 and continue to the next iteration. S625. Find the node x_min closest to x_new in the tree T and calculate the cost cost(x_min, x_new) from x_min to x_new. S4626. For nodes x_near_neighbors in tree T that are within a certain distance from x_new, calculate the cost cost(x_near, x_new) of connecting to x_new via x_near, and select the connection method with the lowest cost. S627, add x_new to the tree T, and update the connection relationship and cost information between nodes; S628. Optimize some nodes in the tree T with the target point goal as the goal, improve the quality of the path by adjusting the connection relationship and cost information, and repeat steps S622-S628 until a path is found or the maximum number of iterations is reached; S629: If the newly added node x_new is close to the target point goal, a feasible path is found and the algorithm ends.

Citation Information

Patent Citations

  • Path planning method and system for intelligent mowing robot

    CN116736845A

  • Autonomous robot system for steep terrain farming operations

    US20220151135A1