Agricultural multi-robot task allocation method and system based on reinforcement learning

Through a reinforcement learning-based method and combined with path planning and attention mechanism strategies to optimize the network, the problem of single goals and poor practicality in agricultural multi-robot task allocation is solved, and efficient and flexible task allocation and resource utilization are achieved in the agricultural environment.

CN120069407APending Publication Date: 2025-05-30INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510117544.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has problems such as single goals and poor practicality in the allocation of agricultural multi-robot tasks, especially in dynamic and complex agricultural environments, with high computational costs and uneven resource allocation.

Method used

A multi-robot scheduling and task allocation method based on reinforcement learning is proposed. The task path cost is calculated through the path planning algorithm, the task allocation objective function is established, and the attention mechanism strategy optimization network is used to optimize the distribution probability between nodes and vehicles, and the model is trained using the strategy gradient method to output the task allocation plan.

Benefits of technology

The agricultural scheduling needs are achieved in terms of shortest total path and balanced workload. The running time of the optimized model in actual use can meet real-time applications, improving the efficiency and flexibility of agricultural multi-robot task allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069407A_ABST
    Figure CN120069407A_ABST
Patent Text Reader

Abstract

The invention discloses an agricultural multi-robot task allocation method based on reinforcement learning, and the method comprises the steps: calculating the path cost of a farmland task based on a path planning algorithm of the farmland task for a plurality of to-be-allocated farmland task plots, and obtaining the path cost of the farmland task according to a path planning result; establishing a task allocation objective function of the agricultural vehicle cluster by taking the workload balance of task allocation and the minimum total path cost as constraints; based on an attention mechanism strategy optimization network of reinforcement learning, determining a distribution probability between nodes and vehicles, formulating a reward function with node workload balance and total path cost minimum constraint according to a target function, using a strategy gradient method to complete task distribution model training, and outputting a task distribution scheme of an agricultural vehicle cluster; and each vehicle in the agricultural vehicle cluster traverses and executes agricultural operation on a task plot according to a given task allocation scheme. The method and the system are used for an agricultural multi-robot and multi-task job scheduling demand scene, and an instant and reasonable robot and task allocation scheme is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to intelligent agricultural machinery and multi-robot scheduling methods, and particularly to an agricultural multi-robot task allocation method and system based on reinforcement learning. Background Art

[0002] Currently, with the development of global agricultural intelligence and digitization, agriculture has entered the fields of "unmanned farms" and "smart agriculture". These operations usually require multiple vehicles to cooperate to improve efficiency, which requires considering how to efficiently complete vehicle scheduling, reasonably allocate operation sequences, and improve the utilization rate of agricultural vehicle resources. To solve these problems, the issue of agricultural multi-robot task allocation has been proposed, which is the core technology of multi-robot systems and an important application in smart agriculture and large-scale farm management.

[0003] In the process of agricultural multi-robot task allocation methods, the completion time of the robot and workload balance are usually two key factors concerned in the actual operation scenario. The former is usually represented by the running time when the last robot stops working, which is directly related to the path cost. Traditional agricultural task allocation methods mainly rely on human experience for decision-making and lack effective allocation means.

[0004] The essence of the agricultural multi-robot task allocation method is to map multiple agricultural vehicles to task plots under specific operation constraints, which is a typical combinatorial optimization problem. Given that the multi-traveling salesman problem is the basic model for various combinatorial optimization scenarios, including vehicle routing planning (VRP), hot rolling scheduling, and global navigation satellite system measurement network design. Therefore, an attempt is made to solve the agricultural multi-robot task allocation method by referring to the solution method of MTSP under the constraints of actual agricultural operation requirements. Currently, the exploration of the agricultural multi-robot task allocation problem includes deterministic methods, heuristic algorithms, market-based, and learning-based strategies. However, the existing technologies have the following disadvantages:

[0005] 1) Deterministic methods aim to minimize the waiting time of working robots and agricultural energy consumption while ensuring compliance with the speed limits of the robots and the correct sequencing of tasks. Although these methods are very precise, due to their dependence on accurate mathematical modeling and environmental prediction, they often lack flexibility in dynamic and complex agricultural environments. Market-based methods, including dynamic task allocation methods for homogeneous agricultural robots, such as using the enhanced contract network algorithm, can reduce the time cost of job clusters by 30.20% to 34.09%. In addition, the heuristic-based clustering auction (HBCA) method introduces PFCI into the auction mechanism to achieve effective and efficient farmland task allocation. However, market-based methods usually have a high computational cost and may lead to uneven resource allocation in some cases, where some vehicles may be overloaded while others are idle.

[0006] 2) Heuristic algorithms, regarded as an effective and fast method, are widely used in agricultural multi-robot scheduling and allocation related work. Commonly used algorithms include genetic algorithms (GA), ant colony optimization (ACO), simulated annealing (SA), and artificial bee colony (ABC). According to the task allocation model, some literature has established a task allocation process based on the improved ACO algorithm considering factors such as supply-demand matching, agricultural vehicle operation ability, operation cycle, and path cost. Currently, some researchers have also completed the task allocation of multiple vehicles from the perspective of path planning using an optimized ACO algorithm. In addition, the agricultural multi-robot task allocation has been transformed into a multi-objective MTSP and solved using the NSGA-II algorithm. Or, based on NSGA-III and the improved ant colony algorithm, an intelligent scheduling method for agricultural multi-vehicle collaborative command has been proposed. However, the performance of heuristic algorithms often depends on the precise adjustment of parameters to adapt to the current scenario conditions and usually requires multiple iterations when applied.

[0007] In summary, the above defects not only lead to the consideration of a single objective function in agricultural multi-robot scheduling and task allocation methods and do not deeply combine with the actual needs of agriculture; the methods of the existing technology take too long time when facing large-scale operations.

[0008] Therefore, it is urgent to reformulate the agricultural multi-robot task allocation as an NWC-MTSP problem (multi-traveling salesman problem with workload constraints). According to the actual needs of agricultural vehicle operation, design and implement the required path planning algorithm, and establish a task allocation objective function based on the path planning results. And optimize the model, and the running time of the optimized model in actual use can meet real-time applications. Summary of the Invention

[0009] To solve the above problems in the prior art such as single objective and poor practicability, a method for agricultural multi-robot scheduling and task allocation based on reinforcement learning is proposed.

[0010] In a first aspect, an embodiment of the present application provides a method for agricultural multi-robot task allocation based on reinforcement learning. The method includes:

[0011] For multiple farmland task plots to be allocated, based on the path planning algorithm of the farmland tasks, calculate the path cost of the farmland tasks, and according to the path planning results, taking the workload balance of task allocation and the minimum total path cost as constraints, establish a task allocation objective function for the agricultural vehicle cluster;

[0012] Based on the attention mechanism policy optimization network of reinforcement learning, determine the allocation probability between nodes and vehicles, formulate a reward function according to the objective function, and use the policy gradient method to complete the training of the task allocation model, and output the task allocation plan for the agricultural vehicle cluster;

[0013] Each vehicle in the agricultural vehicle cluster traverses and performs agricultural operations on the task plot according to the given task allocation plan.

[0014] In a specific embodiment of the present invention, the above path planning algorithm based on farm tasks calculates the path cost of farm tasks, including:

[0015] Obtain the transfer path cost between plots, convert the map information of the farmland into a road network topology map in vector data format, perform transfer path search, use the starting point of the agricultural vehicle cluster as the starting task, and use the algorithm to search for the actual shortest paths of all task plots to construct the shortest path length matrix;

[0016] Obtain the working path cost within the plot, use the full-coverage path planning algorithm to calculate the path trajectories of the turning area and the actual operation area respectively, and sum the two to calculate the working path cost within the plot;

[0017] Calculate the workload wl assigned to each task according to the transfer path cost between plots and the working path cost within the plot.

[0018] In a specific embodiment of the present invention, the above attention mechanism policy optimization network based on reinforcement learning determines the allocation probability between nodes and vehicles, including:

[0019] Define a graph attention network, add the workload attribute wl to each network node according to the requirements of the task allocation plan, and generate an attention mechanism policy optimization network based on reinforcement learning;

[0020] After being processed by the policy optimization network, output the node feature matrix F M and the global feature representation G of the graph M ;

[0021] Based on the node feature matrix F M and the global feature representation G of the graph M , generate the embedding of each agricultural vehicle, perform the allocation of vehicles and plot task nodes, and calculate the probability that an agricultural vehicle selects each plot node.

[0022] In a specific embodiment of the present invention, the above reward function with node workload constraints is formulated according to the objective function, and the task allocation model training is completed using the policy gradient method, including:

[0023] Based on the allocation probability of selecting each plot node given by the policy optimization network, perform random action sampling;

[0024] Construct a reward function \(R(\theta,\lambda)\) based on the action sampling results, and use the classical Traveling Salesman Problem algorithm (LKH3) to complete the minimization optimization of the transfer distance of sub-vehicles, determine the transfer order between each node, calculate the parameter \(\theta\) using the policy gradient algorithm, and train to obtain the optimal policy; where \(\theta\) is the trained policy and \(\lambda\) is the Lagrangian relaxation factor.

[0025] Construct a loss function according to \(R(\theta,\lambda)\), and the loss function is the sum of the products of the negative logarithm probabilities of all samples and the rewards.

[0026] In a specific embodiment of the present invention, the above-mentioned node feature matrix \(F\) M and the global feature representation \(G\) of the graph M , generate the embedding of each agricultural vehicle, including:

[0027] Use the multi-head attention mechanism to generate the embedding for each vehicle, and input the node feature matrix \(F\) through a linear transformation M and the global graph feature \(G\) M , to obtain the query vector \(Q\), key vector \(K\), and value vector \(V\) of each head;

[0028] Calculate the attention scores through the dot product of the key vector \(K\) and the query vector \(Q\), weight each head, and obtain the embedding of each agricultural vehicle;

[0029] By concatenating the cumulative weighted aggregated value vector \(V\) and the attention values of all heads, obtain the feature dimension of the embedding of each agricultural vehicle, form the final \(V\_Embedding\), and dynamically adjust the merging weight values according to the importance of each element.

[0030] In a specific embodiment of the present invention, the above-mentioned calculation of the probability that an agricultural vehicle selects each plot node includes:

[0031] Adopt the generated node feature \(F\) M and \(V\_Embedding\) to calculate the probability that the vehicle selects each node;

[0032] The linear layer receives \(V\_Embedding\) and constructs the updated query vector \(Q'\) and key vector \(K'\);

[0033] Calculate the attention scores through the attention mechanism, adjust the attention scores through the scaled dot product attention mechanism, and convert the attention scores into the probability distribution of each plot node through the softmax function.

[0034] In a second aspect, an embodiment of the present application provides an agricultural multi-robot task allocation system based on reinforcement learning, which adopts the above-mentioned agricultural multi-robot task allocation method based on reinforcement learning. The system includes:

[0035] Objective function establishment module: For multiple farmland task plots to be allocated, based on the path planning algorithm of the farmland tasks, calculate the path cost of the farmland tasks, and according to the path planning results, with the workload balance of task allocation and the minimum total path cost as constraints, establish the task allocation objective function of the agricultural vehicle cluster;

[0036] Task allocation scheme output module: Based on the attention mechanism policy optimization network of reinforcement learning, determine the allocation probability between nodes and vehicles, formulate a reward function according to the objective function, and use the policy gradient method to complete the training of the task allocation model, and output the task allocation scheme of the agricultural vehicle cluster;

[0037] Task allocation scheme execution module: Used for each vehicle in the agricultural vehicle cluster to traverse and execute agricultural operations on the task plots according to the given task allocation scheme.

[0038] In a third aspect, an embodiment of the present application provides an agricultural multi-robot task allocation system, which includes: a server, a client, and an agricultural vehicle cluster;

[0039] When the server executes the program, it implements the steps of the above-mentioned agricultural multi-robot task allocation method based on reinforcement learning, and issues a control instruction;

[0040] The agricultural vehicle cluster is communicatively connected to the server, receives the control instruction issued by the server, and completes the steps of the above-mentioned agricultural multi-robot task allocation method based on reinforcement learning;

[0041] The client is communicatively connected to the server and the agricultural vehicle cluster, and is used to receive control instructions and monitor and manage the agricultural vehicle cluster.

[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the above-mentioned agricultural multi-robot task allocation method based on reinforcement learning.

[0043] In a fifth aspect, an embodiment of the present application provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the above-mentioned agricultural multi-robot task allocation method based on reinforcement learning.

[0044] Compared with the related prior art, it has the following outstanding beneficial effects:

[0045] 1) The method of the present invention proposes to set up two objective functions according to the actual needs of agricultural multi-robot task allocation; at the same time, it ensures the agricultural scheduling requirements in terms of both the shortest total path and workload balance;

[0046] 2) The method of the present invention proposes to optimize the attention mechanism policy network to determine the allocation probability between nodes and vehicles. A reward function with node workload constraints is formulated according to the objective function, and the model training is completed using the policy gradient method. The running time of the optimized model in actual use can meet real-time applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0048] Figure 1 It is a schematic flow chart of the agricultural multi-robot task allocation method of the present invention;

[0049] Figure 2 It is a schematic flow chart of the working path calculation method within the plot in the embodiment of the present invention;

[0050] Figure 3 It is a schematic diagram of the agricultural multi-robot task allocation system in the embodiment of the present invention;

[0051] Figure 4 It is a schematic diagram of the agricultural multi-robot task allocation system in the embodiment of the present invention;

[0052] Figure 5 It is a schematic diagram of the computer hardware of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] It should be noted that the processor described in the present invention is the control center of the electronic device, which can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0054] Optionally, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0055] In a specific implementation, as an embodiment, the processor may include one or more CPUs. Each of these processors may be a single-core processor or a multi-core processor. The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The electronic device may include: a server, a desktop computer, a laptop computer, a smart phone, a tablet computer, an embedded computer, etc., where the embedded computer includes vehicles and robots, etc.

[0056] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation manner may refer to the above method embodiment and will not be elaborated here.

[0057] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto. The actual knowledge structure recognition device may include more or fewer components than those shown in the drawings, or combine certain components, or have different component arrangements.

[0058] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other arbitrary combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on the computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that the computer can access or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium may be a solid-state drive.

[0059] It should also be understood that the term "and / or" in this text is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this text generally represents an "or" relationship between the preceding and following associated objects, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0060] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following items (pieces)" or similar expressions refer to any combination of these items, including any combination of single items (pieces) or plural items (pieces). For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0061] It should also be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0062] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0063] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0064] In addition, in each embodiment of the present invention, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0065] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0066] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are hereinafter given and detailed descriptions are made in conjunction with the accompanying drawings of the specification. The present specification discloses one or more embodiments including the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.

[0067] The following is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they are not elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0068] The method of the present invention aims to reformulate the agricultural multi-robot task allocation as a NWC-MTSP problem (multi-traveling salesman problem with node workload constraints). According to the actual requirements of the operation of agricultural vehicles, design and implement the required path planning algorithm, and establish a task allocation objective function based on the path planning results; optimize the attention mechanism policy network to determine the allocation probability between nodes and vehicles. Develop a reward function with node workload constraints according to the objective function, and use the policy gradient method to complete the model training. The running time of the optimized model in actual use can meet the requirements of real-time applications.

[0069] The method of the present invention analyzes the actual agricultural requirements, takes agricultural plots as task nodes to be executed, and considers the workload attributes of the task plots. With minimizing the maximum sub-robot time cost as the main objective function and workload balance as the constraint, the task allocation problem is transformed into a multi-traveling salesman problem with node workload constraints (NWC-MTSP). Then, an attention mechanism policy optimization network based on reinforcement learning (NWC-APONet) is proposed to minimize the total task completion time while avoiding vehicle overloading or excessive idling.

[0070] The agricultural multi-robot scheduling and task allocation system proposed by the present invention is mainly used for the operation scheduling requirements scenarios of agricultural multi-robots and multi-tasks, and provides an immediate and reasonable robot and task allocation scheme for farm administrators. This system applies a centralized control architecture, that is, it consists of a main processor and each sub-robot.

[0071] The main scenario structure of the same type of agricultural multi-robot task allocation includes three main participants: the server, the client, and the agricultural vehicle. Among them, the task allocation and path planning methods are deployed on the server, and the agricultural robot will be equipped with an on-vehicle controller with high computing power, a positioning module, and a communication module. In the process of agricultural task allocation, it is required to give an allocation plan P for n task plots T: {t 1 , t 2 ,......., t n} and m robots of the same type R: {R 1 , R 2 ,......, R M}, where the allocation plan P includes the mapping relationship and execution order between the agricultural vehicle and the task target plot, and is sent to the agricultural vehicle to complete the specific execution to improve the operation efficiency. The agricultural task is described as t = {x, y, wl}, where x and y represent the two-dimensional coordinates of the plot, and wl represents the workload of the task plot.

[0072] The task allocation of the agricultural vehicle cluster requires each vehicle to start from the same starting point (hangar), traverse and execute operations on the task plots according to the given allocation plan P, and finally return to the starting point. The allocation plan P is as follows:

[0073]

[0074] Among them, a 0 is the starting point, a (i,j) is the jth task assigned to robot i, and t num is the total number of tasks assigned to robot i.

[0075] Agricultural operations usually focus on the operation time and the utilization rate of agricultural robots. For multiple robots performing operation tasks, the task completion time usually refers to the time taken by the last robot. Therefore, to ensure the shortest completion time of agricultural operations, we consider these issues from the perspective of "path cost", with the objective of minimizing the longest path cost of sub-vehicles, and at the same time achieving the effect of balancing the operation time or workload distribution of each robot as much as possible.

[0076] The following describes the method of the embodiments of the present application in detail with specific embodiments:

[0077] Embodiment 1

[0078] As Figure 1 shown, the embodiment of the present application provides a method for task allocation of agricultural multi-robots based on reinforcement learning, and the method includes:

[0079] Step 101: For multiple farmland task plots to be allocated, based on the path planning algorithm of the farmland tasks, calculate the path cost of the farmland tasks, and according to the path planning result, with the workload balance of task allocation and the minimum total path cost as constraints, establish the task allocation objective function of the agricultural vehicle cluster;

[0080] Step 102: Based on the attention mechanism policy optimization network of reinforcement learning, determine the allocation probability between nodes and vehicles, formulate a reward function according to the objective function, and use the policy gradient method to complete the training of the task allocation model, and output the task allocation scheme of the agricultural vehicle cluster;

[0081] Step 103: Each vehicle in the agricultural vehicle cluster traverses and performs agricultural operations on the task plot according to the given task allocation scheme.

[0082] In the specific embodiment of the present invention, in the above step 101: based on the path planning algorithm of the farmland tasks, calculating the path cost of the farmland tasks includes:

[0083] 1) Obtain the transfer path cost between plots, convert the map information of the farmland into a road network topology map in vector data format, perform transfer path search, use the starting point of the agricultural vehicle cluster as the starting task, and use the algorithm to search for the actual shortest paths of all task plots, and construct the shortest path length matrix;

[0084] In the specific embodiment of the present invention, obtaining the path cost includes:

[0085] First, it is necessary to calculate the path cost, and then obtain the target time cost result according to the speed of the robot. In the current operation scenario, the path cost is mainly divided into two parts, the path cost of transfer between plots and the path cost of operation within the plot. It is necessary to calculate all the path costs of each robot separately.

[0086] In the specific embodiments of the present invention, obtaining the transfer path cost between plots includes: calculating the transfer path cost between plots based on farmland map data, which has high-precision GPS positioning data and road network information, and the positioning accuracy can reach the centimeter level. To save storage space and calculation amount, we obtained the GIS map of the farmland and converted it into a road network topology map in vector data format for transfer path search.

[0087] Here, the starting point (warehouse) is regarded as the 0th task, and the central coordinates of the plot are regarded as the coordinate position of the plot. The Floyd algorithm is used to search for the actual shortest path that can be taken by n + 1 task plots, and a shortest path length matrix Dist is constructed. Dist is shown as follows:

[0088]

[0089] Among them, d (i,j) is the actual drivable distance between the ith task plot and the jth task plot, in km, i, j = 0, 1,..., n.

[0090] When initializing the distance matrix, if the fields of the two task plots of i and j are adjacent, then d (i,j) = 0; if the two task plots are not adjacent and need to pass through more than 3 plots, a relatively large integer I (such as I = 1000) is used to replace the distance. This measure aims to increase the transfer cost between non-connected plots to avoid generating overly complex transfer paths for plot tasks.

[0091] 2) Obtain the working path cost within the plot. The full-coverage path planning algorithm is used to calculate the path trajectories of the turning area and the actual operation area respectively, and the two are summed to calculate the working path cost within the plot;

[0092] 3) Calculate the workload wl assigned to each task according to the transfer path cost between plots and the working path cost within the plot.

[0093] As Figure 2 shown, according to the needs of farmland operations, the full-coverage polygon path planning algorithm is mainly used for farmland operations, and the specific embodiments of the present invention implement the full-coverage path planning algorithm. At the same time, turning space needs to be reserved for the operation of agricultural vehicles. Therefore, the path within the plot includes two parts: the turning area and the actual operation area. The turning area is used to meet the turning and row-changing of agricultural vehicles; the operation area consists of a series of parallel trajectory lines, where agricultural vehicles perform operations such as plowing and sowing. The parallel trajectory lines in the operation area are composed of a series of discrete path points (path_no i, consisting of i = {1, 2, …, k}). To obtain the working path points, we first need to get the two endpoints of each route, namely the turning points. When obtaining the turning points of each route, the initial edge is used as the base edge, and the working area is translated and segmented at intervals of the tillage width (w). The base edge point is the starting point of the path point sequence (path_no 0 , path_no 1 ). Each time the base edge is translated parallelly, two intersection points (a, b) with the boundary are obtained. Compare the Euclidean distances between the two intersection points and the last point in the sequence. The calculation formula is:

[0094]

[0095] where path_no i is the last point saved to the path point sequence, (path_no ix , path_no iy ) represents the x, y coordinate values of this point. And a, b are the intersection points with the boundary after translating the base edge, (a x , a y ), (b x , b y ) are the intersection points of the parallel line with the boundary, that is, the two endpoints of the parallel line segment.

[0096] Among the two new intersection points, the point (path_no i ) with the shorter distance from the previous point will be the point first added to the sequence. That is, if d 1 < d 2 , then path_no i+1 = a, path_no i+2 = b; if d 2 < d 1 , then path_no i+1 = b, path_no i+2 = a. According to the above process, the path points are arranged and collected to obtain the complete working path point sequence of the task plot. The cumulative sum of the Euclidean distances between two adjacent points in the sequence is the working distance of the task plot:

[0097]

[0098] In the turning area, since the kinematic model of the agricultural vehicle is the Ackerman structure, the vehicle mainly uses a fishtail route to complete the turn. Combining the width h of the U-turn area, the minimum turning radius r of the vehicle, and the working width w, the U-turn route is planned. The U-turn path is divided into the path l 1 leaving the working area, the reverse path l 2 and the path l 3 entering the working area, and their relationship is as follows:

[0099]

[0100] The calculation formula for the turning distance is:

[0101] plot turnaround =(path num - 1)*(l 1 +l 2 +l 3 )

[0102] The path cost in each task plot is the sum of the transfer distance and the actual working distance:

[0103] cover_path = plot work +plot turnaround

[0104] The length of the working path assigned to each task is the workload (wl) of the task assignment, which is used to calculate the subsequent task assignment process.

[0105] In the specific embodiment of the present invention, in step 102 above: Based on the attention mechanism policy optimization network of reinforcement learning, determining the allocation probability between nodes and vehicles includes:

[0106] 1) Define a graph attention network. According to the requirements of the task assignment scheme, a workload attribute wl is added to each network node to generate an attention mechanism policy optimization network based on reinforcement learning;

[0107] 2) After being processed by the policy optimization network, output the node feature matrix F M and the global feature representation G M ;

[0108] 3) Based on the node feature matrix F M and the global feature representation G M , generate the embedding of each agricultural vehicle, perform the allocation of vehicles and plot task nodes, and calculate the probability that the agricultural vehicle selects each plot node.

[0109] The graph attention network (GAT) is a special type of graph neural network, and its main improvement is the message passing method. It introduces a learnable attention mechanism that can assign weights between each source node and target node, enabling the node to aggregate the information of neighbor nodes and determine which neighbor node's information is more important, rather than aggregating the information of all neighbor nodes with the same weight. In the graph attention network defined in the present invention, according to the actual needs, a workload attribute (wl) is added to each node. By stacking three convolutional layers, each convolutional layer is followed by a batch normalization layer to stabilize and accelerate the training of the deep network. The new feature h' of node i iis the weighted sum of the features of all neighbor nodes, with the weight being the attention factor α ij For each node i, its updated feature h' i can be expressed as:

[0110]

[0111] where N(i) is the set of neighbor nodes of node i, W is the learnable weight matrix, and α ij is the attention factor, representing the importance of node j to node i, and h i is the feature vector before update.

[0112] In addition, to accelerate the training process and improve the generalization ability of the model, we apply batch normalization after each convolutional layer. At the same time, global average pooling is applied to all node features to obtain the global representation of the entire graph. After being processed by GAT, the node feature matrix F M and the global feature representation G M of the graph are given.

[0113] In the specific embodiment of the present invention, in step 102 above: formulating a reward function with node workload constraints according to the objective function and using the policy gradient method to complete the training of the task allocation model includes:

[0114] 1) Conducting random action sampling based on the allocation probability of selecting each plot node given by the policy optimization network;

[0115] 2) Constructing the reward function R(θ,λ) based on the action sampling result, using the traveling salesman problem algorithm to complete the minimization optimization of the transfer distance of the sub-vehicle, determining the transfer order between each node, calculating the parameter θ using the policy gradient algorithm, and training to obtain the optimal policy; where θ is the trained policy and λ is the Lagrange relaxation factor;

[0116] 3) Constructing a loss function according to R(θ,λ), and the loss function is the sum of the product of the negative log probability of all samples and the reward.

[0117] In the specific embodiment of the present invention, in step 102 above: generating the embedding of each agricultural vehicle based on the node feature matrix F M and the global feature representation G M of the graph includes:

[0118] 1) Using the multi-head attention mechanism to generate the embedding for each vehicle, and obtaining the query vector Q, key vector K, and value vector V of each head by linearly transforming the input node feature matrix F M and the global graph feature G M ;

[0119] 2) Calculate the attention scores through the dot product of the key vector K and the query vector Q, and weight each head to obtain the embedding of each agricultural vehicle;

[0120] 3) By concatenating the aggregated weighted value vector V and the attention values of all heads, obtain the feature dimension of the embedding of each agricultural vehicle, form the final V_Embedding, and dynamically adjust the merging weight values according to the importance of each element.

[0121] The graph attention policy network determines the allocation of vehicles to task nodes through a two-stage process. The first stage is to generate embeddings for each vehicle, and the second stage is to perform the allocation of vehicles and task nodes, giving the probability that each vehicle selects each node.

[0122] Among them, embedding refers to the process of converting a certain type of input data (such as text, image, sound, etc.) into a dense numerical vector.

[0123] These vectors usually contain many dimensions, and each dimension represents some abstract feature or attribute of the input data.

[0124] The purpose of embedding is to convert the actual input into a format that enables the computer to process and learn more effectively;

[0125] Among them, V_Embedding refers to the embedding value of the vehicle;

[0126] First, use the multi-head attention mechanism to generate embeddings for each vehicle. The query vector (Q i ), key vector (K i ), and value vector (V i ) of each head are input through linear transformation of the node feature F M and the global graph feature G M to obtain:

[0127]

[0128] Here, W i Q , W i k , are learnable weight matrices, b i Q , b i K , is the bias term, a parameter used to adjust the weights during the learning process. ReLU is a non-linear activation function used to increase the model's expressive power. Each head focuses on different aspects of the input data, enabling the model to capture information in multiple representation subspaces.

[0129] Then, the attention scores are calculated through the dot product of the key K and the query Q. Subsequently, each head is weighted to obtain the embedding of each vehicle. By aggregating the cumulative weighted values of V i and the attention values VE of all heads i are concatenated. Finally, the feature dimension of the vehicle embedding is obtained, thereby forming the final V_Embedding. This process is not just simply merging information but also dynamically adjusting the merging weights according to the importance of each element (represented by the attention scores). This mechanism enables the model to focus on the most relevant parts of the graph.

[0130]

[0131] V_Embedding = W o (Concat(VE 1 , …, VE h )) + b o

[0132] where d ki is the dimension of the key K, the key vector K i and the query vector Q i , T is the transpose, is the learnable weight matrix, are the weights and biases of the projection layer, h is the number of heads, and Concat d represents the concatenation operation on the feature dimension.

[0133] In the specific embodiments of the present invention, calculating the probability that the agricultural vehicle selects each plot node includes:

[0134] 1) Using the generated node feature F M and V_Embedding, calculate the probability that the vehicle selects each node;

[0135] 2) The linear layer receives V_Embedding and constructs the updated query vector Q' and key vector K';

[0136] 3) Calculate the attention scores through the attention mechanism, adjust the attention scores through the scaled dot product attention mechanism, and convert the attention scores into the probability distribution of each plot node through the softmax function.

[0137] I) In the second stage of the policy network, the specific embodiments of the present invention use the node feature F generated in the first stageM and V_Embedding to calculate the probability of the vehicle selecting each node. Specifically, the linear layer receives V_Embedding to construct the query vector Q' and the key vector K'. We still calculate the attention scores through the attention mechanism, adjust the attention scores through the scaled dot-product attention mechanism, and then convert the attention scores into a probability distribution through the softmax function. This process can be expressed as:

[0138]

[0139] where Q' represents the query vector, K' represents the key vector, V' represents the value vector, and d' k is the dimension of the key, and T represents the transpose of these vectors. The above formula calculates the similarity between the query vector and the key vector, and then scales it by d' k to avoid the problem of gradient disappearance caused by an overly large dot product.

[0140] To consider the actual node distance or other factors, we adjust the attention scores in the network:

[0141] Adjusted(U') = U' – dist_factor * Dist

[0142] where dist_factor is a hyperparameter used to adjust the influence of distance on the attention scores.

[0143] Finally, the adjusted attention scores are converted into a probability distribution through the activation function and softmax:

[0144] π θ = softmax(clip_f * tanh(Adjusted(U'))

[0145] where clip_f is a clipping parameter used to adjust the output range of the activation function, thereby affecting the decision-making tendency of the model. π θ reflects the probability that the vehicle selects the current node as the next target, representing the final output of selecting this node in the current state. This attention-based method enables the model to flexibly adapt to different graph structures and the effective allocation of vehicles to nodes.

[0146] II) Training based on policy gradient: The training process is achieved by optimizing the reward function, aiming to minimize the maximum path cost of the sub-vehicles while trying to maintain the balance of the workload distribution.

[0147] Based on the allocation probability given by the policy network, random action sampling is performed. The sampling result (a sample)Indicates which vehicles the task nodes are assigned to. The sampling process ensures that task nodes with high probability have a greater chance of being selected, while allowing a certain degree of exploration, that is, occasionally selecting task nodes with low probability.

[0148] Construct a reward function R(θ,λ) based on the action sampling results, and use LKH3 (a classical traveling salesman problem handling algorithm) to complete the minimization optimization of the sub-vehicle transfer distance. Based on the allocation results, use the classical traveling salesman processing algorithm lkh3 to further optimize the traversal sequence of a single robot after allocation, and determine the transfer order between each node. Construct a loss function according to R(θ,λ). The loss function is the sum of the product of the negative logarithm probability of all samples and the reward. These sums are accumulated in batches. To achieve the training goal, this paper uses the policy gradient algorithm to calculate the parameter θ and train to obtain the optimal policy θ * .

[0149]

[0150] Here, η is the learning rate, θ * and θ represent the parameters before and after update respectively, is the gradient of the loss function with respect to θ.

[0151] Data verification is performed after every five iterations to check the performance of the network. The validation dataset contains 512 batches, and the data in each batch consists of random decimals between 0 and 1. This randomness helps to ensure the generalization ability of the model. At each validation, calculate the average reward value of all validation batches passing through the network as the result of this validation. This reward value reflects the matching degree between the network output and the expected output. If the average reward value of a certain validation is lower than the previous result, it indicates that the performance of the model has improved in the most recent iteration. In this case, the system will automatically save the current network parameters. Through repeated iteration and validation, the optimal network parameter settings (θ * ) can be effectively found and saved, and at the same time the model has better generalization ability.

[0152] Embodiment 2

[0153] As Figure 3 shown, the embodiment of the present application provides an agricultural multi-robot task allocation system based on reinforcement learning, adopting the agricultural multi-robot task allocation method based on reinforcement learning as described above. The system includes:

[0154] Objective function establishment module 201: For multiple farmland task plots to be allocated, based on the path planning algorithm of the farmland tasks, calculate the path cost of the farmland tasks, and according to the path planning results, with the workload balance of task allocation and the minimum total path cost as constraints, establish the task allocation objective function of the agricultural vehicle cluster;

[0155] Task Assignment Scheme Output Module 202: Based on the attention mechanism policy optimization network of reinforcement learning, determine the assignment probability between nodes and vehicles, formulate a reward function according to the objective function, and use the policy gradient method to complete the training of the task assignment model, and output the task assignment scheme of the agricultural vehicle cluster;

[0156] Task Assignment Scheme Execution Module 203: For each vehicle in the agricultural vehicle cluster, traverse and execute agricultural operations on the task plot according to the given task assignment scheme.

[0157] Embodiment III

[0158] As Figure 4 shown, an embodiment of the present application provides an agricultural multi-robot task assignment system, which includes: a server 301, a client 302, and an agricultural vehicle cluster 303;

[0159] When the server 301 executes the program, it implements the steps of the above-mentioned agricultural multi-robot task assignment method based on reinforcement learning, and issues a control instruction;

[0160] The agricultural vehicle cluster 303 is communicatively connected to the server, receives the control instruction issued by the server, and completes the steps of the above-mentioned agricultural multi-robot task assignment method based on reinforcement learning;

[0161] The client 302 is communicatively connected to the server and the agricultural vehicle cluster, and is used to receive the control instruction and monitor and manage the agricultural vehicle cluster.

[0162] Embodiment IV

[0163] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the above-mentioned agricultural multi-robot task assignment method based on reinforcement learning.

[0164] Embodiment V

[0165] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the above-mentioned agricultural multi-robot task assignment method based on reinforcement learning.

[0166] In addition, the agricultural multi-robot task assignment method based on reinforcement learning described in Figure 1 the embodiments of the present application can be implemented by an electronic device, such as a computer device. Figure 5 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application.

[0167] In some of these embodiments, the computer device may further include a communication interface 83 and a bus 80. Among them, as Figure 5 shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other.

[0168] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.

[0169] The memory 82 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.

[0170] The processor 81 reads and executes the computer program instructions stored in the memory 82 to implement any one of the above-mentioned reinforcement learning-based agricultural multi-robot task allocation methods.

[0171] The technical features of the above-mentioned embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0172] The above-mentioned embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A method for allocating agricultural multi-robot tasks based on reinforcement learning, characterized in that: The method comprises: For multiple farmland task plots to be assigned, based on the path planning algorithm of the farmland task, the path cost of the farmland task is calculated, and according to the path planning result, the task allocation objective function of the agricultural vehicle cluster is established with the workload balance of task allocation and the minimum total path cost as constraints; The attention mechanism strategy based on reinforcement learning optimizes the network, determines the allocation probability between nodes and vehicles, formulates a reward function according to the objective function, and uses the policy gradient method to complete the task allocation model training, and outputs the task allocation plan of the agricultural vehicle cluster; Each vehicle of the agricultural vehicle cluster traverses and performs agricultural operations on the task plot according to the given task allocation plan.

2. The agricultural multi-robot task allocation method based on reinforcement learning according to claim 1 is characterized in that: The path planning algorithm based on the farmland task calculates the path cost of the farmland task, including: Obtain the transfer path cost between plots, convert the map information of the farmland into a road network topology in vector data format, perform transfer path search, take the starting point of the agricultural vehicle cluster as the starting point task, use the algorithm to search for the actual shortest path of all task plots, and construct the shortest path length matrix; Obtain the cost of the working path within the plot, use the full coverage path planning algorithm, calculate the path trajectories of the turning area and the actual working area respectively, and sum the two to calculate the cost of the working path within the plot; The workload wl assigned to each task is calculated based on the transfer path cost between the plots and the work path cost within the plot.

3. The agricultural multi-robot task allocation method based on reinforcement learning according to claim 2 is characterized in that: The reinforcement learning-based attention mechanism strategy optimizes the network and determines the allocation probability between nodes and vehicles, including: Define the graph attention network, add the workload attribute wl to each network node according to the task allocation scheme requirements, and generate the attention mechanism strategy optimization network based on reinforcement learning; After the strategy optimization network is processed, the output node feature matrix F M The global feature representation G of the graph M ; Based on the node feature matrix F M The global feature representation G of the graph M , generate embedding for each agricultural vehicle, allocate vehicles and plot task nodes, and calculate the probability of the agricultural vehicle selecting each plot node.

4. The agricultural multi-robot task allocation method based on reinforcement learning according to claim 2 is characterized in that: According to the objective function, a reward function with node workload balance and total path cost minimum constraints is formulated, and the task allocation model training is completed using the policy gradient method, including: Based on the allocation probability of selecting each plot node given by the strategy optimization network, random action sampling is performed; Based on the action sampling results, a reward function R(θ,λ) is constructed, and the classic traveling salesman problem algorithm is used to minimize the transfer distance of the sub-vehicles, determine the transfer order between nodes, use the policy gradient algorithm to calculate the parameter θ, and train to obtain the optimal strategy; where θ is the trained strategy and λ is the Lagrangian relaxation factor; The loss function is constructed based on R(θ,λ), which is the sum of the products of the negative log probability of all samples and the reward.

5. The agricultural multi-robot task allocation method based on reinforcement learning according to claim 3 is characterized in that: Based on the node feature matrix F M The global feature representation G of the graph M , generate embedding for each agricultural vehicle, including: Using the multi-head attention mechanism, embedding is generated for each car, and the node feature matrix F is input through linear transformation M and the global graph feature G M , obtain the query vector Q, key vector K and value vector V of each head; The attention score is calculated by the dot product of the key vector K and the query vector Q, and each head is weighted to obtain the embedding of each agricultural vehicle; By concatenating the cumulative weighted aggregated value vector V and the attention values ​​of all heads, the feature dimensions of each agricultural vehicle embedding are obtained to form the final V_Embedding, and the combined weight value is dynamically adjusted according to the importance of each element.

6. The agricultural multi-robot task allocation method based on reinforcement learning according to claim 3 is characterized in that: The calculating the probability of the agricultural vehicle selecting each plot node includes: The node feature F generated by M and V_Embedding, which calculates the probability of the vehicle selecting each node; The linear layer receives the V_Embedding and constructs the updated query vector Q' and key vector K'; The attention score is calculated by the attention mechanism, adjusted by the scaled dot product attention mechanism, and converted into a probability distribution for each plot node by the softmax function.

7. An agricultural multi-robot task allocation system based on reinforcement learning, using the agricultural multi-robot task allocation method based on reinforcement learning as described in any one of claims 1-6, characterized in that: The system comprises: Objective function establishment module: for multiple farmland task plots to be assigned, based on the path planning algorithm of the farmland task, the path cost of the farmland task is calculated, and according to the path planning result, the task allocation objective function of the agricultural vehicle cluster is established with the workload balance of task allocation and the minimum total path cost as constraints; Task allocation scheme output module: optimizes the network based on the attention mechanism strategy of reinforcement learning, determines the allocation probability between nodes and vehicles, formulates the reward function according to the objective function, and uses the policy gradient method to complete the task allocation model training, and outputs the task allocation scheme of the agricultural vehicle cluster; Task allocation scheme execution module: each vehicle in the agricultural vehicle cluster traverses and performs agricultural operations on the task plot according to the given task allocation scheme.

8. An agricultural multi-robot task allocation system, characterized in that: The system comprises: a server, a client and an agricultural vehicle cluster; When the server executes the program, the steps of the agricultural multi-robot task allocation method based on reinforcement learning as described in any one of claims 1 to 6 are implemented, and a control instruction is issued; The agricultural vehicle cluster is communicatively connected to the server, receives control instructions issued by the server, and completes the steps of the agricultural multi-robot task allocation method based on reinforcement learning as described in any one of claims 1 to 6; The client is communicatively connected to the server and the agricultural vehicle cluster, and is used to receive the control instruction and monitor and manage the agricultural vehicle cluster.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the agricultural multi-robot task allocation method based on reinforcement learning described in any one of claims 1-6 are implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the agricultural multi-robot task allocation method based on reinforcement learning as described in any one of claims 1 to 7 are implemented.