Multi-agent path planning method and system based on parallel computing and reinforcement learning
By adopting parallel computing and reinforcement learning methods in multi-agent path planning, and using technical means such as digital maps and convolutional neural networks, the problem of excessively long path planning of multi-agent path planning is solved, and efficient and safe path planning is achieved.
Patent Information
- Application Number
- CN202510076021.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, when planning multi-agent paths, the planning time is too long, resulting in increased risks and it is difficult to optimize the path planning process.
The multi-agent path planning method based on parallel computing and reinforcement learning is adopted, and the path planning of multi-agents is optimized through technical means such as digital maps, non-feasible area screening, and convolutional neural network path optimization model.
Reduce the calculation amount, improve planning efficiency, ensure that the generated paths comply with actual environmental constraints, narrow the search range, and significantly accelerate the path planning process.
Smart Images

Figure CN120043526A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of path planning technology, and in particular to a multi-agent path planning method and system based on parallel computing and reinforcement learning. Background Art
[0002] With the continuous development of artificial intelligence and automation technology, the application fields of multi-agent path planning methods are becoming more and more extensive.
[0003] In existing technologies, path planning for a single agent can usually achieve good results, but when it is extended to multi-agent path planning, the planning time will increase dramatically, which will lead to increased risks. Taking the path planning of an autonomous vehicle as an example, if the path planning time is too long, the car may not be able to respond to traffic changes ahead in time, thereby increasing the risk of driving.
[0004] It can be seen that how to optimize the path planning of multiple agents has become a technical problem that needs to be urgently solved by technical personnel in this field. Summary of the invention
[0005] The present invention provides a multi-agent path planning method and system based on parallel computing and reinforcement learning to solve the technical problem of improper multi-agent path planning and to optimize the multi-agent path planning process.
[0006] In order to solve the above technical problems, an embodiment of the present invention provides a multi-agent path planning method based on parallel computing and reinforcement learning, which is applied to the path planning process of multiple agents in a target area, including:
[0007] Determine a digitized map of the target area, and perform coordinate conversion processing on each position point in the digitized map to obtain a coordinate matrix of the target area;
[0008] Determining a non-feasible area of the target area based on the environmental data of the target area;
[0009] Determine all feasible path data of the target area according to the coordinate matrix and the infeasible area, wherein the determination process is configured to start from the terminal coordinates corresponding to the agent in the coordinate matrix, use the infeasible area to sequentially screen neighboring coordinate points circumscribed to the terminal coordinates, and determine a feasible path according to the screening results;
[0010] Constructing a convolutional neural network path optimization model based on a reinforcement learning algorithm, inputting all the feasible path data into the convolutional neural network path optimization model, and obtaining the target optimal path for each of the intelligent agents;
[0011] Execute the path planning solution generated by each of the target optimal paths.
[0012] As one preferred solution, determining the non-feasible area of the target area based on the environmental data of the target area includes:
[0013] Based on the environmental data of the target area, an environmental model of the target area is constructed, and obstacles in the target area are analyzed using the environmental model to determine a non-feasible area of the target area.
[0014] As one of the preferred solutions, inputting all the feasible path data into the convolutional neural network path optimization model to obtain the target optimal path of each intelligent agent includes:
[0015] Inputting all the feasible path data into the convolutional neural network path optimization model to obtain the feature vector and weight parameter corresponding to each of the intelligent agents;
[0016] Using a multi-layer perceptron to calculate the feasible path of each of the intelligent agents to obtain coding information, using the weight parameters of the intelligent agents to calculate the coding information corresponding to the intelligent agents to obtain the overall information of all the intelligent agents;
[0017] According to the feature vector and the overall information of all the intelligent agents, the optimal target path of each of the intelligent agents is determined.
[0018] As one of the preferred solutions, the convolutional neural network path optimization model is constructed based on the reinforcement learning algorithm, and the construction process includes:
[0019] According to the environmental data of the intelligent agent and the target area, the state space, action space and reward function of the intelligent agent reflecting the behavior are determined, and the convolutional neural network path optimization model is constructed according to the state space, the action space and the reward function using the reinforcement learning algorithm.
[0020] As one of the preferred solutions, after obtaining the target optimal path of each of the intelligent agents, the multi-agent path planning method based on parallel computing and reinforcement learning also includes:
[0021] The feasibility of the target optimal path is verified by using a simulated annealing algorithm to optimize the target optimal path of the agent.
[0022] Another embodiment of the present invention provides a multi-agent path planning system based on parallel computing and reinforcement learning, which is applied to the path planning process of multiple agents in a target area, including:
[0023] A conversion module, used to determine a digitized map of the target area, and perform coordinate conversion processing on each position point in the digitized map to obtain a coordinate matrix of the target area;
[0024] A determination module, configured to determine a non-feasible area of the target area based on the environmental data of the target area;
[0025] A calculation module, used for determining all feasible path data of the target area according to the coordinate matrix and the infeasible area, wherein the determination process is configured to start from the terminal coordinates corresponding to the agent in the coordinate matrix, and sequentially screen neighboring coordinate points circumscribed to the terminal coordinates with the infeasible area, and determine a feasible path according to the screening results;
[0026] A construction module is used to construct a convolutional neural network path optimization model based on a reinforcement learning algorithm, input all the feasible path data into the convolutional neural network path optimization model, and obtain the target optimal path of each of the intelligent agents;
[0027] An execution module is used to execute the path planning solution generated by each of the target optimal paths.
[0028] As one preferred solution, determining the non-feasible area of the target area based on the environmental data of the target area includes:
[0029] Based on the environmental data of the target area, an environmental model of the target area is constructed, and obstacles in the target area are analyzed using the environmental model to determine a non-feasible area of the target area.
[0030] As one of the preferred solutions, inputting all the feasible path data into the convolutional neural network path optimization model to obtain the target optimal path of each intelligent agent includes:
[0031] Inputting all the feasible path data into the convolutional neural network path optimization model to obtain the feature vector and weight parameter corresponding to each of the intelligent agents;
[0032] Using a multi-layer perceptron to calculate the feasible path of each of the intelligent agents to obtain coding information, using the weight parameters of the intelligent agents to calculate the coding information corresponding to the intelligent agents to obtain the overall information of all the intelligent agents;
[0033] According to the feature vector and the overall information of all the intelligent agents, the optimal target path of each of the intelligent agents is determined.
[0034] As one of the preferred solutions, the convolutional neural network path optimization model is constructed based on the reinforcement learning algorithm, and the construction process includes:
[0035] According to the environmental data of the intelligent agent and the target area, the state space, action space and reward function of the intelligent agent reflecting the behavior are determined, and the convolutional neural network path optimization model is constructed according to the state space, the action space and the reward function using the reinforcement learning algorithm.
[0036] As one of the preferred solutions, after obtaining the target optimal path of each of the intelligent agents, the multi-agent path planning system based on parallel computing and reinforcement learning further includes:
[0037] The optimization module is used to verify the feasibility of the target optimal path by using a simulated annealing algorithm to optimize the target optimal path of the intelligent agent.
[0038] Compared with the prior art, the embodiments of the present invention have the following advantages:
[0039] (1) The present invention reduces the amount of calculation and improves planning efficiency by optimizing the path planning of multiple agents;
[0040] (2) The present invention can ensure that the generated path is a feasible path that meets the actual environmental constraints by screening neighbor coordinate points, thereby narrowing the search range and improving the efficiency of path planning; using convolutional neural networks for path optimization can make full use of the parallel computing and learning capabilities of neural networks to further accelerate the path planning process. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flowchart of a multi-agent path planning method based on parallel computing and reinforcement learning in one embodiment of the present invention;
[0042] Figure 2 It is a schematic diagram of an algorithm flow for determining a feasible path in a multi-agent path planning method based on parallel computing and reinforcement learning in one embodiment of the present invention;
[0043] Figure 3 It is a schematic diagram of an algorithm for determining a feasible path in a multi-agent path planning method based on parallel computing and reinforcement learning in one embodiment of the present invention;
[0044] Figure 4 It is a schematic diagram of an algorithm flow for determining an optimal path to a target in a multi-agent path planning method based on parallel computing and reinforcement learning in one embodiment of the present invention;
[0045] Figure 5 is a schematic diagram of the advantages of a multi-agent path planning method based on parallel computing and reinforcement learning in one embodiment of the present invention;
[0046] Figure 6 It is a schematic diagram of the structure of a multi-agent path planning system based on parallel computing and reinforcement learning in one embodiment of the present invention.
[0047] Reference numerals:
[0048] Among them, 11, conversion module; 12, determination module; 13, calculation module; 14, construction module; 15, execution module. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0050] In the description of this application, the terms "first", "second", "third", etc. are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second", "third", etc. may explicitly or implicitly include one or more of the feature. In the description of this application, unless otherwise specified, "plurality" means two or more.
[0051] In the description of the present application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a connection between the two elements. The terms "vertical", "horizontal", "left", "right", "upper", "lower" and similar expressions used herein are only for illustrative purposes, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0052] In the description of this application, it should be noted that, unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those commonly understood by those skilled in the art. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood by specific circumstances.
[0053] An embodiment of the present invention provides a multi-agent path planning method based on parallel computing and reinforcement learning. For details, see Figure 1 , Figure 1 The figure shows a flow chart of a multi-agent path planning method based on parallel computing and reinforcement learning in one embodiment of the present invention. The method is applied to the path planning process of multiple agents in a target area, and includes:
[0054] S1: Determine a digitized map of the target area, perform coordinate conversion processing on each position point in the digitized map, and obtain a coordinate matrix of the target area;
[0055] S2: determining a non-feasible area of the target area based on the environmental data of the target area;
[0056] S3: determining all feasible path data of the target area according to the coordinate matrix and the infeasible area, wherein the determination process is configured to start from the terminal coordinates corresponding to the agent in the coordinate matrix, and sequentially screen neighboring coordinate points circumscribed to the terminal coordinates with the infeasible area, and determine a feasible path according to the screening results;
[0057] S4: constructing a convolutional neural network path optimization model based on a reinforcement learning algorithm, inputting all the feasible path data into the convolutional neural network path optimization model, and obtaining the target optimal path of each of the intelligent agents;
[0058] S5: Execute the path planning solution generated by each of the target optimal paths.
[0059] In step S1, in the path planning of multiple agents, a digital map is determined according to the positions of all agents and the destinations to be reached by the agents, and accurate coordinate conversion processing is performed on each position point in the digital map. This step aims to unify the coordinate system and adjust the scale to ensure that all position information is accurate. After this conversion process, all position points in the target area are finally integrated into a structured coordinate matrix, which can intuitively show the precise position of each point in the target area under a specific coordinate system.
[0060] Based on the environmental data of the target area, a comprehensive analysis is conducted to determine the various obstacles, dangerous areas or other areas that are not suitable for activities in the target area. These areas are defined as non-feasible areas.
[0061] In addition, an environmental model of the target area can be constructed based on the environmental data of the target area, and the obstacles in the target area can be analyzed using the environmental model to determine the infeasible area of the target area.
[0062] Specifically, the environmental model simulates the physical and ecological characteristics of the target area based on actual environmental data. The environmental model can be constructed using a variety of technical means such as geographic information system (GIS), remote sensing technology, and three-dimensional modeling.
[0063] In step S3, after clarifying the terminal position of the agent in the coordinate matrix and the area that the agent cannot pass through, starting from the terminal coordinates, its neighbor coordinate points are calculated in turn, and the infeasible neighbor coordinate points are excluded according to the information of the infeasible area; for each feasible neighbor coordinate point, the above steps are repeated recursively until the boundary condition is reached. The boundary condition can be traversing all reachable areas or reaching a preset depth limit. After the screening is completed, the feasible path is determined according to the screening results.
[0064] Specifically, in this step, updating from the end point coordinates can save computing time and improve computing efficiency. If the agent is used as the initial point for updating, then at each time step, when the agent moves once, the distance from each agent to the end point of the agent needs to be calculated once. This process is more time-consuming than updating from the end point coordinates. The neighbor coordinate points are usually in four directions: up, down, left, and right, but can also be in eight directions, including diagonals.
[0065] After determining the feasible path, a convolutional neural network path optimization model is constructed based on the reinforcement learning algorithm. The construction process includes: determining the state space, action space and reward function of the agent according to the environmental data of the agent and the target area, and using the reinforcement learning algorithm to construct the convolutional neural network path optimization model based on the state space, action space and reward function.
[0066] Specifically, the state space is the set of all possible states that the agent may be in. In the path optimization problem, the state includes the agent's current position, speed, and direction; the action space is the set of all possible actions that the agent can take. For example, on a two-dimensional plane, the actions include moving up, down, left, and right; the reward function can calculate a reward value based on the agent's current state and the action taken. The reward value can be positive, which means that the agent is encouraged to take a certain action, or negative, which means that the agent is punished for taking a certain action. In the path optimization problem, the reward function may be designed to give a large reward when reaching the end point and a penalty when encountering an obstacle. For example, if there is no obstacle in the next step, the agent will move forward, and if an obstacle is encountered, the agent will stop moving.
[0067] After the convolutional neural network path optimization model is built, all feasible path data are input into the convolutional neural network path optimization model to obtain the optimal path for each agent. Specifically, it includes:
[0068] All feasible path data are input into the convolutional neural network path optimization model to obtain the feature vector and weight parameters corresponding to each agent; in this process, the convolutional neural network path optimization model will use the input feasible path data to extract features related to path optimization, and generate a feature vector and corresponding weight parameters for each agent.
[0069] A multi-layer perceptron is used to calculate the feasible path of each agent to obtain the encoded information, and the weight parameters of the agent are used to calculate the corresponding encoded information of the agent to obtain the overall information of all agents. In this process, the weight parameters of each agent obtained in the convolutional neural network path optimization model are used to weight the encoded information of these agents to obtain the overall information of all agents. This step is to comprehensively consider the path information of all agents in order to perform global path optimization.
[0070] Based on the feature vector and the overall information of all agents, the optimal path to the target of each agent is determined.
[0071] After obtaining the optimal target path of each agent, the simulated annealing algorithm is used to verify the feasibility of the optimal target path to optimize the optimal target path of the agent, and finally the path planning solution generated by each optimal target path is executed.
[0072] Specifically, in the iterative process of the simulated annealing algorithm, by continuously accepting better neighborhood solutions, the algorithm can explore the path space and find a path that is closer to the global optimum; verify the feasibility of each path, that is, ensure that there are no obstacles on the path and that the agent can smoothly reach the destination along the path. For infeasible paths, the algorithm can continue to search or adjust parameters to optimize.
[0073] Another embodiment of the present invention provides a step of determining a feasible path. For details, see Figure 2 and Figure 3 , Figure 2 FIG. 1 is a schematic diagram of an algorithm flow for determining a feasible path in a multi-agent path planning method based on parallel computing and reinforcement learning in one embodiment of the present invention. Figure 3 The figure shows an algorithm diagram of determining a feasible path in a multi-agent path planning method based on parallel computing and reinforcement learning in one embodiment of the present invention. The specific process is as follows:
[0074] Determine the map of the target area, that is, the 2D grid map, initialize it as a coordinate matrix, set the movable position in the target area to 0, and the immovable position to -1, define the auxiliary map and the output map, where the auxiliary map is used to store the coordinates of the neighbor points of each updated position, and the output map is used to store the movable positions in the path planning process.
[0075] The endpoint coordinates corresponding to the agent in the coordinate matrix are used as the first update position, the position corresponding to the neighbor point coordinates in the auxiliary map is assigned as cost, the neighbor point coordinates of the first update position in the auxiliary map are calculated, and the second update position is determined in combination with the movable position in the output map;
[0076] If the second updated position has neighboring point coordinates, continue to iterate and calculate the next updated position until the updated position has no neighboring point coordinates. All feasible paths can be obtained by combining the position of the agent with the corresponding end point coordinates of the agent in the coordinate matrix.
[0077] exist Figure 3 In the figure, circles represent agents, diamonds represent endpoints, different colors represent different agents, green in the Select graph represents the coordinates of the neighboring points of the endpoint, i.e., the auxiliary map, and green in the Update graph represents the feasible area, i.e., the output map. Combining Select with Update, combining the position of the agent with the corresponding endpoint coordinates of the agent in the coordinate matrix, all feasible paths can be obtained.
[0078] Another embodiment of the present invention provides a method for determining an optimal path to a target. For details, see Figure 4 and Figure 5 , Figure 4 FIG. 1 is a schematic diagram of an algorithm flow for determining an optimal path to a target in a multi-agent path planning method based on parallel computing and reinforcement learning in one embodiment of the present invention. Figure 5 A schematic diagram showing the advantages of a multi-agent path planning method based on parallel computing and reinforcement learning. The specific process is as follows:
[0079] exist Figure 3On this basis, all feasible paths for each agent have been obtained respectively, and a convolutional neural network path optimization model is constructed using reinforcement learning algorithms such as convolutional neural network (CNN) and multi-layer perceptron (MLP). All feasible path data are input into the convolutional neural network path optimization model to obtain the feature vector and weight parameter β corresponding to each agent.
[0080] The information encoder is used to calculate the feasible path of each agent to obtain the encoded information, the weight parameter β of the agent is used to calculate the encoded information corresponding to the agent, and the information decoder is used to obtain the overall information of all agents.
[0081] The feature vector and overall information are input into a multi-layer perceptron (MLP) to obtain the optimal path for each agent.
[0082] exist Figure 5 In this paper, RL-LMAPF is the method of this paper, that is, the multi-agent path planning method based on parallel computing and reinforcement learning, Folloewr is the search algorithm based on A*, and PRIMAL2 is the method based on reinforcement learning.
[0083] exist Figure 5 In (a), as the number of agents increases, the throughput also increases. With the increase in throughput, the system is able to process more data in a shorter time, thereby improving the overall processing efficiency.
[0084] exist Figure 5 In (b), as the number of agents increases, the processing time of RL-LMAPF is much lower than that of Folloewr and PRIMAL2.
[0085] Another embodiment of the present invention provides a multi-agent path planning system based on parallel computing and reinforcement learning. For details, see Figure 6 , Figure 6 The figure shows a schematic diagram of a system structure for determining an optimal target path in a multi-agent path planning method based on parallel computing and reinforcement learning in one embodiment of the present invention. The system is applied to a path planning process of multiple agents in a target area, and includes:
[0086] The conversion module 11 is used to determine a digitized map of the target area, and perform coordinate conversion processing on each position point in the digitized map to obtain a coordinate matrix of the target area;
[0087] A determination module 12, configured to determine a non-feasible area of the target area based on the environmental data of the target area;
[0088] A calculation module 13 is used to determine all feasible path data of the target area according to the coordinate matrix and the infeasible area, wherein the determination process is configured to start from the terminal coordinates corresponding to the agent in the coordinate matrix, and sequentially screen neighboring coordinate points circumscribed to the terminal coordinates with the infeasible area, and determine a feasible path according to the screening results;
[0089] A construction module 14 is used to construct a convolutional neural network path optimization model based on a reinforcement learning algorithm, input all the feasible path data into the convolutional neural network path optimization model, and obtain the target optimal path of each of the intelligent agents;
[0090] The execution module 15 is used to execute the path planning solution generated by each of the target optimal paths.
[0091] Compared with the prior art, the embodiments of the present invention have the following advantages:
[0092] (1) The present invention reduces the amount of calculation and improves planning efficiency by optimizing the path planning of multiple agents;
[0093] (2) The present invention can ensure that the generated path is a feasible path that meets the actual environmental constraints by screening neighbor coordinate points, thereby narrowing the search range and improving the efficiency of path planning; using convolutional neural networks for path optimization can make full use of the parallel computing and learning capabilities of neural networks to further accelerate the path planning process.
[0094] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A multi-agent path planning method based on parallel computing and reinforcement learning, characterized in that: Applied to the path planning process of multiple agents in the target area, including: Determine a digitized map of the target area, and perform coordinate conversion processing on each position point in the digitized map to obtain a coordinate matrix of the target area; Determining a non-feasible area of the target area based on the environmental data of the target area; Determine all feasible path data of the target area according to the coordinate matrix and the infeasible area, wherein the determination process is configured to start from the terminal coordinates corresponding to the agent in the coordinate matrix, use the infeasible area to sequentially screen neighboring coordinate points circumscribed to the terminal coordinates, and determine a feasible path according to the screening results; Constructing a convolutional neural network path optimization model based on a reinforcement learning algorithm, inputting all the feasible path data into the convolutional neural network path optimization model, and obtaining the target optimal path for each of the intelligent agents; Execute the path planning solution generated by each of the target optimal paths.
2. The multi-agent path planning method based on parallel computing and reinforcement learning as claimed in claim 1, characterized in that: The determining of the non-feasible area of the target area based on the environmental data of the target area includes: Based on the environmental data of the target area, an environmental model of the target area is constructed, and obstacles in the target area are analyzed using the environmental model to determine a non-feasible area of the target area.
3. The multi-agent path planning method based on parallel computing and reinforcement learning as claimed in claim 1, characterized in that: The step of inputting all the feasible path data into the convolutional neural network path optimization model to obtain the target optimal path for each of the intelligent agents includes: Inputting all the feasible path data into the convolutional neural network path optimization model to obtain the feature vector and weight parameter corresponding to each of the intelligent agents; Using a multi-layer perceptron to calculate the feasible path of each of the intelligent agents to obtain coding information, using the weight parameters of the intelligent agents to calculate the coding information corresponding to the intelligent agents to obtain the overall information of all the intelligent agents; According to the feature vector and the overall information of all the intelligent agents, the optimal target path of each of the intelligent agents is determined.
4. The multi-agent path planning method based on parallel computing and reinforcement learning as claimed in claim 1, characterized in that: The convolutional neural network path optimization model is constructed based on the reinforcement learning algorithm, and the construction process includes: According to the environmental data of the intelligent agent and the target area, the state space, action space and reward function of the intelligent agent reflecting the behavior are determined, and the convolutional neural network path optimization model is constructed according to the state space, the action space and the reward function using the reinforcement learning algorithm.
5. The multi-agent path planning method based on parallel computing and reinforcement learning as claimed in claim 1, characterized in that: After obtaining the target optimal path of each of the intelligent agents, the multi-agent path planning method based on parallel computing and reinforcement learning further includes: The feasibility of the target optimal path is verified by using a simulated annealing algorithm to optimize the target optimal path of the agent.
6. A multi-agent path planning system based on parallel computing and reinforcement learning, characterized in that: Applied to the path planning process of multiple agents in the target area, including: A conversion module, used to determine a digitized map of the target area, and perform coordinate conversion processing on each position point in the digitized map to obtain a coordinate matrix of the target area; A determination module, configured to determine a non-feasible area of the target area based on the environmental data of the target area; A calculation module, used for determining all feasible path data of the target area according to the coordinate matrix and the infeasible area, wherein the determination process is configured to start from the terminal coordinates corresponding to the agent in the coordinate matrix, and sequentially screen neighboring coordinate points circumscribed to the terminal coordinates with the infeasible area, and determine a feasible path according to the screening results; A construction module is used to construct a convolutional neural network path optimization model based on a reinforcement learning algorithm, input all the feasible path data into the convolutional neural network path optimization model, and obtain the target optimal path of each of the intelligent agents; An execution module is used to execute the path planning solution generated by each of the target optimal paths.
7. The multi-agent path planning system based on parallel computing and reinforcement learning as claimed in claim 6, characterized in that: The determining of the non-feasible area of the target area based on the environmental data of the target area includes: Based on the environmental data of the target area, an environmental model of the target area is constructed, and obstacles in the target area are analyzed using the environmental model to determine a non-feasible area of the target area.
8. The multi-agent path planning system based on parallel computing and reinforcement learning as claimed in claim 6, characterized in that: The step of inputting all the feasible path data into the convolutional neural network path optimization model to obtain the target optimal path for each of the intelligent agents includes: Inputting all the feasible path data into the convolutional neural network path optimization model to obtain the feature vector and weight parameter corresponding to each of the intelligent agents; Using a multi-layer perceptron to calculate the feasible path of each of the intelligent agents to obtain coding information, using the weight parameters of the intelligent agents to calculate the coding information corresponding to the intelligent agents to obtain the overall information of all the intelligent agents; According to the feature vector and the overall information of all the intelligent agents, the optimal target path of each of the intelligent agents is determined.
9. The multi-agent path planning system based on parallel computing and reinforcement learning as claimed in claim 6, characterized in that: The convolutional neural network path optimization model is constructed based on the reinforcement learning algorithm, and the construction process includes: According to the environmental data of the intelligent agent and the target area, the state space, action space and reward function of the intelligent agent reflecting the behavior are determined, and the convolutional neural network path optimization model is constructed according to the state space, the action space and the reward function using the reinforcement learning algorithm.
10. The multi-agent path planning system based on parallel computing and reinforcement learning as claimed in claim 6, characterized in that: After obtaining the optimal target path of each of the agents, the multi-agent path planning system based on parallel computing and reinforcement learning further includes: The optimization module is used to verify the feasibility of the target optimal path by using a simulated annealing algorithm to optimize the target optimal path of the intelligent agent.