Method for planning a path, electronic device, and storage medium

By determining the local map in the environmental map and performing local path planning, the problem of low path planning efficiency in the large environmental map is solved, and efficient path planning is achieved.

CN115061480BActive Publication Date: 2025-05-06CHINA ACADEMY OF INFORMATION & COMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210950097.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-09
Publication Date
2025-05-06
Estimated Expiration
2042-08-09

AI Technical Summary

Technical Problem

In the case of larger environmental maps, the calculation amount of existing path planning algorithms increases, resulting in lower path planning efficiency.

Method used

The local map is determined by the local initial position of the agent, and the local path is planned in the local map until the second target position is the same as the first target position. This method reduces the computational volume and improves the efficiency of path planning.

Benefits of technology

Through local path planning, multiple local paths can be obtained according to changes in the local initial position, and the path planning of the agent to reach the first target position is realized. Compared with direct path planning, the calculation amount is significantly reduced and efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115061480B_ABST
    Figure CN115061480B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of path planning, and discloses a method for planning a path, including: obtaining a first target position of an intelligent agent. Determining a local initial position of the intelligent agent. Determining a local map corresponding to the local initial position. Determining a second target position corresponding to the local initial position in the local map. Obtaining a local path of the intelligent agent from the local initial position to the second target position, until the second target position is the same as the first target position. In this way, multiple local paths can be obtained according to the change of the local initial position, and path planning for the intelligent agent to reach the first target position is realized. Compared with directly performing path planning for the intelligent agent to reach the first target position, using a local map for local path planning reduces the amount of calculation and improves the efficiency of path planning. The present application also discloses an electronic device and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of path planning, for example, to a method for planning a path, an electronic device, and a storage medium. Background Art

[0002] At present, with the development of artificial intelligence technology, artificial intelligence technology has been widely used in various fields. The problem of intelligent agents using artificial intelligence for path planning has also attracted more and more attention. Among them, the intelligent agent is an intelligent robot, such as a sweeping robot, a food delivery robot, etc. Path planning refers to the intelligent agent planning an optimal path from the starting point to the target point in the map without touching obstacles. Make the time and energy spent on the optimal path the least. The grid method is usually used for path planning in related technologies. First, the map is divided into many small grids, and the obstacle area, covered area and uncovered area are defined in the map, and a breadth-first or depth-first search is performed to obtain the optimal path.

[0003] In the process of implementing the embodiments of the present disclosure, it was found that there are at least the following problems in the related art: in the related art, the larger the environment map, the greater the amount of algorithm calculation, resulting in lower efficiency of path planning. Summary of the invention

[0004] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0005] The embodiments of the present disclosure provide a method, an electronic device, and a storage medium for planning a path, so as to improve the efficiency of path planning.

[0006] In some embodiments, the method includes: obtaining a first target position of the agent; determining a local initial position of the agent; determining a local map corresponding to the local initial position; determining a second target position corresponding to the local initial position in the local map; and obtaining a local path of the agent from the local initial position to the second target position until the second target position is the same as the first target position.

[0007] In some embodiments, the electronic device includes a processor and a memory storing program instructions, and the processor is configured to execute the above-mentioned method for planning a path when running the program instructions.

[0008] The method, electronic device, and storage medium for planning a path provided by the embodiments of the present disclosure can achieve the following technical effects: determine a local map through the local initial position of the agent, then determine the second target position corresponding to the local initial position in the local map, and obtain the local path of the agent from the local initial position to the second target position until the second target position is the same as the first target position. In this way, it is only necessary to plan the local path from the local initial position to the second target position in the determined local map, and multiple local paths can be obtained according to the change of the local initial position, thereby realizing the path planning for the agent to reach the first target position. Compared with directly planning the path for the agent to reach the first target position, using a local map for local path planning reduces the amount of calculation and improves the efficiency of path planning.

[0009] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements, and the drawings do not constitute a scale limitation, and wherein:

[0011] Figure 1 is a schematic diagram of a method for planning a path provided by an embodiment of the present disclosure;

[0012] Figure 2 is a schematic diagram of a method for constructing a local map provided by an embodiment of the present disclosure;

[0013] Figure 3 is a schematic diagram of another method for planning a path provided by an embodiment of the present disclosure;

[0014] Figure 4 is a schematic diagram of a method for obtaining an action selection model provided by an embodiment of the present disclosure;

[0015] Figure 5 is a schematic diagram of a device for planning a path provided by an embodiment of the present disclosure;

[0016] Figure 6 It is a schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0017] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.

[0018] The terms "first", "second", etc. in the specification and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.

[0019] Unless otherwise stated, the term "plurality" means two or more.

[0020] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B indicates: A or B.

[0021] The term "and / or" is a description of the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.

[0022] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.

[0023] In the disclosed embodiment, an electronic device is used to obtain the first target position of the intelligent agent, and to determine the local initial position of the intelligent agent. Then, a local map corresponding to the local initial position is determined, and the second target position corresponding to the local initial position is determined in the local map. The local path of the intelligent agent from the local initial position to the second target position is obtained until the second target position is the same as the first target position. In this way, by planning the local path from the local initial position to the second target position in the local map, multiple local paths can be obtained according to the change of the local initial position, thereby realizing the path planning for the intelligent agent to reach the first target position. Compared with directly planning the path for the intelligent agent to reach the first target position, using a local map for local path planning reduces the amount of calculation and improves the efficiency of path planning.

[0024] Combination Figure 1 As shown, the embodiment of the present disclosure provides a method for planning a path, including:

[0025] Step S101: The electronic device obtains a first target position of an intelligent agent, wherein the intelligent agent is an intelligent robot.

[0026] Step S102: the electronic device determines the local initial position of the agent.

[0027] Step S103: the electronic device determines a local map corresponding to the local initial position.

[0028] Step S104: the electronic device determines a second target position corresponding to the local initial position in the local map.

[0029] Step S105, the electronic device obtains a local path of the agent from the local initial position to the second target position until the second target position is the same as the first target position.

[0030] Using the method for planning a path provided by an embodiment of the present disclosure, a local map is determined by the local initial position of the agent, and then the second target position corresponding to the local initial position is determined in the local map, and the local path of the agent from the local initial position to the second target position is obtained until the second target position is the same as the first target position. In this way, the local path from the local initial position to the second target position is planned in the determined local map, and multiple local paths can be obtained according to the change of the local initial position, thereby realizing the path planning for the agent to reach the first target position. Compared with directly planning the path for the agent to reach the first target position, using the local map for local path planning reduces the amount of calculation and improves the efficiency of path planning.

[0031] The environmental map includes the initial moving position of the agent, the first target position of the agent, and obstacles. Path planning is to plan an optimal path from the initial moving position to the first target position in the environmental map without touching obstacles. Among them, the initial moving position of the agent is the starting point of the path planned in the environmental map. The first target position is the end point of the path planned in the environmental map. The size of the local map is smaller than the size of the environmental map. In the environmental map, the local map is determined according to the local initial position of the agent, and then the second target position of the initial position is determined in the local map, and a local path is planned in the local map. The local initial position is the starting point of the local path, and the second target position is the end point of the local path.

[0032] Furthermore, the electronic device determines the local map corresponding to the local initial position, including: the electronic device establishes a local map with the local initial position of the agent as the center or corner, so that the local initial position of the agent is located at the center or a corner of the local map. The local map is a grid map. The size of the local map is n×n. n is a positive integer. The local initial position occupies a grid in the local map. In this way, the local map is established with the initial position of the agent as a reference. In order to perform local path planning in the local map and obtain the local path, the amount of calculation is reduced and the efficiency of path planning is improved.

[0033] In some embodiments, n is equal to 4. n is determined by the size of the environment map and / or the complexity of the obstacles. When the environment map is large, that is, the larger the area of ​​the environment map, the larger n is, so that the efficiency of the planned path can be increased. When the complexity of the obstacles is high, that is, the more obstacles there are, the smaller n is, thereby improving the local planning precision. In this way, when the complexity of the environment map is high, by using the local map to decompose the environment map layer by layer and then perform path planning, it simplifies the complexity, reduces the amount of calculation, and improves the efficiency of path planning.

[0034] Furthermore, the electronic device determines the local map corresponding to the local initial position, further comprising: when the local map exceeds the boundary of the environment map, determining that an area in the local map that exceeds the boundary of the environment map is filled with obstacles.

[0035] Optionally, the electronic device determines the local initial position of the intelligent body, including: when the local initial position of the intelligent body is determined for the first time, the electronic device obtains the mobile initial position of the intelligent body. The electronic device determines the mobile initial position as the local initial position. When the local initial position of the intelligent body is determined for the mth time, the third target position is determined as the local initial position, and the third target position is the second target position corresponding to the local initial position determined for the m-1th time, m>2, and is a positive integer. In this way, the end point of the last local path planning, that is, the third target position, can be determined as the starting point of the next local planning path. So as to obtain a coherent path with the mobile initial position and the first target position as the end point. At the same time, since a coherent path with the mobile initial position and the first target position as the end point is not planned at one time, the amount of calculation is reduced and the efficiency of path planning is improved.

[0036] In some embodiments, the electronic device obtains the first target position of the agent in the environmental map. Then, when the local initial position of the agent is determined for the first time, the electronic device obtains the mobile initial position of the agent in the environmental map, and determines the mobile initial position as the local initial position. Determine the local map corresponding to the local initial position, determine the second target position corresponding to the local initial position in the local map, and then obtain the local path of the agent from the local initial position to the second target position. When the electronic device determines the local initial position of the agent for the mth time, the second target position corresponding to the local initial position determined for the m-1th time is determined as the new local initial position, and the local map corresponding to the local initial position is re-determined. Determine the second target position corresponding to the local initial position in the local map, and obtain the local path of the agent from the local initial position to the second target position until the second target position is the same as the first target position.

[0037] In some embodiments, in combination Figure 2 As shown, Figure 2 A method for constructing a local map. The first target position 3 and the moving initial position 1 of the agent are both in the environment map. The environment map is a grid map. When the local initial position of the agent is determined for the first time, the electronic device obtains the moving initial position of the agent in the environment map. The agent is moved to the moving initial position, and then the moving initial position, that is, the current position of the agent, is determined as the local initial position. The electronic device constructs a local map of size 4×4 with the local initial position as the lower left corner of the local map. The constructed local map is Figure 2 The area enclosed by the solid line in the middle. In the local map, determine the second target position 2 corresponding to the local initial position, and then obtain the local path of the agent from the local initial position to the second target position 2. Then control the electronic device to move from the local initial position to the second target position according to the local path. The second target position, that is, the current position of the agent, is re-determined as the initial position, and the initial position is used as the lower left corner of the local map to reconstruct a local map of size 4×4. The newly constructed local map is Figure 2 The area enclosed by the dotted line in the middle. The local path planning is performed again in the local map. In this way, the local map flows in the environment map according to the change of the initial position, and the environment map is decomposed and planned. It is possible to plan multiple local paths connected end to end in the environment map, and finally obtain a coherent path with the initial position of the movement and the first target position as the end point. This facilitates the rapid dynamic path planning in the large environment map.

[0038] Optionally, the electronic device obtains a local path of the agent from the local initial position to the second target position, including: the electronic device obtains the local path of the agent from the local initial position to the second target position using a preset path planning algorithm. The preset path planning algorithm is a D star algorithm. The D star algorithm is a heuristic algorithm.

[0039] In this way, the D star algorithm is used to implement local path planning in the local map, which reduces the amount of calculation and improves the efficiency of path planning. At the same time, by using the D star algorithm to implement local path planning in the local map, the problem of slow search when the D star algorithm is used for path planning in large maps or maps with hundreds of obstacles is overcome, and it is not easy to fall into the local minimum, thereby increasing the success rate and efficiency of completing the path planning.

[0040] In some embodiments, the D star algorithm takes the target point, i.e., the second target position, as the starting point, and searches by placing the target point in a priority queue. Until the node corresponding to the local initial position is dequeued from the queue, the path between the local initial position and the second target position is obtained. In this way, when the state of a certain position in the middle of the path changes, it is only necessary to replan the path from the target node, i.e., the second target position to the position where the state changes. There is no need to replan the path from the local initial position to the second target position as in the A star algorithm, thereby improving the efficiency of path planning.

[0041] Optionally, the electronic device determines the second target position corresponding to the local initial position in the local map, including: the electronic device obtains the initial state of the intelligent body according to the first target position and the local initial position. The electronic device inputs the initial state of the intelligent body into a preset action selection model to obtain a decision action corresponding to the initial state. The electronic device determines the position corresponding to the decision action as the second target position. Among them, the second target position is located in a grid in the local map. In this way, the second target position corresponding to the local initial position can be accurately obtained in the local map through the action selection model, so as to facilitate local path planning in the local map. When the local initial position changes, multiple local paths can be obtained according to the change of the local initial position, so as to realize the path planning of the intelligent body to reach the first target position. The amount of calculation is reduced and the efficiency of path planning is improved.

[0042] Furthermore, the electronic device obtains the initial state of the intelligent body according to the first target position and the local initial position, including: the electronic device obtains the speed of the intelligent body and the obstacle position of the obstacle in the local map. The electronic device obtains the distance difference between the initial position and the first target position and the deflection difference between the initial position and the first target position. The electronic device obtains the initial state of the intelligent body according to the speed of the intelligent body, the local initial position, the obstacle position of the obstacle in the local map, the distance difference between the local initial position and the first target position, and the deflection difference between the local initial position and the first target position.

[0043] Optionally, the electronic device obtains the initial state of the intelligent body according to the speed of the intelligent body, the local initial position, the position of the obstacle in the local map, the distance difference between the local initial position and the first target position, and the deflection difference between the local initial position and the first target position, including: the electronic device calculates Get the initial state of the agent. Where S(m) is the initial state of the agent, and m is used to represent the agent. is the speed of the agent. O(m) is the local initial position of the agent. O(obstacle) is the obstacle position of the obstacle in the local map. D(m) is the distance difference between the local initial position and the first target position. U(m) is the deflection difference between the local initial position and the first target position, and T is the transposition sign.

[0044] Combination Figure 3 As shown, the embodiment of the present disclosure provides a method for planning a path, including:

[0045] Step S301, the electronic device obtains the first target position of the agent.

[0046] Step S302: When the electronic device determines the local initial position of the agent for the first time, it obtains the moving initial position of the agent.

[0047] Step S303: The electronic device determines the movement initial position as the local initial position.

[0048] Step S304: the electronic device determines a local map corresponding to the local initial position.

[0049] Step S305: the electronic device determines a second target position corresponding to the local initial position in the local map.

[0050] Step S306, the electronic device uses a preset path planning algorithm to obtain a local path of the agent from the local initial position to the second target position until the second target position is the same as the first target position. The preset path planning algorithm is a D star algorithm.

[0051] The method for planning a path provided by the embodiment of the present disclosure is adopted to determine a local map through the local initial position of the agent, then determine the second target position corresponding to the local initial position in the local map, and use the D star algorithm to obtain the local path of the agent from the local initial position to the second target position, until the second target position is the same as the first target position. In this way, it is only necessary to plan the local path from the local initial position to the second target position in the determined local map, and then the D star algorithm can be used to obtain multiple local paths according to the change of the local initial position, so as to realize the path planning of the agent to the first target position. Compared with directly planning the path for the agent to reach the first target position, using the local map for local path planning reduces the amount of calculation and improves the efficiency of path planning.

[0052] Optionally, the action selection model is obtained by the following method: the electronic device obtains the first sample target position of the agent. The electronic device determines the first sample initial position of the agent. The electronic device determines the first sample local map corresponding to the first sample local initial position. The electronic device obtains the first sample initial state of the agent based on the first sample target position and the first sample local initial position. The electronic device obtains the first sample alternative decision action corresponding to the first sample local map. The electronic device inputs the first sample initial state and the first sample alternative decision action into a preset deep neural network model for training to obtain the action selection model. In this way, the first sample local map is determined by the first sample local initial position of the agent, and the first sample initial state and the first sample alternative decision action of the agent are obtained, and then the preset deep neural network model is trained using the first sample initial state and the first sample alternative decision action to obtain the action selection model. In this way, the action selection model can perform local path planning in the local map, reduce the amount of calculation, and improve the efficiency of path planning. At the same time, compared with the reinforcement learning method based on grid maps, the agent updates its actions in a grid at each step, and the training convergence speed in a large map will be greatly reduced. This application adopts the local map first sample initial state and the first sample alternative decision action training mode, which reduces the amount of calculation and is more suitable for large environment maps.

[0053] The sample environment map includes the agent's sample movement initial position, the agent's first sample target position, and sample obstacles. Path planning is to plan an optimal path from the sample movement initial position to the first sample target position in the environment map without touching obstacles. Among them, the agent's sample movement initial position is the starting point of the path planned in the sample environment map. The first sample target position is the end point of the path planned in the sample environment map. The size of the first sample local map is smaller than that of the sample environment map. In the sample environment map, the first sample local map is determined according to the agent's first sample local initial position, and then the first sample alternative decision action corresponding to the first sample local map is determined.

[0054] Furthermore, the electronic device determines the first sample local initial position of the agent, including: the electronic device obtains the sample movement initial position of the agent, and the electronic device determines the sample movement initial position as the first sample local initial position.

[0055] In some embodiments, after the electronic device obtains the first sample target position of the agent, the electronic device obtains the sample movement initial position of the agent and determines the sample movement initial position as the first sample local initial position, and then determines the first sample local map corresponding to the first sample local initial position.

[0056] In some embodiments, after the electronic device obtains the first sample target position of the agent, the electronic device obtains the sample movement initial position of the agent and moves the agent to the sample movement initial position. The electronic device determines the sample movement initial position, that is, the first sample current position of the agent as the first sample local initial position. And determines the first sample local map corresponding to the first sample local initial position.

[0057] Further, the electronic device determines the first sample local map corresponding to the first sample local initial position, including: the electronic device takes the first sample local initial position of the agent as the center or corner, establishes the first sample local map, and makes the first sample local initial position of the agent located at the center or a corner of the first sample local map. The first sample local map is a grid map. The size of the first sample local map is n×n.

[0058] Furthermore, the electronic device determines the first sample local map corresponding to the first sample local initial position, and also includes: when the first sample local map exceeds the boundary of the sample environment map, determining that the area in the first sample local map that exceeds the boundary of the sample environment map is filled with obstacles.

[0059] Further, the electronic device obtains the first sample initial state of the agent according to the first sample target position and the first sample local initial position, including: the electronic device obtains the first sample speed of the agent and the first sample obstacle position of the obstacle in the first sample local map. The electronic device obtains the first sample distance difference between the first sample local initial position and the first sample target position and the first sample deflection difference between the first sample local initial position and the first sample target position. The electronic device obtains the first sample initial state of the agent according to the first sample speed of the agent, the first sample local initial position, the first sample obstacle position, the first sample distance difference and the first sample deflection difference.

[0060] Optionally, the electronic device obtains the first sample initial state of the agent according to the first sample speed, the first sample local initial position, the first sample obstacle position, the first sample distance difference and the first sample deflection difference of the agent, including: the electronic device calculates The first sample initial state of the agent is obtained. Wherein, S'(m) is the first sample initial state of the agent, and m is used to represent the agent. is the first sample speed of the agent. O'(m) is the first sample local initial position of the agent. O'(obstacle) is the first sample obstacle position. D'(m) is the first sample distance difference between the first sample local initial position and the first sample target position. U'(m) is the first sample deflection difference between the first sample local initial position and the first sample target position.

[0061] Further, the electronic device obtains the first sample candidate decision action corresponding to the first sample local map, including: the electronic device obtains the sequence number corresponding to each grid in the first sample local map. The sequence number is obtained by numbering each grid in the first sample local map. The electronic device determines the action performed by the intelligent agent when it reaches the grid corresponding to each sequence number from the first sample initial state as the first sample candidate decision action. Further, the sequence number of the grid is 0, 1, 2, ..., n 2 -1.

[0062] Optionally, the electronic device inputs the first sample initial state and the first sample candidate decision action into a preset deep neural network model for training to obtain an action selection model, including: the electronic device inputs the first sample initial state and the first sample candidate decision action into a preset deep neural network model, and uses the first current network of the deep neural network model to select a first sample candidate decision action as the first sample decision action. The electronic device determines the position corresponding to the first sample decision action as the second sample target position. The electronic device obtains the first sample local path of the agent from the first sample local initial position to the second sample target position. The electronic device scores the first sample local path using the second current network of the deep neural network model to obtain the first sample path score. The electronic device obtains the first candidate action selection model corresponding to the first sample path score. The electronic device optimizes the first candidate action selection model according to the first sample path score to obtain the second candidate action selection model. The electronic device obtains the action selection model according to the second candidate action selection model. Among them, the preset deep neural network model is a DDPG (deep deterministic policy gradient, deep deterministic policy gradient algorithm) model. The first current network is an Actor current network. The second current network is a Critic current network. In this way, a first sample candidate decision action is selected as the first sample decision action in the first current network of the deep neural network model to obtain the second sample target position corresponding to the first sample decision action. Then the first sample local path of the agent from the first sample local initial position to the second sample target position is obtained, and the path is scored. The first alternative action selection model is optimized using the score to obtain the second alternative action selection model. Then the action selection model is obtained based on the second alternative action selection model. The action selection model is enabled to obtain the second target position in the local map according to the initial state of the agent, realize local path planning, reduce the amount of calculation, and improve the efficiency of path planning.

[0063] Further, the electronic device selects a first sample candidate decision action as the first sample decision action using the first current network of the deep neural network model, including: the electronic device selects a first sample candidate decision action using the first current network of the deep neural network model through a preset strategy algorithm. The electronic device determines the selected first sample candidate decision action as the first sample decision action. The preset strategy algorithm is an ε-greedy strategy algorithm.

[0064] Furthermore, after the electronic device selects a first sample candidate decision action as the first sample decision action using the first current network of the deep neural network model, the method further includes: the electronic device obtains a first experience trajectory corresponding to the first sample decision action. The electronic device obtains a state-action function corresponding to the first sample decision action according to the first experience trajectory.

[0065] Further, the electronic device determines the second sample local initial position of the agent according to the second sample target position, including: the electronic device determines the second sample target position as the second sample local initial position of the agent. Or, the electronic device moves the agent to the second sample target position, then obtains the second sample current position of the agent, and determines the second sample current position, i.e., the second sample target position, as the second sample local initial position of the agent.

[0066] Furthermore, the electronic device obtains a first sample local path of the agent from the first sample local initial position to the second sample target position, including: the electronic device uses a preset path planning algorithm to obtain the first sample local path of the agent from the first sample local initial position to the second sample target position.

[0067] Furthermore, the electronic device optimizes the first candidate action selection model according to the first sample path score to obtain the second candidate action selection model, including: the electronic device uses a preset optimization algorithm to adjust the parameters in the first candidate action selection model according to the first sample path score to obtain the second candidate action selection model. The preset optimization algorithm is a TD (Temporal Difference) update algorithm.

[0068] Combination Figure 4 As shown, the embodiment of the present disclosure provides a method for obtaining an action selection model, including:

[0069] Step S401: the electronic device obtains a first sample target position of the agent.

[0070] Step S402: the electronic device determines a first sample local initial position of the agent.

[0071] Step S403: the electronic device determines a first sample local map corresponding to the first sample local initial position.

[0072] Step S404: the electronic device obtains a first sample initial state of the agent according to the first sample target position and the first sample local initial position.

[0073] Step S405: the electronic device obtains a first sample candidate decision action corresponding to the first sample local map.

[0074] In step S406, the electronic device inputs the first sample initial state and the first sample candidate decision action into a preset deep neural network model, and selects a first sample candidate decision action through a preset strategy algorithm using the first current network of the deep neural network model.

[0075] Step S407: The electronic device determines the selected first sample candidate decision action as the first sample decision action.

[0076] Step S408: The electronic device determines the position corresponding to the first sample decision action as the second sample target position.

[0077] Step S409: the electronic device uses a preset path planning algorithm to obtain a first sample local path of the agent from the first sample local initial position to the second sample target position.

[0078] Step S410: The electronic device scores the first sample local path using the second current network of the deep neural network model to obtain a first sample path score.

[0079] Step S411: the electronic device obtains a first candidate action selection model corresponding to the first sample path score.

[0080] Step S412: the electronic device optimizes the first candidate action selection model according to the first sample path score to obtain a second candidate action selection model.

[0081] Step S413: the electronic device obtains an action selection model according to the second candidate action selection model.

[0082] The method for obtaining an action selection model provided by the embodiment of the present disclosure is adopted. By inputting the first sample initial state obtained and the first sample alternative decision action corresponding to the first sample local map into a preset deep neural network model, a first sample alternative decision action is selected through a preset strategy algorithm to determine the second sample target position. And the first sample local path of the agent from the first sample local initial position to the second sample target position is obtained. Then the first sample local path is scored to obtain the first sample path score. And the first alternative action selection model is optimized using the obtained first sample path score, so as to obtain the action selection model according to the second alternative action selection model. So that the action selection model. It is possible to obtain the second target position in the local map according to the initial state of the agent, realize local path planning, reduce the amount of calculation, and improve the efficiency of path planning.

[0083] Optionally, the electronic device obtains an action selection model based on the second alternative action selection model, including: the electronic device obtains the current training times. When the current training times are greater than or equal to a preset times threshold, the first alternative action selection model is determined as the action selection model. The current training times are the number of times the agent reaches the first sample target position or the number of times the first sample target position is the same as the second sample target position. In this way, by determining the first alternative action selection model as the action selection model when the current training times are greater than or equal to a preset times threshold, the obtained action selection model can collect the sample state of the agent multiple times through training of the times threshold, so that the action selection model can select the second target position according to the current state information, so as to realize local path planning, reduce the amount of calculation, and improve the efficiency of path planning.

[0084] Optionally, the electronic device acquires the action selection model according to the second alternative action selection model, including: the electronic device acquires the current training number, and when the current training number is less than the preset number threshold, the electronic device determines the second sample local initial position of the agent according to the second sample target position. The electronic device determines the second sample local map corresponding to the second sample local initial position. The electronic device acquires the second sample initial state of the agent according to the first sample target position and the second sample local initial position. The electronic device acquires the second sample alternative decision action corresponding to the second sample local map. The electronic device inputs the second sample initial state and the second sample alternative decision action into the second alternative action selection model for training to obtain the action selection model. In this way, by acquiring the second sample initial state according to the second sample local initial position of the agent when the current training number is less than the preset number threshold, and acquiring the second sample alternative decision action corresponding to the second sample local map, and then inputting the second sample initial state and the second sample alternative decision action into the second alternative action selection model for training to obtain the action selection model. The action selection model can be fully trained for the number threshold times, so that the trained action selection model can select the second target position according to the current state information, so as to realize local path planning, reduce the amount of calculation, and improve the efficiency of path planning.

[0085] Further, the electronic device determines the second sample local initial position of the agent according to the second sample target position, including: the electronic device determines the second sample target position as the second sample local initial position of the agent. Or, the electronic device moves the agent to the second sample target position, then obtains the second sample current position of the agent, and determines the second sample current position as the second sample local initial position of the agent. Wherein, the current position of the agent is the second sample target position.

[0086] Further, the electronic device determines a second sample local map corresponding to the second sample local initial position, including: the electronic device establishes the second sample local map with reference to the second sample local initial position of the agent. The second sample local map is a grid map. The size of the second sample local map is n×n.

[0087] Furthermore, the electronic device determines the second sample local map corresponding to the second sample local initial position, and further includes: when the second sample local map exceeds the boundary of the environment map, determining that the area in the second sample local map that exceeds the boundary of the environment map is filled with obstacles.

[0088] Furthermore, the electronic device obtains the second sample initial state of the agent according to the first sample target position and the second sample local initial position, including: the electronic device obtains the second sample speed of the agent and the second sample obstacle position of the obstacle in the second sample local map. The electronic device obtains the second sample distance difference between the second sample local initial position and the first sample target position and the second sample deflection difference between the second sample local initial position and the first sample target position. The electronic device obtains the second sample initial state of the agent according to the second sample speed of the agent, the second sample local initial position, the second sample obstacle position, the second sample distance difference and the second sample deflection difference.

[0089] Optionally, the electronic device obtains the second sample initial state of the agent according to the second sample speed, the second sample local initial position, the second sample obstacle position, the second sample distance difference and the second sample deflection difference of the agent, including: the electronic device calculates The second sample initial state of the agent is obtained. Wherein, S" (m) is the second sample initial state of the agent, and m is used to represent the agent. is the second sample speed of the agent. O”(m) is the second sample local initial position of the agent. O”(obstacle) is the second sample obstacle position. D”(m) is the second sample distance difference. U”(m) is the second sample deflection angle difference.

[0090] Further, the electronic device obtains the second sample candidate decision action corresponding to the second sample local map, including: the electronic device obtains the sequence number corresponding to each grid in the second sample local map. The sequence number is obtained by numbering each grid in the second sample local map. The electronic device determines the action performed by the agent when it reaches the grid corresponding to each sequence number from the second sample initial state as the second sample candidate decision action.

[0091] Further, the electronic device inputs the second sample initial state and the second sample alternative decision action into the second alternative action selection model for training to obtain the action selection model, including: the electronic device inputs the second sample initial state and the second sample alternative decision action into the second alternative action selection model, and uses the third current network of the second alternative action selection model to select a second sample alternative decision action as the second sample decision action. The electronic device determines the position corresponding to the second sample decision action as the third sample target position. The electronic device obtains the second sample local path of the agent from the second sample local initial position to the third sample target position. The electronic device uses the third current network of the second alternative action selection model to score the second sample local path to obtain the second path score. The electronic device obtains the third alternative action selection model corresponding to the second path score. The electronic device optimizes the third alternative action selection model according to the second path score to obtain the fourth alternative action selection model. The electronic device determines the fourth alternative action selection model as the second alternative action selection model. The electronic device obtains the action selection model according to the second alternative action selection model. Among them, the third current network is the Actor current network. The fourth current network is the Critic current network. Further, the electronic device selects a second sample candidate decision action as the second sample decision action using the third current network of the second candidate action selection model, including: the electronic device selects a second sample candidate decision action through a preset strategy algorithm using the third current network of the second candidate action selection model. The electronic device determines the selected second sample candidate decision action as the sample decision action.

[0092] Furthermore, after the electronic device selects a second sample candidate decision action as the second sample decision action using the third current network of the second candidate action selection model, the method further includes: the electronic device obtains a second experience trajectory corresponding to the second sample decision action. The electronic device obtains a second state-action function corresponding to the second sample decision action according to the second experience trajectory.

[0093] Further, the electronic device obtains a second experience trajectory corresponding to the second sample decision action, including: the electronic device obtains a third sample initial state and a third sample decision action. The electronic device obtains the second experience trajectory corresponding to the second sample decision action based on the third sample initial state, the third sample decision action, the second sample initial state, and the second sample decision action.

[0094] In some embodiments, the electronic device obtains the number of path planning times for obtaining a sample local path in a sample local map. The sample local map includes a first sample local map and a first sample local map. The sample local path includes a first sample local path and a second sample local path. After one training is completed, the path planning times are reset. When the number of path planning times is 2, the third sample initial state is the first sample initial state of the agent obtained by the current training times. When the number of path planning times is 2, the third sample decision action is the first sample decision action of the agent obtained by the current training times. When the number of path planning times is greater than 2, the third sample initial state is the second sample initial state of the agent obtained last time. When the number of path planning times is greater than 2, the third sample decision action is the second sample decision action of the agent obtained last time.

[0095] Further, the electronic device obtains a second experience trajectory corresponding to the second sample decision action according to the third sample initial state, the third sample decision action, the second sample initial state and the second sample decision action, including: the electronic device calculates C=<S”'(m),a”',S”(m),a”> , obtain the second experience trajectory corresponding to the second sample decision action. Wherein, C is the second experience trajectory corresponding to the second sample decision action. S"'(m) is the initial state of the third sample. a"' is the third sample decision action. S"(m) is the initial state of the second sample. a" is the second sample decision action.

[0096] Further, the electronic device obtains a second state-action function corresponding to the second sample decision action according to the second experience trajectory, including: the electronic device calculates

[0097]

[0098] The second state-action function corresponding to the second sample decision action is obtained. θ [S”(m),a”] is the second state-action function corresponding to the second sample decision action. Q θ [S”'(m),a”'] is the state-action function corresponding to the third sample decision action. α is the learning rate, α∈(0,1). R is the reward function corresponding to the third sample decision action. γ is the discount factor, γ∈(0,1). is the maximum value of the state-action function corresponding to each action selected in the second sample initial state. θ is a model parameter. Further, θ is obtained by copying in the third target network of the second alternative action selection model. is the time difference error, which is also the TD error. In this way, the second state-action function corresponding to the second sample decision action is obtained through the TD error, and the state-action function is updated using the TD update algorithm to obtain an accurate state-action function.

[0099] Furthermore, the reward function corresponding to the third sample decision action is obtained by the following method: the electronic device obtains the reward function corresponding to the third sample decision action by calculating R=w1O”'(obstacle)+w2D”'(m)+w3U”'(m). Among them, w1 is the preset first weight. w2 is the preset second weight. w3 is the preset third weight. O”'(obstacle) is the sample obstacle position obtained by the electronic device last time. D”'(m) is the distance difference between the local initial position of the third sample and the target position of the first sample. U”'(m) is the angular difference between the local initial position of the third sample and the target position of the first sample.

[0100] In some embodiments, when the third sample initial state is the first sample initial state of the agent, the third sample local initial position of the agent is the first sample local initial position. When the third sample decision action is the second sample decision action of the agent obtained last time, the third sample local initial position is the second sample local initial position determined last time.

[0101] Further, the electronic device optimizes the third candidate action selection model according to the second path score to obtain a fourth candidate action selection model, including: the electronic device uses a preset optimization algorithm to adjust the parameters in the third candidate action selection model according to the second sample path score to obtain the fourth candidate action selection model. The preset optimization algorithm is a TD update algorithm.

[0102] In this way, the Markov decision process is introduced through the state-action function, and mathematical models such as the state function, action function and reward function in the reinforcement learning optimization algorithm are established, so that the D star algorithm can be used to establish a local path, so that the complex dynamic environment can be simulated in the simulation environment, and the intelligent agent can collect and train samples in the simulation environment. By adjusting the algorithm model hyperparameters, it can be fully trained to achieve a more robust path planning algorithm.

[0103] In some embodiments, the action selection model includes an Actor current network, an Actor target network, a Critic current network, and a Critic target network. Each network is composed of a fully connected neural network. The Actor current network is configured to iteratively update the parameter θ; select a decision action according to the initial state, and interact with the environment to obtain the next initial state and the reward corresponding to the decision action. The Actor target network is configured to copy θ from the Actor current network to obtain θ′ to mitigate the overestimation of the action-state function. The structure of the Actor target network is the same as that of the Actor current network. The Critic current network is configured to score the first sample local path. The structure of the Critic current network is similar to that of the Actor current network. The output of the Critic current network is a continuous value, which represents a score for the current trajectory. The Critic current network is updated using the TD update method. The structure of the Critic target network is the same as that of the Critic current network to reduce the overestimation of the value score.

[0104] Combination Figure 5 As shown, an embodiment of the present disclosure provides a device for planning a path, comprising: a first acquisition module 1, a first determination module 2, a second determination module 3, a third determination module 4 and a second acquisition module 5. The first acquisition module 1 is configured to acquire a first target position of an intelligent agent. The first determination module 2 is configured to determine a local initial position of the intelligent agent. The second determination module 3 is configured to determine a local map corresponding to the local initial position. The third determination module 4 is configured to determine a second target position corresponding to the local initial position in the local map. The second acquisition module 5 is configured to acquire a local path of the intelligent agent from the local initial position to the second target position, until the second target position is the same as the first target position.

[0105] The device for planning a path provided by this embodiment is used to determine a local map through the local initial position of the agent, and then determine the second target position corresponding to the local initial position in the local map, and obtain the local path of the agent from the local initial position to the second target position until the second target position is the same as the first target position. In this way, it is only necessary to plan the local path from the local initial position to the second target position in the determined local map, and multiple local paths can be obtained according to the change of the local initial position, thereby realizing the path planning for the agent to reach the first target position. Compared with directly planning the path for the agent to reach the first target position, using the local map for local path planning reduces the amount of calculation and improves the efficiency of path planning.

[0106] Combination Figure 6As shown, an embodiment of the present disclosure provides an electronic device, which includes a processor (processor) 6 and a memory (memory) 7. Optionally, the device may also include a communication interface (CommunicationInterface) 8 and a bus 9. Among them, the processor 6, the communication interface 8, and the memory 7 can communicate with each other through the bus 9. The communication interface 8 can be used for information transmission. The processor 6 can call the logic instructions in the memory 7 to execute the method for planning a path in the above embodiment.

[0107] In addition, the logic instructions in the above-mentioned memory 7 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0108] The memory 7, as a storage medium, can be used to store software programs, computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 6 executes the program instructions / modules stored in the memory 7 to execute functional applications and data processing, that is, to implement the method for planning a path in the above embodiment.

[0109] The memory 7 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 7 may include a high-speed random access memory and may also include a non-volatile memory.

[0110] The electronic device includes a server or an intelligent agent. The intelligent agent obtains the first target position, the initial position of movement and the obstacle position of the intelligent agent by monitoring the environment map. The server obtains the first target position, the initial position of movement and the obstacle position of the intelligent agent through the intelligent agent uploading.

[0111] By using this electronic device, a local map is determined by the local initial position of the agent, and then the second target position corresponding to the local initial position is determined in the local map, and the local path of the agent from the local initial position to the second target position is obtained until the second target position is the same as the first target position. In this way, it is only necessary to plan the local path from the local initial position to the second target position in the determined local map, and multiple local paths can be obtained according to the change of the local initial position, so as to realize the path planning for the agent to reach the first target position. Compared with directly planning the path for the agent to reach the first target position, using the local map for local path planning reduces the amount of calculation and improves the efficiency of path planning.

[0112] An embodiment of the present disclosure provides a storage medium storing computer executable instructions, wherein the computer executable instructions are configured to execute the above method for planning a path.

[0113] The above-mentioned storage medium may be a transient computer-readable storage medium or a non-transient computer-readable storage medium. Non-transient storage media include: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks, and other media that can store program codes, which may also be transient storage media.

[0114] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible changes. Unless explicitly required, separate components and functions are optional, and the order of operation may vary. The parts and features of some embodiments may be included in or replace the parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates, the singular forms of "a", "an" and "the" are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of listings containing one or more associated ones. In addition, when used in the present application, the term "comprise" and its variants "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the presence of other identical elements in the process, method or device comprising the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.

[0115] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods for each specific application to implement the described functions, but such implementations should not be considered to exceed the scope of the embodiments of the present disclosure. The technicians may clearly understand that, for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above may refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.

[0116] The flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and the block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A method for planning a path, characterized in that: include: Get the first target position of the agent; Determining a local initial position of the agent; Determine a local map corresponding to the local initial position; Determining a second target position corresponding to the local initial position in the local map; Acquire a local path of the agent from a local initial position to a second target position until the second target position is the same as the first target position; Wherein, determining the second target position corresponding to the local initial position in the local map includes: obtaining the initial state of the intelligent agent according to the first target position and the local initial position; inputting the initial state into a preset action selection model to obtain a decision action corresponding to the initial state; and determining the position corresponding to the decision action as the second target position.

2. The method according to claim 1, characterized in that Determining a local initial position of the agent includes: When the local initial position of the agent is determined for the first time, obtaining the moving initial position of the agent; and determining the moving initial position as the local initial position; When the local initial position of the agent is determined for the mth time, the third target position is determined as the local initial position, and the third target position is the second target position corresponding to the local initial position determined for the m-1th time, m>2, and is a positive integer.

3. The method according to claim 1, characterized in that Obtaining a local path of the agent from a local initial position to a second target position, comprising: A preset path planning algorithm is used to obtain a local path of the agent from a local initial position to a second target position.

4. The method according to claim 1, characterized in that: The action selection model is obtained by the following method: Get the first sample target position of the agent; Determining a first sample local initial position of the agent; Determine a first sample local map corresponding to the first sample local initial position; Acquire a first sample initial state of the agent according to the first sample target position and the first sample local initial position; Obtaining a first sample candidate decision action corresponding to the first sample local map; The first sample initial state and the first sample alternative decision action are input into a preset deep neural network model for training to obtain an action selection model.

5. The method according to claim 4, characterized in that Inputting the first sample initial state and the first sample candidate decision action into a preset deep neural network model for training to obtain an action selection model, including: Inputting the first sample initial state and the first sample candidate decision action into a preset deep neural network model, and using a first current network of the deep neural network model to select a sample candidate decision action as the first sample decision action; Determine the position corresponding to the first sample decision action as the second sample target position; Obtaining a sample local path of the agent from a first sample local initial position to a second sample target position; Scoring the sample local path using the second current network of the deep neural network model to obtain a path score; Obtaining a first candidate action selection model corresponding to the path score; Optimizing the first candidate action selection model according to the path score to obtain a second candidate action selection model; The action selection model is acquired according to the second candidate action selection model.

6. The method according to claim 5, characterized in that Acquiring the action selection model according to the second candidate action selection model includes: Obtaining a current number of training times; and determining the first candidate action selection model as the action selection model when the current number of training times is greater than or equal to a preset number threshold.

7. The method according to claim 5, characterized in that Acquiring the action selection model according to the second candidate action selection model includes: Obtaining a current training number, and when the current training number is less than a preset number threshold, determining a second sample local initial position of the agent according to the second sample target position; Determine a second sample local map corresponding to the second sample local initial position; Acquire a second sample initial state of the agent according to the first sample target position and the second sample local initial position; Obtaining a second sample candidate decision action corresponding to the second sample local map; The second sample initial state and the second sample candidate decision action are input into the second candidate action selection model for training to obtain an action selection model.

8. An electronic device comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the method for planning a path according to any one of claims 1 to 7 when running the program instructions.

9. A storage medium storing program instructions, characterized in that: When the program instructions are executed, the method for planning a path according to any one of claims 1 to 7 is executed.

Citation Information

Patent Citations

  • Navigation method and device, storage medium and equipment

    CN113108796A

  • Path planning method and device, robot and storage medium

    CN113791616A