Route planning system, route planning method, roadmap construction device, model generation device, and model generation method
The route planning system optimizes multi-agent path planning in continuous spaces by constructing individual roadmaps for each agent, considering attributes and obstacles, enhancing path finding and reducing search costs.
Patent Information
- Application Number
- JP2021169361
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-10-15
AI Technical Summary
Conventional multi-agent path planning methods in continuous spaces face challenges in finding optimal paths while minimizing computational and time costs, as densely placed nodes increase path combinations but also increase search costs, and sparsely placed nodes reduce the likelihood of finding optimal paths.
A route planning system that constructs individual roadmaps for each agent using a trained roadmap construction model, considering agent attributes and environmental obstacles, to optimize path selection and reduce search costs.
The system enhances the likelihood of finding optimal paths for agents with diverse attributes while reducing search costs by arranging nodes in appropriate ranges, suitable for each agent's path, thus improving path planning efficiency.
Smart Images

Figure 0007707846000004 
Figure 0007707846000005 
Figure 0007707846000006
Abstract
Description
Technical Field
[0001] The present invention relates to a route planning system, a route planning method, a roadmap construction device, a model generation device, and a model generation method.
Background Art
[0002] There exists a multi-agent route planning problem of planning a route for each of a plurality of agents to move to a destination. An agent is, for example, a moving body that autonomously moves (e.g., a transport robot, a cleaning robot, an autonomous driving vehicle, a drone, etc.), a person, a device operated by a person, a manipulator, or the like. Conventionally, the multi-agent route planning problem has been solved on a predetermined grid map (grid space). For example, in Non-Patent Document 1, a method of performing route planning for each agent on a grid map using a model trained by machine learning has been proposed.
[0003] According to the method using this grid map, since the movable positions of each agent are defined in advance, it is possible to find the route of each agent relatively easily. However, the movement of each agent is restricted by the predetermined grid. Therefore, there is a limit to obtaining a better route (e.g., the shortest route in real space) for each agent.
[0004] Therefore, a method of solving the multi-agent route planning problem on a continuous space where each agent can move to a free position instead of a predetermined grid space has been studied. When solving the route planning problem on a continuous space, an approach of constructing a roadmap and searching for the route of each agent on the constructed roadmap is often adopted (e.g., Non-Patent Document 2).
[0005] The roadmap is composed of nodes and edges, and specifies the range for exploring the paths of each agent. The nodes indicate movable positions, and the edges connect the nodes to indicate that movement is possible between the connected nodes. When constructing this roadmap, nodes can be placed at arbitrary positions according to the continuous space to be explored for paths. Therefore, compared with the method using a grid map where the movable positions are fixed in advance, the method of exploring paths on this continuous space may obtain better paths for each agent.
Prior Art Documents
Non-Patent Documents
[0006]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0007] The inventors of the present invention have found that the above conventional methods for solving the multi-agent path planning problem in a continuous space have the following problems. That is, in the conventional methods, a common road map is constructed for each agent by placing nodes throughout the continuous space. In one example, the nodes are randomly placed throughout the continuous space. As another example, in the method proposed in Non-Patent Document 2, by using a model trained by machine learning, nodes are placed at positions where other agents can be avoided near obstacles. In either method, basically, the nodes are placed throughout the continuous space.
[0008] At this time, by densely placing nodes on the continuous space, the number of combinations of paths available to each agent increases. Therefore, there is a higher possibility of finding a more optimal path for each agent by that road map. However, when nodes are densely placed, the number of nodes increases accordingly, so the cost (computational cost and time cost) required for path search increases. Therefore, when nodes are sparsely placed, the cost required for search can be reduced, but the possibility of finding an optimal path decreases. In the worst case, a situation may occur where a path that can bypass an obstacle or another agent cannot be found. Therefore, the conventional methods have a problem that it is difficult to achieve both finding a more optimal path and reducing the search cost.
[0009] Note that this problem does not occur only in the scenario of searching for the movement paths of multi-agents. Each node may be configured to indicate other states (for example, speed, direction, etc.) in addition to the position. In this case, the same problem can occur in the scenario of searching for the transition path from the start state to the goal state of each agent in the continuous state space.
[0010] In one aspect, the present invention has been made in view of such circumstances, and an object thereof is to provide a technique for increasing the possibility of finding a more optimal path for an agent and reducing the cost required for search when solving the multi-agent path planning problem in a continuous state space.
Means for Solving the Problem
[0011] In order to solve the above-described problems, the present invention adopts the following configuration.
[0012] That is, a route planning system according to one aspect of the present invention includes an information acquisition unit, a map construction unit, and a search unit. The information acquisition unit is configured to acquire target information including a start state and a goal state in the continuous state space of each of a plurality of agents. The map construction unit is configured to construct a roadmap for each agent from the acquired target information using a trained roadmap construction model. The search unit is configured to search for a path of each agent from the start state to the goal state on the roadmap constructed for each agent.
[0013] The roadmap construction model includes a first processing module, a second processing module, and an estimation module. The first processing module is configured to generate first feature information from target agent information including a goal state of a target agent and candidate states at a target time step. The second processing module is configured to generate second feature information from other agent information including a goal state of an agent other than the target agent and candidate states at the target time step. The estimation module is configured to estimate one or more candidate states at the next time step of the target time step of the target agent from the generated first feature information and the second feature information. The trained roadmap construction model is generated by machine learning using learning data obtained from correct paths of a plurality of learning agents.
[0014] Constructing the roadmap for each agent includes treating any one of the plurality of agents as the target agent, treating at least a part of the remaining agents among the plurality of agents as the other agents, designating the start state of any one of the agents indicated by the acquired target information as the candidate state at the first target time step of the target agent, estimating one or more candidate states at the next time step by the trained roadmap construction model, and designating each of the one or more candidate states estimated at the next time step as the candidate state at a new target time step until the goal state or its vicinity state of any one of the agents is included in the one or more candidate states estimated at the next time step, and repeating the estimation of the candidate state at the next time step by the trained roadmap construction model. This process is configured by individually designating each of the plurality of agents as the target agent and executing it.
[0015] The candidate state corresponds to the node constituting the roadmap. That is, the estimation module is configured to estimate the arrangement of one or more nodes that can transition to the next time step on the continuous state space from each piece of feature information of the target time step for the target agent. The candidate state at the target time step and the candidate state at the next time step obtained by estimation are connected by an edge. The edge indicates that it is possible to transition from one candidate state (node) to another candidate state (node).
[0016] In this configuration, a load map for each agent is constructed using a trained load map construction model. The trained load map construction model used is generated by machine learning using learning data obtained from the correct paths of a plurality of learning agents. According to this machine learning, the load map construction model can acquire the ability to construct a load map in accordance with an appropriate path from the start to the goal of the agent. Therefore, according to this configuration, it is possible to construct for each agent a load map in which nodes are arranged in an appropriate path (i.e., the path to be searched) from the start state to the goal state of each agent and in the surrounding area thereof. That is, according to this configuration, by using a trained load map construction model that has learned the correct path, a load map suitable for each agent can be obtained. On the other hand, in the above conventional method, a common load map for all agents is constructed by arranging nodes throughout the continuous state space. Compared with this conventional method, in the load map obtained by this configuration, the range in which nodes are arranged can be narrowed down to an appropriate range for each agent. That is, it is possible to omit the arrangement of nodes at positions that are wasteful for the search path of each agent. Therefore, even if nodes are densely arranged on the load map, an increase in the number of nodes can be suppressed. Thus, according to this configuration, when solving the multi-agent path planning problem on the continuous state space, the possibility of finding a more optimal path for the agent can be increased, and the cost required for the search can be reduced.
[0017] In the path planning system according to the above aspect, the roadmap construction model may further include a third processing module configured to generate third feature information from environmental information including information about obstacles. The estimation module may be configured to estimate one or more candidate states at the next time step from the generated first feature information, second feature information, and third feature information. The acquired target information may be configured to further include information about the obstacles existing in the continuous state space. Using the trained roadmap construction model may include constructing the environmental information from the information included in the acquired target information and providing the constructed environmental information to the third processing module. In this configuration, considering the situation of the environment including obstacles, the arrangement of the nodes constituting the roadmap of each agent can be estimated. Therefore, even in an environment where obstacles exist, a suitable roadmap can be constructed for each agent. Thus, according to this configuration, the possibility of finding an optimal path for the agent on the continuous state space can be increased, and the cost related to the search can be further reduced.
[0018] In the path planning system according to the above aspect, the target agent information may be configured to further include the candidate states of the target agent at time steps before the target time step. The other agent information may be configured to further include the candidate states of the other agents at time steps before the target time step. In this configuration, considering the state transition of each agent in time series, the arrangement of the nodes constituting the roadmap of each agent can be estimated. Therefore, a suitable roadmap can be constructed for an agent that reaches the goal from the start through a plurality of time steps. Thus, according to this configuration, the possibility of finding an optimal path for the agent on the continuous state space can be increased, and the cost related to the search can be further reduced.
[0019] In the route planning system according to the above aspect, the target agent information may be configured to further include the attributes of the target agent.
[0020] In the conventional method, since a common roadmap is used for path search for each agent, it is difficult to consider the individual differences of each agent (that is, to handle agents with different attributes). As an example, assume a scenario of searching for the movement path of each agent that is a moving object on a continuous space where there are narrow passages. In this scenario, when the sizes of the agents may be different, some agents may be able to pass through the passage, but there may be cases where the remaining agents cannot pass through the passage. In such a case, in the conventional method, if the nodes constituting the roadmap are arranged on the passage, there is a possibility of searching for a path that is actually impassable for agents that cannot pass through the passage. On the other hand, if the nodes constituting the roadmap are not arranged on the passage, it is highly likely that an optimal path cannot be searched for agents that can pass through the passage. Thus, in the conventional method, it has been difficult to handle agents with different attributes.
[0021] In contrast, in this configuration, while considering the attributes of each agent, the placement of the nodes constituting the roadmap is estimated for each agent. Thereby, even when there are agents with different attributes, an appropriate roadmap can be constructed for each agent. In the above case, for agents that can pass through the passage (when the passage exists on or around the optimal path), a roadmap with nodes arranged on the passage can be constructed, and for agents that cannot pass through the passage, a roadmap with no nodes arranged on the passage can be constructed. Therefore, according to this configuration, even when agents with different attributes are mixed, the multi-agent path planning problem on the continuous state space can be appropriately solved.
[0022] In the path planning system according to the above aspect, the attributes of the target agent may include at least any one of size, shape, maximum speed, and weight. According to this configuration, even when agents with different sizes, shapes, maximum speeds, and weights are mixed, the multi-agent path planning problem in the continuous state space can be appropriately solved.
[0023] In the path planning system according to the above aspect, the other agent information may be configured to further include the attributes of the other agents. In this configuration, considering the attributes of other agents, the arrangement of the nodes constituting the roadmap of the target agent can be estimated. Therefore, even in an environment where agents with various attributes exist, a roadmap suitable for each agent can be constructed. Thus, according to this configuration, the possibility of finding an optimal path for the agent in the continuous state space can be further enhanced, and the cost related to the search can be further reduced.
[0024] In the path planning system according to the above aspect, the attributes of the other agents may include at least any one of size, shape, maximum speed, and weight. According to this configuration, even in an environment where agents with different sizes, shapes, maximum speeds, and weights exist, a roadmap suitable for each agent can be constructed. Thereby, the possibility of finding an optimal path for the agent in the continuous state space can be further enhanced, and the cost related to the search can be further reduced.
[0025] In the path planning system according to the above aspect, the target agent information may be configured to further include a direction flag indicating the direction in which the target agent transitions in the continuous state space. When the transition direction of the correct path obtained from the training data used for machine learning is biased, the trained roadmap construction model may construct a roadmap in which nodes are arranged only in the biased direction for each agent. If the nodes are arranged in a biased direction, the selection width of the state transition becomes narrow, so it may become impossible to find an optimal path for at least one of the agents. On the other hand, in this configuration, the arrangement direction of the nodes constituting the roadmap of each agent can be controlled by the direction flag given to the roadmap construction model. As a result, it is possible to suppress the nodes from being arranged in a biased direction in the roadmap of each agent. As a result, the possibility of finding an optimal path for the agent in the continuous state space can be further increased.
[0026] In the path planning system according to the above aspect, each of the plurality of agents may be a moving body configured to move autonomously. According to this configuration, in a scenario of solving the path planning problem of a plurality of moving bodies, it is possible to increase the possibility of finding a more optimal path for each moving body and reduce the cost of search. Note that the moving body may be any device configured to be able to move autonomously by machine control. The moving body may be, for example, a mobile robot, an autonomous driving vehicle, a drone, or the like.
[0027] The embodiment of the present invention is not limited to a path planning system configured to construct a roadmap for each agent and search for the path of each agent using a trained roadmap construction model. One aspect of the present invention may be a roadmap construction device configured to construct a roadmap of each agent in any of the above embodiments, or a model generation device configured to generate a trained roadmap construction model used in any of the above embodiments.
[0028] For example, a roadmap construction apparatus according to an aspect of the present invention includes an information acquisition unit configured to acquire target information including start states and goal states in the continuous state spaces of a plurality of agents, and a map construction unit configured to construct a roadmap for each of the agents from the acquired target information using a trained roadmap construction model. According to this configuration, by using a trained roadmap construction model that has learned a correct path, a suitable roadmap can be obtained for each agent. As a result, when solving the multi-agent path planning problem on a continuous state space, the possibility of finding a more optimal path for the agents can be increased, and the cost required for the search can be reduced.
[0029] For example, a model generation device according to an aspect of the present invention includes a data acquisition unit configured to acquire learning data generated from correct paths of a plurality of learning agents, and a learning processing unit configured to perform machine learning of a roadmap construction model using the acquired learning data. The learning data includes a goal state and a plurality of data sets in the correct paths of the respective learning agents. Each of the plurality of data sets is composed of a combination of a state at a first time step and a state at a second time step of each learning agent. The second time step is the next time step after the first time step. The machine learning of the roadmap construction model treats any one of the plurality of learning agents as the target agent, treats at least a part of the remaining learning agents among the plurality of learning agents as the other agents, and for each of the data sets, gives the state at the first time step of any one of the learning agents as a candidate state at the target time step of the target agent to the first processing module, and gives the state at the first time step of at least a part of the remaining learning agents as a candidate state at the target time step of the other agents to the second processing module, so as to train the roadmap construction model such that the candidate state at the next time step of the target agent estimated by the estimation module matches the state at the second time step of any one of the learning agents. According to this configuration, it is possible to generate a trained roadmap construction model that has acquired the ability to construct a suitable roadmap for each agent. By using this trained roadmap construction model, when solving the multi-agent path planning problem on a continuous state space, it is possible to increase the possibility of finding a more optimal path for the agent and reduce the cost required for exploration.
[0030] In the model generation device according to the above aspect, the target agent information may be configured to further include a direction flag indicating a direction in which the target agent transitions in the continuous state space. Each of the data sets may be configured to further include a training flag indicating a direction from the state at the first time step to the state at the second time step in the continuous state space. The machine learning of the roadmap construction model may include providing, to the first processing module, the training flag of any one of the learning agents as the direction flag of the target agent when estimating a candidate state of the target agent at the next time step for each of the data sets.
[0031] In this configuration, based on the training flags included in each data set, the transition direction of each data set used for machine learning can be managed. By using each data set for machine learning so as to train the transitions in each direction evenly with this training flag, it is possible to generate a trained roadmap construction model in which the direction of placing nodes is less likely to be biased. Further, since the target agent information includes an item for the direction flag, it is possible to generate a trained roadmap construction model that has acquired the ability to control the direction of placing nodes with the direction flag. By using this trained roadmap construction model, the direction in which nodes are placed in the roadmap of each agent can be prevented from being biased, and as a result, the selection width of state transitions can be prevented from becoming narrow. As a result, the possibility of finding an optimal path for the agent on the continuous state space can be further increased.
[0032] As another form of the route planning system, roadmap construction device, and model generation device according to each of the above aspects, one aspect of the present invention may be an information processing method for realizing all or a part of each of the above configurations, or an information processing system, or a program, or a computer-readable storage medium storing such a program, such as a computer or other device or machine. Here, a computer-readable storage medium is a medium that stores information such as a program by an electrical, magnetic, optical, mechanical, or chemical action.
[0033] For example, a route planning method according to one aspect of the present invention includes steps in which a computer acquires target information including start states and goal states in the continuous state spaces of a plurality of agents, constructs a roadmap for each agent from the acquired target information using a trained roadmap construction model, and searches for a path for each agent from the start state to the goal state on the roadmap constructed for each agent.
[0034] For example, a route planning program according to one aspect of the present invention is a program for causing a computer to execute steps of acquiring target information including start states and goal states in the continuous state spaces of a plurality of agents, constructing a roadmap for each agent from the acquired target information using a trained roadmap construction model, and searching for a path for each agent from the start state to the goal state on the roadmap constructed for each agent.
[0035] For example, a roadmap construction method according to one aspect of the present invention is an information processing method in which a computer executes: a step of acquiring target information including start states and goal states in the continuous state spaces of a plurality of agents; and a step of constructing a roadmap for each of the agents from the acquired target information using a trained roadmap construction model.
[0036] For example, a roadmap construction program according to one aspect of the present invention is a program for causing a computer to execute: a step of acquiring target information including start states and goal states in the continuous state spaces of a plurality of agents; and a step of constructing a roadmap for each of the agents from the acquired target information using a trained roadmap construction model.
[0037] For example, a model generation method according to one aspect of the present invention is an information processing method in which a computer executes: a step of acquiring learning data generated from correct paths of a plurality of learning agents; and a step of performing machine learning of a roadmap construction model using the acquired learning data.
[0038] For example, a model generation program according to one aspect of the present invention is a program for causing a computer to execute: a step of acquiring learning data generated from correct paths of a plurality of learning agents; and a step of performing machine learning of a roadmap construction model using the acquired learning data.
Advantages of the Invention
[0039] According to the present invention, when solving the multi-agent path planning problem on a continuous state space, it is possible to increase the possibility of finding a more optimal path for an agent and reduce the cost required for search.
Brief Description of the Drawings
[0040]
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16A
Figure 16B
Figure 17
Embodiments for Carrying Out the Invention
[0041] Hereinafter, embodiments according to one aspect of the present invention (hereinafter, also referred to as "the present embodiment") will be described with reference to the drawings. However, the present embodiment described below is merely an exemplification of the present invention in every respect. Needless to say, various improvements and modifications can be made without departing from the scope of the present invention. That is, in carrying out the present invention, a specific configuration according to the embodiment may be appropriately adopted. Note that the data appearing in the present embodiment is described in natural language, but more specifically, it is specified by a pseudo language, command, parameter, machine language, etc. recognizable by a computer.
[0042] §1 Application Example FIG. 1 schematically shows an example of a scene to which the present invention is applied. As shown in FIG. 1, the information processing system according to the present embodiment includes a model generation device 1 and a path planning system 2.
[0043] [Model Generation Device] The model generation device 1 according to the present embodiment is at least one computer configured to generate a trained roadmap construction model 5.
[0044] The model generation device 1 acquires learning data 3 generated from the correct paths of a plurality of learning agents. The model generation device 1 performs machine learning on the load map construction model 5 using the acquired learning data 3.
[0045] (Load map construction model) FIG. 2A schematically illustrates an example of the configuration of the load map construction model 5 according to the present embodiment. The load map construction model 5 according to the present embodiment includes a first processing module 51, a second processing module 52, a third processing module 53, and an estimation module 55.
[0046] The first processing module 51 is configured to generate first feature information from the target agent information 41. The target agent information 41 includes information about the target agent that is useful for constructing the load map of the target agent. The target agent is an agent that is focused on as a target for constructing a load map by the load map construction model 5 among a plurality of agents. The target agent information 41 is configured to include the goal state of the target agent in the continuous state space and the candidate state at the target time step.
[0047] The second processing module 52 is configured to generate second feature information from the other agent information 43. The other agent information 43 includes information about one or more other agents other than the target agent that is useful for constructing the load map of the target agent. The other agents are agents other than the agent that is focused on as the target agent among a plurality of agents. The other agent information 43 is configured to include the goal states of one or more other agents in the continuous state space and the candidate state at the target time step.
[0048] The third processing module 53 is configured to generate third feature information from the environmental information 45. The environmental information 45 is configured to include information about obstacles. An obstacle is an object that can prevent the state transition (e.g., movement) of an agent such as a wall, a package, an animal (human, non-human animal), a step (staircase), etc. The obstacle may be an actually existing object or a virtual object. The number of obstacles may be arbitrary. The obstacle may be an object that can undergo a state transition (e.g., an object that can move) or an object that does not undergo a state transition (e.g., a stationary object). The information about the obstacle may include information indicating attributes of the obstacle such as the presence or absence, position, size, existence range, shape, etc. of the obstacle. The environmental information 45 may further include other information other than the information about the obstacle, which is other information about the environment in which the agent undergoes a state transition (e.g., the situation of the environment that can affect the state transition of the agent, the rules provided in the environment, etc.). As an example, when assuming a moving body that moves outdoors (real space) as the agent, the environmental information 45 may further include information about the weather, road surface conditions, etc. For example, when assuming a moving body that travels on a road such as a general vehicle or an autonomous vehicle as the agent, the environmental information 45 may further include information about the congestion situation, traffic regulations, traffic rules (e.g., the number of traffic lights, one-way traffic, overtaking prohibition, legal maximum speed, etc.). For example, when assuming a flying moving body such as a drone as the agent, the environmental information 45 may further include information about the weather conditions, aviation rules, etc.
[0049] As a result of the operations of the above-described processing modules 51-53, first feature information, second feature information, and third feature information are generated from the input data (target agent information 41, other agent information 43, and environmental information 45). Each feature information may be represented as data in an arbitrary format that can be handled by a computer (e.g., a fixed-length sequence of numbers).
[0050] The estimation module 55 is configured to estimate one or more candidate states of the target agent at the next time step of the target time step from the generated first feature information, second feature information, and third feature information. The configuration of the estimation module 55 for obtaining one or more candidate states may be appropriately determined according to the embodiment. As an example, the estimation module 55 may be configured to be able to change the output (estimation result) each time a trial is performed. In an example of FIG. 2A, the estimation module 55 is configured to be able to change the estimation result by further receiving an input of a multi-dimensional random number vector in addition to each feature information. The value of the random number vector may be obtained by any method such as a known method, for example. According to this configuration, by changing the value of the random number vector given to the estimation module 55, the estimation result obtained from the estimation module 55 can be changed. Therefore, by changing the value of the given random number vector and executing the process of estimating the candidate state from each feature information (the arithmetic process of the estimation module 55) a plurality of times, a plurality of different candidate states can be obtained. Note that the random number vector may be referred to by another expression such as noise, for example.
[0051] The roadmap construction model 5 according to the present embodiment includes the above-described respective modules (51-53, 55), and is configured to estimate one or more candidate states of the target agent at the next time step from the information of the target time step. Therefore, according to the roadmap construction model 5 according to the present embodiment, by repeating the process of estimating the candidate state at the next time step from the candidate state at the target time step while replacing the target time step until reaching the goal state from the start state, the roadmap of the target agent can be constructed.
[0052] The target time step is the time step that is focused on when estimating the placement (candidate state) of nodes in the process of constructing the roadmap. In the candidate state at the first target time step, the start state of each agent is specified. In the candidate state at the second and subsequent target time steps, the candidate state estimated by the roadmap construction model 5 in the estimation process of the immediately preceding target time step is specified. That is, the candidate state at the next time step obtained by the estimation process of the roadmap construction model 5 is specified as the candidate state at the new target time step, and the estimation process of the roadmap construction model 5 is repeated. Thereby, a roadmap from the start state to the goal state can be constructed. For convenience, when the target time step is expressed as the "t-th time step", the next time step may be expressed as the "(t + 1)-th time step" (t is a natural number). The first target time step may be expressed as the "first time step".
[0053] The roadmap defines the range of state transitions that an agent can take in the continuous state space. The constructed roadmap is composed of nodes and edges. Nodes indicate candidate states where the target agent can transition in the continuous state space. That is, the candidate states correspond to the nodes that make up the roadmap. In the process of constructing the roadmap, the node indicating the candidate state at the target time step and the node indicating the estimated candidate state at the next time step are connected by an edge. The edge indicates that the connected nodes (candidate states) can transition in one time step on the continuous state space. The time interval of the time step may be appropriately determined according to the embodiment.
[0054] A continuous state space is a space in which the values of states can take continuous values. States (start state, goal state, candidate state) relate to the dynamic characteristics of an agent that can change over time, such as position, velocity, orientation, etc. The characteristics adopted as states may be appropriately selected according to the embodiment. A state may include multiple types of characteristics. As an example, a state may include at least any one of position, velocity, and orientation. When the characteristic adopted as a state includes position, the continuous state space corresponds to the space indicating the position where the agent exists, the transition on the continuous state space corresponds to movement, and the multi-agent path planning problem corresponds to the problem of planning the movement path of each agent.
[0055] The start state is the state (e.g., the starting point of movement, the current location) that starts a transition on the continuous state space. The goal state is the state (e.g., the destination of movement) that is the target of the transition on the continuous state space. The range of values that a state can take may be arbitrarily defined. In one example, the range of values that a state can take may be defined within a finite range, such as from velocity 0 to the maximum velocity.
[0056] (Various information) The target agent information 41 is configured to include information indicating the goal state of the target agent in the continuous state space and the candidate state at the target time step. In addition to these information, the target agent information 41 may further include other information that can be used for constructing the roadmap of the target agent (specifically, estimating the node placement). In this embodiment, the target agent information 41 may be further configured to include the candidate state of the target agent at a time step earlier than the target time step, the attributes of the target agent (attribute information), the direction flag, and the Cost-to-go feature.
[0057] The time steps that can be treated as the previous time step (hereinafter also referred to as the "past time step") depend on the time step focused on as the target time step. Each time the estimation process by the roadmap construction model 5 is repeated, the number of time steps that can be treated as the past time step increases. The number of time steps to be treated as the past time step does not have to be particularly limited and may be appropriately determined according to the embodiment. When there are a plurality of time steps before the target time step, at least a part of the plurality of time steps may be treated as the past time step.
[0058] In one example, all time steps before the target time step may be treated as the past time step. However, the information of the past time step is referred to in order to estimate the suitable next node arrangement from the arrangement history of the nodes of the target agent. If there is at least a part of the information in the arrangement history of the nodes, that information can be used as a clue to estimate the suitable arrangement of the nodes in the next time step. Therefore, it is not necessarily the case that all time steps before the target time step must be treated as the past time step. In another example, any number of time steps before the target time step may be treated as the past time step. Typically, any number of time steps immediately before the target time step may be treated as the past time step.
[0059] Also, in one example, one or more past time steps may include the time step immediately before the target time step. However, the past time step does not necessarily have to be the immediately previous time step. The past time step may be configured not to include the time step immediately before the target time step. When there is no time step before the target time step, the information indicating the candidate state in the past time step may be omitted.
[0060] Note that providing the first processing module 51 with information indicating candidate states in past time steps may include, for example, providing information indicating candidate states at the target time step to a model having a recursive structure, such as a recurrent neural network, in a time series. In this case, even if the information indicating candidate states in past time steps is not explicitly included in the target agent information 41, the information may be treated as being included in the target agent information 41.
[0061] The attributes relate to static characteristics of an agent that basically do not change over time, such as size, shape, maximum speed, weight, etc. The attribute information included in the target agent information 41 is appropriately configured to indicate the attributes of the target agent. The characteristics adopted as attributes may be appropriately selected according to the embodiment. The attributes may include multiple types of characteristics. As an example, the attributes of the target agent may include at least any one of size, shape, maximum speed, and weight. The maximum speed indicates the maximum amount of movement per unit time. The maximum speed may be determined by any index. As an example, the maximum speed may be the upper limit value of the agent's ability or the regulated value (e.g., legal speed).
[0062] The direction flag is configured to indicate the direction in which the target agent transitions (promotes the transition) in the continuous state space from the target time step to the next time step. The Cost-to-go feature is a feature vector that numerically represents whether it approaches the goal state for each transition direction (e.g., in the case of two dimensions, each point in a K×K rectangular region centered on the candidate state) from the candidate state (node) at the target time step in the continuous state space.
[0063] Other agent information 43 is configured to include information indicating goal states of one or more other agents in a continuous state space and candidate states at a target time step. Similar to the target agent information 41, the other agent information 43 may further include other information about other agents that can be used to construct the roadmap of the target agent in addition to these information. In the present embodiment, the other agent information 43 may be further configured to include candidate states of other agents at a time step prior to the target time step (past time step), attributes (attribute information) of other agents, and Cost-to-go features.
[0064] The information indicating the candidate states of past time steps, the attribute information, and the Cost-to-go feature included in the other agent information 43 may be the same as the target agent information 41, except for aspects regarding other agents. That is, regarding other agents, if there are a plurality of time steps before the target time step, at least a part of the plurality of time steps may be treated as the past time steps of the other agents. If there is no time step before the target time step, in the other agent information 43, the information indicating the candidate states of past time steps may be omitted. Providing the information indicating the candidate states in the past time steps to the second processing module 52 may include, for example, providing the information indicating the candidate states at the target time step to a model having a recursive structure, such as a recurrent neural network, in a time series. In this case, even if the information indicating the candidate states in the past time steps is not explicitly included in the other agent information 43, it may be treated as if the information is included in the other agent information 43. The attribute information included in the other agent information 43 is appropriately configured to indicate the attributes of other agents. The characteristics adopted as attributes may be appropriately selected according to the embodiments. The attributes may include multiple types of characteristics. As an example, the attributes of other agents may include at least any one of size, shape, maximum speed, and weight. The Cost-to-go feature included in the other agent information 43 is, regarding other agents, in the continuous state space, a feature vector that numerically represents whether it approaches the goal state for each arbitrary transition direction from the candidate state at the target time step.
[0065] In the process of repeating the estimation process by the roadmap construction model 5, various information included in the target agent information 41, other agent information 43, and environmental information 45 may be appropriately updated so as to conform to the time step that is the target time step in the estimation process. For example, information indicating candidate states of the target agent and other agents at the target time step included in the target agent information 41 and other agent information 43 may be updated based on the estimation results of the candidate states obtained in the estimation process for the immediately preceding time step. On the other hand, information that does not change during the repetition of the estimation process (for example, the goal state of each agent) may be continuously used as it is. The data formats of various information included in the target agent information 41, other agent information 43, and environmental information 45 do not have to be particularly limited and may be appropriately selected according to the embodiment.
[0066] In an example of FIG. 2A, the roadmap construction model 5 has a separable structure for each module (51-53, 55). That is, in the roadmap construction model 5, each module (51-53, 55) is arranged so as to be clearly distinguishable. However, the structure of the roadmap construction model 5 does not have to be limited to such an example. The roadmap construction model 5 only needs to be configured to receive the input of each information (41, 43, 45) at the target time step and output the estimation results of one or more candidate states (nodes) of the target agent at the next target time step. In another example, each module (51-53, 55) may be at least partially integrally configured, and thereby may be arranged so as not to be clearly distinguishable. Roughly the part related to the input of each information (41, 43, 45) may be regarded as each processing module 51-53, and roughly the part related to the output may be regarded as the estimation module 55. Each processing module 51-53 may be regarded as an encoder that encodes data including information indicating candidate states at the target time step into each feature information, and the estimation module 55 may be regarded as a decoder that restores candidate states at the next time step from each feature information.
[0067] (Processing example of the second processing module) When constructing the roadmap of the target agent, the number of other agents to be processed by the second processing module 52 may be appropriately selected according to the embodiment. In one example, among the plurality of existing agents, all agents except the target agent may be treated as other agents. However, other agents that have little influence on the transition of the target agent are less likely to affect the path of the roadmap of the target agent. Therefore, in another example, the efficiency of processing may be improved by omitting the handling of such other agents with little influence. As a specific example, any number of agents existing in the vicinity of the target agent may be treated as other agents. The number of other agents processed by the second processing module 52 may be a fixed number or a variable number.
[0068] FIG. 2B schematically shows an example of the arithmetic processing of the second processing module 52 according to the present embodiment. In the example of FIG. 2B, the second processing module 52 individually receives the input of the other agent information 43 of other agents, and the feature information (F i ) and the priority (w i ) for the other agent. The priority indicates the degree of importance attached to the information of the corresponding other agent when constructing the roadmap of the target agent. The priority may be read as weight, importance, etc. As an example, when there are M - 1 other agents, the other agent information 43 of each of the other agents is input to the second processing module 52, and the arithmetic processing of the second processing module 52 is repeated M - 1 times, so that M - 1 pieces of feature information (F i ) and the priority (w i ) can be obtained (i is an integer from 1 to M - 1, and M is a natural number of 2 or more). The second processing module 52 is configured to include an arithmetic module that executes a process of integrating the estimation results (outputs) obtained by the following formula 1.
[0069]
Equation
[0070] Note that the configuration of the second processing module 52 for handling information of any number of other agents is not necessarily limited to such an example. In another example, the maximum number of other agents to be handled may be defined in advance. The second processing module 52 may be configured to receive inputs of information of the maximum number of other agents and output respective feature information.
[0071] (Machine learning processing) Returning to FIG. 1, the learning data 3 is configured to include a goal state and a plurality of data sets in the correct path of each learning agent. The learning agent is an agent used to obtain the correct path. The learning agent may be the same as or different from the agent to be inferred (i.e., the agent targeted for path planning). The correct path indicates a path that transitions from the start state to the goal state of the learning agent. The correct path may be obtained manually or from the calculation result of an arbitrary algorithm. Also, the correct path may be obtained based on an arbitrary index. As an example, when handling the position as a state, the correct path may be a path with a short distance, a path that can be reached quickly, a path with low cost, etc.
[0072] Each data set is composed of a combination of the state of each learning agent at the first time step and the state at the second time step. The first time step may be any time step in the correct path. When the first time step is selected as the first time step, the state at the first time step may be composed of the start state in the correct path. The second time step is the next time step after the first time step.
[0073] The machine learning of the roadmap construction model 5 is configured by the following process. That is, in the machine learning of the roadmap construction model 5, the model generation device 1 treats any one of the plurality of learning agents as the target agent, and treats at least a part of the remaining learning agents other than any one of the plurality of learning agents as other agents. Then, for each dataset, the model generation device 1 gives the state of any learning agent at the first time step to the first processing module 51 as the candidate state of the target agent at the target time step, and gives the state of at least a part of the remaining learning agents at the first time step to the second processing module 52 as the candidate state of the other agents at the target time step, so as to train the roadmap construction model 5 so that the candidate state of the target agent at the next time step estimated by the estimation module 55 matches the state of any learning agent at the second time step. As a result of this machine learning, a trained roadmap construction model 5 can be generated that has acquired the ability to estimate an appropriate candidate state of the target agent at the next time step from the candidate states of the target agent and other agents at the target time step.
[0074] [Route Planning System] On the other hand, the route planning system 2 according to the present embodiment is at least one computer configured to solve the multi-agent route planning problem in the continuous state space using the trained roadmap construction model 5 generated by the above machine learning.
[0075] The route planning system 2 acquires target information 221 including the start state and the goal state of each of the plurality of agents in the continuous state space. The route planning system 2 uses the trained roadmap construction model 5 to construct a roadmap 225 for each agent from the acquired target information 221.
[0076] The process of constructing the roadmap 225 for each agent is configured by individually specifying each of the plurality of agents as a target agent and executing the following processes. That is, the route planning system 2 treats any one of the plurality of agents as a target agent and treats at least a part of the remaining agents among the plurality of agents as other agents. The route planning system 2 designates the start state of any agent indicated by the acquired target information 221 as a candidate state at the first target time step of the target agent, and estimates one or more candidate states at the next time step by the trained roadmap construction model 5. The route planning system 2 designates each of the one or more candidate states estimated at the next time step as a candidate state at a new target time step until the goal state or its vicinity state of any agent is included in the one or more candidate states estimated at the next time step, and repeats the estimation of the candidate state at the next time step by the trained roadmap construction model 5. The route planning system 2 individually designates each of the plurality of agents as a target agent and executes these processes. As a result, the roadmap 225 can be constructed for each agent.
[0077] The route planning system 2 searches for the route of each agent from the start state to the goal state on the roadmap 225 constructed for each agent. Then, the route planning system 2 outputs information indicating the searched route.
[0078] [Advantages and effects] As described above, in the model generation device 1 according to the present embodiment, the trained roadmap construction model 5 is generated by machine learning using the learning data 3 obtained from the correct paths of a plurality of learning agents. According to this machine learning, the roadmap construction model 5 can acquire the ability to construct a roadmap in accordance with an appropriate path from the start state to the goal state of the agent. Therefore, in the path planning system 2, it is possible to construct a roadmap 225 in which nodes are arranged in an appropriate path (that is, the path to be searched) from the start state to the goal state of each agent and in its peripheral range for each agent. Thereby, compared with the conventional roadmap construction method of arranging nodes throughout the space, in the constructed roadmap 225 for each agent, the range of arranging nodes can be narrowed to a range suitable for path search of each agent. That is, it is possible to omit the arrangement of nodes at positions that are wasteful for the search path of each agent. Therefore, even if nodes are densely arranged on the roadmap 225, an increase in the number of nodes can be suppressed. Thus, according to the present embodiment, when solving the multi-agent path planning problem on the continuous state space, the possibility of finding a more optimal path for the agent can be increased, and the cost required for the search can be reduced.
[0079] In addition, in one example, as shown in FIG. 1, the model generation device 1 and the path planning system 2 may be connected to each other via a network. The type of the network may be appropriately selected, for example, from the Internet, a wireless communication network, a mobile communication network, a telephone network, a dedicated network, etc. However, the method of exchanging data between the model generation device 1 and the path planning system 2 may not be limited to such an example and may be appropriately selected according to the embodiment. In another example, data may be exchanged between the model generation device 1 and the path planning system 2 using a storage medium.
[0080] Also, in the example of FIG. 1, the model generation device 1 and the route planning system 2 are separate computers respectively. However, the configuration of the information processing system according to the present embodiment may not be limited to such an example and may be appropriately determined according to the embodiment. In another example, the model generation device 1 and the route planning system 2 may be an integrated computer. In still another example, at least one of the model generation device 1 and the route planning system 2 may be configured by a plurality of computers.
[0081] §2 Configuration Example [Hardware Configuration] <Model Generation Device> FIG. 3 schematically shows an example of the hardware configuration of the model generation device 1 according to the present embodiment. As shown in FIG. 3, the model generation device 1 according to the present embodiment is a computer in which a control unit 11, a storage unit 12, a communication interface 13, an input device 14, an output device 15, and a drive 16 are electrically connected. In FIG. 3, the communication interface is described as "communication I / F". The same notation will be used in FIG. 4 described later.
[0082] The control unit 11 includes a hardware processor such as a CPU (Central Processing Unit), a RAM (Random Access Memory), a ROM (Read Only Memory), etc., and is configured to execute information processing based on programs and various data. The control unit 11 (CPU) is an example of a processor resource. The storage unit 12 is an example of a memory resource and is composed of, for example, a hard disk drive, a solid state drive, etc. In the present embodiment, the storage unit 12 stores various information such as a model generation program 81, learning data 3, and learning result data 125.
[0083] The model generation program 81 is a program for causing the model generation device 1 to execute information processing (Figure 10) of machine learning described later to generate the trained roadmap construction model 5. The model generation program 81 includes a series of instructions for the information processing. The learning data 3 is used for the machine learning of the roadmap construction model 5. The learning result data 125 indicates information regarding the trained roadmap construction model 5 generated by machine learning. In the present embodiment, the learning result data 125 is generated as a result of executing the model generation program 81.
[0084] The communication interface 13 is, for example, a wired LAN (Local Area Network) module, a wireless LAN module, etc., and is an interface for performing wired or wireless communication via a network. The model generation device 1 can perform data communication with other computers via the communication interface 13.
[0085] The input device 14 is, for example, a device for performing input such as a mouse, a keyboard, etc. The output device 15 is, for example, a device for performing output such as a display, a speaker, etc. The operator can operate the model generation device 1 by using the input device 14 and the output device 15. The input device 14 and the output device 15 may be integrally configured by, for example, a touch panel display.
[0086] The drive 16 is, for example, a CD drive, a DVD drive, etc., and is a drive device for reading various information such as programs stored in the storage medium 91. At least one of the model generation program 81 and the learning data 3 may be stored in the storage medium 91.
[0087] The storage medium 91 is a medium that stores information such as programs by means of electrical, magnetic, optical, mechanical, or chemical actions so that various devices such as computers and other apparatuses and machines can read the stored programs and other information. The model generation device 1 may acquire at least one of the model generation program 81 and the learning data 3 from this storage medium 91.
[0088] Here, in FIG. 3, as an example of the storage medium 91, a disk-type storage medium such as a CD or a DVD is illustrated. However, the type of the storage medium 91 is not limited to the disk type and may be other than the disk type. Examples of storage media other than the disk type include semiconductor memories such as flash memories. The type of the drive 16 may be appropriately selected according to the type of the storage medium 91.
[0089] Regarding the specific hardware configuration of the model generation device 1, components may be omitted, replaced, or added as appropriate according to the embodiment. For example, the control unit 11 may include a plurality of hardware processors. The hardware processors may be composed of a microprocessor, an FPGA (field-programmable gate array), a DSP (digital signal processor), or the like. The storage unit 12 may be composed of a RAM and a ROM included in the control unit 11. At least one of the communication interface 13, the input device 14, the output device 15, and the drive 16 may be omitted. The model generation device 1 may be composed of a plurality of computers. In this case, the hardware configurations of the respective computers may or may not match. Further, the model generation device 1 may be a general-purpose server device, a general-purpose PC (Personal Computer), an industrial PC, or the like in addition to an information processing device designed specifically for the provided service.
[0090] <Route Planning System> FIG. 4 schematically shows an example of the hardware configuration of the route planning system 2 according to the present embodiment. As shown in FIG. 4, the route planning system 2 according to the present embodiment is a computer in which a control unit 21, a storage unit 22, a communication interface 23, an input device 24, an output device 25, and a drive 26 are electrically connected.
[0091] The control unit 21 to the drive 26 and the storage medium 92 of the route planning system 2 may each be configured in the same manner as the control unit 11 to the drive 16 and the storage medium 91 of the model generation device 1 respectively. The control unit 21 includes a CPU, a RAM, a ROM, etc., which are hardware processors, and is configured to execute various information processes based on programs and data. The control unit 21 (CPU) is an example of a processor resource. The storage unit 22 is an example of a memory resource and is composed of, for example, a hard disk drive, a solid state drive, etc. In the present embodiment, the storage unit 22 stores various information such as a route planning program 82 and learning result data 125.
[0092] The route planning program 82 is a program for causing the route planning system 2 to execute information processing (FIG. 11) described later, which constructs a roadmap 225 for each agent using the trained roadmap construction model 5 and searches for the route of each agent on the constructed roadmap 225. The route planning program 82 includes a series of instructions for the information processing. At least one of the route planning program 82 and the learning result data 125 may be stored in the storage medium 92. The route planning system 2 may acquire at least one of the route planning program 82 and the learning result data 125 from the storage medium 92.
[0093] The route planning system 2 can perform data communication with other computers via the communication interface 23. The operator can operate the route planning system 2 by using the input device 24 and the output device 25.
[0094] Regarding the specific hardware configuration of the route planning system 2, depending on the embodiment, components can be omitted, replaced, or added as appropriate. For example, the control unit 21 may include a plurality of hardware processors. The hardware processors may be composed of a microprocessor, FPGA, DSP, etc. The storage unit 22 may be composed of a RAM and a ROM included in the control unit 21. At least one of the communication interface 23, the input device 24, the output device 25, and the drive 26 may be omitted. The route planning system 2 may be composed of a plurality of computers. In this case, the hardware configurations of the respective computers may or may not match. Also, the route planning system 2 may be a general-purpose server device, a general-purpose PC, an industrial PC, etc., in addition to an information processing device designed specifically for the provided service.
[0095] [Software Configuration] <Model Generation Device> FIG. 5 schematically shows an example of the software configuration of the model generation device 1 according to the present embodiment. The control unit 11 of the model generation device 1 expands the model generation program 81 stored in the storage unit 12 into the RAM. Then, the control unit 11 executes the instructions included in the model generation program 81 expanded in the RAM by the CPU. As a result, as shown in FIG. 5, the model generation device 1 according to the present embodiment is configured to include a data acquisition unit 111, a learning processing unit 112, and a storage processing unit 113 as software modules. That is, in the present embodiment, each software module of the model generation device 1 is realized by the control unit 11 (CPU).
[0096] The data acquisition unit 111 is configured to acquire learning data 3 generated from the correct paths of a plurality of learning agents. The learning processing unit 112 is configured to perform machine learning of the roadmap construction model 5 using the acquired learning data 3. The storage processing unit 113 is configured to generate information regarding the trained roadmap construction model 5 generated by machine learning as learning result data 125, and store the generated learning result data 125 in a predetermined storage area. The learning result data 125 may be appropriately configured to include information for reproducing the trained roadmap construction model 5.
[0097] (Example of acquisition of learning data) FIG. 6 schematically shows an example of a method for acquiring the learning data 3 according to the present embodiment. In an example of FIG. 6, a scenario is assumed in which two learning agents (A, B) acquire learning data 3 from one situation in which they transition from different start states to different goal states. In an example of this situation, the correct path of the learning agent A is one that transitions from the upper left start state to the lower right goal state in 7 time steps, and the correct path of the learning agent B is one that transitions from the lower left start state to the upper right goal state in 7 time steps. Note that the various information shown in FIG. 6, such as the number of learning agents, the start state of each learning agent, the goal state, each state constituting the correct path, and the number of time steps from the start state to the goal state, is merely an example for convenience of explanation. The situation for obtaining the learning data 3 is not limited to such an example.
[0098] The learning data 3 may be appropriately generated from the information indicating the correct paths of the respective learning agents so as to include information capable of constituting a plurality of sets of combinations of training data and correct data (teacher signals, labels) in machine learning. In machine learning, the training data is given to each processing module 51 - 53 of the roadmap construction model 5. The correct data is used as the true value of the candidate state in the next time step of the target agent estimated by the estimation module 55 by giving the training data.
[0099] In this embodiment, the target agent information 41 provided to the first processing module 51 may be configured to include the goal state of the target agent, the candidate state at the target time step, the candidate state at the past time step, attributes, a direction flag, and Cost-to-go features. The other agent information 43 provided to the second processing module 52 may be configured to include the goal state of the other agent, the candidate state at the target time step, the candidate state at the past time step, attributes, and Cost-to-go features. The environmental information 45 provided to the third processing module 53 may be configured to include information about obstacles.
[0100] Accordingly, the training data 3 may be configured to include the goal state 31 in the correct path of each training agent, a plurality of data sets 32, the training attribute information 33 of each training agent, and the training environmental information 35. Each data set 32 may be configured to include the state 321 at the first time step, the state 323 at the second time step, a training flag 325, and Cost-to-go features (not shown) for training. The training flag 325 is configured to indicate the direction from the state 321 at the first time step to the state 323 at the second time step in the continuous state space. The training flag 325 is used as training data for the direction flag.
[0101] As shown in FIG. 6, the information indicating the goal state 31 of each training agent (A, B) can be obtained from the information of the final state (goal state) in the correct path of each training agent (A, B).
[0102] Each dataset 32 of each learning agent (A, B) can be obtained from information on combinations of two consecutive time steps (first time step, second time step) in the correct path of each learning agent (A, B). For the first time step, any time step from the start state to the state immediately before the goal state may be selected, and the second time step may be the time step next to the time step selected as the first time step. In an example of FIG. 6, in the correct path of each learning agent (A, B), seven combinations of the first time step and the second time step can be formed respectively.
[0103] As an example of a method for obtaining each dataset 32, in the correct path of each learning agent (A, B), a combination of two consecutive time steps is formed. The states of the two consecutive time steps of the formed combination can be obtained as the state 321 of the first time step and the state 323 of the second time step. Subsequently, by generating information indicating the direction from the state 321 of the first time step to the state 323 of the second time step, the training flag 325 can be obtained. By generating a feature vector that numerically represents whether or not it approaches the goal state 31 for each transition direction from the state 321 of the first time step, the Cost-to-go feature for training can be obtained. Then, by associating the obtained state 321 of the first time step, the state 323 of the second time step, the training flag 325, and the Cost-to-go feature for training, each dataset 32 of each learning agent (A, B) can be obtained. In FIG. 6, as an example, a scenario of obtaining two datasets 32 for each of the learning agents (A, B) is illustrated. The state 321 of the first time step, the training flag 325, and the Cost-to-go feature for training constitute training data of information of the target time step, and the state 323 of the second time step constitutes correct data (teacher signal, label) of the candidate state of the next time step.
[0104] Note that at least one of the training flag 325 and the training Cost-to-go feature may be pre-generated before being used in machine learning, or may be generated when used in machine learning. Also, in one example, each dataset 32 may be configured to further include information indicating a state (not shown) at a time step before the first time step (past time step). However, if information indicating the time series of the correct path is retained, the information indicating the state at the past time step can be obtained from the dataset 32 generated by a combination of time steps before the target dataset 32. Therefore, in another example, each dataset 32 may not include information indicating the state at the past time step. Similarly, when the first processing module 51 and the second processing module 52 are configured by a model having a recursive structure and information on the target agent 41 and other agent information 43 are given in time series so that information at past time steps is referred to, each dataset 32 may not include information indicating the state at the past time step.
[0105] The training attribute information 33 of each learning agent (A, B) can be appropriately obtained from the information of each learning agent (A, B). In one example, the training attribute information 33 may be configured to indicate at least one of the size, shape, maximum speed, and weight of the learning agent. The training attribute information 33 of each learning agent (A, B) may be obtained from each learning agent (A, B), may be obtained manually, or may be obtained from other computers, external storage devices, etc. The method for obtaining the training attribute information 33 may not be particularly limited and may be appropriately selected according to the embodiment.
[0106] The training environment information 35 can be appropriately obtained from the information of the environment of the situation where the correct path has been obtained. The training environment information 35 may include information about obstacles. The obstacles may be, for example, walls, packages, animals (humans, animals other than humans), steps (stairs), etc. The information about obstacles may include, for example, information indicating attributes of obstacles such as the presence or absence, position, size, existence range, shape, etc. of the obstacles. In the environment of the situation where the correct path is obtained, any obstacles may or may not exist. The information about obstacles can be appropriately obtained from the information of the obstacles in this environment.
[0107] As described above, various information constituting the learning data 3 (the goal state 31 in the correct path of each learning agent, the plurality of data sets 32, the training attribute information 33 of each learning agent, and the training environment information 35) can be obtained. These information can be obtained for each situation where the correct path of each learning agent is obtained. The number of situations for obtaining these information does not need to be particularly limited and may be appropriately determined according to the embodiment. The situations for obtaining these information may be realized in the real space or simulated in the virtual space. By collecting these information from each situation, the learning data 3 can be generated.
[0108] Note that each learning agent is an agent used for obtaining the learning data 3. Each learning agent may be an agent existing in the real space (real agent) or a virtual agent (virtual agent). The correct path of each learning agent may be obtained manually or from the calculation result by any algorithm (for example, a known search algorithm). The correct path may be obtained from the result of the state transition performed by any agent in the past. The data set 32 among the above various information constituting the learning data 3 may be obtained from at least a part of the correct path of each obtained learning agent.
[0109] (An example of a machine learning method) The roadmap construction model 5 is composed of a machine learning model including one or more operation parameters for executing an operation (estimation process) to solve a task, where the values of the one or more operation parameters are adjusted by machine learning. If it is possible to execute an operation process for estimating a candidate state at the next target time step, the type, configuration, and structure of the machine learning model adopted in the roadmap construction model 5 do not have to be particularly limited and may be appropriately determined according to the embodiment. Each processing module 51-53 and the estimation module 55 may be configured as a part of the machine learning model adopted in the roadmap construction model 5.
[0110] As an example, the roadmap construction model 5 may be composed of a neural network. In this case, the weights of the connections between each node, the threshold values of each node, etc. are examples of operation parameters. The first processing module 51 and the second processing module 52 may be composed of, for example, a multi-layer perceptron, a recurrent neural network, etc. The third processing module may be composed of, for example, a multi-layer perceptron, a convolutional neural network, etc. The estimation module 55 may be composed of a neural network, etc. The type of layer, the number of layers, the number of nodes in each layer, and the connection relationship of the nodes of the roadmap construction model 5 (each processing module 51-53 and the estimation module 55) may be appropriately determined according to the embodiment.
[0111] In the machine learning of the roadmap construction model 5, the learning processing unit 112 treats any one of a plurality of learning agents as a target agent and treats at least a part of the remaining learning agents as other agents. Hereinafter, for convenience of explanation, the learning agent selected as the target agent is also referred to as the "learning target agent", and the learning agent selected as the other agent is also referred to as the "learning other agent".
[0112] The learning processing unit 112 constructs a combination of training data and correct answer data (teacher signal, label) from the learning data 3, and trains the roadmap construction model 5 so that the output obtained by providing the training data matches the correct answer data. In the present embodiment, among the various types of information included in the learning data 3, the learning processing unit 112 uses the goal state 31 of the learning target agent, the state 321 at the first time step of each data set 32, the training flag 325, the Cost-to-go feature for training, and the training attribute information 33 to construct the training data (training target agent information) provided to the first processing module 51. The learning processing unit 112 uses the goal state 31 of the other learning agents, the state 321 at the first time step of each data set 32, the Cost-to-go feature for training, and the training attribute information 33 among the various types of information included in the learning data 3 to construct the training data (training other agent information) provided to the second processing module 52. The learning processing unit 112 constructs the training data provided to the third processing module 53 using the training environment information 35 among the various types of information included in the learning data 3. The learning processing unit 112 constructs the correct answer data based on the state 323 at the second time step of each data set 32 of the learning target agent. A combination of training data and correct answer data can be generated for each sample for each data set 32. By providing the configured training data to each of the processing modules 51-53, an estimation result of the candidate state of the target agent at the next time step can be obtained from the estimation module 55. The learning processing unit 112 trains the roadmap construction model 5 so that, for each data set 32, the candidate state of the target agent estimated by the estimation module 55 at the next time step matches the correct answer data (the state 323 at the second time step of the corresponding learning target agent) when the above training data is provided to each of the processing modules 51-53.
[0113] Training the roadmap construction model 5 is configured by adjusting (optimizing) the values of the arithmetic parameters of the roadmap construction model 5 so that the output obtained for the training data conforms to the correct data. The method for solving the optimization problem may be appropriately selected according to the type, configuration, structure, etc. of the machine learning model adopted in the roadmap construction model 5.
[0114] As an example of a machine learning method when the roadmap construction model 5 is configured by a neural network, the learning processing unit 112 gives the training data composed of the learning data 3 to each processing module 51 - 53, and executes the forward propagation arithmetic processing of each processing module 51 - 53. As a result of this arithmetic processing, the learning processing unit 112 acquires feature information from each processing module 51 - 53. The learning processing unit 112 gives the obtained feature information and the random vector to the estimation module 55, and executes the forward propagation arithmetic processing of the estimation module 55. The value of the random vector may be obtained as appropriate. As a result of this arithmetic processing, the learning processing unit 112 obtains the result of estimating the candidate state of the target agent at the next time step. The learning processing unit 112 calculates the error between the obtained estimation result and the corresponding correct data, and further calculates the gradient of the calculated error. The learning processing unit 112 calculates the error of the values of the arithmetic parameters of the roadmap construction model 5 (each processing module 51 - 53 and the estimation module 55) by backpropagating the calculated gradient of the error by the error backpropagation method. Then, the learning processing unit 112 updates the values of the arithmetic parameters based on the calculated error. The learning processing unit 112 adjusts the values of the arithmetic parameters of the roadmap construction model 5 so that the sum of the errors between the estimation result obtained by giving the training data and the correct data becomes small by this series of update processes. The adjustment of the values of these arithmetic parameters may be repeated, for example, until a predetermined condition such as being executed a specified number of times or the sum of the calculated errors becoming equal to or less than a threshold value is satisfied. Also, for example, the conditions of machine learning such as the loss function and the learning rate may be appropriately set according to the embodiment. By this machine learning process, a trained roadmap construction model 5 can be generated.
[0115] FIG. 7 schematically shows another example of the machine learning method of the roadmap construction model 5 according to the present embodiment. In an example of the machine learning method shown in FIG. 7, for machine learning of the roadmap construction model 5, a first processing module 51Z, a second processing module 52Z, and a third processing module 53Z, each having the same configuration as the first processing module 51, the second processing module 52, and the third processing module 53, are prepared. For convenience of explanation, the first processing module 51, the second processing module 52, and the third processing module 53 are collectively referred to as the first encoder E1, and the first processing module 51Z, the second processing module 52Z, and the third processing module 53Z are collectively referred to as the second encoder E2. In FIG. 7, "SA" indicates a learning agent that is treated as a target agent, and "A" is further appended to various information of the learning target agent SA in the learning data 3. Similarly, "SB" indicates a learning agent that is treated as another agent, and "B" is further appended to various information of the learning other agent SB in the learning data 3.
[0116] In an example of FIG. 7, the learning processing unit 112 constructs training data to be given to the first processing module 51 based on the goal state 31A of the learning target agent SA, the state 321A at the first time step of each dataset 32A, the training flag 325A and the Cost-to-go feature for training, and the training attribute information 33A. The learning processing unit 112 constructs training data to be given to the second processing module 52 based on the goal state 31B of the other learning agent SB, the state 321B at the first time step of each dataset 32B and the Cost-to-go feature for training, and the training attribute information 33B. The learning processing unit 112 constructs training data to be given to the third processing module 53 based on the training environment information 35. On the other hand, the learning processing unit 112 replaces the state 321A at the first time step of each dataset 32A of the learning target agent SA with the state 323A at the second time step, and constructs training data to be given to the first processing module 51Z with information similar to the training data given to the first processing module 51 except for this. The training data given to the second processing module 52Z and the third processing module 53Z is the same as that given to the second processing module 52 and the third processing module 53. The learning processing unit 112 constructs correct answer data based on the state 323A at the second time step of each dataset 32A of the learning target agent SA.
[0117] The learning processing unit 112 gives the training data composed of the learning data 3 to the first encoder E1 and executes the forward propagation arithmetic processing of the first encoder E1. As a result of this arithmetic processing, the learning processing unit 112 acquires feature information from the first encoder E1. Similarly, the learning processing unit 112 gives the training data composed of the learning data 3 to the second encoder E2 and executes the forward propagation arithmetic processing of the second encoder E2. As a result of this arithmetic processing, the learning processing unit 112 acquires feature information from the second encoder E2.
[0118] The learning processing unit 112 calculates the distribution error between the feature information obtained from each encoder (E1, E2), and further calculates the gradient of the calculated distribution error. The learning processing unit 112 calculates the error of the value of the operation parameter of the first encoder E1 by backpropagating the calculated error gradient to the first encoder E1 by the error backpropagation method.
[0119] Also, the learning processing unit 112 gives the feature information and the random number vector obtained from the second encoder E2 to the estimation module 55, and executes the forward propagation operation process of the estimation module 55. The value of the random number vector may be obtained as appropriate. As a result of this operation process, the learning processing unit 112 obtains from the estimation module 55 the result of estimating the candidate state of the target agent at the next time step. The learning processing unit 112 calculates the reconstruction error between the obtained estimation result and the corresponding correct data (the state 323A at the second time step), and further calculates the gradient of the calculated reconstruction error. The learning processing unit 112 calculates the error of the value of the operation parameter of each encoder (E1, E2) and the estimation module 55 by backpropagating the calculated error gradient to each encoder (E1, E2) and the estimation module 55 by the error backpropagation method.
[0120] Then, the learning processing unit 112 updates the values of the operation parameters of each encoder (E1, E2) and the estimation module 55 based on the calculated error. The learning processing unit 112 adjusts the value of the operation parameter of the roadmap construction model 5 so that the sum of the calculated distribution error and reconstruction error becomes smaller by this series of update processes. The adjustment of the value of this operation parameter may be repeated, for example, until a predetermined condition such as being executed a specified number of times or the sum of the calculated errors becoming less than or equal to a threshold value is satisfied. Also, for example, conditions of machine learning such as a loss function and a learning rate may be appropriately set according to the embodiment. By this machine learning process, a trained roadmap construction model 5 (the first encoder E1 and the estimation module 55) can be generated.
[0121] Note that, as described above, information indicating the state of past time steps may be appropriately acquired, and the training data provided to the first processing module 51 and the second processing module 52 may further include the information indicating the state of the obtained past time steps. Alternatively, by providing the training data in time series, the information indicating the state of past time steps may be appropriately reflected in the roadmap construction model 5 (the first processing module 51 and the second processing module 52). The same applies to the second encoder E2 when the method illustrated in FIG. 7 is adopted. Also, the learning target agent may be arbitrarily selected from a plurality of learning agents. In the process of solving the optimization problem, for example, the learning target agent may be changed at any timing such as each time a series of update processes is repeated.
[0122] The save processing unit 113 generates learning result data 125 for reproducing the trained roadmap construction model 5 generated by the above machine learning. If the trained roadmap construction model 5 can be reproduced, the configuration of the learning result data 125 does not particularly need to be limited and may be appropriately determined according to the embodiment. As an example, the learning result data 125 may include information indicating the values of each operation parameter obtained by adjusting the above machine learning. In some cases, the learning result data 125 may further include information indicating the structure of the roadmap construction model 5. The structure may be specified by, for example, the number of layers from the input layer to the output layer, the type of each layer, the number of neurons included in each layer, the connection relationship between adjacent neurons, and the like. The save processing unit 113 saves the generated learning result data 125 in a predetermined storage area.
[0123] <Route Planning System> FIG. 8 schematically shows an example of the software configuration of the route planning system 2 according to the present embodiment. The control unit 21 of the route planning system 2 expands the route planning program 82 stored in the storage unit 22 into the RAM. Then, the control unit 21 executes the instructions included in the route planning program 82 expanded in the RAM by the CPU. As a result, as shown in FIG. 8, the route planning system 2 according to the present embodiment is configured to include an information acquisition unit 211, a map construction unit 212, a search unit 213, and an output unit 214 as software modules. That is, in the present embodiment, similar to the model generation device 1, each software module of the route planning system 2 is also realized by the control unit 21 (CPU).
[0124] The information acquisition unit 211 is configured to acquire target information 221 including the start state and the goal state of the transition in the continuous state space of each of the plurality of agents. The map construction unit 212 includes a trained road map construction model 5 by holding the learning result data 125. The map construction unit 212 is configured to construct a road map 225 for each agent from the acquired target information 221 using the trained road map construction model 5. The search unit 213 is configured to search for a path for each agent from the start state to the goal state on the road map 225 constructed for each agent by a predetermined path search method. The output unit 214 is configured to output information indicating the searched path.
[0125] (An example of the process of constructing a road map) The target information 221 is used to constitute the target agent information 41, other agent information 43, and environment information 45 given to each processing module 51-53 when constructing the roadmap 225 of each agent. In the present embodiment, the target information 221 may be configured to further include information on obstacles existing in the continuous state space, information indicating the attributes of each agent (attribute information), and transition flags at each target time step of each agent, in addition to information indicating the start state and goal state of each agent. The attributes of each agent may include at least any one of size, shape, maximum speed, and gravity. The transition flag is configured to indicate the direction of transition (prompting transition) in the continuous state space from the target time step to the next time step when treating each agent as the target agent. The transition flag may be obtained each time a process of estimating candidate states in the next time step is tried.
[0126] The map construction unit 212 treats any one of the plurality of agents as the target agent, and treats at least a part of the remaining agents other than the agent treated as the target agent as other agents. Hereinafter, the agent treated as the target agent among the plurality of agents will also be described as the "tentative target agent", and the agent treated as other agents will also be described as the "tentative other agent". The map construction unit 212 estimates one or more candidate states of the tentative target agent in the next time step by using the trained roadmap construction model 5. By treating each agent individually as the target agent, one or more candidate states of each agent in the next time step can be obtained. The map construction unit 212 repeats the use of the trained roadmap construction model 5 so that the roadmap from the start state to the goal state of each agent is constructed. Thereby, the roadmap 225 for each agent can be constructed.
[0127] In the process of constructing the roadmap 225 for each agent, using the roadmap construction model 5 includes constructing the target agent information 41 based on the information about the tentative target agents included in the target information 221, and providing the constructed target agent information 41 to the first processing module 51. Using the roadmap construction model 5 includes constructing the other agent information 43 based on the information about the tentative other agents included in the target information 221, and providing the constructed other agent information 43 to the second processing module 52. Using the roadmap construction model 5 includes constructing the environmental information 45 from the information about the obstacles included in the target information 221, and providing the constructed environmental information 45 to the third processing module 53.
[0128] Specifically, the map construction unit 212 generates, for each agent, a feature vector that numerically represents whether it approaches the goal state for each transition direction from the start state included in the target information 221, thereby obtaining the Cost-to-go feature to be provided as the information of the first target time step. The map construction unit 212 constructs the target agent information 41 to be provided to the first processing module 51 as the information of the first target time step (the first time step) based on the start state, goal state, attribute information, and transition flag of the tentative target agents included in the target information 221, and the obtained Cost-to-go feature. The map construction unit 212 constructs the other agent information 43 to be provided to the second processing module 52 as the information of the first target time step based on the start state, goal state, and attribute information of the tentative other agents included in the target information 221, and the obtained Cost-to-go feature. That is, the map construction unit 212 designates the start state of the tentative target agent as the candidate state at the first target time step of the target agent, and designates the start state of the tentative other agent as the candidate state at the first target time step of the other agent. The map construction unit 212 constructs the environmental information 45 to be provided to the third processing module 53 from the information about the obstacles included in the target information 221.
[0129] The map construction unit 212 provides the configured target agent information 41, other agent information 43, and environment information 45 to each processing module 51-53, and executes the arithmetic processing of the trained roadmap construction model 5. Thereby, the map construction unit 212 obtains the estimation results of one or more candidate states of the tentative target agent at the next time step (the second time step) from the estimation module 55. In the present embodiment, the map construction unit 212 can obtain a plurality of different estimation results for the candidate states of the tentative target agent at the next time step by changing the value of the random vector given to the estimation module 55 during this estimation process. The map construction unit 212 can obtain one or more candidate states of each agent at the next time step by individually treating each agent as the target agent and executing the above arithmetic processing.
[0130] The map construction unit 212 designates each of one or more candidate states of each agent at the next time step (the (t + 1)-th time step, where t is a natural number of 1 or more) as a candidate state of each agent at the new target time step, and repeats the estimation of the candidate state at the next time step by the trained roadmap construction model 5. Specifically, the map construction unit 212 generates, for each agent, a feature vector that numerically represents whether each of the one or more estimated candidate states approaches the goal state in any transition direction, thereby obtaining a Cost-to-go feature to be given as information on the new target time step for each of the one or more estimated candidate states. The map construction unit 212 constructs target agent information 41 to be given to the first processing module 51 as information on the new target time step based on each of the one or more estimated candidate states of the tentative target agent, the goal state, attribute information, a transition flag, and the obtained Cost-to-go feature. The map construction unit 212 constructs other agent information 43 to be given to the second processing module 52 as information on the new target time step based on each of the one or more estimated candidate states of the tentative other agents (the estimation results obtained by treating the tentative other agents as target agents), the goal state, attribute information, and the obtained Cost-to-go feature. The map construction unit 212 constructs environment information 45 to be given to the third processing module 53 from the information on obstacles included in the target information 221. The map construction unit 212 gives the constructed target agent information 41, other agent information 43, and environment information 45 (information on the new target time step) to each of the processing modules 51-53 to execute the arithmetic processing of the trained roadmap construction model 5. As a result, the map construction unit 212 obtains, from the estimation module 55, the estimation results of one or more candidate states of the tentative target agent at the time step next to the new target time step. The map construction unit 212 can obtain one or more candidate states of each agent at the time step next to the new target time step by treating each agent as a target agent individually and executing the above arithmetic processing.
[0131] That is, the map construction unit 212 updates the information indicating the candidate states of each agent and the Cost-to-go feature given as the target agent information 41 and the other agent information 43 based on the estimation results of one or more candidate states in the next time step obtained for each agent, and re-executes the estimation process of the candidate states in the next time step by the trained roadmap construction model 5. The map construction unit 212 repeats the estimation process of the candidate states in the next time step for the tentative target agent by the trained roadmap construction model 5 until the goal state or its neighboring state of the tentative target agent is included in one or more candidate states estimated in the next time step. The map construction unit 212 individually designates each agent as the target agent and executes the repetition of the above series of estimation processes. Thereby, the roadmap 225 for each agent can be constructed.
[0132] Note that the transition flag may be obtained at an arbitrary timing in each time step. As an example, the transition flag may be obtained when obtaining different estimation results of candidate states by changing the value of the random vector for each time step and each agent. Thereby, the direction of arranging the node (candidate state) in the next time step with respect to the node (candidate state) in the target time step can be controlled. Also, the transition flag may be obtained by an arbitrary method. The transition flag may be obtained, for example, by a method such as an instruction by an operator, random, selection based on an arbitrary index / a predetermined rule. As an example of a method for obtaining the transition flag based on an arbitrary index or a predetermined rule, among a predetermined range centered on the direction approaching the goal state from the candidate state in the target time step of the tentative target agent, the direction in which the node (estimation result of the candidate state in the next target time step) does not exist or is few (for example, the density is below the threshold) may be selected as the direction indicated by the transition flag. Thereby, it is possible to prevent the direction in which the nodes are arranged in the constructed roadmap 225 from being biased.
[0133] Also, the target agent information 41 and the other agent information 43 given to the first processing module 51 and the second processing module 52 at each time step may each include information indicating the candidate states of the respective agents at past time steps. That is, using the trained roadmap construction model 5 may involve, when repeating the process of estimating the candidate state of the next target time step, providing the candidate state of the tentative target agent at time steps prior to the target time step as the candidate state at the past time step to the first processing module 51, and providing the candidate state of the tentative other agent at time steps prior to the target time step as the candidate state at the past time step to the second processing module 52. The information indicating the candidate states of the respective agents at past time steps may be appropriately obtained from the estimation results of the candidate states obtained by previous estimation processes and the start state. Alternatively, since the first processing module 51 and the second processing module 52 have a recursive structure, information from past time steps may be reflected by executing the above-described estimation process in time series.
[0134] FIG. 9 schematically shows an example of the process of constructing a roadmap 225 for each agent by the trained roadmap construction model 5 according to the present embodiment. Specifically, FIG. 9 shows a part of the process of executing the above-described estimation process assuming that there are M agents, and when any one of the agents is treated as the target agent, all the remaining M−1 agents are treated as other agents. However, as described above, the number of agents and the treatment of other agents need not be limited to such an example.
[0135] In an example of FIG. 9, in the process of constructing the roadmap 225 of the first agent, the first agent is treated as the target agent, and the second to M-th agents are treated as other agents (the top line of FIG. 9). At the stage where the T-th time step is the target time step, the target agent information 41 is constructed from the information about the first agent at the T-th time step, and the other agent information 43 is constructed from the information about each of the second to M-th agents at the T-th time step. Thus, as described above, at the stage where the T-th time step is the target time step, the input information given to the roadmap construction model 5 when treating the first agent as the target agent is obtained.
[0136] The same applies when treating each of the second to M-th agents as the target agent. The target agent information 41 is constructed from the information about the tentative target agent at the T-th time step, and the other agent information 43 is constructed from the information about the tentative other agent at the T-th time step. When the T-th time step is the first target time step (T = 1), the start state of each is specified as the candidate state of the tentative target agent and the tentative other agent that constitute the target agent information 41 and the other agent information 43. In other cases (T is 2 or more), any one of the one or more candidate states estimated immediately before is specified as the candidate state of the tentative target agent and the tentative other agent that constitute the target agent information 41 and the other agent information 43.
[0137] Then, by providing the obtained input information to the roadmap construction model 5 and executing the arithmetic processing of the roadmap construction model 5, it is possible to obtain the estimation results of one or more candidate states of the tentative target agent at the next time step. When the T-th time step is the target time step, in the process of treating the first agent as the target agent, it is possible to obtain the estimation results of one or more candidate states of the first agent at the (T + 1)-th time step. In order to obtain the input information to be provided to the roadmap construction model 5 in the next stage (i.e., the stage of treating the (T + 1)-th time step as the new target time step), based on this estimation result, the information indicating the candidate state of the first agent at the target time step and the Cost-to-go feature are updated. Thereby, information regarding the first agent at the (T + 1)-th time step can be obtained. The same applies to the second to M-th agents. In the process of treating each agent as the target agent, it is possible to obtain the estimation results of one or more candidate states of each agent at the (T + 1)-th time step. Based on the obtained estimation results, by updating the information indicating the candidate state at the target time step and the Cost-to-go feature, information regarding each agent at the (T + 1)-th time step can be obtained.
[0138] Using the information of each agent at the obtained (T + 1)-th time step, by respectively executing the estimation process on the tentative target agent by the above-mentioned roadmap construction model 5, it is possible to obtain the estimation results of one or more candidate states of each agent at the (T + 2)-th time step. The estimation process by the trained roadmap construction model 5 is repeatedly executed for each agent until the goal state or its neighboring state is included in one or more candidate states at the estimated next time step. Thereby, a roadmap 225 for each agent (in the example of FIG. 9, M roadmaps) can be constructed.
[0139] <Others> In this embodiment, an example in which each software module of the model generation device 1 and the route planning system 2 is realized by a general-purpose CPU is described. However, part or all of the above software modules may be realized by one or more dedicated processors (for example, a graphics processing unit). Each of the above modules may be realized as a hardware module. Further, with respect to the software configurations of the model generation device 1 and the route planning system 2, omission, substitution, and addition of software modules may be appropriately performed according to the embodiment.
[0140] §3 Operation Example [Model Generation Device] FIG. 10 is a flowchart showing an example of the processing procedure of the model generation device 1 according to this embodiment. The following processing procedure of the model generation device 1 is an example of the model generation method. However, the following processing procedure of the model generation device 1 is merely an example, and each step may be changed as much as possible. Further, with respect to the following processing procedure of the model generation device 1, omission, substitution, and addition of steps are possible as appropriate according to the embodiment.
[0141] (Step S101) In step S101, the control unit 11 operates as a data acquisition unit 111 and acquires the learning data 3 generated from the correct routes of a plurality of learning agents.
[0142] The learning data 3 may be appropriately generated from the correct routes of the learning agents by the method exemplified in FIG. 6. In this embodiment, the learning data 3 may be configured to include the goal state 31 in the correct route of each learning agent, a plurality of data sets 32, training attribute information 33 for each learning agent, and training environment information 35. Each data set 32 may be configured to include the state 321 at the first time step, the state 323 at the second time step, a training flag 325, and a Cost-to-go feature for training.
[0143] The learning data 3 may be automatically generated by the operation of a computer, or may be manually generated by including at least partially the operations of an operator. Also, the generation of the learning data 3 may be performed by the model generation device 1, or may be performed by another computer other than the model generation device 1. The control unit 11 may automatically or manually generate the learning data 3 by the operation of the operator via the input device 14. Alternatively, the control unit 11 may obtain the learning data 3 generated by another computer via, for example, a network, a storage medium 91, etc. A part of the learning data 3 may be generated by the model generation device 1 and the other part may be generated by one or more other computers.
[0144] The amount of various information acquired as the learning data 3 may not be particularly limited and may be appropriately determined so that machine learning can be performed. When the learning data 3 is acquired, the control unit 11 proceeds to the next step S102.
[0145] (Step S102) In step S102, the control unit 11 operates as the learning processing unit 112 and performs machine learning of the load map construction model 5 using the acquired learning data 3.
[0146] As an example, the control unit 11 performs initial setting of the machine learning model constituting the load map construction model 5. The structure of the machine learning model and the initial values of the calculation parameters may be given by a template, or may be determined by the input of an operator. When performing additional learning or re-learning, the control unit 11 may perform initial setting of the load map construction model 5 based on the learning result data obtained by past machine learning.
[0147] Next, the control unit 11 treats any one of the plurality of learning agents as a target agent, and treats at least a part of the remaining learning agents as other agents, and constructs a combination of training data and correct answer data from the learning data 3. The training data given to the first processing module 51 is composed of the goal state 31 of the learning target agent, the state 321 at the first time step of each data set 32, the training flag 325 and the Cost-to-go feature for training, and the training attribute information 33. The training data given to the second processing module 52 is composed of the goal state 31 of the learning other agent, the state 321 at the first time step of each data set 32 and the Cost-to-go feature for training, and the training attribute information 33. The training data given to the third processing module 53 is composed of the training environment information 35. The correct answer data is composed of the state 323 at the second time step of each data set 32 of the learning target agent. The control unit 11 adjusts the value of the calculation parameter of the roadmap construction model 5 so that the candidate state at the next time step of the target agent estimated by the estimation module 55 conforms to the correct answer data by giving the above training data to each processing module 51-53 for each data set 32.
[0148] When adopting the machine learning method illustrated in FIG. 7 and adopting a neural network for the roadmap construction model 5, the control unit 11 further prepares a second encoder E2 having the same configuration as the first encoder E1 in the above initial setting. The configuration of the training data and the correct answer data given to each encoder (E1, E2) is as described above. The training data given to the first processing module 51Z of the second encoder E2 is configured by replacing the state 321A at the first time step of each data set 32A with the state 323A at the second time step among the information included in the training data given to the first processing module 51Z of the first encoder E1. Except for this point, the training data given to the second encoder E2 is the same as the training data given to the first encoder E1.
[0149] The control unit 11 provides each encoder (E1, E2) with its respective training data, and acquires feature information from each encoder (E1, E2) by executing forward propagation arithmetic processing. The control unit 11 calculates the error (inter-distribution error) between the feature information obtained from each encoder (E1, E2). The control unit 11 adjusts the values of the arithmetic parameters of the first encoder E1 (each processing module 51-53) so that the sum of the calculated inter-distribution errors becomes small by the error backpropagation method. At the same time, the control unit 11 gives the feature information obtained from the second encoder E2 and the random number vector to the estimation module 55, and acquires the estimation result of the candidate state of the target agent at the next time step from the estimation module 55 by executing the forward propagation arithmetic processing of the estimation module 55. The control unit 11 calculates the error (reconstruction error) between the obtained estimation result and the corresponding correct data (the state 323A at the second time step). The control unit 11 adjusts the values of the arithmetic parameters of the first encoder E1, the second encoder E2, and the estimation module 55 so that the sum of the calculated reconstruction errors becomes small by the error backpropagation method. As a result of this machine learning process, a trained roadmap construction model 5 can be generated that has acquired the ability to estimate an appropriate candidate state of the target agent at the next time step from the candidate states of the target agent and other agents at the target time step.
[0150] In this embodiment, each sample of the training data includes a training flag 325. According to the training flag 325, the transition direction of the samples used for machine learning can be managed. Thereby, the training data can be used for machine learning so as to train the transitions in each direction evenly. As an example, the control unit 11 may count the number of samples in each transition direction by referring to the training flag 325. Then, in the above machine learning, the control unit 11 may extract (sample) the samples in each transition direction with a probability that is the reciprocal of the obtained number of samples. The transition direction may be defined to have a certain width, for example, up, down, left, right, etc. As another example, the control unit 11 may adjust the number of samples in each transition direction so that the number of samples in each transition direction is uniform (the same or the difference is below a threshold). As an example of the adjustment method, for a transition direction with a smaller number of samples than other transition directions, the control unit 11 may execute a process of increasing the number of samples by methods such as replication and data augmentation. By training the transitions in each direction evenly with the training flag 325, a trained load map construction model 5 in which the direction of arranging nodes is less likely to be biased can be generated.
[0151] When the process of machine learning is completed as described above, the control unit 11 proceeds to the next step S103.
[0152] (Step S103) In step S103, the control unit 11 operates as a storage processing unit 113 and generates information regarding the trained load map construction model 5 generated by machine learning as learning result data 125. The control unit 11 stores the generated learning result data 125 in a predetermined storage area.
[0153] The predetermined storage area may be, for example, the RAM in the control unit 11, the storage unit 12, an external storage device, a storage medium, or a combination thereof. The storage medium may be, for example, a CD, a DVD, etc. The control unit 11 may store the learning result data 125 in the storage medium via the drive 16. The external storage device may be, for example, a data server such as a NAS (Network Attached Storage). In this case, the control unit 11 may store the learning result data 125 in the data server via the network using the communication interface 13. Also, the external storage device may be, for example, an external storage device connected to the model generation device 1.
[0154] When the storage of the learning result data 125 is completed, the control unit 11 ends the processing procedure of the model generation device 1 according to this operation example.
[0155] Note that the generated learning result data 125 may be provided to the route planning system 2 at an arbitrary timing. For example, the control unit 11 may transfer the learning result data 125 to the route planning system 2 as the processing of step S103 or separately from the processing of step S103. The route planning system 2 may obtain the learning result data 125 by receiving this transfer. Also, for example, the route planning system 2 may obtain the learning result data 125 by accessing the model generation device 1 or the data server via the network using the communication interface 23. Also, for example, the route planning system 2 may obtain the learning result data 125 via the storage medium 92. Also, for example, the learning result data 125 may be pre - incorporated into the route planning system 2.
[0156] Furthermore, the control unit 11 may update or newly generate the learning result data 125 by repeatedly performing the processes of the above steps S101 to S103 regularly or irregularly. During this repetition, at least some of the learning data 3 used for machine learning may be appropriately changed, modified, added, deleted, etc. Then, the control unit 11 may update the learning result data 125 held by the route planning system 2 by providing the updated or newly generated learning result data 125 to the route planning system 2 by any method.
[0157] [Route Planning System] FIG. 11 is a flowchart showing an example of the processing procedure of the route planning system 2 according to the present embodiment. The processing procedure of the route planning system 2 described below is an example of a route planning method. However, the processing procedure described below is only an example, and each step may be changed as much as possible. Also, regarding the following processing procedure, steps may be appropriately omitted, replaced, and added according to the embodiment.
[0158] (Step S201) In step S201, the control unit 21 operates as an information acquisition unit 211 and acquires target information 221 including the start state and goal state of the transition in the continuous state space of each of the plurality of agents.
[0159] In the present embodiment, the control unit 21 may acquire, as the target information 221, information indicating the start state and goal state of each agent, information regarding obstacles existing in the continuous state space, information indicating the attributes of each agent (attribute information), and a transition flag at each target time step of each agent. The various types of information may be acquired by any method such as, for example, information processing, operator input, acquisition from the storage unit 22, or acquisition from an external device (e.g., another computer, storage medium, external storage device, etc.). At least a part of the various types of information may be held in advance in the route planning system 2 or may be acquired by any method when used. When the target information 221 is acquired, the control unit 21 proceeds to the next step S202.
[0160] (Step S202) In step S202, the control unit 21 operates as a map construction unit 212, and sets the trained load map construction model 5 by referring to the learning result data 125. Then, the control unit 21 constructs a load map 225 for each agent from the acquired target information 221 using the trained load map construction model 5.
[0161] As described above, the control unit 21 treats any one of the plurality of agents as a target agent, and treats at least a part of the remaining agents other than the agent treated as the target agent as other agents. The control unit 21 constructs target agent information 41 and other agent information 43 from the information of each agent included in the target information 221. The control unit 21 constructs environment information 45 from the information regarding obstacles included in the target information 221. The configurations of the target agent information 41, the other agent information 43, and the environment information 45 are as described above. When the environment information 45 is configured to further include information other than the information regarding obstacles, the control unit 21 may acquire the information in the environment to be subjected to path search as the target information 221, and construct the environment information 45 further including the acquired information. The information regarding the environment to be subjected to path search including the said information and the information regarding obstacles may be referred to as "target environment information". The control unit 21 may acquire the target environment information as the target information 221, and construct the environment information 45 provided to the third processing module 53 based on the acquired target environment information.
[0162] The control unit 21 provides the configured target agent information 41, other agent information 43, and environment information 45 to each processing module 51 - 53, and executes the arithmetic processing of the trained roadmap construction model 5. Thereby, the control unit 21 obtains the estimation results of one or more candidate states of the tentative target agent at the next time step. In the configuration example of FIG. 2A, during the estimation process, the control unit 21 can obtain a plurality of different estimation results for the candidate states of the tentative target agent at the next time step by changing the value of the random vector given to the estimation module 55. The control unit 21 treats each agent as a target agent individually, and by executing the above arithmetic processing, obtains one or more candidate states of each agent at the next time step.
[0163] If the obtained candidate state (node) is unreachable from the candidate state (node) at the target time step for some reason, such as colliding with an obstacle, the control unit 21 may discard or modify the obtained candidate state (node). As an example, the control unit 21 may discard the obtained candidate state and obtain a new candidate state at the next time step by another method such as random sampling. As another example, the control unit 21 may modify the obtained candidate state by moving the node corresponding to the obtained candidate state within a neighborhood range in the continuous state space. The moving direction may be determined arbitrarily. Also, the neighborhood range may be appropriately determined by a threshold value or the like.
[0164] Also, in order to evaluate whether to adopt the obtained candidate state (node), the control unit 21 may determine whether there is a candidate state (node) that can be replaced with the candidate state (node) obtained by the current stage of estimation processing among the candidate states (nodes) obtained by the previous estimation processing. The index for evaluating whether replacement is possible may be appropriately set according to the embodiment. As an example, a candidate state that can be connected to other candidate states (nodes) in the same manner as the candidate state (node) obtained at the current stage and is arranged in the vicinity range on the continuous state space with the candidate state (node) obtained at the current stage may be evaluated as a replaceable candidate state. The vicinity range in the evaluation may be appropriately determined by a threshold value or the like. If there is no candidate state that can be replaced with the candidate state obtained at the current stage, the control unit 21 may adopt the obtained candidate state as it is. On the other hand, if there is a replaceable candidate state, the control unit 21 may defer the adoption of the obtained candidate state. Alternatively, the control unit 21 may integrate (e.g., average) the replaceable candidate state and the obtained candidate state. The control unit 21 connects, with an edge, the candidate state (node) reachable from the candidate state obtained at the current stage among the candidate states (nodes) obtained up to the previous stage and the candidate state obtained at the current stage.
[0165] When the estimation processing for the target time step is completed, the control unit 21 uses each of the one or more candidate states of each agent at the next time step estimated as the candidate state of each agent at the new target time step. That is, the control unit 21 updates the information indicating the candidate state of each agent and the Cost-to-go feature given as the target agent information 41 and the other agent information 43 according to the estimation results of the one or more candidate states at the next time step obtained for each agent. Then, the control unit 21 executes the estimation processing by the trained roadmap construction model 5 again.
[0166] The control unit 21 repeatedly executes the estimation process by the trained roadmap construction model 5 until the goal state or a state in the vicinity thereof of the tentative target agent is included in one or more candidate states at the estimated next time step. Whether the estimated candidate state is a state in the vicinity (arranged in the vicinity of the goal state) may be appropriately evaluated by a threshold value or the like. In the present embodiment, in the process of repeating the estimation process, information indicating candidate states of each agent at past time steps may be given to the trained roadmap construction model 5 as the target agent information 41 and the other agent information 43. Alternatively, since the trained roadmap construction model 5 (the first processing module 51 and the second processing module 52) has a recursive structure, information of past time steps may be appropriately reflected in the estimation process. Thereby, not only the information of the candidate state at the target time step but also the information of the candidate state at the past time steps may be referred to for estimating the candidate state at the next time step. Further, when each agent is treated as a target agent and the estimation process of the candidate state is executed, the transition flag may be newly obtained at an arbitrary timing such as each time the process of estimating the candidate state is tried at each time step. The transition flag may be obtained by an arbitrary method such as an instruction by an operator, random, selection based on an arbitrary index / a predetermined rule, or the like. Through the above processing, the control unit 21 can construct the roadmap 225 for each agent.
[0167] Note that, regarding the estimation process at each time step of each agent, the control unit 21 may generate candidate states at the next time step by another algorithm (for example, randomly transitioning states) without executing the above load map construction model 5 with a predetermined probability. The predetermined probability may be set as appropriate. Thereby, even if it is difficult to appropriately estimate candidate states by the trained load map construction model 5 due to a large difference in the possible paths of the agent between the path planning problem for which learning data 3 is acquired in the learning stage and the path planning problem in the inference stage, the candidate states (nodes) generated by another algorithm are included in the load map, so that path planning can be stably performed.
[0168] FIG. 12 schematically shows an example of load maps (225X, 225Y) for each agent (X, Y) constructed by the trained load map construction model 5 according to the present embodiment. In the example of FIG. 12, there are two agents (X, Y). Agent X plans a transition from the upper left to the lower right in the continuous state space, and agent Y plans a transition from the lower left to the upper right. In this scenario, the load map 225X (solid line) shows an example of the load map 225 constructed for agent X, and the load map 225Y (dotted line) shows an example of the load map 225 constructed for agent Y.
[0169] The trained load map construction model 5 has acquired the ability to construct the load map 225 in accordance with an appropriate path from the start state to the goal state of the agent by machine learning using the learning data 3 obtained from the correct paths of a plurality of learning agents. Therefore, as illustrated in FIG. 12, in the load maps (225X, 225Y) for each agent (X, Y) to be constructed, nodes are likely to be arranged in a range with a high probability of the existence of an optimal path from the start state to the goal state of each agent (X, Y), and it is possible to make it difficult for nodes to be arranged in other ranges.
[0170] When the construction of the roadmap 225 for each agent is completed, the control unit 21 proceeds to the next step S203.
[0171] (Step S203) Returning to FIG. 11, in step S203, the control unit 21 operates as a search unit 213 and searches for the path of each agent from the start state to the goal state on the roadmap 225 constructed for each agent by a predetermined path search method. The path search method may be appropriately selected according to the embodiment. In one example, known path search methods such as CBS (Conflict-based Search), prioritized planning, etc. may be used to search for the path of each agent. When the search for the path of each agent is completed, the control unit 21 proceeds to the next step S204.
[0172] (Step S204) In step S204, the control unit 21 operates as an output unit 214 and outputs information indicating the searched path (hereinafter, also referred to as "search result information").
[0173] The output destination and the content of the information to be output may be appropriately determined according to the embodiment. As an example, the control unit 21 may output the information indicating the searched path of each agent as it is. The output information may be utilized for each agent to perform a state transition. For example, when the agent is a human or a device operated by a human, the human may perform his or her own state transition or the operation of the device based on the output information. As another example, when each agent can be directly or indirectly controlled via a control device, the control unit 21 may output, as search result information, an instruction to instruct the operation of each agent according to the search result of the path (that is, to prompt each agent to perform a state transition according to the search result). The output destination may be, for example, a RAM, a storage unit 22, an output device 25, another computer, an agent, etc.
[0174] Note that the number of agents targeted for path planning does not have to be particularly limited and may be appropriately selected according to the embodiment. When repeating path planning, the number of agents may vary. Also, each agent targeted for path planning may be an agent existing in the real space (real agent) or a virtual agent (virtual agent). Each agent may be the same as or different from the learning agent.
[0175] As long as each agent can transition states, its type does not have to be particularly limited and may be appropriately selected according to the embodiment. Each agent may be, for example, a moving body, a human, a device operated by a human, a manipulator, etc. The device operated by a human may be, for example, a general vehicle, etc. When the agent is a manipulator, solving the path planning problem may correspond to solving the motion planning problem of the manipulator. All agents targeted for path planning may be of the same type, or at least some agents may be of a different type from other agents. As an example, each agent may be a moving body configured to move autonomously. The moving body may be, for example, a mobile robot, an autonomous driving vehicle, a drone, etc. The mobile robot may be configured to move, for example, for the purpose of transporting, guiding, patrolling, cleaning, etc. of articles.
[0176] FIG. 13 schematically shows an example of an agent according to this embodiment. In the scene of FIG. 13, there are two moving bodies (6X, 6Y). The moving body 6X located in the upper left is planning to move to the goal GX in the lower right, and the moving body 6Y located in the lower left is planning to move to the goal GY in the upper right. Assume that the continuous state space indicates the positions of the respective moving bodies (6X, 6Y). Each of the moving bodies (6X, 6Y) is an example of an agent.
[0177] In this scenario, the control unit 21 may treat each moving body (6X, 6Y) as an agent and execute the processes of steps S201 to S203 to search for the movement paths to the respective goals (GX, GY) of each moving body (6X, 6Y). And as an example of the process of this step S204, the control unit 21 may instruct each moving body (6X, 6Y) to move according to the search result by outputting the obtained search result information to each moving body (6X, 6Y).
[0178] When the output of the search result information is completed, the control unit 21 ends the processing procedure of the path planning system 2 according to this operation example.
[0179] Note that the situation where path planning is performed may occur or have occurred in the real space, or may be realized by a simulation in the virtual space. Accordingly, the control unit 21 may execute a series of information processes of steps S201 to S204 at an arbitrary timing. The timing for performing path planning may be online or offline.
[0180] As an example, in a scenario where each agent is a moving body and each moving body performs the planning of a movement path (for example, FIG. 13), the path planning system 2 may acquire the information on the current location and the destination of each moving body in real time and online. The path planning system 2 may treat the current location as the start state and the destination as the goal state, and execute the processes of steps S201 to S204. Thereby, the path planning system 2 may plan the movement paths of each moving body in real time. Also, the path planning system 2 may acquire the information on the current location and the destination of each moving body periodically or irregularly, and update the planning of the movement paths of each moving body by executing the series of information processes of steps S201 to S204 again. When repeating the series of information processes of steps S201 to S204, the number of moving bodies (agents) may vary. Thereby, appropriate path information may be continuously created for each moving body until each moving body reaches the destination.
[0181] [Features] As described above, in this embodiment, in the process of step S102, the trained roadmap construction model 5 is generated by machine learning using the learning data 3 obtained from the correct paths of a plurality of learning agents. According to this machine learning, it is possible to generate a trained roadmap construction model 5 that has acquired the ability to construct a roadmap in accordance with an appropriate path from the start state to the goal state of the agent. In the process of step S202, by using this trained roadmap construction model 5, as illustrated in FIG. 12, it is possible to construct a roadmap 225 in which nodes are arranged focusing on a range with a high probability of the existence of an optimal path for each agent. As a result, in the constructed roadmap 225 for each agent, it is possible to narrow down the range in which nodes are arranged to a range suitable for path search of each agent. Therefore, even if nodes are densely arranged on the roadmap 225, an increase in the number of nodes can be suppressed. Thus, according to this embodiment, in the process of step S203, when solving the multi-agent path planning problem on the continuous state space, the possibility of finding a more optimal path for the agent can be increased, and the cost required for the search can be reduced.
[0182] Also, in this embodiment, the roadmap construction model 5 includes a third processing module 53 that processes information on an environment including obstacles. As a result, the trained roadmap construction model 5 can estimate the arrangement of nodes constituting the roadmap 225 for each agent in consideration of the situation of the environment including obstacles. Therefore, in the process of step S202, it is possible to construct a roadmap 225 suitable for the environment for each agent. As a result, in the process of step S203, the possibility of finding an optimal path for the agent on the continuous state space can be further increased, and the cost related to the search can be further reduced.
[0183] Also, in the present embodiment, the first processing module 51 and the second processing module 52 may be configured to further process the candidate states in the past time steps as the target agent information 41 and the other agent information 43. Thereby, the trained roadmap construction model 5 can estimate the arrangement of the nodes constituting the roadmap 225 for each agent while considering the state transition of each agent in time series. Therefore, in the process of step S202, a roadmap 225 suitable for the agents that reach the goal state from the start state through a plurality of time steps can be constructed. As a result, in the process of step S203, the possibility of finding an optimal path for the agents on the continuous state space can be further increased, and the cost related to the search can be further reduced.
[0184] Also, in the present embodiment, the first processing module 51 may be configured to further process the attribute information of the target agent as the target agent information 41. Thereby, the trained roadmap construction model 5 can estimate the arrangement of the nodes constituting the roadmap 225 for each agent while considering the attributes of each agent. Therefore, even when there are agents with different attributes, in the process of step S202, a suitable roadmap 225 can be constructed for each agent. As a result, even when agents with different attributes are mixed, in the process of step S203, the multi-agent path planning problem on the continuous state space can be appropriately solved. In one example, the attributes of the target agent may include at least any one of size, shape, maximum speed, and weight. Thereby, in the process of step S202, a suitable roadmap 225 can be constructed for each agent with respect to the constraints regarding the outer shape, speed, and weight.
[0185] Also, in the present embodiment, the second processing module 52 may be configured to further process the attribute information of other agents as the other agent information 43. Thereby, the trained roadmap construction model 5 can estimate the arrangement of nodes suitable for the target agent in consideration of the attributes of other agents. Therefore, even in an environment where agents with various attributes exist, in the process of step S202, a roadmap 225 suitable for each agent can be constructed. As a result, in the process of step S203, the possibility of finding an optimal path for the agent in the continuous state space can be further enhanced, and the cost related to the search can be further reduced. Note that, in one example, the attributes of other agents may include at least any one of size, shape, maximum speed, and weight. Thereby, even when agents with different sizes, speeds, and weights are mixed, in the process of step S203, the possibility of finding an optimal path for each agent can be further enhanced, and the cost related to the search can be further reduced.
[0186] Also, in the present embodiment, the first processing module 51 may be configured to further process the information of the direction flag indicating the transition direction as the target agent information 41. Thereby, in the process of step S102, based on the training flag 325, the training data can be used for machine learning to train the transitions in each direction evenly. As a result, a trained roadmap construction model 5 in which the direction of arranging nodes is less likely to be biased can be generated. Further, since the item of the direction flag is included in the target agent information 41, a trained roadmap construction model 5 that has acquired the ability to control the direction of arranging nodes by the direction flag can be generated.
[0187] Accordingly, in the process of step S202, by using this trained roadmap construction model 5, it is possible to prevent the direction in which nodes are arranged in the roadmap 225 of each agent from being biased. When nodes are arranged in a biased direction, the transition flag given as a direction flag can be used to control the direction in which the nodes are arranged, and the direction in which the nodes are arranged can be controlled. As a result, in the constructed roadmap 225, it is possible to prevent the selection width of state transitions from becoming narrow. As a result, in the process of step S203, it is possible to further increase the possibility of finding an optimal path for the agent on the continuous state space.
[0188] §4 Modifications Although the embodiments of the present invention have been described in detail above, the description so far is merely an exemplification of the present invention in every aspect. Needless to say, various improvements or modifications can be made without departing from the scope of the present invention. For example, the following changes are possible. In the following, the same reference numerals are used for the same components as in the above embodiment, and the description of the same points as in the above embodiment is omitted as appropriate. The following modifications can be combined as appropriate.
[0189] <4.1> In the above embodiment, the path planning system 2 is configured to execute a series of processes of acquiring the target information 221, constructing the roadmap 225 of each agent, and solving the path planning problem of each agent. However, the device configuration for executing each process does not have to be limited to such an example. As another example, the device for constructing the roadmap 225 of each agent and the device for searching for the path of each agent on the continuous state space based on the obtained roadmap 225 may be configured by one or a plurality of separate and independent computers.
[0190] FIG. 14 schematically shows an example of the configuration of the route planning system according to this modified example. In this modified example, the route planning system is composed of a roadmap construction device 201 and a route search device 202. The roadmap construction device 201 is at least one computer configured to construct a roadmap 225 for each agent using a trained roadmap construction model 5. The route search device 202 is at least one computer configured to search for the route of each agent on the constructed roadmap 225.
[0191] The hardware configurations of the roadmap construction device 201 and the route search device 202 may be the same as those of the hardware configuration of the above route planning system 2. The roadmap construction device 201 and the route search device 202 may be directly connected or may be connected via a network. In one example, the roadmap construction device 201 and the route search device 202 may perform data communication with each other through these connections. In another example, data may be exchanged between the roadmap construction device 201 and the route search device 202 via a storage medium or the like.
[0192] The roadmap construction program may be configured to include instructions up to constructing the roadmap 225 among the instructions included in the route planning program 82. The roadmap construction program may further include an instruction to output the constructed roadmap 225. The roadmap construction device 201 may be configured to include an information acquisition unit 211, a map construction unit 212, and an output unit 216 as software modules by executing this roadmap construction program. The output unit 216 may be configured to output the constructed roadmap 225 for each agent.
[0193] The route search program may be configured to include, among the instructions included in the route planning program 82, the instructions for performing route search until the search results are output. The route search program may further include an instruction to acquire the road map 225 constructed for each agent. The route search device 202 may be configured to include the map acquisition unit 218, the search unit 213, and the output unit 214 as software modules by executing this route search program. The map acquisition unit 218 may be configured to acquire the road map 225 constructed for each agent.
[0194] Note that, similar to the above embodiment, part or all of the software modules of the road map construction device 201 and the route search device 202 may be realized by one or more dedicated processors. Each module of the road map construction device 201 and the route search device 202 may be realized as a hardware module. Also, regarding the software configuration of each of the road map construction device 201 and the route search device 202, omission, substitution, and addition of software modules may be appropriately performed according to the embodiment.
[0195] In the road map construction device 201 according to this modification example, the control unit operates as the information acquisition unit 211 and executes the process of step S201. Next, the control unit operates as the map construction unit 212 and executes the process of step S202. Thereby, the road map 225 for each agent is constructed. Then, the control unit operates as the output unit 216 and outputs the constructed road map 225 for each agent. The output destination and the output method may be appropriately selected according to the embodiment. The output destination may include the route search device 202. The constructed road map 225 for each agent may be provided to the route search device 202 at an arbitrary timing and by an arbitrary method.
[0196] In the path search device 202 according to this modification example, the control unit operates as a map acquisition unit 218 and acquires the load map 225 constructed for each agent. The path for acquiring the load map 225 of each agent may not be particularly limited and may be appropriately selected according to the embodiment. Next, the control unit operates as a search unit 213 and executes the process of step S203 above. Then, the control unit operates as an output unit 214 and executes the process of step S204 above.
[0197] According to the load map construction device 201 according to this modification example, by using the trained load map construction model 5, it is possible to construct, for each agent, a load map 225 in which nodes are arranged focusing on a range with a high probability of the existence of an optimal path. Therefore, even if nodes are densely arranged on the load map 225, an increase in the number of nodes can be suppressed. Thus, in the path search device 202, when solving the multi-agent path planning problem on the continuous state space, it is possible to increase the possibility of finding a more optimal path for the agent and reduce the cost required for the search. The load map construction device 201 according to this modification example can enable the path search device 202 to enjoy such effects.
[0198] <4.2> The machine learning method of the load map construction model 5 may not be limited to the method described above. As long as it is possible to acquire the ability to estimate the candidate states of the target agent at the next time step from various information at the target time step, the machine learning method of the load map construction model 5 may not be particularly limited and may be appropriately selected according to the embodiment. As an example, adversarial learning may be adopted as the machine learning method of the load map construction model 5.
[0199] FIG. 15 schematically shows an example of a machine learning method of the roadmap construction model 5 by adversarial learning according to this modified example. In an example of FIG. 15, in order to execute adversarial learning, a discriminator 59 is prepared for the roadmap construction model 5. A machine learning model having calculation parameters is used for the discriminator 59. The discriminator 59 is configured to execute a process of receiving an input of a sample and discriminating whether the origin of the input sample is correct data or an estimation result by the roadmap construction model 5. If this discrimination process can be executed, the type of the machine learning model used for the discriminator 59 does not particularly need to be limited and may be appropriately selected according to the embodiment. In an example, a neural network may be adopted for the discriminator 59. The discriminator 59 may be further configured to receive an input of various information given to the roadmap construction model 5. That is, in adversarial learning, at least a part of the information (training data) given to the roadmap construction model 5 may also be input to the discriminator 59.
[0200] The adversarial learning is configured to include a first training step and a second training step. The first training step is configured by training the discriminator 59 to use the estimation result of the candidate state in the next time step by the roadmap construction model 5 and the correct data as input data and to discriminate the origin of the input data. The estimation result of the candidate state is obtained by giving the roadmap construction model 5 training data composed of information indicating the state 321 at the first time step of each data set 32 and executing the calculation process of the roadmap construction model 5. The correct data is composed of information indicating the state 323 at the second time step of each data set 32. The second training step is configured by training the roadmap construction model 5 so that the discrimination performance of the discriminator 59 decreases when the estimation result of the candidate state in the next time step by the roadmap construction model 5 is input to the discriminator 59. In FIG. 15, that the correct data is the origin is expressed as "true", and that the estimation result by the roadmap construction model 5 is the origin is expressed as "false". The expression method of the origin of each sample may be appropriately changed.
[0201] As an example of the training process, the control unit 11 prepares the roadmap construction model 5 and the discriminator 59 through the initial setting process. The control unit 11 prepares the training data and the correct answer data from the learning data 3. The composition of the training data may be the same as the training data given to the first encoder E1. That is, the training data given to the first processing module 51 may be composed of the goal state 31A of the learning target agent SA, the state 321A at the first time step of each data set 32A, the training flag 325A, the Cost-to-go feature for training, and the training attribute information 33A. The training data given to the second processing module 52 may be composed of the goal state 31B of the learning other agent SB, the state 321B at the first time step of each data set 32B, the Cost-to-go feature for training, and the training attribute information 33B. The training data given to the third processing module 53 may be composed of the training environment information 35. The correct answer data may be composed of the state 323A at the second time step of each data set 32A of the learning target agent SA.
[0202] The control unit 11 inputs each sample of the training data into the roadmap construction model 5 and executes the arithmetic processing of the roadmap construction model 5. That is, the control unit 11 gives each sample of the training data to each processing module 51-53 and executes the arithmetic processing of each processing module 51-53 to obtain the feature information from each processing module 51-53. The control unit 11 gives the obtained feature information and the random number vector to the estimation module 55 and executes the arithmetic processing of the estimation module 55. Through this arithmetic processing, the control unit 11 obtains each sample of the result of estimating the candidate state at the next time step from the roadmap construction model 5.
[0203] The control unit 11 can obtain, from the discriminator 59, the result of identifying the origin of each input sample by inputting each sample of the estimation result and the correct data to the discriminator 59 and executing the arithmetic processing of the discriminator 59. The control unit 11 adjusts the value of the arithmetic parameter of the discriminator 59 so that the error between this identification result and the true value (true / false) of the identification becomes small. The method of adjusting the value of the arithmetic parameter may be appropriately selected according to the type of the machine learning model. As an example, when a neural network is adopted in the configuration of the discriminator 59, the value of the arithmetic parameter of the discriminator 59 may be adjusted by the error backpropagation method. A known optimization method may be adopted as the method of adjusting the arithmetic parameter of the discriminator 59. Thereby, the discriminator 59 can be trained to acquire the ability to identify the origin of the input sample.
[0204] Also, the control unit 11 inputs each sample of the training data to the roadmap construction model 5 and executes the arithmetic processing of the roadmap construction model 5. By this arithmetic processing, the control unit 11 obtains, from the roadmap construction model 5, each sample of the result of estimating the candidate state at the next time step. The control unit 11 can obtain, from the discriminator 59, the result of identifying the origin of each input sample by inputting each sample of the estimation result to the discriminator 59 and executing the arithmetic processing of the discriminator 59. The control unit 11 calculates the error so that this identification result is incorrect (that is, the error becomes as small as misidentifying the origin of the input sample as the correct data), and adjusts the value of the arithmetic parameter of the roadmap construction model 5 so that the calculated error becomes small. The method of adjusting the value of the arithmetic parameter of the roadmap construction model 5 may be the same as that of the above embodiment. Thereby, the roadmap construction model 5 can be trained to acquire the ability to generate an estimation result that degrades the identification performance of the discriminator 59 (that is, misidentifies with the correct data).
[0205] In one example of FIG. 15, in the training step of the discriminator 59, the control unit 11 adjusts the values of the arithmetic parameters of the discriminator 59 while fixing the values of the arithmetic parameters of the load map construction model 5. On the other hand, in the training step of the load map construction model 5, the control unit 11 adjusts the values of the arithmetic parameters of the load map construction model 5 while fixing the values of the arithmetic parameters of the discriminator 59. The control unit 11 alternately and repeatedly executes the training processes of the discriminator 59 and the load map construction model 5. As a result, corresponding to the improvement in the discrimination performance of the discriminator 59, the load map construction model 5 acquires the ability to generate an estimation result that approximates the correct data (the state 323A at the second time step of each data set 32A), that is, the discrimination by the discriminator 59 fails. Therefore, also by the machine learning method according to this modification example, similarly to the above-described embodiment, a trained load map construction model 5 that has acquired the ability to estimate an appropriate candidate state at the next time step of the target agent from the candidate states at the target time steps of the target agent and other agents can be generated.
[0206] Note that the method of adversarial learning does not have to be limited to such an example. In another example, a gradient reversal layer may be arranged between the load map construction model 5 and the discriminator 59. The gradient reversal layer is configured to pass the value as it is during the forward propagation operation and reverse the value during the backpropagation. With this gradient reversal layer, the control unit 11 may execute the training process of the discriminator 59 and the training process of the load map construction model 5 in the adversarial learning at once.
[0207] <4.3> In the above-described embodiment, the Cost-to-go feature may be omitted from the target agent information 41 and the other agent information 43. Accordingly, the process of acquiring the Cost-to-go feature may be omitted. The information process related to the Cost-to-go feature may be omitted from the information process of the load map construction model 5. The Cost-to-go feature for training may be omitted from the learning data 3. The training process related to the Cost-to-go feature may be omitted from the machine learning process of the load map construction model 5.
[0208] In the above embodiment, information regarding the direction flag may be omitted from the target agent information 41. Accordingly, the process of obtaining the transition flag may be omitted. Information processing regarding the direction flag may be omitted from the information processing of the roadmap construction model 5. The training flag 325 may be omitted from the learning data 3. Training processing regarding the training flag 325 may be omitted from the machine learning processing of the roadmap construction model 5.
[0209] In the above embodiment, the attribute information of the target agent may be omitted from the target agent information 41. Accordingly, the process of obtaining the attribute information of the target agent may be omitted. Information processing regarding the attribute information of the target agent may be omitted from the information processing of the roadmap construction model 5. Regarding the learning target agent, the training attribute information 33 may be omitted from the learning data 3. Training processing regarding the training attribute information 33 of the learning target agent may be omitted from the machine learning processing of the roadmap construction model 5.
[0210] In the above embodiment, the attribute information of other agents may be omitted from the other agent information 43. Accordingly, the process of obtaining the attribute information of other agents may be omitted. Information processing regarding the attribute information of other agents may be omitted from the information processing of the roadmap construction model 5. Regarding the learning other agents, the training attribute information 33 may be omitted from the learning data 3. Training processing regarding the training attribute information 33 of the learning other agents may be omitted from the machine learning processing of the roadmap construction model 5.
[0211] In the above-described embodiment, information regarding candidate states in past time steps may be omitted from the target agent information 41 and the other agent information 43. Accordingly, the process of obtaining information on candidate states in past time steps may be omitted. Information processing regarding candidate states in past time steps may be omitted from the information processing of the roadmap construction model 5. Information on candidate states in past time steps may be omitted from the training data. Training processing regarding candidate states in past time steps may be omitted from the machine learning processing of the roadmap construction model 5.
[0212] In the above-described embodiment, the third processing module 53 may be omitted from the roadmap construction model 5. Accordingly, the process of obtaining the environmental information 45 may be omitted. Information processing regarding the third processing module 53 may be omitted from the information processing of the roadmap construction model 5. The training environmental information 35 may be omitted from the learning data 3. Training processing regarding the training environmental information 35 may be omitted from the machine learning processing of the roadmap construction model 5.
[0213] In the above-described embodiment, the input / output form of the roadmap construction model 5 may be appropriately changed. The roadmap construction model 5 may be configured to further receive input of information other than the target agent information 41, the other agent information 43, and the environmental information 45. The roadmap construction model 5 may be configured to further output information other than the estimation result of the candidate state of the target agent at the next time step.
[0214] §5 Examples To verify the effectiveness of the present invention, methods according to the following examples and comparative examples were configured. However, the present invention is not limited to the following examples.
[0215] (Problem setting) · Problem: Search for the shortest movement path from the start positions of 21 to 30 agents to the goal position · Environment: 2D plane ([0,1] 2 ) · Continuous state space: A continuous space indicating position (state = position) · Agent - Shape: Circular - Size (radius): The size of the environment × 1 / 32 × magnification factor, where the magnification factor is randomly determined from {1, 1.25, 1.5} - Maximum speed: The size of the environment × 1 / 64 × magnification factor, where the magnification factor is randomly determined from {1, 1.25, 1.5} - Number: Randomly determined between 21 and 30 - Start position / Goal position: Randomly determined · Obstacle - Shape: Circular - Size (radius): The size of the environment × magnification factor, where the magnification factor is randomly determined between 0.05 and 0.08 - Position: Randomly determined
[0216] First, based on the above problem settings, 1100 instances of the path planning problem were created. Among them, 1000 instances were used as the training data for the roadmap construction model according to the examples. By constructing a roadmap through random sampling and performing prioritized planning (reference: David Silver, "Cooperative Pathfinding", Proceedings of the Artificial Intelligence for Interactive Digital Entertainment Conference (AIIDE) (2005), 117 - 122.) using the obtained roadmap, the correct paths of each agent in the 1000 instances were obtained (that is, the correct paths of each agent in each instance were obtained by the same method as in the comparative example described later). On the other hand, the remaining 100 instances were used as the evaluation data for the examples and the comparative example respectively.
[0217] <Example> In the example, a roadmap construction model having the same configuration as the above embodiment (Figure 2A) was generated. The conditions of the roadmap construction model and the input information (target agent information, other agent information, environmental information) are shown below. Learning data was obtained from the above 1000 types of instances, and machine learning was performed by the method shown in Figure 7 to generate a trained roadmap construction model. The Adam (learning rate = 0.001) was adopted as the optimization algorithm for machine learning, the batch size was set to 50, and the number of epochs was set to 1000. The trained model with the minimum loss for the evaluation data was used for evaluation.
[0218] (Conditions of the roadmap construction model) · Encoder (the first processing module - the third processing module) - Configuration: fully connected layer -> batch normalization -> ReLU -> fully connected layer -> batch normalization -> ReLU - Dimension of the output (feature information): 64 dimensions · Estimation module - Configuration: fully connected layer -> batch normalization -> ReLU -> fully connected layer -> batch normalization -> ReLU - Dimension of the output (estimation result): 3 dimensions, the distance (1 dimension) and direction (represented by a unit vector, 2 dimensions) from the candidate position of the target time step to the candidate position of the next time step (Conditions of the input information) · Candidate state and goal state at the target time step: The distance and direction from the candidate position at the target time step to the goal position are represented by 2D unit vectors · Candidate state at the past time step: The distance and direction from the candidate position one time step before to the candidate position at the target time step are represented by 2D unit vectors · Information about obstacles: After representing the environment as a 160×160 grid, a binary feature vector representing the presence or absence of obstacles in the 19×19 cells around the candidate position at the target time step in 0 / 1 · Cost-to-go feature: A binary feature vector representing whether each point in the 19×19 cells around the candidate position at the target time step approaches the goal state in 0 / 1 · Direction flag: The direction to the candidate position in the next time step is expressed as a three - dimensional vector of [1, 0, 0], [0, 1, 0], and [0, 0, 1] depending on whether it is "left (more than 30 degrees to the left as seen from the goal direction)", "front (within 30 degrees to the left and right of the goal direction)", or "right (more than 30 degrees to the right as seen from the goal direction)". · Attribute information is omitted · Other agents: In the target time step, the five agents from the agent located fifth closest to the target agent among the agents located closest to the target agent · Processing method of the second processing module: The method shown in Figure 2B was adopted to integrate the information of each agent
[0219] Next, for each of the above 100 instances (evaluation data), a roadmap for each agent was constructed using the generated trained roadmap construction model. When constructing the roadmap, at time step t, with the probability of the following formula 2, without using the trained roadmap construction model, a node (candidate position in the next time step) was placed at a position randomly moved from the candidate position in the target time step.
[0220]
Number
[0221] Then, for each instance, on the roadmap obtained for each agent, the shortest path for each agent was searched. The path search method adopted prioritized planning.
[0222] <Comparative example> On the other hand, in the comparative example, for each of the above 100 types of instances (evaluation data), a roadmap for each agent was constructed by randomly arranging 5,000 nodes for each agent. Then, for each instance, the shortest path of each agent was searched on the obtained roadmap. The same method (prioritized planning) as in the example was adopted for the path search method.
[0223] <Result> FIG. 16A shows the roadmap constructed for a certain agent by the method according to the example and the result of path search. FIG. 16B shows the roadmap constructed for a certain agent by the method according to the comparative example and the result of path search. FIG. 17 shows the path planning results of each agent obtained by the method according to the example for a certain instance (evaluation data). Table 1 below shows the performance results of the methods according to the example and the comparative example for 100 types of instances (evaluation data).
[0224] [Table 1] Note that for the path planning according to the example and the comparative example, a commercially available PC equipped with a CPU (Intel Core i7-7800) and a RAM (32 GB) (without using a GPU) was used. The processing time of each part for the evaluation data was obtained by measuring the time taken from the start to the end of the processing of each part using the said PC.
[0225] As shown by an example in Fig. 17, by the path planning method according to the embodiment, the path planning of each agent could be appropriately performed. Regarding this obtained path planning, as shown in Table 1, in the embodiment, in terms of the success rate and whether the obtained path is optimal, the performance was almost equivalent to that of the comparative example. On the other hand, as shown by an example in Figs. 16A and 16B, in the embodiment, compared with the comparative example, the range where the nodes of the roadmap are arranged could be narrowed down to an appropriate range. As a result, in the embodiment, compared with the comparative example, the time required for path search could be significantly reduced. From these results, according to the present invention, when solving the multi-agent path planning problem on a continuous state space, it was verified that it is possible to increase the possibility of finding a more optimal path for the agent and to reduce the cost required for search.
Explanation of Signs
[0226] 1... Model generation device, 11... Control unit, 12... Storage unit, 13... Communication interface, 14... Input device, 15... Output device, 16... Drive, 81... Model generation program, 91... Storage medium, 111... Data acquisition unit, 112... Learning processing unit, 113... Saving processing unit, 125... Learning result data, 2... Path planning system, 21... Control unit, 22... Storage unit, 23... Communication interface, 24... Input device, 25... Output device, 26... Drive, 82... Path planning program, 92... Storage medium, 211... Information acquisition unit, 212... Map construction unit, 213... Search unit, 214... Output unit, 221... Target information, 225... Roadmap, 3... Learning data, 31... Goal state, 32... Data set, 321... State at the first time step, 323... State at the second time step, 325… Training flag, 33… Attribute information for training, 35… Environmental information for training, 41… Target agent information, 43… Other agent information, 45… Environmental information, 5… Roadmap construction model, 51… First processing module, 52… Second processing module, 53… Third processing module, 55… Estimation module
Claims
1. An information acquisition unit configured to acquire target information including start states and goal states in the continuous state spaces of a plurality of agents; A map construction unit configured to construct a roadmap for each agent from the acquired target information using a trained roadmap construction model; A search unit configured to search for a path of each agent from the start state to the goal state on the roadmap constructed for each agent; A path planning system comprising: The roadmap construction model includes: A first processing module configured to generate first feature information from target agent information including the goal state of a target agent and candidate states at a target time step; A second processing module configured to generate second feature information from other agent information including the goal states of agents other than the target agent and candidate states at the target time step; and An estimation module configured to estimate one or more candidate states at the next time step of the target time step of the target agent from the generated first feature information and second feature information; The trained roadmap construction model is generated by machine learning using learning data obtained from correct paths of a plurality of learning agents. Constructing the roadmap for each agent includes: Treating any one of the plurality of agents as the target agent; Treating at least some of the remaining agents among the plurality of agents as the other agents; Designating the start state of any one of the agents indicated by the acquired target information as a candidate state at the first target time step of the target agent, estimating one or more candidate states at the next time step by the trained roadmap construction model, and Repeating the estimation of candidate states at the next time step by the trained roadmap construction model by designating each of the one or more candidate states at the estimated next time step as a candidate state at a new target time step until the goal state or a state in the vicinity thereof of any one of the agents is included in the one or more candidate states at the estimated next time step. A path planning system configured by specifying and executing the process for each of the plurality of agents individually for the target agent. Path planning system.
2. The roadmap construction model further includes a third processing module configured to generate third feature information from environmental information including information on obstacles. The estimation module is configured to estimate one or more candidate states at the next time step from the generated first feature information, second feature information, and third feature information. The target information to be acquired is further configured to include information on the obstacles existing in the continuous state space. Using the trained roadmap construction model includes constructing the environmental information from the information included in the acquired target information and providing the constructed environmental information to the third processing module. The path planning system according to claim 1.
3. The target agent information is further configured to include candidate states of the target agent at time steps earlier than the target time step. The other agent information is further configured to include candidate states of the other agents at time steps earlier than the target time step. The path planning system according to claim 1 or 2.
4. The target agent information is further configured to include the attributes of the target agent. The path planning system according to any one of claims 1 to 3.
5. The attributes of the target agent include at least any one of size, shape, maximum speed, and weight. The path planning system according to claim 4.
6. The other agent information is further configured to include the attributes of the other agents. The path planning system according to any one of claims 1 to 5.
7. The attributes of the other agents include at least any one of size, shape, maximum speed, and weight. The path planning system according to claim 6.
8. The target agent information is further configured to include a direction flag indicating the direction in which the target agent transitions in the continuous state space. The path planning system according to any one of claims 1 to 7.
9. Each of the plurality of agents is a moving body configured to move autonomously. The path planning system according to any one of claims 1 to 8.
10. A computer performs the steps of: obtaining target information including start states and goal states in the continuous state spaces of a plurality of agents; constructing a roadmap for each agent from the obtained target information using a trained roadmap construction model; searching for a path for each agent from the start state to the goal state on the roadmap constructed for each agent; A path planning method for performing the above steps, wherein the roadmap construction model includes: a first processing module configured to generate first feature information from target agent information including the goal state of a target agent and candidate states at a target time step; a second processing module configured to generate second feature information from other agent information including the goal states of agents other than the target agent and candidate states at the target time step; and an estimation module configured to estimate one or more candidate states at the next time step of the target time step of the target agent from the generated first feature information and the second feature information; The trained roadmap construction model is generated by machine learning using learning data obtained from correct paths of a plurality of learning agents. Constructing the roadmap for each agent includes: treating any one of the plurality of agents as the target agent, treating at least some of the remaining agents among the plurality of agents as the other agents, designating the start state of the any agent indicated by the obtained target information as a candidate state at the first target time step of the target agent, estimating one or more candidate states at the next time step by the trained roadmap construction model, and repeating the estimation of candidate states at the next time step by the trained roadmap construction model by designating each of the one or more candidate states at the estimated next time step as a candidate state at a new target time step until the goal state or a state in the vicinity thereof of the any agent is included in the one or more candidate states at the estimated next time step. A path planning method configured by specifying and executing the process for each of the plurality of agents individually for the target agent. Path planning method.
11. An information acquisition unit configured to acquire target information including a start state and a goal state in the continuous state space of each of a plurality of agents, A map construction unit configured to construct a roadmap for each agent from the acquired target information using a trained roadmap construction model, A roadmap construction apparatus comprising: The roadmap construction model includes: A first processing module configured to generate first feature information from target agent information including a goal state of a target agent and candidate states at a target time step, A second processing module configured to generate second feature information from other agent information including a goal state of an agent other than the target agent and candidate states at the target time step, and An estimation module configured to estimate one or more candidate states at the next time step of the target time step of the target agent from the generated first feature information and the second feature information, Comprising The trained roadmap construction model is generated by machine learning using learning data obtained from the correct paths of a plurality of learning agents, Constructing the roadmap for each agent includes: Treating any one of the plurality of agents as the target agent, Treating at least a part of the remaining agents among the plurality of agents as the other agents, Designating the start state of any one of the agents indicated by the acquired target information as a candidate state at the first target time step of the target agent, and estimating one or more candidate states at the next time step by the trained roadmap construction model, and Repeating the estimation of candidate states at the next time step by the trained roadmap construction model by designating each of the one or more candidate states at the estimated next time step as a candidate state at a new target time step until the goal state or a state in the vicinity thereof of any one of the agents is included in the one or more candidate states at the estimated next time step. A roadmap construction device configured by executing the process by individually designating each of the plurality of agents to the target agent. Roadmap construction device.
12. A data acquisition unit configured to acquire learning data generated from the correct paths of a plurality of learning agents, A learning processing unit configured to perform machine learning of a roadmap construction model using the acquired learning data, A model generation device comprising: The roadmap construction model includes: A first processing module configured to generate first feature information from target agent information including the goal state of the target agent in the continuous state space and candidate states in the target time step; A second processing module configured to generate second feature information from other agent information including the goal state of other agents other than the target agent and candidate states in the target time step, and An estimation module configured to estimate one or more candidate states in the next time step of the target time step of the target agent from the generated first feature information and second feature information. Comprising: The learning data includes the goal state in the correct path of each learning agent and a plurality of data sets, Each of the plurality of data sets is composed of a combination of the state of each learning agent at the first time step and the state at the second time step, The second time step is the next time step of the first time step, The machine learning of the roadmap construction model includes: Treating any one of the plurality of learning agents as the target agent, Treating at least a part of the remaining learning agents among the plurality of learning agents as the other agents, and For each of the data sets, the state of any of the learning agents at the first time step is provided to the first processing module as a candidate state of the target agent at the target time step, and the states of at least some of the remaining learning agents at the first time step are provided to the second processing module as candidate states of the other agents at the target time step, so that the candidate state of the target agent at the next time step estimated by the estimation module conforms to the state of any of the learning agents at the second time step, training the roadmap construction model; Composed of; Model generation device.
13. The target agent information is configured to further include a direction flag indicating the direction in which the target agent transitions in the continuous state space, Each of the data sets is configured to further include a training flag indicating the direction from the state at the first time step to the state at the second time step in the continuous state space, The machine learning of the roadmap construction model includes, for each of the data sets, providing the training flag of any of the learning agents to the first processing module as the direction flag of the target agent when estimating the candidate state of the target agent at the next time step. The model generation device according to claim 12.
14. A computer, obtaining learning data generated from the correct paths of a plurality of learning agents; performing machine learning of a roadmap construction model using the obtained learning data; A model generation method for executing, The roadmap construction model includes a first processing module configured to generate first feature information from target agent information including a goal state of a target agent in a continuous state space and a candidate state at a target time step; a second processing module configured to generate second feature information from other agent information including a goal state of an agent other than the target agent and a candidate state at the target time step, and An estimation module configured to estimate one or more candidate states of the target agent at the next time step after the target time step from the generated first feature information and the second feature information. Comprising The training data includes a goal state and a plurality of data sets in the correct path of each training agent. Each of the plurality of data sets is composed of a combination of the state of each training agent at the first time step and the state at the second time step. The second time step is the next time step after the first time step. The machine learning of the roadmap construction model Treat any one of the plurality of training agents as the target agent. Treat at least a part of the remaining training agents among the plurality of training agents as the other agents, and For each data set, the state of any one of the training agents at the first time step is given to the first processing module as a candidate state of the target agent at the target time step, and the state of at least a part of the remaining training agents at the first time step is given to the second processing module as a candidate state of the other agents at the target time step, so that the candidate state of the target agent at the next time step estimated by the estimation module conforms to the state of any one of the training agents at the second time step, training the roadmap construction model. Composed by Model generation method.
Citation Information
Patent Citations
Path planning method based on multi-agent enhanced learning
CN109059931A
Vehicle path planning method based on reinforcement learning
CN111415048A
System and method for adaptive route planning
JP2007531110A
Method and system for controlling a vehicle
US20190361452A1
Apparatus, method and article to facilitate motion planning in an environment having dynamic objects
WO2020117958A1