Mobile control device, mobile body, learning device, learning method, and program
The movement control device and learning method address overfitting by using a simulator system with varied obstacle densities to learn adaptive movement paths, ensuring accurate control in environments with different congestion levels.
Patent Information
- Application Number
- JP2021162069
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-09-30
AI Technical Summary
Conventional methods for determining movement paths in complex environments with varying congestion levels often result in overfitting, leading to inappropriate path determination in environments with fewer moving objects.
A movement control device and learning method that utilize a simulator system with multiple environments of varying obstacle densities to learn a policy that maximizes cumulative rewards, allowing for adaptive movement paths based on congestion levels.
Enables flexible and appropriate movement control in environments with different congestion levels by preventing overfitting and ensuring accurate path determination across varying densities of obstacles.
Smart Images

Figure 0007716951000003 
Figure 0007716951000004 
Figure 0007716951000005
Abstract
Description
Technical Field
[0001] The present invention relates to a movement control device, a moving body, a learning device, a learning method, and a program.
Background Art
[0002] In recent years, attempts have been made to determine the movement route of a moving body by AI (artificial intelligence) generated by machine learning. In addition, research and practical application of reinforcement learning, which determines actions based on observed values and calculates rewards based on feedback obtained from the real environment or simulator to optimize model parameters, have also been advanced.
[0003] In relation to this, in order to take safe and secure avoidance actions against human movement, an invention of a route determination device that determines the route when a self-driving robot moves to a destination under the condition that traffic participants including pedestrians exist in the traffic environment to the destination is disclosed (see Patent Document 1). This route determination device includes a predicted route determination unit that determines a predicted route, which is a predicted value of the robot's route, so as to avoid interference between the robot and traffic participants using a predetermined prediction algorithm, and a route determination unit that determines the robot's route using a predetermined control algorithm so that an objective function including the distance to the traffic participant closest to the robot and the robot's speed as independent variables becomes the maximum value when it is assumed that the robot moves from the current position along the predicted route.
[0004] In addition, Non-Patent Document 1 describes multi-stage training in which reinforcement learning is performed while gradually increasing the number of agents for decentralized motion planning in a high-density and dynamic environment.
[0005] In addition, Non-Patent Document 2 describes a multi-scenario multi-stage training framework as a method for learning a policy that can appropriately determine the operation of a moving body.
Prior Art Documents
Patent Documents
[0006] [Patent Document 1] International Publication No. 2020 / 136977 [Non-Patent Document]
[0007] [Non-Patent Document 1] Samaneh Hosseini Semnani, Hugh Liu, Michael Everett, Anton de Ruiter, and Jonathan P How. Multi-agent motion planning for dense and dynamic environments via deep reinforcement learning. IEEE Robotics and Automation Letters, 5(2):3221-3226, 2020. [Non-Patent Document 2] P. Long, T. Fan, X. Liao, W. Liu, H. Zhang, and J. Pan. Towards optimally decentralized multi-robot collision avoidance via deep reinforcement learning. In 2018 IEEE International Conference on Robotics and Automation (ICRA). [Summary of the Invention] [Problems to be Solved by the Invention]
[0008] However, in the conventional method, as a result of learning an environment with a larger number of moving objects in order to cope with a complex environment, overfitting occurs, and an inappropriate movement path may be determined in an environment with a small number of existing moving objects. Thus, in the prior art, it may not be possible to appropriately determine a movement path according to the degree of congestion of the environment.
[0009] The present invention has been made in consideration of such circumstances, and one of its objectives is to provide a movement control device, a moving body, a learning device, a learning method, and a program that can determine an appropriate movement mode according to the degree of environmental congestion.
Means for Solving the Problems
[0010] The movement control device, the moving body, the learning device, the learning method, and the program according to this invention adopt the following configurations.
[0011] (1): The movement control device according to one aspect of this invention includes a route determination unit that determines the route of the moving body according to the number of obstacles existing around the moving body, and a control unit that moves the moving body along the route determined by the route determination unit.
[0012] (2): In the aspect of (1) above, the route determination unit determines the route of the moving body based on the policy of the operation learned by the simulator and the learning unit, and the policy of the operation is such that the simulator simultaneously executes the simulation of the operations of the moving body and the obstacles in a plurality of environments with different numbers of obstacles, and the learning unit updates the policy so that the reward obtained by applying the reward function to the processing result of the simulator is maximized.
[0013] (3): In the aspect of (2) above, the policy of the operation is learned based on the processing results of a plurality of the simulators, the number of obstacles in the environment is different for each of the plurality of simulators, and the learning unit updates the policy of the operation so that the cumulative sum of each reward obtained by applying the reward function to each processing result of the plurality of simulators is maximized.
[0014] (4): The mobile body according to one aspect of the present invention includes any of the above-described mobile body control devices, a working unit for providing a predetermined service to a user, and a driving device for moving the mobile body. The driving device drives the mobile body to move in a movement mode determined by the mobile body control device.
[0015] (5): The learning device according to one aspect of the present invention is a simulator that executes a simulation of the operation of a mobile body, and includes a plurality of the simulators in which the number of the existing mobile bodies or obstacles is different for each simulator, and a learning unit that learns the policy of the operation so that the cumulative sum of each reward obtained by applying a reward function to each processing result of the plurality of simulators is maximized.
[0016] (6): In the aspect of (5) above, the plurality of simulators are executed by separate processors associated with each of them.
[0017] (7): In the aspect of (5) or (6) above, different maximum numbers of the mobile bodies or the obstacles are set for the plurality of simulators, and the plurality of simulators execute the simulation while gradually increasing the number of the mobile bodies or the obstacles from a specified minimum number to their respective maximum numbers.
[0018] (8): In any of the aspects of (5) to (7) above, the plurality of simulators execute the simulation in parallel for a plurality of environments in which the number of the mobile bodies or the obstacles is the same in each stage of the simulation.
[0019] (9): In any of the aspects of (5) to (8) above, the reward function includes at least one of the degree of reaching the target of the mobile body, the number of collisions of the mobile body, and the moving speed of the mobile body as a variable.
[0020] (10): In any of the aspects (5) to (9) above, the reward function includes, as an independent variable, the change in the movement vector of the moving body or the obstacle existing around the autonomous mobile body.
[0021] (11): The learning method according to one aspect of the present invention is such that a computer executes a simulation of the operation of a moving body by a plurality of the simulators in which the number of existing moving bodies or obstacles is different for each simulator, and learns the policy of the operation so that the cumulative sum of each reward obtained by applying a reward function to each processing result of the plurality of the simulators is maximized.
[0022] (12): The program according to one aspect of the present invention causes a computer to execute a simulation of the operation of a moving body by a plurality of the simulators in which the number of existing moving bodies or obstacles is different for each simulator, and causes the computer to learn the policy of the operation so that the cumulative sum of each reward obtained by applying a reward function to each processing result of the plurality of the simulators is maximized.
Advantages of the Invention
[0023] (1) - (4) According to the above, by providing a route determination unit that determines the route of the moving body according to the number of obstacles existing around the moving body, and a control unit that moves the moving body along the route determined by the route determination unit, it is possible to determine an appropriate moving mode according to the degree of congestion of the environment.
[0024] Also, according to (5) - (12), by providing a simulator that executes a simulation of the operation of a moving body, a plurality of the simulators in which the number of existing moving bodies or obstacles is different for each simulator, and a learning unit that learns the policy of the operation so that the cumulative sum of each reward obtained by applying a reward function to each processing result of the plurality of the simulators is maximized, it is possible to determine an appropriate moving mode according to the degree of congestion of the environment.
Brief Description of the Drawings
[0025]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Mode for Carrying Out the Invention
[0026] Hereinafter, embodiments of the movement control device, the moving body, the learning device, the learning method, and the program of the present invention will be described with reference to the drawings.
[0027] <First Embodiment> FIG. 1 is a schematic diagram of the configuration of the movement control system 1 according to the embodiment. The movement control system 1 includes a learning device 100 and a moving body 200. The learning device 100 is realized by one or more processors. The learning device 100 is a device that determines actions by computer simulation for a plurality of moving bodies, derives or obtains rewards based on state changes and the like caused by those actions, and learns actions (operations) that maximize the rewards. An operation is, for example, movement within a simulation space. Although operations other than movement may be the learning targets, in the following description, an operation means movement. The simulator that determines the movement may be executed in a device different from the learning device 100, but in the following description, the simulator is assumed to be executed by the learning device 100. The learning device 100 stores in advance environment information that is a premise for simulation, such as map information. The learning result of the learning device 100 is mounted on the moving body 200 as a policy PL.
[0028] [Learning device] FIG. 2 is a diagram showing a configuration example of the learning device 100 according to the embodiment. The learning device 100 includes, for example, a learning unit 110, a plurality of simulators 120, and an experience accumulation unit 130. These components are realized, for example, by a hardware processor such as a CPU (Central Processing Unit) executing a program (software). Some or all of these components may be realized by hardware (including a circuit unit; circuitry) such as LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), or GPU (Graphics Processing Unit), or may be realized by cooperation between software and hardware. The program may be stored in advance in a storage device (a storage device including a non-transitory storage medium) such as an HDD (Hard Disk Drive), a flash memory, or an SSD (Solid State Drive), or may be stored in a removable storage medium (a non-transitory storage medium) such as a DVD or a CD-ROM, and may be installed by mounting the storage medium on a drive device.
[0029] The learning unit 110 updates the policy according to various reinforcement learning algorithms based on the evaluation information indicating the result of the evaluation by the experience accumulation unit 130 of the state changes generated by the plurality of simulators 120. The learning unit 110 repeatedly executes until learning is completed the operation of outputting the updated policy to the plurality of simulators 120. The policy is, for example, a neural network having parameters (hereinafter also simply referred to as a "network"), which probabilistically outputs actions (operations) that an agent can take in response to the input of environmental information. Here, the agent is a moving body existing in the simulation space (environment), and is the moving body that is the target of learning the operation. The agent is an example of a "self-moving body". The environmental information is information indicating the state of the environment. The policy may be a rule-based function having parameters. The learning unit 110 updates the policy by updating the parameters based on the evaluation information. The learning unit 110 supplies the updated parameters to each simulator 120.
[0030] The simulator 120 inputs the operation target and the current state (the initial state if it is immediately after the start of the simulation) into the policy, and derives the state change that is the result of the operations of the self-agent and other agents. The policy may be, for example, a DNN (Deep Neural Network), but may also be a policy of other forms such as a rule-based policy. The policy derives the occurrence probability for each of the assumed multiple types of operations. For example, in a simple example, assuming that the assumed plane extends vertically and horizontally, the results such as moving right 80%, moving left 10%, moving up 10%, and moving down 0% are output. The simulator 120 applies a random number to this result, and if the random number value is 0% or more and less than 80%, it moves right, if the random number value is 80% or more and less than 90%, it moves left, and if the random number value is 90% or more, it moves up, and thus derives the state change of the agent.
[0031] A plurality of simulators 120 use the policy (network) updated by the learning unit 110 to execute simulations for environments with different numbers of agents and in which a plurality of agents exist, thereby determining the actions of the agents in each environment. Here, the determination of the action means deriving the above-described state change for the agent. In this embodiment, four simulators are assumed as the plurality of simulators 120. For example, in this embodiment, the first to fourth simulators 120A to 120D determine the movement of 2 agents, 4 agents, 8 agents, and 10 agents, respectively. Note that the environment may include moving bodies that do not depend on policies other than the agents. For example, the environment may include, in addition to agents that move based on a policy, stationary moving bodies, moving bodies that operate according to an operation model different from the policy, and the like.
[0032] Specifically, each simulator 120 updates the policy (network) with the parameters supplied from the learning unit 110, inputs the current state obtained from the simulation result of the previous time (the previous sampling period) into the updated network, and applies a random number to the output result to determine the action of each agent this time (the current sampling period). By inputting the determined action into the environment EV by each simulator 120, the updated state and reward are generated by the environment EV. The reward is generated when the environment EV inputs the determined action into the reward function. Each simulator 120 supplies the experience accumulation unit 130 with experience information based on the actions determined for each agent. For example, the experience information includes information on the actions determined for the agent, the state before the action and the state after the action, and the reward obtained by that action.
[0033] The experience accumulation unit 130 accumulates the experience information supplied from each simulator 120, and samples the experience information with high priority from the accumulated experience information and supplies it to the learning unit 110. The priority is a priority based on the height of the learning effect in the learning of the network NW, and is determined by, for example, a TD (Temporal Difference) error. Note that the priority may be appropriately updated based on the learning result of the learning unit 110.
[0034] Based on the experience information supplied from the experience accumulation unit 130, the learning unit 110 updates the parameters of the network NW so that the reward obtained by the movement of each agent is maximized. The learning unit 110 supplies the updated parameters to each simulator 120. Each simulator 120 updates the network NW with the parameters supplied from the learning unit 110.
[0035] The learning unit 110 may use any of various reinforcement learning algorithms. By repeatedly executing such parameter updates, the learning unit 110 learns the appropriate movement of the agents in an environment where a plurality of agents exist. The network learned in this way is supplied to the mobile body 200 as a policy.
[0036] Note that the reward function used when the environment EV calculates the reward may be any function as long as it gives a larger reward as the agent makes a more appropriate movement. For example, as shown in Equation (1), a reward function R1 given when the own agent arrives at the destination, a reward function R2 given when the own agent achieves smooth movement, a reward function R3 that becomes smaller when the own agent affects the movement vector of another agent, and a reward function R4 in which the distance to be maintained when the own agent approaches another agent is made variable according to the direction in which the other agent is facing may be included in the function R as the reward function. Further, the reward function R may be a function including at least one of R1, R2, R3, and R4.
[0037]
Number
[0038] For example, the reward function R1 becomes a positive fixed value when the destination is reached, and becomes a value proportional to the distance change to the destination (positive if the distance change is in the decreasing direction and negative if it is in the increasing direction) when the destination has not been reached. The reward function R1 is an example of the "first function".
[0039] For example, the reward function R2 is a function such that the smaller the third derivative of the position of the agent in the two-dimensional plane, that is, the jerk, the larger the value. The reward function R2 is an example of the "second function".
[0040] For example, the reward function R3 is a function that returns a low evaluation value when the self-agent enters a predetermined area. According to such a reward function R3, for example, a low evaluation can be given to an action in which the self-agent passes through the area (predetermined area) in front of other agents, and a not-too-low evaluation can be given to an action in which the self-agent passes through the side or the back. The reward function R3 is an example of the "third function".
[0041] Figure 3 is a diagram for explaining the reward function R4. Figure 3 shows an environment in which people P1, P4, and P5 and robots R2, R3, and R5 are mixed as an example of a simulation environment. In Figure 3, the points D1 to D5 are the destination points of each moving object. Specifically, the point D1 is the destination point of the person P1, the point D2 is the destination point of the robot R2, the point D3 is the destination point of the robot R3, the point D4 is the destination point of the person P4, and the point D5 is the destination point of the person P5.
[0042] Here, taking the robot R5 as the target robot, as the reward function R4 for learning a movement method that does not obstruct the movement of people by the target robot, it can be defined, for example, as in the following formula (2).
[0043]
Number
[0044] (2) In the formula, R4 is a reward function for learning a movement method that does not inhibit human movement, and is a function that gives a greater reward to a movement that does not inhibit human movement. i is the identification number of a moving object such as a person or a robot existing in the environment, and N is the maximum number thereof. Also, a i represents an action (hereinafter referred to as "the first action") determined by the state of the environment including the target robot R5 for each moving object, and b i represents an action (hereinafter referred to as "the second action") determined by the state of the environment excluding (ignoring) the target robot R5. w is a coefficient that takes the difference between the first action and the second action for each moving object and converts a value corresponding to the sum thereof into a negative reward value as a penalty. That is, formula (2) calculates a reward that becomes smaller as the difference between the first action and the second action becomes larger. According to such a reward function, for example, the target robot R5 can learn a movement method such that its own movement does not affect the movement of other moving objects. The reward function R4 is an example of "the fourth function".
[0045] The learning operation of the network described above explains the operation when each simulator 120 performs a simulation with a predetermined number of agents. The learning device 100 of the present embodiment is configured to learn the operations of moving objects in a plurality of environments with different numbers of agents in parallel by executing the above-described reinforcement learning while gradually increasing the number of agents in the simulation. This method of learning the policy of the environment with the final number of agents while gradually increasing the number of agents (hereinafter referred to as "stepwise reinforcement learning") is known as one of the methods for improving the accuracy of reinforcement learning (see, for example, Non-Patent Document 1).
[0046] Figure 4 is a diagram showing an example of the effect of stepwise reinforcement learning. In Figure 4, the horizontal axis represents the progress of learning at each step, and the vertical axis represents the accuracy of learning. According to Figure 4, it can be seen that learning while gradually increasing the number of agents step by step as 2, 4, 8, 10 enables learning of actions that obtain higher rewards compared to starting learning with 10 agents from the beginning.
[0047] However, in an environment where there are multiple agents, the policy learned with 10 agents does not necessarily determine appropriate movements in all environments. This is because, in learning movements, although determining a destination that does not contact other moving objects or obstacles (i.e., learning as an action that can obtain a high reward) is prioritized, depending on the state of the environment (e.g., the density of agents existing in the environment), the priority of other matters may become higher. That is, the learning result of movements in an environment with a larger number of agents may be overlearning when determining movements in an environment with a smaller number of agents.
[0048] Figures 5 and 6 are diagrams showing an example of overlearning of a policy. Figure 5 shows an example of movement based on a policy learned with 2 agents, and Figure 6 shows an example of movement based on a policy learned with 10 agents. Both Figures 5 and 6 show the movement route determined for one agent A to start from the starting point B, avoid the obstacle C, and arrive at the destination D. From Figures 5 and 6, it can be seen that in the policy learned in the 2-agent environment, agent A starts the avoidance action of the obstacle C promptly after leaving the starting point B, while in the policy learned in the 10-agent environment, agent A starts the avoidance action at a position closer to the obstacle C.
[0049] Such differences in avoidance behavior can be considered to be the result of learning, for example, in an environment with a larger number of agents, it is easier to interfere with other agents, so in order not to interfere with other agents, the avoidance behavior starts at a position closer to the obstacle C. Also, for example, such differences in avoidance behavior can be considered to be the result of learning that in an environment with a smaller number of agents, it is less likely to interfere with other agents, so in order to improve the safety of movement, the direction of movement is changed more gently.
[0050] In any case, in the conventional stepwise reinforcement learning, when learning is sequentially performed individually from an environment with a small number of agents to an environment with a large number of agents, the learning result in the last learning environment dominates in the determination of the movement pattern by the policy. Therefore, even if the movement in an environment with a large number of agents can be accurately learned, the policy generated by the learning will be optimized for an environment with a large number of agents, and there may be cases where appropriate actions cannot be determined in environments with different numbers of agents. Therefore, in the learning device 100 of the present embodiment, a configuration is adopted in which a plurality of simulators 120 are operated in parallel to learn environments with different numbers of agents in parallel.
[0051] FIG. 7 is a diagram showing a state in which the learning device 100 learns the operation for environments with different numbers of agents using a plurality of simulators 120. As described above, in the learning device 100 of the present embodiment, the simulators 120A, 120B, 120C, and 120D determine the operation of each agent for environments with 2 agents, 4 agents, 8 agents, and 10 agents, respectively. Specifically, each simulator 120 starts the simulation with the specified minimum number of agents, and executes the simulation while gradually increasing the number of agents up to the maximum number of each simulator 120.
[0052] For example, in this embodiment, since the maximum number of agents in simulator 120B is 4, the simulation first starts with 2 agents. When the learning with 2 agents has progressed to a certain extent, the simulation shifts to 4 agents. Similarly, since the maximum number of agents in simulator 120C is 8, the simulation first starts with 2 agents. When the learning with 2 agents has progressed to a certain extent, the simulation shifts to 4 agents. When the learning with 4 agents has progressed to a certain extent, the simulation shifts to 8 agents. Similarly, since the maximum number of agents in simulator 120D is 10, the simulation first starts with 2 agents. When the learning with 2 agents has progressed to a certain extent, the simulation shifts to 4 agents. When the learning with 4 agents has progressed to a certain extent, the simulation shifts to 8 agents. When the learning with 8 agents has progressed to a certain extent, the simulation shifts to 10 agents. When the agents reach the maximum number in simulators 120B, 120C, and 120D, the simulation at the maximum number is continued until the learning ends. Note that since the maximum number of agents in simulator 120A is 2, the simulation is executed with 2 agents from the beginning to the end of the learning.
[0053] Note that in FIG. 7, for simplicity, the environments with the same number of agents in each learning stage are shown in the same state. However, this means that the simulation of the environments with the same number of agents is executed in consecutive learning stages, and does not mean that exactly the same simulation is repeatedly executed. Also, in each simulator, the representation of the simulation of the environments with the same number of agents for each learning stage means that the simulation with the same number of agents is performed in consecutive learning stages, and does not necessarily mean that the start and end of the simulation are performed for each learning stage. When the number of agents does not change, the start and end of the simulation may be performed for each learning stage, or may be continuously performed in consecutive learning stages.
[0054] According to such a configuration, learning in environments with different numbers of agents can be advanced evenly, so it becomes possible to flexibly respond to environments with any number of agents. That is, by using the policy learned in such a method, the movement control device 250 can control the moving body 200 so that the moving body 200 moves in an appropriate manner according to the number of surrounding moving bodies. Further, by using the policy learned in such a method, the movement control unit 252 of the movement control device 250 can determine the path of the moving body 200 according to the number of obstacles existing around the moving body 200. The movement control unit 252 is an example of a "path determination unit".
[0055] Specifically, different maximum agent numbers are preset in each simulator 120, and each simulator 120 executes simulations while gradually increasing the number of agents from a small number of agents to its respective maximum agent number. Note that the learning device 100 may be configured to allocate computing resources to each simulator 120 in a time-sharing manner, or may be configured to allocate computing resources that each simulator 120 can use in parallel. For example, the learning device 100 may include CPUs more than the number of simulators 120 and may be configured to allocate a separate CPU to each simulator 120 as computing resources. FIG. 7 shows an example in which the first to fourth CPUs #1 to #4 are allocated to the simulators 120A to 120D. The computing resources allocated to each simulator 120 may be in units of physical cores of the CPU or may be in units of virtual cores realized by a technology such as SMT (Simultaneous Multithreading Technology).
[0056] According to the learning device 100 described above, learning of the operation of the agent by reinforcement learning can be distributed and executed in parallel on a plurality of simulators 120 corresponding to each environment with different numbers of agents. As a result, the movement control device 250 to which the policy that is the learning result of the learning device 100 is applied can determine an appropriate movement mode according to the congestion level of the environment.
[0057] [Moving body] FIG. 8 is a diagram showing a configuration example of the moving body 200. The moving body 200 includes, for example, a movement control device 250, a peripheral detection device 210, a moving body sensor 220, a work unit 230, and a drive device 240. The moving body 200 may be a vehicle or a device such as a robot. The movement control device 250, the peripheral detection device 210, the moving body sensor 220, the work unit 230, and the drive device 240 are connected to each other by a multiplex communication line such as a CAN (Controller Area Network) communication line, a serial communication line, a wireless communication network, or the like.
[0058] The peripheral detection device 210 is a device for detecting the environment around the moving body 200 and the operations of other moving bodies in the periphery. The peripheral detection device 210 includes, for example, a positioning device including a GPS receiver and map information, and an object recognition device such as a radar device and a camera. The positioning device measures the position of the moving body 200 and matches the position with the map information. The radar device emits radio waves such as millimeter waves around the moving body 200 and detects radio waves (reflected waves) reflected by an object to detect at least the position (distance and azimuth) of the object. The radar device may detect the position and movement vector of the object. The camera is, for example, a digital camera using a solid-state imaging device such as a CCD (Charge Coupled Device) or a CMOS (Complementary Metal Oxide Semiconductor), and is provided with an image processing device for recognizing the position of an object from a captured image. The peripheral detection device 210 outputs information such as the position of the moving body 200 on the map and the position of an object (including other moving bodies corresponding to the other agents described above) existing around the moving body 200 to the movement control device 250.
[0059] The mobile body sensor 220 includes, for example, a speed sensor that detects the speed of the mobile body 200, an acceleration sensor that detects acceleration, a yaw rate sensor that detects the angular velocity around the vertical axis, an azimuth sensor that detects the orientation of the mobile body 200, and the like. The mobile body sensor 220 outputs the detected result to the mobile body control device 250.
[0060] The working unit 230 is, for example, a device that provides a predetermined service to the user. The service here is, for example, work such as loading or unloading goods onto a transportation device. The working unit 230 includes, for example, a magic arm, a loading platform, an HMI (Human Machine Interface) such as a microphone and a speaker, and the like. The working unit 230 operates according to the content instructed by the mobile body control device 250.
[0061] The driving device 240 is a device for moving the mobile body 200 in a desired direction. When the mobile body 200 is a robot, the driving device 240 includes, for example, two or more legs and actuators. When the mobile body 200 is a vehicle, a micromobility device, or a robot that moves on wheels, the driving device 240 includes wheels (steering wheels, drive wheels) and motors, engines, etc. for rotating the wheels.
[0062] The mobile body control device 250 includes, for example, a mobile body control unit 252 and a storage unit 256. The mobile body control unit 252 is realized, for example, by a hardware processor such as a CPU executing a program (software). The program may be stored in advance in a storage device (non-transitory storage medium) such as an HDD or a flash memory, or may be stored in a removable storage medium (non-transitory storage medium) such as a DVD or a CD-ROM, and may be installed by mounting the storage medium on a drive device. Some or all of these components may be realized by hardware (including a circuit unit; circuitry) such as an LSI, an ASIC, an FPGA, or a GPU, or may be realized by the cooperation of software and hardware.
[0063] The storage unit 256 is, for example, an HDD, a flash memory, a RAM, a ROM, or the like. Information such as a policy 256A is stored in the storage unit 256. The policy 256A is a policy PL generated by the learning device 100 and is based on the policy at the end point of the learning stage processing.
[0064] The movement control unit 252 inputs information such as the position of the moving body 200 on the map detected by the peripheral detection device 210, the position of an object existing around the moving body 200, and the information of the destination input by the user into the policy 256A, for example. Thereby, the position (movement mode) where the moving body 200 should proceed next is determined, and the determined position is output to the drive device 240. By repeating this, the route of the moving body 200 is sequentially determined.
[0065] According to the movement control device 250 described above, by applying the policy that is the learning result of the learning device 100 of the embodiment, while moving the moving body 200 in a manner according to the environmental congestion degree, a predetermined service can be provided to the user.
[0066] <Second Embodiment> In the movement control system 1 of the second embodiment, similar to the movement control system 1 of the first embodiment, the learning device 100 simulates the movement of agents in environments with different numbers of agents by a plurality of simulators 120, the experience accumulation unit 130 generates evaluation information based on the simulation results, and the learning unit 110 updates the parameters of the network based on the evaluation information.
[0067] On the other hand, in the movement control system 1 of the first embodiment, in the learning device 100, each simulator 120 executes stepwise reinforcement learning in one environment (see FIG. 7), whereas in the movement control system 1 of the second embodiment, the movement control system 1 is different from the learning device 100 of the first embodiment in that each simulator 120 executes simulations at each learning stage in a plurality of environments with the same number of agents. Other configurations are the same as those of the movement control system 1 of the first embodiment (see FIGS. 1, 2, 7, etc.).
[0068] FIG. 9 is an image diagram showing a state in which a plurality of simulators 120 execute simulations in a plurality of environments with the same number of agents in the learning device 100 of the second embodiment. Also in the second embodiment, as in the first embodiment, different CPUs are assigned to each simulator 120 as computable resources that can be used simultaneously. For example, FIG. 9 shows an example in the case where the maximum number of agents that can be calculated simultaneously by one CPU (hereinafter referred to as "maximum parallel number") is 40.
[0069] Here, since the maximum number of agents of the first simulator 120A is set to 2, the simulation is always executed in a 2-agent environment from the first stage to the fourth stage. In this case, since the maximum parallel number per CPU is 40, the first simulator 120A executes the simulation in parallel for 20 two-agent environments.
[0070] Similarly, since the maximum number of agents of the second simulator 120B is set to 4, first, in the first stage, the simulation in a 2-agent environment is executed, and in the second stage, the simulation shifts to a 4-agent environment with the maximum number of agents, and in the second to fourth stages, the simulation in a 4-agent environment is executed. In this case, since the maximum number of agents per CPU is 40, the second simulator 120B executes the simulation in parallel for 20 two-agent environments in the first stage, similar to the first simulator 120A, and executes the simulation in parallel for 9 four-agent environments in the second to fourth stages. Here, for the sake of easy understanding of the image, 9 (= 3 × 3) four-agent environments (the total number of agents is 36 = 9 × 4 < 40) are shown, but the second simulator 120B may be configured to execute the simulation in parallel for 10 four-agent environments, which is the maximum parallel number.
[0071] Similarly, since the maximum number of agents in the third simulator 120C is set to 8, first in the first stage, it executes simulations in a 2-agent environment, then in the second stage, it moves on to simulations in a 4-agent environment, in the third stage, it moves on to simulations in an 8-agent environment with the maximum number of agents, and in the third to fourth stages, it executes simulations in an 8-agent environment. In this case, since the maximum number of agents per CPU is 40, the third simulator 120C, in the first stage, similar to the first simulator 120A, executes simulations in parallel for 20 2-agent environments, in the second stage, similar to the second simulator 120B, executes simulations in parallel for 10 4-agent environments, and in the third to fourth stages, executes simulations in parallel for 4 8-agent environments. Here, for the sake of easy understanding of the image, 4 (=2×2) 8-agent environments (total number of agents is 32 = 8×4 < 40) are shown, but the third simulator 120C may be configured to execute simulations in parallel for 5 8-agent environments, which is the maximum parallel number.
[0072] Similarly, since the maximum number of agents in the fourth simulator 120D is set to 10, first in the first stage, it executes simulations in a 2-agent environment, then in the second stage, it moves on to simulations in a 4-agent environment, in the third stage, it moves on to simulations in an 8-agent environment, and in the fourth stage, it moves on to simulations in a 10-agent environment with the maximum number of agents. In this case, since the maximum number of agents per CPU is 40, the fourth simulator 120D, in the first stage, similar to the first simulator 120A, executes simulations in parallel for 20 2-agent environments, in the second stage, similar to the second simulator 120B, executes simulations in parallel for 9 4-agent environments, in the third stage, similar to the third simulator 120C, executes simulations in parallel for 4 8-agent environments, and in the fourth stage, executes simulations in parallel for 4 10-agent environments.
[0073] Note that in Fig. 9, the number of agents in a plurality of environments generated by each CPU is unified in each learning stage, but this is not essential. As long as the maximum parallelism is not exceeded and it conforms to the gradual increase in the number of agents, the number of agents does not need to be unified across multiple environments. For example, in the final stage of learning, the number of moving bodies in each environment can be set to 2 for CPU#1, 2 to 6 (decrease the number of moving bodies and increase the number of environments) for CPU#2, 2 to 6 (increase the number of moving bodies and decrease the number of environments) for CPU#3, and 2 to 10 (increase the number of moving bodies and decrease the number of environments) for CPU#4.
[0074] Also, in Fig. 9, for simplicity, environments with the same number of agents in each learning stage are represented in the same state, but this means that simulations of environments with the same number of agents are executed simultaneously, not that exactly the same simulation is executed simultaneously.
[0075] Also, in Fig. 9, similar to Fig. 7, for simplicity, environments with the same number of agents in each learning stage are represented in the same state, but this means that simulations of environments with the same number of agents are executed in consecutive learning stages, not that exactly the same simulation is repeatedly executed. Also, in each simulator, the representation of simulations of environments with the same number of agents for each learning stage means that simulations with the same number of agents are carried out in consecutive learning stages, not necessarily that the start and end of the simulation are carried out for each learning stage. When the number of agents does not change, the start and end of the simulation may be carried out for each learning stage, or may be continuously carried out in consecutive learning stages.
[0076] In the movement control system 1 of the second embodiment configured as described above, the learning device 100 can execute simulations in parallel for a plurality of environments with the same number of agents. With such a configuration, the movement control system 1 of the embodiment can efficiently learn the movement of each agent in an environment where a plurality of agents exist.
[0077] Further, in the movement control system 1 of the second embodiment, each of the simulators for each of the plurality of CPUs virtually forms a plurality of environments, the total value of the moving bodies for each CPU is unified by the plurality of CPUs, and the number of agents corresponding to the number of environments is generated in each environment. According to such a configuration, the movement control system 1 of the embodiment can prevent a bias for each CPU from occurring in the collected experience and can more efficiently learn the movement of each agent.
[0078] In the present embodiment, it is assumed that the policy is updated only in the learning stage and not after being mounted on the moving body, but learning may be continued even after being mounted on the moving body.
[0079] As described above, the embodiments for carrying out the present invention have been described using the embodiments, but the present invention is not limited to such embodiments, and various modifications and substitutions can be made without departing from the gist of the present invention.
[0080] The above-described embodiments can be expressed as follows. A storage device storing a program; A hardware processor, and by the hardware processor executing the program stored in the storage device, simulating the operation of the moving body by a plurality of the simulators in which the number of existing moving bodies or obstacles is different for each simulator; learning the policy of the operation so that the cumulative sum of each reward obtained by applying a reward function to each processing result of the plurality of simulators is maximized. A learning device configured as described above.
[0081] The above-described embodiment can be expressed as follows. A storage device that stores a program, And a hardware processor, By the hardware processor executing the program stored in the storage device, Determine the path of the moving body according to the number of obstacles existing around the moving body, Move the moving body along the determined path, A movement control device configured as described above.
[0082] The above-described embodiment can be expressed as follows. A storage device that stores a program, And a hardware processor, By the hardware processor executing the program stored in the storage device, The working unit provides a predetermined service to the user, The driving device drives the self-moving body to move in the movement mode determined by the above movement control device, A moving body configured as described above.
Explanation of Signs
[0083] 1... Movement control system, 100... Learning device, 110... Learning unit, 120... Simulator, 120A... First simulator, 120B... Second simulator, 120C... Third simulator, 120D... Fourth simulator, 130... Experience accumulation unit, 200... Moving body, 210... Peripheral detection device, 220... Moving body sensor, 230... Working unit, 240... Driving device, 250... Movement control device, 252... Movement control unit, 254... Control unit, 256... Storage unit
Claims
1. A path determination unit that determines a path along which a mobile object to be controlled moves according to the number of obstacles existing around the mobile object to be controlled; A control unit that moves the mobile object to be controlled along the path determined by the path determination unit; comprising: The path determination unit is a plurality of simulators corresponding to each of a plurality of environments with different numbers of obstacles, the plurality of simulators that execute simulations of the operations of the mobile object in the corresponding environments, and determines the path of the mobile object to be controlled based on the policy of the operations learned by a learning unit; The policy of the operations is learned by the learning unit updating the current policy of the operations so that the reward obtained by applying a reward function to the execution results of the simulations simultaneously performed by the plurality of simulators for the plurality of environments is maximized. A mobile body control device.
2. The policy of the operations is learned by the learning unit updating the policy of the operations so that the cumulative sum of each reward obtained by applying the reward function to the plurality of execution results of the simulations by the plurality of simulators is maximized. The mobile body control device according to Claim 1.
3. The mobile body control device according to any one of Claims 1 or 2, a working unit for providing a predetermined service to a user, a driving device for moving the mobile body, comprising: The driving device drives the mobile body to move in the movement mode determined by the mobile body control device. A mobile body.
4. A plurality of simulators corresponding to each of a plurality of environments with different numbers of obstacles, the plurality of simulators that execute simulations of the operations of the mobile object in the corresponding environments, and a learning unit that learns the policy of the operations by updating the current policy of the operations so that the cumulative sum of each reward obtained by applying a reward function to the execution results of the simulations simultaneously performed by the plurality of simulators for the plurality of environments is maximized; A learning device comprising.
5. The plurality of simulators are executed by separate processors associated with each of them, and execute the simulations related to the plurality of environments in parallel. The learning device according to Claim 4.
6. The plurality of simulators are For the number of moving objects or obstacles in the corresponding environment, different maximum numbers are set, Execute the simulation while gradually increasing the number of moving objects or obstacles from the specified minimum number to the maximum number set for each, The learning device according to claim 4 or 5.
7. Each of the plurality of simulators executes simulations in parallel for a plurality of environments with the same number of moving objects or obstacles at each stage of the simulation. The learning device according to any one of claims 4 to 6.
8. The reward function includes at least one of the degree of reaching the target of the moving object, the number of collisions of the moving object, and the moving speed of the moving object as a variable. The learning device according to any one of claims 4 to 7.
9. The reward function includes the change in the movement vector of the moving objects or obstacles existing around the self-moving object as an independent variable. The learning device according to any one of claims 4 to 8.
10. A computer Executing, by a plurality of simulators corresponding to each of a plurality of environments with different numbers of obstacles, simulations of the operations of the moving objects in the environments corresponding to each of the plurality of simulators; Learning the policy of the operation by updating the current operation policy so that the cumulative sum of each reward obtained by applying the reward function to the execution results of the simulations simultaneously performed by the plurality of simulators for the plurality of environments is maximized. A learning method having.
11. On a computer Executing, by a plurality of simulators corresponding to each of a plurality of environments with different numbers of obstacles, simulations of the operations of the moving objects in the environments corresponding to each of the plurality of simulators; Learning the policy of the operation by updating the current operation policy so that the cumulative sum of each reward obtained by applying the reward function to the execution results of the simulations simultaneously performed by the plurality of simulators for the plurality of environments is maximized. A program for causing the above to be executed.
Citation Information
Patent Citations
Navigation obstacle avoidance method, device and system based on learning and fusion
CN113253733A
Information providing method, information providing system, and computer program
JP2019106114A
Robot control model learning method, robot control model learning apparatus, robot control model learning program, robot control method, robot control apparatus, robot control program, and robot
JP2021077286A
Path planning in mobile robots
US20210191404A1
Path determination device, robot, and path determination method
WO2020136977A1