Mobile body control device, mobile body, learning device and method, and storage medium
By simulating environments with varying numbers of obstacles using multiple simulators and optimizing the path of the moving body using reinforcement learning algorithms, the problem of environmental congestion affecting path decision-making in existing technologies is solved, and precise movement control in different environments is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies struggle to determine appropriate movement paths based on environmental congestion in complex environments, leading to inappropriate movement paths being determined in environments with fewer moving bodies.
Multiple simulators are used to simulate environments with different numbers of obstacles. The path determination of the moving body is optimized by reinforcement learning algorithm, and the reward function maximization strategy is used to learn to adapt to environments with different levels of congestion.
It enables the determination of the appropriate mode of movement based on the degree of environmental congestion, thereby improving the accuracy and safety of path decision-making for mobile entities in different environments.
Smart Images

Figure CN115903773B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a mobile body control device, a mobile body, a learning device, a learning method, and a storage medium. BACKGROUND
[0002] In recent years, attempts to determine a movement path of a mobile body using AI (artificial intelligence) generated through machine learning are being made. Further, research and practical use of reinforcement learning that determines an action based on an observation value and optimizes model parameters to maximize a reward obtained based on feedback from an actual environment or a simulator are also being advanced.
[0003] In connection therewith, in order to take a safe and secure evasive action against a person's movement, an invention of a path determination device that determines a path for an autonomous mobile robot to move to a destination under the condition that a traffic participant including a pedestrian exists in a traffic environment to the destination is disclosed (see Patent Literature 1). The path determination device is provided with: a prediction path determination section that determines a prediction path that is a predicted value of the path of the robot using a prescribed prediction algorithm so as not to interfere with the traffic participant; and a path determination section that determines the path of the robot using a prescribed control algorithm in such a way that a target function becomes a maximum value, the target function including a distance between the robot and the nearest traffic participant and a speed of the robot as arguments when the robot is assumed to move from a current position to the prediction path.
[0004] Further, in Non Patent Literature 1, it is described that, regarding decentralized motion planning in a high-density dynamic environment, multi-stage training is performed while the number of agents is increased in stages.
[0005] Further, in Non Patent Literature 2, as a method of learning a policy that can appropriately determine the action of a mobile body, a multi-scenario-multi-stage-training framework is described.
[0006] PRIOR ART DOCUMENTS
[0007] PATENT LITERATURE
[0008] PATENT LITERATURE 1: International Publication No. 2020 / 136977
[0009] NON PATENT LITERATURE
[0010]
Non Patent Literature 1
[0011]
Non Patent Literature 2
[0012] Problems to be Solved by the Invention
[0013] However, in the conventional method, in order to cope with a complex environment, an environment in which a large number of moving bodies exist is learned, and as a result, sometimes over-learning occurs, and an inappropriate moving path is determined in an environment in which a small number of moving bodies exist. In this way, in the conventional technology, sometimes a moving path cannot be appropriately determined according to the congestion degree of the environment.
[0014] The present invention was completed in consideration of such a situation, and one of the objects thereof is to provide a moving body control device, a moving body, a learning device, a learning method, and a storage medium that can determine an appropriate moving method according to the congestion degree of the environment.
[0015] Means for Solving the Problems
[0016] The moving body control device, the moving body, the learning device, the learning method, and the storage medium of the present invention adopt the following structure.
[0017] (1) : The mobile body control device of one aspect of the present invention includes a path decision section that decides a path of a mobile body in accordance with a number of obstacles existing in the periphery of the mobile body, and a control section that moves the mobile body along the path decided by the path decision section.
[0018] (2) : In the aspect of the above (1), the path decision section decides a path of a mobile body based on a policy of action learned by a simulator and a learning section, the policy of action being a policy of action learned by the learning section so as to maximize a cumulative sum of rewards obtained by applying a reward function to each processing result of a plurality of simulators, the simulators performing simulation of actions of the mobile body and obstacles in a plurality of environments in which the number of obstacles differs for each of the simulators.
[0019] (3) : In the aspect of the above (2), the policy of action is learned based on processing results of a plurality of simulators, the number of obstacles in the environment differing for each of the simulators, the learning section updating the policy of action so as to maximize a cumulative sum of rewards obtained by applying a reward function to each processing result of the plurality of simulators, thereby learning the policy of action.
[0020] (4) : The mobile body of one aspect of the present invention includes any of the mobile body control devices described above, a work section for providing a prescribed service to a user, and a drive device for moving the mobile body, the drive device being driven so as to move the mobile body in a manner decided by the mobile body control device.
[0021] (5) : The learning device of one aspect of the present invention includes a plurality of simulators that perform simulation of actions of a mobile body, and in the plurality of simulators, the number of mobile bodies or obstacles existing differs for each of the simulators, and a learning section that learns a policy of action so as to maximize a cumulative sum of rewards obtained by applying a reward function to each processing result of the plurality of simulators.
[0022] (6) : In the aspect of the above (5), the plurality of simulators are executed by separate processors that have a corresponding relationship with the plurality of simulators, respectively.
[0023] (7) : In the aspect of the above (5) or (6), a maximum number of the mobile bodies or the obstacles is set to differ for each of the plurality of simulators, and the plurality of simulators perform simulation while increasing the number of the mobile bodies or the obstacles from a prescribed minimum number to the maximum number for each of the plurality of simulators in stages.
[0024] (8) In any of the aspects (5) to (7) described above, the plurality of simulators perform simulation in parallel with respect to a plurality of environments in which the number of mobile bodies or obstacles is the same in each stage of simulation.
[0025] (9) In any of the aspects (5) to (8) described above, the reward function includes, as a variable, at least one of the degree of arrival of a mobile body to a goal, the number of collisions of a mobile body, and the moving speed of a mobile body.
[0026] (10) In any of the aspects (5) to (9) described above, the reward function includes, as an argument, a change in the moving vector of a mobile body or an obstacle existing in the vicinity of the mobile body.
[0027] (11) The learning method of an aspect of the present application causes a computer to perform simulation of the behavior of a mobile body using a plurality of simulators in which the number of mobile bodies or obstacles existing differs for each simulator, and to learn a policy for the behavior so as to maximize the cumulative sum of each reward obtained by applying a reward function to each processing result of the plurality of simulators.
[0028] (12) The storage medium of an aspect of the present application causes a computer to perform simulation of the behavior of a mobile body using a plurality of simulators in which the number of mobile bodies or obstacles existing differs for each simulator, and to learn a policy for the behavior so as to maximize the cumulative sum of each reward obtained by applying a reward function to each processing result of the plurality of simulators.
[0029] Effects of Invention
[0030] According to (1) to (4), there are provided a path decision section that decides a path of a mobile body in accordance with the number of obstacles existing in the vicinity of the mobile body, and a control section that moves the mobile body along the path decided by the path decision section, whereby an appropriate moving method can be decided in accordance with the degree of congestion of the environment.
[0031] Further, according to (5) to (12), there are provided a plurality of simulators that perform simulation of the behavior of a mobile body, and in which the number of mobile bodies or obstacles existing differs for each simulator, and a learning section that learns a policy for the behavior so as to maximize the cumulative sum of each reward obtained by applying a reward function to each processing result of the plurality of simulators, whereby an appropriate moving method can be decided in accordance with the degree of congestion of the environment. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 is a schematic diagram of the structure of the mobile body control system of the embodiment.
[0033] Figure 2 is a diagram showing a configuration example of the learning device.
[0034] Figure 3 is a diagram explaining a reward function R4.
[0035] Figure 4 is a diagram showing an example of an effect of reinforcement learning of the stage.
[0036] Figure 5 is a first diagram showing an example of overlearning of the network.
[0037] Figure 6 is a second diagram showing an example of overlearning of the network.
[0038] Figure 7 is a diagram showing a case where the learning device learns actions with respect to environments different in the number of agents respectively using a plurality of simulators.
[0039] Figure 8 is a diagram showing a configuration example of the mobile body.
[0040] Figure 9 is a diagram showing a case where a plurality of simulators in the learning device execute simulation under a plurality of environments of the same number of agents.
[0041] Diagram text translation:
[0042] 1 … mobile body control system, 100 … learning device, 110 … learning unit, 120 … simulator, 120A … first simulator, 120B … second simulator, 120C … third simulator, 120D … fourth simulator, 130 … experience accumulation unit, 200 … mobile body, 210 … surrounding recognition device, 220 … mobile body sensor, 230 … work unit, 240 … drive device, 250 … mobile body control device, 252 … movement control unit, 254 … control unit, 256 … storage unit. DETAILED DESCRIPTION
[0043] Hereinafter, embodiments of the mobile body control device, the mobile body, the learning device, the learning method, and the storage medium of the present application will be described with reference to the drawings.
[0044] <First Embodiment>
[0045] Figure 1is a diagram showing the structure of the mobile body control system 1 of the embodiment. The mobile body control system 1 is provided with a learning device 100 and a mobile body 200. The learning device 100 is implemented by one or more processors. The learning device 100 is a device that determines actions of a plurality of mobile bodies by computer simulation, derives or acquires a reward based on a change in a state resulting from the actions, and learns actions (movements) that maximize the reward. The actions refer to, for example, movements within a simulation space. Actions other than movements can also be learning targets, but in the following description, the actions refer to movements. The simulator that determines the movements can also be executed in a device different from the learning device 100, but in the following description, the simulator is executed in the learning device 100. The learning device 100 is pre-stored with environmental information that becomes a premise of the simulation, such as map information. The learning result of the learning device 100 is mounted on the mobile body 200 as a policy PL.
[0046] [Learning device]
[0047] Figure 2 is a diagram showing an example of the structure of the learning device 100 of the embodiment. The learning device 100 is provided with, for example, a learning unit 110, a plurality of simulators 120, and an experience accumulation unit 130. These constituent elements are implemented by, for example, a hardware processor such as a CPU (Central Processing Unit) executing a program (software). Part or all of these constituent elements can also be implemented by a hardware (including circuitry) such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), a GPU (Graphics Processing Unit), and can also be implemented by a cooperative operation of software and hardware. The program can be pre-stored in a storage device (a storage device provided with a non-transitory storage medium) such as an HDD (Hard Disk Drive), a flash memory, an SSD (Solid State Drive), and can also be stored in a removable storage medium (a non-transitory storage medium) such as a DVD, a CD-ROM, and installed by mounting the storage medium on a drive device.
[0048] The learning unit 110 updates the policy in accordance with various reinforcement learning algorithms based on evaluation information indicating the results of the experience accumulation unit 130 evaluating the state changes generated by the plurality of simulators 120. The learning unit 110 repeatedly performs the process of outputting the updated policy to the plurality of simulators 120 until learning is completed. The policy refers to, for example, a neural network having parameters (hereinafter, also simply referred to as "network"), and outputs an action that the agent can take in a probabilistic manner with respect to the input of environment information. Here, the agent refers to a mobile body that exists within a simulation space (environment) and is the object of learning actions. The agent is an example of "this mobile body". The environment information is information indicating the state of the environment. The policy can also be a rule-based function having parameters. The learning unit 110 updates the policy by updating the parameters based on the evaluation information. The learning unit 110 supplies the updated parameters to each of the simulators 120.
[0049] The simulator 120 inputs the action target and the current state (if it is the initial state after starting the simulation) to the policy, and derives a state change as a result of the actions of the agent and other agents. The policy is, for example, a DNN (Deep Neural Network), but can also be a rule-based policy or another type of policy. The policy derives the occurrence probability for each of a plurality of types of actions assumed. For example, in a simple example, it is assumed that a plane extends in up, down, left, and right directions, and outputs a result of 80% right movement, 10% left movement, 10% up movement, and 0% down movement. The simulator 120 applies a random number to the result, and derives the state change of the agent in such a manner that if the random number value is 0% or more and less than 80%, it is right movement, if the random number value is 80% or more and less than 90%, it is left movement, and if the random number value is 90% or more, it is up movement.
[0050] The plurality of simulators 120 use the policy (network) updated by the learning unit 110 to perform simulations with respect to environments in which the number of agents is different for each environment and in which a plurality of agents exist, and thereby determine the actions of the agents in each environment. Note that the determination of the action referred to here refers to the derivation of the above-described state change with respect to the agent. In the present embodiment, for example, four simulators are assumed as the plurality of simulators 120. For example, in the present embodiment, it is assumed that the first simulator 120A to the fourth simulator 120D determine the movement of 2 agents, 4 agents, 8 agents, and 10 agents, respectively. Note that the environment can also include mobile bodies other than agents that do not depend on the policy. For example, the environment can include, in addition to agents that move based on the policy, a mobile body that is stationary, a mobile body that moves in a different action model from the policy, and the like.
[0051] Specifically, each simulator 120 updates the parameter update policy (network) supplied from the learning unit 110, and inputs the current state obtained from the simulation result of the last time (the sampling period before the present sampling period) to the updated network, and applies a random number to the output result, thereby deciding the action of each agent for the present time (the present sampling period). Each simulator 120 generates the updated state and the reward from the environment EV by inputting the decided action to the environment EV. The reward is generated by inputting the decided action to the reward function through the environment EV. Each simulator 120 supplies the experience information obtained based on the action decided for each agent to the experience accumulation unit 130. For example, the experience information includes information on the action decided by the agent, the state before the action, the state after the action, and the reward obtained by the action.
[0052] The experience accumulation unit 130 accumulates the experience information supplied from each simulator 120, and samples the experience information with high priority from the accumulated experience information and supplies it to the learning unit 110. The priority is the priority obtained based on the degree of learning effect in the learning of the network NW, and is decided by, for example, the TD (Temporal Difference) error. Note that the priority can be appropriately updated based on the learning result of the learning unit 110.
[0053] The learning unit 110 updates the parameters of the network NW based on the experience information supplied from the experience accumulation unit 130, so as to maximize the reward obtained by the movement of each agent. The learning unit 110 supplies the updated parameters to each simulator 120. Each simulator 120 updates the network NW by the parameters supplied from the learning unit 110.
[0054] The learning unit 110 can use any of various reinforcement learning algorithms. The learning unit 110 learns the appropriate movement of the agent in the environment in which a plurality of agents exist by repeatedly performing the update of the parameters. The network learned in this way is supplied to the mobile body 200 as a policy.
[0055] Note that the reward function used by the environment EV when calculating the reward can be an arbitrary function as long as the agent is given a greater reward the more appropriate movement it takes. For example, a function R including a reward function Rl, a reward function R2, a reward function R3, and a reward function R4 can be used as the reward function, where the reward function Rl is a function given when the agent reaches the destination, the reward function R2 is a function given when the agent smoothly achieves movement, the reward function R3 is a function that becomes smaller when the agent causes a change in the movement vector of another agent, and the reward function R4 is a function that makes the distance that should be maintained when the agent approaches another agent variable depending on the direction in which the other agent is facing. In addition, the reward function R can be a function including at least one of Rl, R2, R3, and R4.
[0056] R = Rl + R2 + R3 + R4... (1)
[0057] For example, the reward function Rl is a function that becomes a positive fixed value when reaching the destination and becomes a value proportional to the change in distance from the destination (positive if the distance changes in the decreasing direction and negative if it changes in the increasing direction) when not reaching the destination. The reward function Rl is an example of a "first function".
[0058] For example, the reward function R2 is a function that becomes a value that becomes larger the smaller the third derivative of the position on the two-dimensional plane of the agent, i.e., the jerk. The reward function R2 is an example of a "second function".
[0059] For example, the reward function R3 is a function that returns a low evaluation value when the agent enters a prescribed area. According to such a reward function R3, for example, a low evaluation can be given to the agent for the action in the area (prescribed area) that is in front of another agent, and a less low evaluation can be given to the action in the side or back. The reward function R3 is an example of a "third function".
[0060] Figure 3 is a graph for illustrating the reward function R4. Figure 3 An environment in which the person Pl, P4, and P5, and the robots R2, R3, and R5 are mixed together is shown as an example of a simulation environment. In the environment shown in Figure 3 , the locations Dl to D5 are the destination locations of the respective moving bodies. Specifically, the location Dl is the destination location of the person Pl, the location D2 is the destination location of the robot R2, the location D3 is the destination location of the robot R3, the location D4 is the destination location of the person P4, and the location D5 is the destination location of the person P5.
[0061] Here, the robot R5 is the target robot, and as a reward function R4 for learning a movement method in which the target robot does not obstruct the movement of a person, it can be defined as, for example, the following (2).
[0062]
[0063] In (2), R4 is a reward function for learning a movement method in which the target robot does not obstruct the movement of a person, and is a function that gives a larger reward the more the movement is one in which the target robot does not obstruct the movement of a person. i is the identification number of a moving body such as a person or a robot present in the environment, and N is the maximum number thereof. Further, a i represents an action determined in accordance with the state of the environment including the target robot R5 (hereinafter referred to as "first action"). i represents an action determined in accordance with the state of the environment not including the target robot R5 (in a case in which the target robot R5 is ignored) (hereinafter referred to as "second action"). w is a coefficient that takes the difference between the first action and the second action with respect to each moving body, and transforms a value corresponding to the sum thereof into a negative reward value as a penalty. That is, (2) is an expression in which the larger the difference between the first action and the second action is, the smaller the reward is calculated to be. According to such a reward function, for example, the target robot R5 can learn a movement method in which the action of the target robot R5 does not have an influence on the movement of other moving bodies. The reward function R4 is an example of a "fourth function".
[0064] The learning action of the network described above describes the action when each of the simulators 120 performs simulation with a predetermined number of agents. The learning device 100 of the present embodiment is configured to learn the action of a moving body in a plurality of environments different in the number of agents in parallel by performing the reinforcement learning described above while gradually increasing the number of agents in the simulation. As one of the methods for improving the accuracy of reinforcement learning, there is known a method of learning a policy of an environment with a final number of agents while gradually increasing the number of agents (hereinafter referred to as "phased reinforcement learning").
[0065] Figure 4 is a graph indicating an example of the effect of the phased reinforcement learning. In Figure 4 , the horizontal axis indicates the degree of progress of learning in each phase, and the vertical axis indicates the accuracy of learning. According to Figure 4 it is known that, compared with a case in which learning is advanced with 10 agents from the beginning, a case in which learning is advanced while the number of agents is increased in stages by 2, 4, 8, and 10 is able to learn an action with a higher reward.
[0066] However, in an environment in which a plurality of agents exist, the policy learned with 10 agents does not necessarily determine appropriate movement in all environments. This is because, in learning of movement, while the movement destination in which priority is given to determination of non-contact with other moving bodies, obstacles, and the like (i.e., learning as an action that can obtain a high reward) is determined, the priority of other matters is sometimes high depending on the state of the environment (e.g., the density of agents existing in the environment, and the like). That is, the learning result of movement in an environment in which the number of agents is larger sometimes becomes overlearning when movement in an environment in which the number of agents is smaller is determined.
[0067] Figure 5 and Figure 6 is a graph indicating an example of overlearning of a policy. Figure 5 indicates an example of movement based on a policy learned with 2 agents, Figure 6 indicates an example of movement based on a policy learned with 10 agents. Figure 5 and Figure 6 both indicate a movement path determined so that one agent A departs from a departure place B and reaches a destination D while avoiding an obstacle C. From Figure 5 and Figure 6 It is understood that, in the policy learned in an environment with 2 agents, the agent A starts the avoidance action of the obstacle C soon after departure from the departure place B, in contrast to this, in the policy learned in an environment with 10 agents, the agent A starts the avoidance action at a position closer to the obstacle C.
[0068] Such a difference in avoidance action can be considered, for example, to be caused as a result of learning that, since the more the number of agents, the more likely it is to interfere with other agents, the avoidance action is started at a position closer to the obstacle C in order to avoid interference with other agents. In addition, such a difference in avoidance action can be considered, for example, to be caused as a result of learning that, since the less the number of agents, the less likely it is to interfere with other agents, the direction of travel is changed more gently in order to improve safety of movement.
[0069] In summary, in the reinforcement learning of the related art stage, in a case in which learning is individually performed in order from an environment in which the number of agents is small to an environment in which the number of agents is large, in determination of a movement pattern based on a policy, the learning result based on the last learning environment dominates. Therefore, even if movement in an environment in which a plurality of agents exist can be learned with good accuracy, the policy generated by the learning is optimized in an environment in which a plurality of agents exist, and sometimes cannot determine appropriate action in an environment in which the number of agents is different. Therefore, in the learning device 100 of the present embodiment, a structure in which environments in which the number of agents is different are learned in parallel by causing a plurality of simulators 120 to act in parallel is provided.
[0070] Figure 7 is a diagram indicating a case where the learning device 100 learns actions with respect to environments in which the number of agents is different for each of the plurality of simulators 120. As described above, in the learning device 100 of the present embodiment, the simulators 120A, 120B, 120C, 120D determine actions of each agent with respect to environments in which there are 2 agents, 4 agents, 8 agents, and 10 agents, respectively. Specifically, each of the simulators 120 starts simulation with a predetermined minimum number of agents, and performs simulation while gradually increasing the number of agents until the maximum number of agents for each of the simulators 120.
[0071] For example, in the present embodiment, the simulator 120B, since the maximum number of agents is 4, first starts simulation with 2 agents, and switches to simulation with 4 agents when learning with 2 agents has progressed to a certain extent. Similarly, the simulator 120C, since the maximum number of agents is 8, first starts simulation with 2 agents, switches to simulation with 4 agents when learning with 2 agents has progressed to a certain extent, and switches to simulation with 8 agents when learning with 4 agents has progressed to a certain extent. Similarly, the simulator 120D, since the maximum number of agents is 10, first starts simulation with 2 agents, switches to simulation with 4 agents when learning with 2 agents has progressed to a certain extent, switches to simulation with 8 agents when learning with 4 agents has progressed to a certain extent, and switches to simulation with 10 agents when learning with 8 agents has progressed to a certain extent. The simulators 120B, 120C, 120D, when the number of agents reaches the maximum number, continue simulation with the maximum number until the end of learning. Note that the simulator 120A, since the maximum number of agents is 2, performs simulation with 2 agents from the beginning to the end of learning.
[0072] Note that, in Figure 7 , for simplicity, the same number of agents in the environment is expressed with the same state in each learning stage, but this means that simulation of the same number of agents in the environment is performed in consecutive learning stages, and does not mean that exactly the same simulation is repeatedly performed. In addition, in each of the simulators, expressing simulation of the same number of agents in the environment by each learning stage means that simulation of the same number of agents is performed in consecutive learning stages, and does not necessarily mean that the start and end of simulation are performed by each learning stage. In the case where the number of agents does not change, the start and end of simulation can be performed by each learning stage, or can be continued in consecutive learning stages.
[0073] According to such a structure, learning of environments different in the number of agents can be generally performed, and thus any environment in the number of agents can be flexibly coped with. That is, by using a policy learned in such a method, the mobile body control device 250 can control the mobile body 200 to move in an appropriate manner in accordance with the number of surrounding mobile bodies. Further, by using a policy learned in such a method, the movement control section 252 of the mobile body control device 250 can determine a path of the mobile body 200 in accordance with the number of obstacles existing in the periphery of the mobile body 200. The movement control section 252 is an example of a "path determination section".
[0074] Specifically, each of the simulators 120 is set in advance with a different maximum number of agents, and each of the simulators 120 performs simulation from a small number of agents while increasing the number of agents in stages until the respective maximum number of agents. Note that the learning device 100 can be configured to allocate computing resources to each of the simulators 120 in a time-divided manner, or can be configured to allocate computing resources that can be used in parallel by each of the simulators 120. For example, the learning device 100 can have a number of CPUs greater than the number of simulators 120, and can be configured to allocate a separate CPU as a computing resource to each of the simulators 120. Figure 7 An example in which the first to fourth CPUs #1 to #4 are allocated to the simulators 120A to 120D is shown. The computing resources allocated to each of the simulators 120 can be a physical core unit of a CPU, or a virtual core unit realized by a technology such as SMT (Simultaneous Multithreading Technology).
[0075] According to the learning device 100 described above, learning of actions of agents based on reinforcement learning can be dispersed to and performed in parallel by a plurality of simulators 120 corresponding to environments different in the number of agents. Thus, the mobile body control device 250 to which a policy that is a result of learning by the learning device 100 is applied can determine an appropriate movement manner in accordance with the congestion degree of an environment.
[0076] [Mobile Body]
[0077] Figure 8 is a diagram showing an example of a structure of the mobile body 200. The mobile body 200 has, for example, a mobile body control device 250, a periphery sensing device 210, a mobile body sensor 220, a work section 230, and a drive device 240. The mobile body 200 can be a vehicle, or a device such as a robot. The mobile body control device 250, the periphery sensing device 210, the mobile body sensor 220, the work section 230, and the drive device 240 are connected to each other through a multiway communication line such as a CAN (Controller Area Network) communication line, a serial communication line, a wireless communication network, or the like.
[0078] The surrounding recognition device 210 is a device for recognizing the environment of the surrounding of the mobile body 200, and the action of other mobile bodies in the surrounding. The surrounding recognition device 210 is provided with, for example, a position measurement device including a GPS receiver, map information, and the like, and an object recognition device such as a radar device, a camera, and the like. The position measurement device detects the position of the mobile body 200, and matches the position with the map information. The radar device radiates electric waves such as millimeter waves to the surrounding of the mobile body 200, and detects the electric waves (reflected waves) reflected by objects to detect at least the position (distance and direction) of the objects. The radar device can also detect the position and movement vector of the objects. The camera is, for example, a digital camera using a solid-state imaging element such as a CCD (Charge Coupled Device), a CMOS (Complementary Metal Oxide Semiconductor), and the like, and is attached with an image processing device that recognizes the position of the objects from the captured image. The surrounding recognition device 210 outputs information such as the position on the map of the mobile body 200, the position of the objects (including other mobile bodies corresponding to the aforementioned other agents) existing in the surrounding of the mobile body 200, and the like, to the mobile body control device 250.
[0079] The mobile body sensor 220 includes, for example, a speed sensor that detects the speed of the mobile body 200, an acceleration sensor that detects acceleration, a yaw rate sensor that detects the angular velocity around the vertical axis, a direction sensor that detects the orientation of the mobile body 200, and the like. The mobile body sensor 220 outputs the detected results to the mobile body control device 250.
[0080] The work section 230 is, for example, a device that provides a prescribed service to the user. The service here refers to, for example, work such as loading and unloading of goods to and from a transport device. The work section 230 includes, for example, a magic arm, a loading platform, a microphone, a speaker, and the like, HMI (Human machine Interface), and the like. The work section 230 performs the action in accordance with the instruction from the mobile body control device 250.
[0081] The drive device 240 is a device for moving the mobile body 200 in a desired direction. In the case where the mobile body 200 is a robot, the drive device 240 includes, for example, two or more leg sections and actuators. In the case where the mobile body 200 is a vehicle, a micro mobile body, or a robot that moves using wheels, the drive device 240 includes wheels (steering wheels, drive wheels) and a motor, an engine, or the like for rotating the wheels.
[0082] The mobile body control device 250 includes, for example, a movement control section 252 and a storage section 256. The movement control section 252 is implemented, for example, by a hardware processor such as a CPU executing a program (software). The program can be stored in advance in a storage device (non-transitory storage medium) such as an HDD or a flash memory, or in a removable storage medium (non-transitory storage medium) such as a DVD or a CD-ROM, and installed by mounting the storage medium in a drive device. Part or all of these components can also be implemented by hardware (including circuitry) such as an LSI, an ASIC, an FPGA, or a GPU, and can also be implemented in cooperation with software and hardware.
[0083] The storage section 256 is, for example, an HDD, a flash memory, a RAM, a ROM, or the like. The storage section 256 stores, for example, information such as a policy 256A. The policy 256A is a policy PL generated by the learning device 100, and is a policy based on the final point in time of the processing of the learning phase.
[0084] The movement control section 252 inputs, for example, the position of the mobile body 200 on a map detected by the surrounding detection device 210, information on the positions of objects existing in the surroundings of the mobile body 200, and information on a destination input by a user, to the policy 256A, thereby determining the position (movement method) in which the mobile body 200 should proceed next, and outputs the determined position to the drive device 240. The path of the mobile body 200 is determined in sequence by repeating this processing.
[0085] According to the mobile body control device 250 described above, by applying the policy that is the result of learning by the learning device 100 of the embodiment, it is possible to move the mobile body 200 in a manner corresponding to the congestion of the environment while providing a prescribed service to a user.
[0086] <Second Embodiment>
[0087] The mobile body control system 1 of the second embodiment, like the mobile body control system 1 of the first embodiment, has the learning device 100 simulate the movement of an agent in environments in which the number of agents differs, using a plurality of simulators 120, the experience accumulation section 130 generate evaluation information based on the results of the simulation, and the learning section 110 update the parameters of the network based on the evaluation information.
[0088] On the other hand, the mobile body control system 1 of the first embodiment has each simulator 120 perform reinforcement learning of a phase in one environment in the learning device 100 (see Figure 7), and in contrast, the mobile body control system 1 of the second embodiment differs from the learning device 100 of the first embodiment in that each simulator 120 performs simulation in each learning stage in a plurality of environments with the same number of agents. The other structures are the same as those of the mobile body control system 1 of the first embodiment (see FIG. 1). Figure 1 、 Figure 2 、 Figure 7 and the like).
[0089] Figure 9 is an image diagram showing a case where a plurality of simulators 120 perform simulation in a plurality of environments with the same number of agents in the learning device 100 of the second embodiment. In the second embodiment, each simulator 120 is also assigned a different CPU as a computing resource that can be used simultaneously, as in the first embodiment. For example, Figure 9 is an example of a case where the maximum number of agents that can be simultaneously and in parallel computed with one CPU (hereinafter referred to as "maximum parallel number") is 40.
[0090] Here, the first simulator 120A performs simulation in a 2-agent environment from the first stage to the fourth stage because the maximum number of agents is set to 2. In this case, because the maximum parallel number per 1 CPU is 40, the first simulator 120A performs simulation in parallel with respect to 20 2-agent environments.
[0091] Similarly, the second simulator 120B performs simulation in a 2-agent environment in the first stage first, and in the second stage, the simulation switches to a 4-agent environment with the maximum number of agents, and performs simulation in a 4-agent environment in the second to fourth stages. In this case, because the maximum number of agents per 1 CPU is 40, the second simulator 120B performs simulation in parallel with respect to 20 2-agent environments in the first stage as in the first simulator 120A, and performs simulation in parallel with respect to 9 4-agent environments in the second to fourth stages. Note that here, 9 (= 3 x 3) 4-agent environments (the total number of agents is 36 = 9 x 4 < 40) are shown in order to make the image easy to understand, but the second simulator 120B can also be configured to perform simulation in parallel with respect to 10 4-agent environments as the maximum parallel number.
[0092] Similarly, since the maximum number of agents is set to 8, the third simulator 120C first performs simulation in a 2-agent environment in the first stage, transitions to a 4-agent environment in the second stage, and then transitions to an 8-agent environment (the maximum number of agents) in the third stage. Simulations in the 8-agent environment are performed in the third and fourth stages. In this case, since the maximum number of agents per CPU is 40, the third simulator 120C performs simulations in the first stage with 20 2-agent environments in parallel, similar to the first simulator 120A. In the second stage, it performs simulations in the second stage with 10 4-agent environments in parallel, similar to the second simulator 120B. In the third and fourth stages, it performs simulations in the third and fourth stages with 4 8-agent environments in parallel. It should be noted that, for ease of visual representation, 4 (=2×2) 8-agent environments are shown here (total number of agents is 32 = 8×4 < 40), but the third simulator 120C can also be configured to perform simulations in the third stage with 5 8-agent environments (the maximum number of parallel environments).
[0093] Similarly, since the maximum number of agents is set to 10, the fourth simulator 120D first performs simulation in a 2-agent environment in the first stage, then transitions to a 4-agent environment in the second stage, an 8-agent environment in the third stage, and finally a 10-agent environment in the fourth stage. In this case, since the maximum number of agents per CPU is 40, the fourth simulator 120D performs simulations in parallel with the first simulator 120A in the first stage for 20 2-agent environments, in parallel with the second simulator 120B in the second stage for 9 4-agent environments, in parallel with the third simulator 120C in the third stage for 4 8-agent environments, and in parallel with the fourth simulator 120C in the fourth stage for 4 10-agent environments.
[0094] It should be noted that, in Figure 9 In the learning process, the number of agents in the multiple environments generated by each CPU is consistent across all learning stages. However, this is not mandatory. As long as the maximum number of agents does not exceed the maximum number of parallel agents and the increase in the number of agents is appropriate for each stage, the number of agents does not need to be consistent across multiple environments. For example, in the final stage of learning, the number of agents in each environment can be set to 2 in CPU #1, 2 to 6 in CPU #2 (reducing the number of agents while increasing the number of environments), 2 to 6 in CPU #3 (increasing the number of agents while decreasing the number of environments), and 2 to 10 in CPU #4 (increasing the number of agents while decreasing the number of environments).
[0095] In addition, Figure 9In the above, for simplicity, the same number of agents of the same environment is expressed in the same state at each learning stage, but this means that the simulation of the same number of agents of the same environment is performed at the same time, and does not mean that the same simulation is performed at the same time.
[0096] In addition, in the above Figure 9 In the above, for simplicity, the same number of agents of the same environment is expressed in the same state at each learning stage, but this means that the simulation of the same number of agents of the same environment is performed at the same time, and does not mean that the same simulation is performed at the same time. Figure 7 In the above, for simplicity, the same number of agents of the same environment is expressed in the same state at each learning stage, but this means that the simulation of the same number of agents of the same environment is performed at the same time, and does not mean that the same simulation is performed at the same time.
[0097] In the mobile body control system 1 of the second embodiment configured as described above, the learning device 100 can perform the simulation in parallel with respect to a plurality of environments of the same number of agents. With this configuration, the mobile body control system 1 of the embodiment can efficiently learn the movement of each agent in an environment in which a plurality of agents exist.
[0098] In addition, in the mobile body control system 1 of the second embodiment, a plurality of environments are hypothetically formed by the simulators of each of the plurality of CPUs respectively, the aggregate values of the mobile bodies of each CPU are unified in the plurality of CPUs, and the number of agents corresponding to the number of environments is generated in each environment. With this configuration, the mobile body control system 1 of the embodiment can prevent bias of each CPU from occurring in the collected experience, and can more efficiently learn the movement of each agent.
[0099] In the present embodiment, it is assumed that the update of the policy is performed only at the learning stage, and is not performed after being mounted on the mobile body, but the learning can also be continued after being mounted on the mobile body.
[0100] The above describes the specific embodiments of the present application using the embodiments, but the present application is not at all limited by such embodiments, and various modifications and substitutions can be made within the scope of the gist of the present application.
[0101] The above-described embodiments can be expressed as follows.
[0102] A learning device configured to include:
[0103] a storage device in which a program is stored; and
[0104] a hardware processor,
[0105] by the hardware processor executing a program stored in the storage device, the following processing is performed:
[0106] simulation of the movement of the mobile body is performed using a plurality of the simulators that differ in the number of mobile bodies or obstacles present per simulator;
[0107] a policy of the movement is learned so as to maximize a cumulative sum of each reward obtained by applying a reward function to each processing result of the plurality of the simulators.
[0108] The embodiments described above can be expressed as follows.
[0109] A mobile body control device is configured to include:
[0110] a storage device in which a program is stored; and
[0111] a hardware processor,
[0112] by the hardware processor executing a program stored in the storage device, the following processing is performed:
[0113] a path of the mobile body is decided in accordance with the number of obstacles present in the periphery of the mobile body;
[0114] the mobile body is caused to move along the decided path.
[0115] The embodiments described above can be expressed as follows.
[0116] A mobile body is configured to include:
[0117] a storage device in which a program is stored; and
[0118] a hardware processor,
[0119] by the hardware processor executing a program stored in the storage device, the following processing is performed:
[0120] a prescribed service is provided to a user using a job unit;
[0121] driving is performed using a driving device so that the mobile body moves in a movement manner decided by the mobile body control device.
Claims
1. A mobile body control device, wherein, The moving body control device includes: A path determination unit determines the path of movement of the controlled object's moving body based on the number of obstacles present around the moving body; and The control unit causes the moving body of the controlled object to move along the path determined by the path determination unit. The path determination unit determines the path of the controlled object's moving body based on the action strategies learned by multiple simulators and the learning unit. The multiple simulators correspond to multiple environments with different numbers of obstacles, and each simulator performs a simulation of the moving body's actions in its corresponding environment. The action strategy refers to the strategy learned by updating the current action strategy through the learning unit to maximize the reward obtained by applying the reward function to the execution results of the simulations performed simultaneously by the multiple simulators with respect to the multiple environments.
2. The moving body control device according to claim 1, wherein, The strategy of the action refers to the strategy learned by updating the strategy of the action through the learning unit to maximize the accumulation and sum of the rewards obtained by applying the reward function to the multiple execution results of the simulation performed by the multiple simulators.
3. A mobile body, wherein, The mobile body has: The moving body control device according to claim 1 or 2; The operations department, which provides prescribed services to users; and A drive mechanism, used to move the moving body. The driving device drives the mobile body to move in a manner determined by the mobile body control device.
4. A learning device, wherein, The learning device includes: Multiple simulators, each corresponding to a different number of environments with varying numbers of obstacles, simulate the actions of moving bodies within their respective environments; and The learning unit learns the action strategy by updating the current action strategy to maximize the cumulative sum of the rewards obtained by applying a reward function to the execution results of the simulations performed simultaneously by the plurality of simulators with respect to the plurality of environments.
5. The learning device according to claim 4, wherein, The plurality of simulators are executed by separate processors that have established corresponding relationships with each of the plurality of simulators, and the simulations involving the plurality of environments are executed in parallel.
6. The learning device according to claim 4, wherein, Each of the multiple simulators has a different maximum number set for the number of moving objects or obstacles in its corresponding environment. The multiple simulators perform the simulation by progressively increasing the number of moving objects or obstacles from a predetermined minimum number to a separately set maximum number.
7. The learning device according to claim 4, wherein, Each of the plurality of simulators performs simulations in parallel with respect to multiple environments having the same number of moving bodies or obstacles at each stage of the simulation.
8. The learning device according to any one of claims 4 to 7, wherein, The reward function includes, as a variable, at least one of the following: the degree of arrival of the moving body at the target, the number of collisions of the moving body, and the moving speed of the moving body.
9. The learning device according to any one of claims 4 to 7, wherein, The reward function includes, as an independent variable, the changes in the movement vectors of moving bodies or obstacles existing around the moving body.
10. A learning method, wherein, The learning method enables the computer to perform the following processes: Simulations of the actions of a moving body in the environment corresponding to the multiple environments, each with a different number of obstacles, are performed by multiple simulators corresponding to the multiple simulators respectively. as well as The strategy for the action is learned by updating the current action strategy to maximize the cumulative sum of the rewards obtained by applying a reward function to the execution results of the simulations performed simultaneously by the multiple simulators with respect to the multiple environments.
11. A storage medium storing a program, wherein, The program is used to enable the computer to perform the following steps: Simulations of the actions of a moving body in the environment corresponding to the multiple environments, each with a different number of obstacles, are performed by multiple simulators corresponding to the multiple simulators respectively. as well as The strategy for the action is learned by updating the current action strategy to maximize the cumulative sum of the rewards obtained by applying a reward function to the execution results of the simulations performed simultaneously by the multiple simulators with respect to the multiple environments.
Citation Information
Patent Citations
Path determination device, robot, and path determination method
WO2020136977A1
Route determination method
CN111673731A
Navigation obstacle avoidance method, device and system based on learning and fusion
CN113253733A
Information providing method, information providing system, and computer program
JP2019106114A