Mobile body, control device and control method thereof, storage medium, and learning device

By combining the collaborative efforts of the path determination unit and the control unit with reward function evaluation, the learning device optimizes the movement strategy of the mobile body, solving the problem that the mobile body has difficulty predicting the actions of other mobile bodies with high accuracy in a crowd environment. This achieves high-affinity action, reduces interference, and improves safety and coordination.

CN115903774BActive Publication Date: 2026-02-13HONDA MOTOR CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211186194.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-09-30
Filing Date
2022-09-27
Publication Date
2026-02-13
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

In existing technologies, it is difficult for mobile entities to accurately predict the movements of other mobile entities in crowd environments, leading to frequent interference with the crowd and an inability to take actions that are highly compatible with surrounding mobile entities.

Method used

The path determination unit determines the path of a mobile body in a way that reduces the change in the movement vector of other mobile bodies, and the control unit makes the mobile body move along the path. The action of the mobile body is evaluated by combining the reward function. The learning device simulates and optimizes the action strategy of the mobile body to reduce interference with other mobile bodies.

Benefits of technology

It can take actions that are highly compatible with surrounding mobile entities without having to predict their future movements, thereby reducing obstacles to other mobile entities and improving their safety and coordination in crowd environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115903774B_ABST
    Figure CN115903774B_ABST
Patent Text Reader

Abstract

A mobile body control device, a mobile body, a mobile body control method, a storage medium, and a learning device are disclosed. The mobile body control device includes a path decision section that decides a path of a mobile body in a manner that reduces a change in a movement vector of another mobile body existing in a periphery of the mobile body, and a control section that moves the mobile body along the path decided by the path decision section.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a mobile body control device, a mobile body, a mobile body control method, a program, and a learning device. BACKGROUND

[0002] In recent years, with the development of AT (Artificial Intelligence), research is being conducted to determine a path through reinforcement learning in an environment where autonomous mobile bodies coexist with people. However, in a traffic environment where people are crowded, the interference state between the robot and the pedestrian occurs frequently.

[0003] In relation thereto, there is disclosed an invention of a path determination device that determines a path when an autonomous mobile robot moves to a destination under the condition that traffic participants including pedestrians exist in a traffic environment to the destination, in order to take a safe and secure evasive action against the movement of people (refer to International Publication No. 2020 / 136977). The path determination device is provided with a prediction path determination section that determines a prediction path that is a predicted value of the path of the robot using a prescribed prediction algorithm to avoid interference between the robot and the traffic participants, and a path determination section that determines the path of the robot using a prescribed control algorithm so that a target function that includes the distance to the closest traffic participant to the robot and the speed of the robot as independent variables becomes a maximum value when the robot is assumed to move on the prediction path from the current position.

[0004] In addition, in "Socially Aware Motion Planning with Deep Reinforcement Learning", Yu Fan Chen, Michael Everett, Miao Liu, Jonathan P. How, 2017.3.26, <https: / / arxiv.org / pdf / 1703.08862.pdf>, with respect to the reward function, it is described that in order to improve the coordination with the surrounding people, a reward function is created on the basis of consideration of the three modes of crossing, facing, and overtaking, and the robot is caused to learn using a prescribed algorithm.

[0005] In addition, in "Mapless Navigation among Dynamics with Social-safety-awareness: a reinforcement learning approach from 2D laser scans", Jun Jin, Nhat M. Nguyen, Nazmus Sakib, Daniel Graves, Hengshuai Yao, and Martin Jagersand, 2020.3.5., https: / / arxiv.org / pdf / 1911.03074.pdf, regarding the reward function, it is described that a reward function is created in association with the number of people moving in an area overlapping in the respective advancing directions of the robot and the people, and the robot is caused to learn using a prescribed algorithm. SUMMARY

[0006] In the above-described related art, since the influence of the movement of the moving body on the movement of the other moving bodies around the moving body is not taken into consideration, sometimes an action with high affinity to the other moving bodies around the moving body cannot be taken. In addition, in the technology described in PCT International Publication No. 2020 / 136977, although the movement (path of the robot) of the other moving bodies is predicted, it is difficult to predict the movement of the other moving bodies with high accuracy with the current technology.

[0007] One of the objects of the aspects of the present application is to provide a moving body control device, a moving body, a moving body control method, a program, and a learning device that enable a moving body to take an action with high affinity to other moving bodies around the moving body without predicting the future movement of the other moving bodies.

[0008] The moving body control device of the first aspect of the present application includes a path decision section that decides a path of the moving body in such a manner as to reduce changes in movement vectors of other moving bodies present around the moving body, and a control section that causes the moving body to move along the path decided by the path decision section.

[0009] The second aspect can also be the moving body control device of the first aspect described above, in which the path decision section decides the path of the moving body in such a manner as to reduce the sum of changes in movement vectors of a plurality of the other moving bodies.

[0010] The third aspect can also be the moving body control device of the first aspect or the second aspect described above, in which the path decision section decides the path of the moving body in such a manner that the value of a reward function that keeps changes in movement vectors of the other moving bodies as independent variables is a good value.

[0011] The fourth aspect can be the mobile body control device of any one of the first to third aspects, wherein the path determination unit determines the path of the own mobile body so that the own mobile body does not enter a region that is larger in a direction in which the movement vector of the other mobile body is directed and is smaller in a side direction and an opposite direction of the direction in which the movement vector of the other mobile body is directed.

[0012] The fifth aspect of the present application is a mobile body including the mobile body control device of any one of the first to fourth aspects, a surrounding detection unit that detects a surrounding environment, a work unit that provides a predetermined service to a user, and a drive unit that is controlled by the mobile body control device and moves the mobile body, the mobile body control device outputting a control parameter that moves the mobile body by inputting a state of the other mobile body based on the surrounding environment.

[0013] The sixth aspect of the present application is a mobile body control method that causes a computer to perform processes of determining a path of an own mobile body so as to reduce a change in a movement vector of an other mobile body existing in a surrounding of the own mobile body and moving the own mobile body along the path.

[0014] The seventh aspect of the present application is a storage medium storing a program that causes a computer to perform processes of determining a path of an own mobile body so as to reduce a change in a movement vector of an other mobile body existing in a surrounding of the own mobile body and moving the own mobile body along the path.

[0015] The eighth aspect of the present application is a learning device including a simulation unit that simulates movement actions of an own mobile body and an other mobile body, an evaluation unit that applies a reward function to a processing result of the simulation unit and evaluates at least the movement action of the own mobile body, and a learning unit that learns based on an evaluation result of the evaluation unit, the evaluation unit evaluating the movement action of the own mobile body more highly as the change in the movement vector of the other mobile body is smaller.

[0016] The ninth aspect of the present application can be the learning device of the eighth aspect, wherein the evaluation unit evaluates the movement action of the own mobile body more lowly in a case where the own mobile body enters a region that is larger in a direction in which the movement vector of the other mobile body is directed and is smaller in a side direction and an opposite direction of the direction in which the movement vector of the other mobile body is directed.

[0017] According to the first to third aspects, the fifth to seventh aspects, it is possible to move the mobile body so as not to hinder the movement of the other mobile body without predicting the future action of the other mobile body existing in the surrounding. As a result, it is possible to make the mobile body take an action that is highly friendly to the other mobile body existing in the surrounding.

[0018] According to the above-described fourth aspect, it is possible to determine the path of the mobile body while taking into account the personal space.

[0019] According to the above-described first aspect to the seventh aspect, it is possible to cause the mobile body to take an action that is more friendly to the other mobile bodies in the surroundings without predicting the future actions of the other mobile bodies in the surroundings.

[0020] According to the above-described eighth aspect, it is possible to perform learning in such a manner that the movement of the other mobile bodies is not hindered as much as possible without predicting the future actions of the other mobile bodies in the surroundings. As a result, it is possible to generate a policy that causes the mobile body to take an action that is more friendly to the other mobile bodies in the surroundings.

[0021] According to the above-described ninth aspect, it is possible to perform learning while taking into account the personal space.

[0022] According to the above-described eighth aspect to the ninth aspect, it is possible to perform learning for causing the mobile body to take an action that is more friendly to the other mobile bodies in the surroundings without predicting the future actions of the other mobile bodies in the surroundings. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 is a schematic diagram showing the system structure of an embodiment.

[0024] Figure 2 is a block diagram of a learning device.

[0025] Figure 3 is a graph for explaining a reward function R3.

[0026] Figure 4 is a graph for explaining a reward function R4.

[0027] Figure 5 is a flowchart showing an example of the process of the learning process of reinforcement learning performed by the learning device.

[0028] Figure 6 is a block diagram of a mobile body. DETAILED DESCRIPTION

[0029] Hereinafter, with reference to the accompanying drawings, an embodiment of a mobile body control device, a mobile body, a mobile body control method, a program, and a learning device of the present application will be described.

[0030] [Learning Device]

[0031] Figure 1is a schematic diagram showing a system configuration of an embodiment. A mobile body control system 1 is provided with a learning device 100 and a mobile body 200. The learning device 100 is implemented by one or more processors. The learning device 100 is a device that determines an action by computer simulation for a plurality of mobile bodies, derives or acquires a reward based on a change in an environment resulting from the action, and learns an action (movement) that maximizes the reward. The action refers to, for example, movement within a simulation space. Actions other than movement can also be set as learning targets, but in the following description, the action refers to movement. A simulator (simulation section described later) that determines movement can also be executed in a device different from the learning device 100, but in the following description, the simulator is executed by the learning device 100. The learning device 100 is pre-stored with environmental information such as map information that becomes a premise of simulation. The learning result of the learning device 100 is mounted on the mobile body 200 as an action determination model MD.

[0032] Figure 2 is a block diagram of the learning device 100. The learning device 100 is provided with, for example, a learning section 110, a simulation section 120, and an evaluation section 130. The learning device 100 is a device that inputs a movement target generated by a self-agent (becomes a self-mobile body in the mobile body 200) in order to reach a certain destination and a position, a moving direction, and a moving speed of other agents (other mobile bodies) into a policy, performs reinforcement learning of updating the policy based on a result of evaluating a change in a state (a change in an environment) resulting therefrom, and outputs the policy after learning.

[0033] The self-agent refers to a hypothetical movement subject that envisions a mobile body such as a robot or a vehicle. The other agent also similarly refers to a hypothetical movement subject that envisions a mobile body such as a robot or a vehicle. The policy is also used in determination of the movement of the other agent, but the policy of the other agent can or can not be updated.

[0034] The learning unit 110, the simulation unit 120, and the evaluation unit 130 are realized, for example, by a hardware processor such as a CPU (Central Processing Unit) executing a program (software). The program can be stored in advance in a storage device (non-transitory storage medium) such as an HDD (Hard Disk Drive), a flash memory, or the like, or can be stored in a removable storage medium (non-transitory storage medium) such as a DVD (Digital Versatile Disc), a CD-ROM (Read Only Memory), or the like, and installed by mounting the storage medium in a drive device. Some or all of these components can be realized by hardware (including circuitry) such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), a GPU (Graphics Processing Unit), or the like, or can be realized by a combination of software and hardware.

[0035] The learning unit 110 updates the policy in accordance with various reinforcement learning algorithms based on the evaluation result obtained by the evaluation unit 130 evaluating the state change generated by the simulation unit 120 and the collision determination result.

[0036] The simulation unit 120 inputs the action target and the previous state (if it is just after the start of simulation, the initial state) into the policy, and derives a state change as a result of the action of the own agent and other agents. The policy is, for example, a DNN (Deep Neural Network), but can also be a rule-based policy or another form of policy. The policy is a policy that derives a probability of occurrence for each of a plurality of assumed actions. For example, in a simple example, as an assumption that the plane expands upward and downward to the left and right, a result of outputting right movement 80%, left movement 10%, up movement 10%, and down movement 0% is output. The simulation unit 120 causes a random number to act on this result, and derives a state change of the agent so that if the random number value is 0% or more and less than 80%, it is right movement, if the random number value is 80% or more and less than 90%, it is left movement, and if the random number value is 90% or more, it is up movement.

[0037] The evaluation unit 130 calculates a value (reward function value) of a reward function R that evaluates the state change of the own agent output by the simulation unit 120, and evaluates the action of the own agent.

[0038] As shown in equation (1), the reward function R includes: a reward function R1 assigned when the agent reaches its destination; a reward function R2 assigned when the agent successfully moves; a reward function R3 that decreases when the agent causes a change in the movement vector of other agents; and a reward function R4 that varies in distance when the agent approaches other agents based on the direction they are facing. Reward function R3 is an example of the first reward function, and reward function R4 is an example of the second reward function.

[0039] R = R1 + R2 + R3 + R4…(1)

[0040] The reward function R1 is a fixed positive value when the destination is reached, and a value proportional to the change in distance to the destination when the destination is not reached (positive if the distance change is in the direction of decrease, and negative if the distance change is in the direction of increase).

[0041] The reward function R² is a function that is, the smaller the third derivative of the agent's position in the two-dimensional plane, i.e., the jerk (jump), the larger the reward function R².

[0042] Figure 3 This is a diagram used to illustrate the reward function R3. The reward function R3, calculated at time t (control loop), is the following function: a′ represents the movement vector a′ of the other agents from the state of other agents at time t-1 to time t. i,t (The movement vectors of other agents assuming the absence of this agent) and the movement vectors of other agents from the state of other agents at time t-1 to time t. i,t The evaluation function R3 compares the movement vectors of other agents (assuming the presence of this agent) and the larger the difference between them, the lower the evaluation value of the result. In other words, the reward function R3 is a function that evaluates the agent's actions more highly if the agent does not cause changes to the movement vectors of other agents in the vicinity. The reward function R3 is an objective function that retains the changes in the movement vectors of other agents as independent variables; for example, a larger value indicates a better value. The evaluation unit 130 can also derive its own movement vectors a′ of other agents from the state of other agents at time t-1 to time t under the assumption that this agent does not exist. i,t Alternatively, you can request the simulation department 120 to export it.

[0043]

[0044] In equation (2), W is a negative coefficient, or in other words, the reward function R3 is a function that returns a lower evaluation value as the value after ∑ increases. i,tis the movement vector of each other agent from time t-1 to time t (in the presence of the present agent). a' is the movement vector of each other agent from time t-1 to time t (in the absence of the present agent). i is the identification number of the other agent, and N is the number of all other agents present. i,t is the movement vector of each other agent from time t-1 to time t (in the absence of the present agent). i is the identification number of the other agent, and N is the number of all other agents present.

[0045] In Figure 3 , the agent H is the present agent, and the agents A1 to A5 are other agents. For example, at time t, the other agent A1 moves with a movement vector of a 1,t , the other agent A2 moves with a movement vector of a 2,t , the other agent A3 moves with a movement vector of a 3,t , the other agent A4 moves with a movement vector of a 4,t , and the other agent A5 moves with a movement vector of a 5,t . In contrast, returning to the state at time t-1, for the movement vector in the case where the present agent H is assumed to be absent, a' is indicated for the other agent A1, a' is indicated for the other agent A2, a' is indicated for the other agent A3, a' is indicated for the other agent A4, and a' is indicated for the other agent A5. 1,t 2,t 3,t 4,t 5,t

[0046] Figure 4 is a diagram for explaining the reward function R4. The reward function R4 is a function that returns a lower evaluation value in the case where the present agent enters a prescribed region. The region around the other agent A is divided into the following four regions (spaces) for consideration. For example, assume that it is divided into a close space surrounded by the boundary line D1, a personal space surrounded by the boundary line D1 and the boundary line D2, a social space surrounded by the boundary line D2 and the boundary line D3, and a public space surrounded by the boundary line D3 and the boundary line D4.

[0047] In the present embodiment, for example, the reward function R4 is a function that returns a lower evaluation value in the case where the external boundary line, that is, D2, of the personal space among these spaces is entered. The personal space, like the social space and the public space, is wider with respect to the direction (or movement direction) in which the other agent A is directed (or moves) and is narrower with respect to other directions. Thus, a lower evaluation is given to the other agent A for such an action as looking ahead, and a less low evaluation is given to the other agent A for such an action as looking aside or behind.

[0048] ​​​​​The evaluation unit 130 can determine that the self-agent has collided with the other agent when the coordinates of the self-agent and the other agent coincide, and can determine that the self-agent has collided with the other agent when the self-agent has entered the personal space of the other agent. When it is determined that a collision has occurred, the evaluation unit 130 ends the episode, initializes the states of the agents, and starts the next episode. The evaluation unit 130 outputs the collision determination result and the action evaluation result to the learning unit 110. The details are described using a flowchart.

[0049] Figure 5 is a flowchart illustrating an example of the process of the learning process of the reinforcement learning performed by the learning device 100.

[0050] First, the simulation unit 120 receives the action target of the self-agent from the learning device 100 (step S200). Next, the learning device 100 simulates the actions of the agents for one cycle amount using the action target as one of the inputs (step S202).

[0051] Next, the evaluation unit 130 determines whether the self-agent has collided with the other agent (step S204). When it is determined that the self-agent has not collided with the other agent in the surroundings, the evaluation unit 130 evaluates the action of the self-agent using the reward function R (step S206), and outputs the evaluation result to the learning unit 110.

[0052] Next, the learning unit 110 updates the policy in accordance with the reinforcement learning algorithm based on the evaluation result of the evaluation unit 130 (step S208). The updated policy of the learning unit 110 is output to the simulation unit 120, and the simulation unit 120 simulates the actions of the agents for the next cycle using the received policy.

[0053] Next, the learning device 100 determines whether the update amount of the parameters of the policy for each time is below a threshold value based on the change in the state as a result of the actions of the self-agent and the other agents (step S210). The update amount of the parameters herein refers to, for example, the amount of change in the parameters such as the movement vector of the self-agent in the n-th time from the parameters such as the movement vector of the self-agent in the n-1-th time, the sum of the absolute values of the amounts of change in the parameters, or the like. When the update amount of the parameters of the policy is below a certain threshold value m, that is, when the parameters of the policy have not substantially changed, the learning device 100 ends the process of the learning process. When the update amount of the parameters of the policy is not below the certain threshold value m, the learning device 100 returns to step S202.

[0054] Alternatively, the process of the learning process can be ended when the process for a predetermined number of cycles is ended.

[0055] In a case where it is determined in step S204 that the present agent has collided with another agent in the periphery, the evaluation unit 130 outputs the determination result to the learning unit 110, and lowers the evaluation value of the reward function (step S212). Then, the evaluation unit 130 outputs the evaluation result to the learning unit 110, and the learning unit 110 updates the policy based on the evaluation result of the evaluation unit 130 (step S214). Further, the learning device 100 initializes the state of each agent, and returns to step S202.

[0056] According to the learning device 100 described above, it is possible to generate an action decision model (policy) by reinforcement learning in a manner that does not hinder the actions of other moving bodies in the periphery as much as possible. Thereby, in the moving body control device 250 that adopts the action decision model, it is possible to cause the moving body 200 to take an action that is highly compatible with the actions of other moving bodies in the periphery.

[0057] [moving body]

[0058] Figure 6 is a configuration diagram of the moving body 200. The moving body 200, for example, is provided with a moving body control device 250, a periphery sensing device 210, a moving body sensor 220, a work unit 230, and a drive device 240. The moving body 200 can be a vehicle, or a device such as a robot. The moving body control device 250, the periphery sensing device 210, the moving body sensor 220, the work unit 230, and the drive device 240 are connected to each other through a multiplex communication line such as a CAN (Controller Area Network) communication line, a serial communication line, a wireless communication network, or the like.

[0059] The periphery sensing device 210 is a device for sensing the environment of the periphery of the moving body 200, and the actions of other moving bodies in the periphery. The periphery sensing device 210, for example, is provided with a GPS receiver, a positioning device including map information, a radar device, and an object recognition device such as a camera. The positioning device performs positioning of the position of the moving body 200, and matches the position with map information. The radar device radiates an electric wave such as a millimeter wave to the periphery of the moving body 200, and detects an electric wave (reflected wave) reflected by an object to at least detect the position (distance and direction) of the object. The radar device can also detect the position and movement vector of the object. The camera is, for example, a digital camera that uses a solid-state imaging element such as a CCD (Charge Coupled Device) or a CMOS (Complementary Metal Oxide Semiconductor), and is attached with an image processing device that recognizes the position of an object from a captured image. The periphery sensing device 210 outputs information such as the position of the moving body 200 on a map, and the positions of objects (including other moving bodies equivalent to the aforementioned other agents) existing in the periphery of the moving body 200 to the moving body control device 250.

[0060] The mobile body sensor 220 includes, for example, a speed sensor that detects the speed of the mobile body 200, an acceleration sensor that detects acceleration, a yaw rate sensor that detects the angular velocity around the vertical axis, a direction sensor that detects the orientation of the mobile body 200, and the like. The mobile body sensor 220 outputs the detected results to the mobile body control device 250.

[0061] The work section 230 is, for example, a device that provides a prescribed service to a user. The service here refers to, for example, a work such as loading and unloading of goods and the like to and from a conveyance apparatus. The work section 230 includes, for example, a Magic arm, a stage, a microphone, a speaker, and the like HMI (Human machine Interface), and the like. The work section 230 acts in accordance with the content instructed by the mobile body control device 250.

[0062] The drive device 240 (drive section) is a device for moving the mobile body 200 in a desired direction. In the case where the mobile body 200 is a robot, the drive device 240 includes, for example, two or more leg sections and actuators. In the case where the mobile body 200 is a vehicle, a micro mobile body, or a robot that moves by wheels, the drive device 240 includes wheels (steering wheels, drive wheels) and a motor, an engine, or the like for rotating the wheels.

[0063] The mobile body control device 250 includes, for example, a path decision section 252, a control section 254, and a storage section 256. The path decision section 252 and the control section 254 are each realized by, for example, a hardware processor such as a CPU executing a program (software).

[0064] The program can be stored in advance in a storage device (non-transitory storage medium) such as an HDD, a flash memory, or the like, or in a removable storage medium (non-transitory storage medium) such as a DVD, a CD-ROM, or the like, and installed by mounting the storage medium to a drive device. Some or all of these constituent elements can be realized by hardware (including circuitry) such as an LSI, an ASIC, an FPGA, a GPU, or the like, or by a cooperation of software and hardware.

[0065] The storage section 256 is, for example, an HDD, a flash memory, a RAM, a ROM, or the like. Information such as a behavior decision model MD256A is stored in the storage section 256. The behavior decision model MD256A is a model based on the policy at the final time point of the learning phase processing generated by the learning device 100.

[0066] The path decision unit 252, for example, inputs information (state of the object) such as the position of the mobile body 200 on the map, the position of the object existing in the periphery of the mobile body 200, and the like detected by the periphery detection device 210, and furthermore, information of the destination input by the user to the action decision model MD256A, thereby deciding the position where the mobile body 200 should advance next. The path decision unit 252 decides the path of the mobile body 200 in sequence by repeatedly performing the above processing.

[0067] The control unit 254 controls the driving device 240 in such a manner that the mobile body 200 moves on the path decided by the path decision unit 252.

[0068] According to the mobile body control device 250 described above, the path of the mobile body 200 is generated based on the action decision model (policy) generated by reinforcement learning in such a manner as to as little as possible hinder the actions of other mobile bodies in the periphery, and the mobile body 200 is caused to move along the path, so that the mobile body 200 can take an action with high affinity to the actions of other mobile bodies in the periphery.

[0069] In the present embodiment, it is assumed that the update of the policy is performed only during the learning stage, and the update of the policy is not performed after the mobile body is mounted, but learning can be continued after the mobile body is mounted.

[0070] The specific embodiments of the present application are described above using the embodiments, but the present application is not at all limited to such embodiments, and various modifications and substitutions can be added within the scope of the gist of the present application.

[0071] The above-described embodiments can be expressed as follows.

[0072] A mobile body control device includes:

[0073] a storage device that stores a program; and

[0074] a hardware processor connected to the storage device,

[0075] the hardware processor executes the program to decide a path of the present mobile body in such a manner as to reduce a change in a movement vector of other mobile bodies existing in the periphery of the present mobile body, and cause the present mobile body to move along the decided path.

[0076] The above-described embodiments can be expressed as follows.

[0077] A learning device includes:

[0078] a storage device that stores a program; and

[0079] a hardware processor connected to the storage device,

[0080] The program is executed by the hardware processor to simulate movement of the mobile body and movement of other mobile bodies, a reward function is applied to a result of the simulation, movement of the mobile body is evaluated at least based on a result of the application of the reward function, and learning is performed based on a result of the evaluation, and when the evaluation is performed, the smaller the change in the movement vector of the other mobile bodies, the higher the movement of the mobile body is evaluated.

Claims

1. A mobile body control device comprising: a path decision section that decides a path of a self mobile body in a manner to reduce changes in movement vectors of other mobile bodies existing in a periphery of the self mobile body, based on output results output by a model, by inputting at least position information of the other mobile bodies in the periphery of the self mobile body and information indicating a movement direction of the self mobile body to the model; and a control section that moves the self mobile body along the path decided by the path decision section, wherein the model is a model that learns evaluation results obtained by evaluating movement actions of the self mobile body at least by applying a reward function R to processing results obtained by simulating the movement actions of the self mobile body and the other mobile bodies, respectively, whereby the smaller the changes in the movement vectors of the other mobile bodies, the higher the evaluation of the movement actions of the self mobile body, wherein the reward function R is calculated by the following equation: R = Rl + R2 + R3 + R4 wherein Rl is a reward function given in a case where the self mobile body reaches a destination; R2 is a reward function given in a case where the self mobile body successfully achieves movement; R3 is a reward function that becomes smaller in a case where the self mobile body causes changes in the movement vectors of the other mobile bodies; and R4 is a reward function that makes a distance to be maintained by the self mobile body when the self mobile body approaches the other mobile bodies variable according to a direction in which the other mobile bodies are heading.

2. The mobile body control device according to claim 1, wherein the path decision section decides the path of the self mobile body in a manner to reduce a sum of the changes in the movement vectors of a plurality of the other mobile bodies.

3. The mobile body control device according to claim 1 or 2, wherein the path decision section decides the path of the self mobile body in a manner to make a value of the reward function R that holds the changes in the movement vectors of the other mobile bodies as independent variables a good value.

4. The mobile body control device according to claim 1 or 2, wherein the path decision section decides the path of the self mobile body in a manner that the self mobile body does not enter an area that is larger in a direction in which the movement vectors of the other mobile bodies are heading and smaller in a side direction and an opposite direction of the direction in which the movement vectors of the other mobile bodies are heading.

5. A mobile body comprising: the mobile body control device according to any one of claims 1 to 4; a periphery sensing device that senses a periphery environment; a work section for providing a prescribed service to a user; and a drive section that is controlled by the mobile body control device and moves the mobile body, wherein the mobile body control device outputs a control parameter that moves the mobile body by inputting a state of the other mobile bodies based on the periphery environment. wherein 6. A mobile body control method comprising: causing a computer to perform the following processing: deciding a path of a self mobile body in a manner to reduce changes in movement vectors of other mobile bodies existing in a periphery of the self mobile body, based on output results output by a model, by inputting at least position information of the other mobile bodies in the periphery of the self mobile body and information indicating a movement direction of the self mobile body to the model; and moving the self mobile body along the path decided. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ moving the self moving body along the path, the model is a model that learns evaluation results obtained by evaluating at least the movement action of the self moving body by applying a reward function R to processing results obtained by simulating the movement actions of the self moving body and the other moving bodies respectively, whereby the smaller the change in the movement vector of the other moving body, the higher the evaluation of the movement action of the self moving body, the reward function R is calculated by the following equation: where R1 is a reward function given in the case where the self moving body reaches the destination; R2 is a reward function given in the case where the self moving body successfully achieves movement; R3 is a reward function that becomes smaller in the case where the self moving body causes a change in the movement vector of the other moving body; and R4 is a reward function in which the distance to be maintained when the self moving body approaches the other moving body is variable depending on the direction in which the other moving body is facing.

7. A storage medium storing a program, wherein the program causes a computer to perform the following processing: determining a path of a self moving body in a manner to reduce a change in a movement vector of other moving bodies existing in the periphery of the self moving body, based on output results output by a model by at least inputting position information of the other moving bodies in the periphery of the self moving body and information indicating a movement direction of the self moving body to the model; and moving the self moving body along the path, the model is a model that learns evaluation results obtained by evaluating at least the movement action of the self moving body by applying a reward function R to processing results obtained by simulating the movement actions of the self moving body and the other moving bodies respectively, whereby the smaller the change in the movement vector of the other moving body, the higher the evaluation of the movement action of the self moving body, the reward function R is calculated by the following equation: where R1 is a reward function given in the case where the self moving body reaches the destination; R2 is a reward function given in the case where the self moving body successfully achieves movement; R3 is a reward function that becomes smaller in the case where the self moving body causes a change in the movement vector of the other moving body; and R4 is a reward function in which the distance to be maintained when the self moving body approaches the other moving body is variable depending on the direction in which the other moving body is facing.

8. A learning device comprising: a simulation unit that simulates movement actions of a self moving body and other moving bodies respectively; an evaluation unit that evaluates at least the movement action of the self moving body by applying a reward function R to processing results of the simulation unit; and a learning unit that learns based on evaluation results of the evaluation unit, the smaller the change in the movement vector of the other moving body, the higher the evaluation of the movement action of the self moving body by the evaluation unit, the reward function R is calculated by the following equation: where R1 is a reward function given in the case where the self moving body reaches the destination; R2 is a reward function given in the case where the self moving body successfully achieves movement; R3 is a reward function that becomes smaller in the case where the self moving body causes a change in the movement vector of the other moving body; and R4 is a reward function in which the distance to be maintained when the self moving body approaches the other moving body is variable depending on the direction in which the other moving body is facing. wherein R1 is a reward function given in the case where the mobile body reaches a destination; R2 is a reward function given in the case where the mobile body successfully achieves movement; R3 is a reward function that becomes smaller in the case where the mobile body causes a change in the movement vector of another mobile body; and R4 is a reward function in which a distance that the mobile body should maintain when approaching another mobile body is variable according to a direction in which the other mobile body is facing.

9. The learning device according to claim 8, wherein The evaluation section evaluates the movement action of the mobile body to be lower in the case where the mobile body enters a region that is larger in a direction in which the movement vector of the other mobile body is facing and smaller in a side direction and an opposite direction of the direction in which the movement vector of the other mobile body is facing.

Citation Information

Patent Citations

  • Route determining device, robot, and route determining method

    CN111736592A