Behavioral evaluation presentation system, behavioral evaluation presentation method, behavioral evaluation presentation device, and program

The behavior evaluation presentation system addresses the lack of route reasoning clarity by planning and predicting environmental states for autonomous robots, allowing users to understand the robot's actions through mapped evaluations.

JP7794296B2Active Publication Date: 2026-01-06NEC CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024511090
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2026-01-06
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

Users monitoring autonomous mobile robots cannot understand the reasoning behind the selected route, especially when multiple routes are available, as existing technologies do not provide clear insights into the environmental changes and risks over time.

Method used

A behavior evaluation presentation system that plans multiple routes, predicts environmental states based on elapsed time, evaluates the behavior at each area, and superimposes these evaluations on a map for user understanding.

Benefits of technology

Enables users to recognize the basis for the autonomous mobile robot's route selection by presenting the predicted environmental states and risks, enhancing user understanding of the robot's actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007794296000001
    Figure 0007794296000001
  • Figure 0007794296000002
    Figure 0007794296000002
  • Figure 0007794296000003
    Figure 0007794296000003
Patent Text Reader

Abstract

This behavioral assessment presentation system (1) comprises: a route planning means (101) which plans a first route along which a moving body moves to a first area on a map and a second route along which the moving body moves to a second area; a state prediction means (102) which predicts the environmental state of the first area in response to a passage time estimated to be required for the moving body to move along the first route and the environmental state of the second area in response to a passage time estimated to be required for the moving body to move along the second route; an assessment means (103) which assesses, on the basis of the predicted environmental state of the first area, the behavior of the moving body that moves to the first area and assesses, on the basis of the predicted environmental state of the second area, the behavior of the moving body that moves to the second area; and a presentation means (104) which overlaps the assessment in the first area and the assessment in the second area on the map and presents the overlapped result.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a behavior evaluation presentation system, a behavior evaluation presentation method, and a behavior evaluation presentation device. [Background technology]

[0002] Devices (e.g., autonomous mobile robots) that behave based on the results of reinforcement learning have been developed.

[0003] Patent Document 1 describes a predictive behavior decision device that acquires environmental state values ​​and determines its own behavior based on the results of predicting changes in the environmental state. Patent Document 2 describes a device that creates a risk map by applying classified traffic participants to previously learned apparent risks, and controls its own vehicle based on the created risk map. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] International Publication No. 2004 / 068399 [Patent Document 2] Japanese Patent Application Publication No. 2019-106049 Summary of the Invention [Problem to be solved by the invention]

[0005] With the techniques disclosed in Patent Documents 1 and 2, a user monitoring an autonomous mobile robot may not be able to understand the reasoning behind the route the autonomous mobile robot has selected. For example, if there are multiple routes to a destination, the user may not be able to understand the reasoning behind the autonomous mobile robot selecting one of the multiple routes.

[0006] The present disclosure has been made to solve such problems, and aims to provide a behavior evaluation presentation system, method, device, etc. that presents the basis for a moving object's route selection. [Means for solving the problem]

[0007] A behavior evaluation presentation system according to one aspect of the present disclosure includes: a route planning means for planning a first route along which a moving object moves to a first area on a map and a second route along which the moving object moves to a second area on the map; predicting an environmental state of the first area according to an estimated elapsed time required for the moving object to move along the first route; a state prediction means for predicting an environmental state of the second area in accordance with an estimated elapsed time required for the moving object to move along the second route; an evaluation means for evaluating the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, and for evaluating the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area; a presentation means for presenting the evaluation in the first area and the evaluation in the second area by superimposing them on the map; Equipped with.

[0008] A behavioral evaluation presentation method according to one aspect of the present disclosure includes: planning a first route for a moving object to move to a first area on a map and a second route for the moving object to move to a second area on the map; predicting an environmental state of the first area according to an estimated elapsed time required for the moving object to move along the first route; predicting an environmental state of the second area according to an estimated elapsed time required for the moving object to move along the second route; Evaluating the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, and evaluating the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area; The evaluation in the first area and the evaluation in the second area are presented superimposed on the map.

[0009] An action evaluation presentation device according to one aspect of the present disclosure includes: a route planning unit that plans a first route along which a moving object moves to a first area on a map, and a second route along which the moving object moves to a second area on the map; predicting an environmental state of the first area according to an estimated elapsed time required for the moving object to move along the first route; a state prediction means for predicting an environmental state of the second area in accordance with an estimated elapsed time required for the moving object to move along the second route; an evaluation means for evaluating the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, and for evaluating the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area; a presentation means for presenting the evaluation in the first area and the evaluation in the second area by superimposing them on the map; Equipped with. [Effects of the Invention]

[0010] According to the present disclosure, it is possible to provide a behavior evaluation presentation system, method, device, etc. that presents the basis for a moving object's route selection. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a block diagram showing a configuration of a behavior evaluation presentation system according to a first embodiment. [Figure 2] 4 is a flowchart showing the operation of the behavior evaluation presentation system according to the first embodiment. [Figure 3] FIG. 10 is a block diagram showing an example of the configuration of a behavior evaluation presentation system according to a second embodiment. [Figure 4] FIG. 10 is a diagram illustrating a method for visualizing on a map the basis for route selection of a moving object according to the second embodiment. [Figure 5] FIG. 10 is a diagram for explaining an example of creating a judgment criterion map based on evaluation in a toy model according to the second embodiment. [Figure 6] FIG. 10 is a diagram illustrating an example of calculating a state value according to the second embodiment. [Figure 7] FIG. 10 is a diagram illustrating an example of creating a determination criterion map according to the second embodiment. [Figure 8] FIG. 10 is a diagram for explaining a map creation method that takes time into consideration through simulation according to the second embodiment. [Figure 9] FIG. 10 is a diagram illustrating a partial observation system in another embodiment. [Figure 10] 10A and 10B are diagrams illustrating a route planning method by a route planning unit according to another embodiment. [Figure 11] 1 is a block diagram showing an example of the hardware configuration of a behavior evaluation presentation device, etc. DETAILED DESCRIPTION OF THE INVENTION

[0012] Embodiment 1 Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In this embodiment, multiple paths of a moving object are calculated, and an evaluation of the behavior of the moving object is calculated taking into consideration changes in the environment over time as the moving object moves along the paths, and the evaluation is presented to the user. The moving object may be any of a variety of mobile devices, such as an autonomous mobile robot or an autonomous mobile vehicle.

[0013] The behavior evaluation presentation system 1 can be realized by one or more computers. The computer can include a memory, a processor, etc. The behavior evaluation presentation system 1 can be used by a user who is monitoring the autonomous movement of a mobile object. The behavior evaluation presentation system 1 includes a route planning unit 101, a state prediction unit 102, an evaluation unit 103, and a presentation unit 104. Some of the components may be provided on a cloud computer connected via a network.

[0014] The route planning unit 101 is also called a route planning means. The route planning unit 101 plans multiple routes to multiple areas. Specifically, the route planning unit 101 plans a first route along which a mobile object moves to a first area on a map, and a second route along which the mobile object moves to a second area on the map. The map may be an overall map that can indicate areas in which the mobile object can move or areas in which the mobile object cannot move. The map may be provided by a user, or may be generated from information collected by the above-mentioned sensor unit (e.g., a camera, LiDAR).

[0015] The state prediction unit 102 is also called a state prediction means. The state prediction unit 102 executes a simulation regarding the movement of the mobile object according to elapsed time, changes in the state of the environment, etc. The state prediction unit 102 predicts the state of the environment of the first area according to the elapsed time estimated to be required for the movement of the mobile object along the first route, and predicts the state of the environment of the second area according to the elapsed time estimated to be required for the movement of the mobile object along the second route.

[0016] The evaluation unit 103 is also called an evaluation means. The evaluation unit 103 evaluates the behavior of the mobile object that has moved to the first area based on the predicted environmental state of the first area, and evaluates the behavior of the mobile object that has moved to the second area based on the predicted environmental state of the second area.

[0017] The presentation unit 104 is also called presentation means. The presentation unit 104 may be, for example, any display device for presenting a map to a user. The presentation unit 104 presents the evaluation in the first area and the evaluation in the second area by superimposing them on the map. The map showing the evaluation of the behavior of the mobile object in each area is also called a judgment criterion map.

[0018] FIG. 2 is a flowchart illustrating the operation of the behavior evaluation presentation system according to the first embodiment. The route planning unit 101 plans a first route along which a moving object moves to a first area on a map, and a second route along which the moving object moves to a second area (step S101).

[0019] The state prediction unit 102 predicts the environmental state of the first area according to the elapsed time estimated to be required for the moving object to move along the first route, and predicts the environmental state of the second area according to the elapsed time estimated to be required for the moving object to move along the second route (step S102). For example, if a flood is spreading in the environment in which the moving object moves, the state prediction unit 102 can predict whether the environmental state of the first area will become flooded according to the elapsed time estimated to be required for the moving object to move along the first route. Also, the state prediction unit 102 can predict whether the environmental state of the second area will become flooded according to the elapsed time estimated to be required for the moving object to move along the second route.

[0020] The evaluation unit 103 evaluates the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, and evaluates the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area (step S103). The evaluation unit 103 can evaluate the behavior of the moving object that has moved to the first area, for example, using a learned value function and the environmental state of the first area according to elapsed time.

[0021] The presentation unit 104 presents the evaluation in the first area and the evaluation in the second area by superimposing them on the map (step S104).

[0022] According to the first embodiment described above, it is possible to predict the states of the first and second regions, which are the destinations, in consideration of the elapsed time required for the movement of the moving object, and to evaluate the behavior of the moving object to each region. Furthermore, by presenting the evaluation results superimposed on a map, the user can recognize the basis for the autonomous movement of the moving object.

[0023] Embodiment 2 We consider having an autonomous mobile robot (also called a mobile body) perform the task of reconnaissance without being detected by a third party. The autonomous mobile robot generates a route that takes into account the risk of being detected by a third party and performs reconnaissance. In this case, the aim is to present a judgment criteria map (also called a risk map) so that the user monitoring the autonomous mobile robot's behavior can recognize which locations are dangerous and to what extent, as a basis for the autonomous mobile robot's behavior. Note that the risk referred to here does not simply refer to the risk or safety of the robot being in a certain location, but rather to the risk or safety of moving from one location to another.

[0024] For the control of an autonomous mobile robot, reinforcement learning is performed in a simulated environment created by a simulator, with the robot's coordinates and surrounding information as inputs and the robot's speed and direction as outputs. The value function and policy obtained through reinforcement learning can clarify the local risk of performing a certain action, i.e., which direction and speed are most valuable. Methods such as Grad-CAM (Gradient-weighted Class Activation Mapping) can determine which parts of the map are important overall, but because they do not take time changes into account, it is unclear how the important parts affect the overall map. The risk beyond a local area, such as which locations within the map are dangerous, is unclear. Therefore, the reason why the autonomous mobile robot is performing its current action is unclear to the monitoring user. Furthermore, the contribution to the value level and changes in the environment over time are also unclear to the monitoring user. Therefore, the present disclosure relates to improvements in the presentation of mobile object control and path planning learned through reinforcement learning to users.

[0025] FIG. 3 is a block diagram illustrating an example of the configuration of a behavior evaluation presentation system according to the second embodiment. The behavior evaluation presentation system 1 includes a control unit 150 mounted on an autonomous mobile robot and a behavior evaluation presentation device 100. The behavior evaluation presentation device 100 can be used, for example, by a user monitoring the autonomous mobile robot. The behavior evaluation presentation device 100 can be implemented by a computer including a memory, a processor, a display, and the like. The behavior evaluation presentation device 100 includes a path planning unit 101, a state prediction unit 102, an evaluation unit 103, a presentation unit 104, and a map storage unit 110. Note that the configuration of the behavior evaluation presentation device 100 is merely exemplary, and some of the components may be provided outside the device. Some of the components may be provided on different network devices or on a cloud computer connected via a network. Furthermore, the components may be connected to each other via a wired or wireless network.

[0026] The control unit 150 is provided in the mobile object and acquires information about the environment around the mobile object (i.e., the current situation) using a state acquisition unit 151. The control unit 150 is configured by a computer including a memory, a processor, etc. The control unit 150 can also control the movement of the mobile object. The state acquisition unit 151 acquires the position of the mobile object on a map (e.g., coordinates on a map) at a specific time point (e.g., the current time point) and information about the surrounding environment of the mobile object via a sensor unit (e.g., a camera, LiDAR (Light Detection and Ranging or Laser Imaging Detection and Ranging)). The sensor unit may be a camera provided in the mobile object, or a surveillance camera installed in the environment where the mobile object moves (e.g., on the ceiling or wall of a room). When a sensor unit (e.g., a surveillance camera) is installed in the environment where the mobile object moves, the state acquisition unit 151 may acquire information about the environment around the mobile object (i.e., the current situation) from the sensor unit via a wireless network. The acquired environmental information is sent to the behavior evaluation presentation device 100, for example, via a wireless network. In some embodiments, the state acquisition unit 151 may be provided within the behavior evaluation presentation device 100. Although not shown, the mobile object may be an autonomous mobile robot, and may be provided with a drive unit or the like to enable autonomous travel. In this specification, the state acquisition unit is also referred to as a state acquisition means.

[0027] The map storage unit 110 stores an overall map (e.g., a grid map) of the environment around the mobile object. The map may be an overall map that can indicate areas where the mobile object can move or areas where the mobile object cannot move. The map may be provided by a user or may be generated from information collected by the sensor unit (e.g., a camera, LiDAR) described above. In some embodiments, the map storage unit 110 may be stored in a memory unit of the mobile object or a memory unit of the behavior evaluation presentation device 100. The map storage unit 110 is also referred to as a map storage means.

[0028] The route planning unit 101 is also called a route planning means. The route planning unit 101 plans a route for the mobile body from one point to another point based on a map. The route planning unit 101 can also calculate a drive control method for the mobile body to another point. The route planning unit 101 can calculate a route from the current location of the mobile body to a target location using an algorithm. This algorithm may be, for example, A*, but is not limited to this, and various algorithms known to those skilled in the art can be used. The target location may be a goal, or may be each point on a map, each area (for example, each grid on a grid map), etc. The route planning unit 101 can calculate not only the route the mobile body actually travels, but also multiple routes to each point on a map that the mobile body can travel.

[0029] The route planning unit 101 plans a plurality of routes for a moving object to travel. The route planning unit 101 plans a first route for a moving object to travel to a first area on a map at a specific time point, and plans a second route for the moving object to travel to a second area at the specific time point.

[0030] The state prediction unit 102 is also called a state prediction means, an estimation means, or an estimation unit. The state prediction unit 102 performs a simulation of the movement and state changes of the mobile object. Based on map information and information on the surrounding environment of the mobile object, the state prediction unit 102 estimates changes in the environmental state, the position to which the mobile object has moved at that time, and the reward to be obtained along the route.

[0031] When the mobile object moves along a planned first route and arrives at a first area at a first time point, the state prediction unit 102 predicts the environmental state of the first area according to the elapsed time associated with the movement from a specific time point to the first time point. Also, when the mobile object moves along the planned second route and arrives at the second area at a second time point, the state prediction unit 102 predicts the environmental state of the second area according to the elapsed time associated with the movement from the specific time point to the second time point.

[0032] The state prediction unit 102 also predicts the state of each area at each time when the moving object moves along the first route and arrives at each area.The state prediction unit 102 also predicts the state of an area at each time when the moving object moves along the second route and arrives at an area on the route.

[0033] The evaluation unit 103 calculates an evaluation value from the current location of the moving object to a certain point (for example, each grid on a grid map). The evaluation unit 103 is also called evaluation means. The method for calculating the evaluation value is as follows. [maxQ(S',*)+R]-maxQ(S,*) Note that maxf(*) represents the maximum value of the arguments in the * part.

[0034] The evaluation unit 103 evaluates the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, and evaluates the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area.

[0035] The evaluation unit 103 evaluates the behavior of the moving object moving to the first area along the first route based on the predicted environmental state of the first area. The evaluation unit 103 can also evaluate the behavior of the moving object moving to the first area based on the predicted environmental state of each area up to the first area.

[0036] Furthermore, the evaluation unit 103 evaluates the behavior of the moving object moving along the second route to the second area based on the predicted environmental state of the second area. Furthermore, the evaluation unit 103 can evaluate the behavior of the moving object moving along the second route to the second area based on the predicted environmental state of each area up to the second area.

[0037] The evaluation unit 103 calculates an evaluation value of the behavior of the moving object in each area of ​​the route based on route information indicating the route of the moving object, the predicted environmental state, and a criterion function indicating an index of the behavior of the moving object. The evaluation unit 103 calculates the evaluation value of the behavior of the moving object based on the equation [V(S')+R]-V(S) relating to the state S of the moving object, the final state S' of the moving object, the criterion function V, and the reward R obtained during the transition from state S to S' on the simulator.

[0038] These evaluation values, rewards obtained on the routes, and multiple routes to each grid are stored in a storage unit (not shown) of the behavior evaluation presentation device 100, and can be acquired as needed during a simulation by the state prediction unit 102 or when the evaluation unit 103 calculates an evaluation value. This storage unit is also called a state holding unit.

[0039] The presentation unit 104 is, for example, a display device, and presents the evaluation results calculated by the evaluation unit 103 to the user by superimposing them on a map. In this specification, the map on which the evaluation results are superimposed is also referred to as a judgment criterion map or a risk map. This judgment criterion map can indicate the degree of risk of any point or area on the map. A user monitoring the moving object can recognize the basis for the autonomous movement of the moving object based on this judgment criterion map.

[0040] Here, with reference to FIG. 4, a method for visualizing on a map the moving objects moving on the map and the basis for route selection will be described. In this embodiment, a value function is required instead of a policy gradient system such as Actor-Critic or Q-learning. The mobile object moves so as to maximize this value function Q. The simulator environment is (S, a → S', r), the value function Q(S, a) obtained by learning, and the algorithm A(s, s') gives the route from the current location s to the position s' on the map.

[0041] In Figure 4, consider reinforcement learning for agent AG, which is currently in state S. For each point s' on the map, algorithm A is used to find each path. Next, agent AG moves along each path on the simulator. The reward obtained during this process is R, and the state after the movement is S'. At this time, the score of each point or area s' on the map is given by equation (1). [maxQ(S',*)+R]-maxQ(S,*)...(1) In other words, [value when moving towards s'] - (current value = value of optimal action) The value of moving towards s' is given as the sum of (the value of the state after the movement) and (the reward obtained up to s'). If the agent AG meets the termination condition during the movement, the score of the point or area where the agent AG is located at the end is given by formula (2). R-maxQ(S,*) (2) In other words, [reward obtained until the end condition is met] - (current value = value of optimal action) The robot is displayed on a map. End conditions can be set arbitrarily based on a variety of conditions, such as when the robot is caught in a flood, when the robot comes into contact with an obstacle, or when the robot comes close to a person.

[0042] Next, with reference to FIG. 5, the creation of a criterion map based on evaluation of a toy model will be described. In this example, for the sake of simplicity, a one-dimensional path along which the autonomous mobile robot can move in only one direction will be used.

[0043] A goal is set at one end of the one-dimensional path. Agent AG, an autonomous mobile robot, can start moving from any grid on the one-dimensional path and move one square to the left or right of its current position on the map in one step. Figure 5 also shows a flood FD on the map. With each step, the flood FD floods one square adjacent to the current position of the flood FD. In other words, the flood FD changes the environment surrounding the autonomous mobile robot as a state change. The reward R obtained by agent AG per step is set to -0.1, and the reward R at the goal is set to +10. The game ends when agent AG is engulfed by the flood FD or when agent AG reaches the goal. These conditions may also be referred to as end conditions in this specification. The end states shown here are merely examples, and various states can be set.

[0044] Under these conditions, sufficient learning eventually converges to the state value shown in the upper diagram of Figure 5. The upper part of Figure 5 shows the coordinates of the moving object acquired by the state acquisition unit 151 and the state value of the current environment. In this case, the value of each grid in the judgment criteria map is calculated using the above equations (1) and (2). Specifically, the aforementioned path planning unit 101 plans a path from the current position to each grid. When the agent AG moves to each grid along the planned path, the state prediction unit 102 estimates the change in the environmental state of each grid, the position of the moving object at that time, and the reward obtained along the path. The evaluation unit 103 calculates the evaluation value using the above equations (1) and (2). The presentation unit 104 presents the map with the evaluation value displayed for each grid to the user.

[0045] The judgment criteria map for the grid where agent AG is currently located is shown in the lower diagram of Figure 5. In this example, the judgment criteria map shows the position of agent AG and the score for each grid. The judgment criteria map may also display the arrival time when agent AG moves to each grid.

[0046] In principle, an autonomous mobile robot moves autonomously toward a goal so as to maximize its own value function. For example, if the autonomous mobile robot selects a route not anticipated by a monitoring user, the criteria map presents the user with the rationale for the robot's route selection. In other words, when the autonomous mobile robot moves from its current location to each grid on the map, the criteria map also indicates risks (e.g., the risk of being swept away by a flood) that arise due to state changes that take into account the elapsed time until the move. Based on the criteria map created in this way, a user monitoring the autonomous mobile robot can recognize the reasons for the autonomous mobile robot's route selection.

[0047] Specifically, in the judgment criteria map for the current position of the autonomous mobile robot shown in Figure 5, grids with a value of 0 indicate a lower risk, while grids with a value of -9.7 indicate a higher risk (i.e., a higher risk of being swamped by floodwaters). A grid with a value of -0.2 indicates that there is no risk of being swamped by floodwaters when the autonomous mobile robot moves to that grid, but that the autonomous mobile robot could be swamped by floodwaters in the next step unless it moves toward the goal. Therefore, a grid with a value of -0.2 indicates a slightly higher risk than a grid with a value of 0.

[0048] Another example of creating a criteria map will now be described with reference to Figures 6 and 7. In this example, a more complex two-dimensional path is shown. The two grids indicated by diagonal lines (Goal_1 and Goal_2 in FIG. 7) are the goals. Flood FD indicates the source of the flood. The movement rules, rewards, and termination conditions are the same as those in the example shown in FIG. 5 above. The monitoring user does not know which of the two goals, Goal_1 or Goal_2, the autonomous mobile robot is moving towards. FIG. 6 shows the state value of each grid on the map when the agent AG, which is an autonomous mobile robot, is at the current position shown in FIG. 6. In other words, FIG. 6 shows the environment in which the state acquisition unit 151 has acquired the position of the mobile body at the current time and the state of the environment.

[0049] Figure 7 shows the criterion map for the current position shown in Figure 6. The value of each grid in the criterion map is calculated using equations (1) and (2) above. In the criterion map at t = 0, the time it takes for an autonomous mobile robot at its current position to reach each grid varies for each grid, so each grid in the criterion map can indicate a different time risk or value. For example, as shown in Figure 7, the grid one grid away from the current position of the agent AG at t = 0 indicates the value that would be expected if the robot were to move to that grid at t = 1. Similarly, the grid two grids away indicates the value that would be expected if the robot were to move to that grid at t = 2. The grid three grids away indicates the value that would be expected if the robot were to move to that grid at t = 3. When calculating this value, the change in state over time due to movement, i.e., the grid being engulfed by the flood FD, is taken into account. For example, at t = 3, the grid three grids away from the current position becomes flooded. The value of [maxQ(S',*)+R] in the above formula (1) becomes almost 0, and as a result, the value of the above formula (1) becomes low. In this way, the values ​​of all grids on the path that the robot takes to the goal are calculated using the above formulas (1) and (2).

[0050] While the values ​​of each grid may be displayed on a map, as in the example shown in FIG. 6, the judgment criteria map in this example displays different patterns on each grid to indicate differences in value or risk. Plain white grids indicate scores near 0, i.e., relatively low risk. On the other hand, the grids shown with the pattern in FIG. 7 indicate scores near -10, i.e., relatively high risk. As described above, the grids shown with the pattern in FIG. 7 indicate that when the robot moves to each grid, taking into account the passage of time associated with the movement, it will be engulfed by the flood FD. In other words, the grids shown with the pattern in FIG. 7 indicate a higher risk than the grids shown with the pattern in FIG. 7. The judgment criteria map shown in FIG. 7 indicates that moving toward Goal_2 rather than Goal_1 is less risky because the autonomous mobile robot can reach Goal_2 via the path consisting of the plain white grids. Note that in some embodiments, the presentation unit 104 can display an arrow on the map along the path with the lowest risk among multiple paths. The presenting unit 104 may also present the route in another manner (for example, a navigation display) so as to guide the user to a route with a lower risk among the multiple routes. In this way, the monitoring user can use the judgment criteria map to recognize the reason why the autonomous mobile robot selects a route toward Goal_2 instead of Goal_1.

[0051] An example in which a moving object moves to a first area (for example, grid G03 in FIG. 7) and a second area (for example, grid G12 in FIG. 7) will be specifically described below. The route planning unit 101 plans a first route (e.g., G00 → G01 → G02 → G03) along which a moving body in a specific area (e.g., G00) moves to a first area on the map (e.g., grid G03 in FIG. 7) at a specific time point (e.g., t = 0). The route planning unit 101 also plans a second route (e.g., G00 → G01 → G12) along which a moving body in a specific area (e.g., G00) moves to a second area on the map (e.g., grid G12 in FIG. 7) at a specific time point (e.g., t = 0).

[0052] When a mobile object moves along a planned first route and arrives at a first region (e.g., grid G03 in FIG. 7) at a first time point (e.g., t=3), the state prediction unit 102 predicts the environmental state of the first region (e.g., flood) according to the elapsed time (e.g., 3 seconds) involved in the movement from a specific time point (e.g., t=0) to the first time point (e.g., t=3). The state prediction unit 102 also predicts the state of each region (e.g., non-flooding) at the time (e.g., t=1, t=2) when the mobile object moves along the first route and arrives at each region (e.g., G01, G02).

[0053] Furthermore, when the moving body moves along the planned second route and arrives at a second area (e.g., grid G12 in Figure 7) at a second time point (e.g., t = 2), the state prediction unit 102 predicts the environmental state of the second area (e.g., non-flooding) according to the elapsed time (2 seconds) associated with the movement from a specific time point (e.g., t = 0) to the second time point (e.g., t = 2).

[0054] The state prediction unit 102 also predicts the state (e.g., non-flood) of an area (e.g., G01) on the route at the time (e.g., t=1) when the moving object moves along the second route and reaches the area (e.g., G01) on the route.

[0055] Furthermore, when the moving object moves along the planned first route to a first region (for example, grid G03 in FIG. 7), the state prediction unit 102 estimates a reward obtained on the first route. Also, when the moving object moves along the planned second route to a second region (for example, grid G12 in FIG. 7), the state prediction unit 102 predicts a reward obtained on the second route.

[0056] Furthermore, when the moving object moves along the first route and reaches each area (e.g., G01, G02), the state prediction unit 102 predicts the reward obtained until the moving object reaches each area. When the moving object moves along the second route and reaches an area (e.g., G01) on the route, the state prediction unit 102 predicts the reward obtained until the moving object reaches the area.

[0057] The evaluation unit 103 evaluates the behavior of the mobile object moving along the first route to the first area (e.g., G03) based on the predicted environmental state (e.g., flood) of the first area. The evaluation unit 103 can also evaluate the behavior of the mobile object moving to the first area based on the environmental state (e.g., non-flood) of each area (e.g., G01, G02) on the way to the predicted first area.

[0058] Furthermore, the evaluation unit 103 evaluates the behavior of the moving object moving to the second area along the second route based on the predicted environmental state (e.g., non-flooding) of the second area (e.g., G12). The evaluation unit 103 can also evaluate the behavior of the moving object moving to the second area along the second route based on the environmental state (e.g., non-flooding) of an area (e.g., G01) on the way to the predicted second area.

[0059] The evaluation unit 103 can evaluate the behavior of the mobile object moving to the first area based on the predicted reward obtained by the mobile object moving on the first route up to the first area. The evaluation unit 103 can evaluate the behavior of the mobile object moving to the second area based on the predicted reward obtained by the mobile object moving on the second route up to the second area.

[0060] The presentation unit 104 presents the evaluation in the first region (e.g., G03) (which is a value near -10 and is therefore shown by the pattern in FIG. 7) and the evaluation in the second region (e.g., G12) (which is a value near 0 and is therefore shown in white in FIG. 7) by superimposing them on the map. In some embodiments, the presentation unit 103 can present the evaluation in the first region and each region up to the first region (e.g., G01, G02) (all shown in white in FIG. 7), and the evaluation in the second region and each region up to the second region (e.g., G0) (white in FIG. 7) by superimposing them on the map.

[0061] In the judgment criteria map shown in FIG. 7, the difference in value or risk of each grid is expressed by the different patterns of each grid, but it may also be expressed by different colors. In this case, for example, a white grid can represent a relatively low risk, and an orange grid can represent a relatively high risk. In this case, the goal can be represented by a red grid and the source of the flood water by a blue grid. Also, if a moving object can take multiple routes at the current time, an arrow can be displayed on the corresponding route on the map to indicate that the route with the lower risk should be selected. In this way, various notations can be used for the judgment criteria map so that the monitoring user can easily determine the risk of the route.

[0062] In the above example, the judgment criterion map at t=0 is shown, but if the state change is as expected at t=1, the judgment criterion map does not change. However, if an unexpected state change occurs at t=1, the judgment criterion map at t=1 may be updated to be different from the judgment criterion map at t=0. In that case, at t=1, the current state after the state change may be recognized by a state acquisition unit (for example, a sensor, etc.), and processing may be executed again by the route planning unit 101, state prediction unit 102, evaluation unit 103, presentation unit 104, etc.

[0063] A method for creating a map that takes time into consideration through simulation will be described with reference to FIG. 8 and the above formula (1). The state prediction unit 102 of the behavior evaluation presentation device 100 simulates the process in which the robot moves to each grid on the map. In the above formula (1), [max Q(S',*)+R], the term max Q(S',*) indicates that the state of the environment at the time (t=k) when the robot reaches that grid from its current position is taken into consideration, and R indicates that the reward obtained during the robot's movement up to that point is taken into consideration. Furthermore, in the above formula (1), the score for each grid is calculated by comparing [max Q(S',*)+R] with the present value max Q(S,*). The present value can be said to be the value of taking the least damaging action in the current state, that is, the value of the optimal action. This score calculation is repeated until the time t=n when the autonomous mobile robot reaches the goal. As a result, the decision criteria map at t=0 can represent the risk taking into account the passage of time.

[0064] <Other embodiments> In the above embodiment, the case where the state-action value function Q(s, a) is learned has been described. However, in some embodiments, the reward variance Va(s, a) can also be incorporated so that it is learned simultaneously during learning. In this case, the index F(s, a) = Q(s, a) - βVa(s, a) can be used. The autonomous mobile robot is made to behave so as to maximize F. To achieve this, the index F(s, a) needs to increase Q while suppressing the reward variance Va when β > 0. In other words, the index F(s, a) = Q(s, a) - βVa(s, a) can be called an index for risk-sensitive reinforcement learning. This embodiment can be directly applied as a value function that includes risk by replacing Q(s, a) with F(s, a).

[0065] Next, the case of a partial observation system will be described with reference to FIG. In reality, there are cases where the state acquisition unit 151 cannot observe the state of all areas. For example, in the example of FIG. 9, consider a case where there is a visibility restriction near Goal_2. That is, the state acquisition unit 151 can grasp the current position of the mobile object and the state of the environment up to the visibility restriction, but cannot grasp the state near Goal_2 due to an obstacle in the view, etc. In this case, the state prediction unit 102 may predict the state using the following method.

[0066] For the observable area, the current state is determined. On the other hand, for the unobservable state of each grid, the state prediction unit 102 creates a distribution that assumes possible states for each grid and performs sampling. For example, as in the example above, in an environment where flooding may occur, the possible states of one square are one of three states: goal, flood state, or non-flood state. If the agent AG is unable to observe the state of an area far from its current location, i.e., the state near Goal_2, the unobservable area is sampled assuming that each possible state is uniformly distributed. In this example, the state of each grid with limited visibility is sampled as one of the three states: goal, flood state, and non-flood state, with one-third of the grid sampled. As described above, the state prediction unit 102 predicts the state of each grid and the reward for the route agent AG takes to grid G21. The state prediction unit 102 repeats these processes a predetermined number of times. The evaluation unit 103 then calculates the sample average of the value [maxQ(S',*)+R]-maxQ(S,*). Similarly, the processing of the state prediction unit 102 and the evaluation unit 103 is executed for the grid G22 and the grid G23.

[0067] Alternatively, the state prediction unit 102 may predict the state of each grid with limited visibility using prior knowledge. For example, as prior knowledge, the behavior evaluation presentation system 1 recognizes the location of the goal on the map. The behavior evaluation presentation system 1 recognizes that in areas with limited visibility, the probability of flooding is lower than the probability of non-flooding. In this case, the state prediction unit 102 assumes that the location of the goal grid is fixed, and samples the other two squares before the goal assuming a 1 / 10 probability of flooding and a 9 / 10 probability of non-flooding. The state prediction unit 102 predicts the state and reward of each grid along the agent AG's path to grid G21, as described above. The state prediction unit 102 repeats these processes a predetermined number of times. Thereafter, the evaluation unit 103 calculates the sample average of the value [maxQ(S',*)+R]-maxQ(S,*). The state prediction unit 102 and evaluation unit 103 similarly perform the processes for grids G22 and G23. In this case, each grid on the criteria map can display the sample mean of the value of [maxQ(S',*)+R]-maxQ(S,*).

[0068] FIG. 10 is a diagram illustrating a route planning method by a route planning unit according to another embodiment. When the environment in which the moving object moves is a continuous space, the path planning unit 101 can divide the continuous space into a grid as shown in Fig. 10 and select a representative point (for example, a central point) from each grid. The path planning unit 101 can plan a drive control method for the moving object (for example, the number of rotations of the wheels of the moving object). Thereafter, as described above, the path planning unit 101, the state prediction unit 102, the evaluation unit 103, and the presentation unit 104 may perform the respective processes.

[0069] In reinforcement learning, the value of the reward obtained n steps from the present is generally evaluated as α^n using a time discount rate α(0-1). The time discount rate can be used, for example, to converge the reward of an infinitely continuing episode. In the case of time discounting, if you receive a reward r0 in state s, and then receive rewards r1, r2..., the evaluation of state s is V(s)=r0+α*r1+α^2*r2+... In the toy model in the above embodiment, it is explained that there is no time discount (α=1), but the time discount rate may be taken into consideration. When there is time discount, R is calculated by multiplying the reward for each time by α^k and adding them together (also referred to as time-discounted reward in this specification). Depending on the elapsed time n [max α^n×Q(S',*)+R]-maxQ(S,*) This becomes:

[0070] FIG. 11 is a block diagram showing a configuration example of the behavior evaluation presentation device 100 and the control unit 150 (hereinafter referred to as the behavior evaluation presentation device 100, etc.). Referring to FIG. 11, the behavior evaluation presentation device 100, etc. includes a network interface 1201, a processor 1202, and a memory 1203. The network interface 1201 is used to communicate with other network node devices that constitute a communication system. The network interface 1201 may be used to perform wireless communication. For example, the network interface 1201 may be used to perform wireless LAN communication defined in the IEEE 802.11 series or mobile communication defined in the 3GPP (3rd Generation Partnership Project). Alternatively, the network interface 1201 may include, for example, a network interface card (NIC) conforming to the IEEE 802.3 series.

[0071] The processor 1202 reads and executes software (computer programs) from the memory 1203 to perform the processing of the behavior evaluation presentation device 100 and the like described using flowcharts or sequences in the above-described embodiments. The processor 1202 may be, for example, a microprocessor, an MPU (Micro Processing Unit), or a CPU (Central Processing Unit). The processor 1202 may include multiple processors.

[0072] The memory 1203 is configured by a combination of volatile memory and non-volatile memory. The memory 1203 may include storage located remotely from the processor 1202. In this case, the processor 1202 may access the memory 1203 via an I / O interface (not shown).

[0073] 11, the memory 1203 is used to store a group of software modules. The processor 1202 reads out and executes these software modules from the memory 1203, thereby performing the processing of the behavior evaluation presentation device 100 and the like described in the above-described embodiment.

[0074] As explained using FIG. 2, each of the processors included in the behavior evaluation presentation device 100 and the like executes one or more programs including a group of instructions for causing a computer to execute the algorithm explained using the drawing.

[0075] In the above examples, the program includes instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments. The program may be stored on a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable medium or tangible storage medium includes random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technology, CD-ROM, digital versatile disc (DVD), Blu-ray® disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The program may also be transmitted on a transitory computer-readable medium or communication medium. By way of example and not limitation, transitory computer-readable medium or communication medium includes electrical, optical, acoustic, or other forms of propagated signals.

[0076] The present invention is not limited to the above-described embodiment, and various modifications can be made without departing from the spirit and scope of the present invention. The above-described examples can also be implemented in combination with each other.

[0077] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. (Appendix 1) a route planning means for planning a first route along which a moving object moves to a first area on a map and a second route along which the moving object moves to a second area on the map; predicting an environmental state of the first area according to an estimated elapsed time required for the moving object to move along the first route; a state prediction means for predicting an environmental state of the second area in accordance with an estimated elapsed time required for the moving object to move along the second route; an evaluation means for evaluating the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, and for evaluating the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area; a presentation means for presenting the evaluation in the first area and the evaluation in the second area by superimposing them on the map; A behavior evaluation presentation system comprising: (Appendix 2) the state prediction means predicts a state of each area at a time when the moving object is estimated to reach each area along the first route on the way to the first area; predicting a state of each area along the second route at a time when the moving object is estimated to arrive at each area along the second route; the evaluation means evaluates the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area and the predicted states of each area up to the first area, and evaluates the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area and the predicted states of each area up to the second area; The behavioral evaluation presentation system described in Appendix 1, wherein the presentation means presents the evaluation in the first area and each area up to the first area, and the evaluation in the second area and each area up to the second area, superimposed on the map. (Appendix 3) The behavior evaluation presentation system described in Appendix 1 or 2, wherein the evaluation means evaluates the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, the predicted states of each area up to the first area, and the reward obtained by the moving object until it moves to the first area, and evaluates the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area, the predicted states of each area up to the second area, and the reward obtained by the moving object until it moves to the second area. (Appendix 4) The behavior evaluation presentation system described in Appendix 1, wherein the evaluation means calculates an evaluation value of the behavior of the moving body in each area of ​​the route based on route information indicating the route of the moving body, the predicted environmental state, and a reference function indicating an indicator of the behavior of the moving body. (Appendix 5) further comprising a state acquisition means for acquiring the position of the moving body and the state of the environment in which the moving body moves, the state prediction means predicts an environmental state of the first area according to an estimated elapsed time required for the moving object to travel from the acquired position of the moving object to the first area; predicting an environmental state of the second area according to an estimated elapsed time required for the moving object to travel from the acquired position of the moving object to the second area; 2. The behavioral evaluation presentation system according to claim 1. (Appendix 6) the state prediction means predicts a state of an area on the map for which the state acquisition means cannot acquire the state of the environment, assuming that each possible state of the area is uniformly distributed; 6. The behavior evaluation presentation system according to claim 5, wherein the evaluation means evaluates the behavior of the moving object that has moved to the first area and the behavior of the moving object that has moved to the second area. (Appendix 7) the state prediction means predicts a state of an area on the map for which the state acquisition means cannot acquire the state of the environment, assuming that each possible state of the area is distributed with a predetermined probability; 6. The behavior evaluation presentation system according to claim 5, wherein the evaluation means evaluates the behavior of the moving object that has moved to the first area and the behavior of the moving object that has moved to the second area. (Appendix 8) planning a first route for a moving object to move to a first area on a map and a second route for the moving object to move to a second area on the map; predicting an environmental state of the first area according to an estimated elapsed time required for the moving object to move along the first route; predicting an environmental state of the second area according to an estimated elapsed time required for the moving object to move along the second route; Evaluating the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, and evaluating the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area; A behavior evaluation presentation method, wherein the evaluation in the first area and the evaluation in the second area are presented by being superimposed on the map. (Appendix 9) predicting a state of each area along the first route at a time when the moving object is estimated to arrive at each area along the first route; predicting a state of each area along the second route at a time when the moving object is estimated to arrive at each area along the second route; Evaluating the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area and the predicted states of each area up to the first area, and evaluating the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area and the predicted states of each area up to the second area; A behavioral evaluation presentation method described in Appendix 8, in which the evaluations in the first area and each area up to the first area, and the evaluations in the second area and each area up to the second area are superimposed and presented on the map. (Appendix 10) A behavior evaluation presentation method as described in Appendix 8 or 9, which evaluates the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, the predicted states of each area up to the first area, and the reward obtained by the moving object until it moves to the first area, and evaluates the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area, the predicted states of each area up to the second area, and the reward obtained by the moving object until it moves to the second area. (Appendix 11) A behavior evaluation presentation method described in Appendix 8, which calculates an evaluation value of the behavior of the moving object in each area of ​​the route based on route information indicating the route of the moving object, the predicted environmental state, and a reference function indicating an indicator of the behavior of the moving object. (Appendix 12) Acquire the position of the moving object and the state of the environment in which the moving object moves; predicting an environmental state of the first area according to an estimated elapsed time required for the moving object to travel from the position of the moving object to the first area; predicting an environmental state of the second area according to an estimated elapsed time required for the moving object to travel from the position of the moving object to the second area; 10. The behavioral assessment presentation method described in Appendix 8. (Appendix 13) For areas on the map where the environmental state cannot be obtained, predicting states assuming a uniform distribution of possible states of the region; 13. The behavior evaluation presentation method according to claim 12, wherein the behavior of the moving object that has moved to the first area and the behavior of the moving object that has moved to the second area are evaluated. (Appendix 14) For areas on the map where the environmental state cannot be obtained, predicting a state assuming that each possible state of the region is distributed with a predetermined probability; 13. The behavior evaluation presentation method according to claim 12, wherein the behavior of the moving object that has moved to the first area and the behavior of the moving object that has moved to the second area are evaluated. (Appendix 15) a route planning means for planning a first route along which a moving object moves to a first area on a map and a second route along which the moving object moves to a second area on the map; predicting an environmental state of the first area according to an estimated elapsed time required for the moving object to move along the first route; a state prediction means for predicting an environmental state of the second area in accordance with an estimated elapsed time required for the moving object to move along the second route; an evaluation means for evaluating the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, and for evaluating the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area; a presentation means for presenting the evaluation in the first area and the evaluation in the second area by superimposing them on the map; A behavior evaluation presentation device comprising: (Appendix 16) the state prediction means predicts a state of each area at a time when the moving object is estimated to reach each area along the first route on the way to the first area; predicting a state of each area along the second route at a time when the moving object is estimated to arrive at each area along the second route; the evaluation means evaluates the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area and the predicted states of each area up to the first area, and evaluates the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area and the predicted states of each area up to the second area; The behavioral evaluation presentation device described in Appendix 15, wherein the presentation means presents the evaluation in the first area and each area up to the first area, and the evaluation in the second area and each area up to the second area, superimposed on the map. (Appendix 17) The behavior evaluation presentation device described in Appendix 15 or 16, wherein the evaluation means evaluates the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, the predicted states of each area up to the first area, and the reward obtained by the moving object until it moves to the first area, and evaluates the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area, the predicted states of each area up to the second area, and the reward obtained by the moving object until it moves to the second area. (Appendix 18) The behavior evaluation presentation device described in Appendix 15, wherein the evaluation means calculates an evaluation value of the behavior of the moving body in each area of ​​the route based on route information indicating the route of the moving body, the predicted environmental state, and a reference function indicating an indicator of the behavior of the moving body. (Appendix 19) further comprising a state acquisition means for acquiring the position of the moving body and the state of the environment in which the moving body moves, the state prediction means predicts an environmental state of the first area in accordance with an estimated elapsed time required for the moving object to travel from a position of the moving object to the first area; predicting an environmental state of the second area according to an estimated elapsed time required for the moving object to travel from the position of the moving object to the second area; 16. The behavioral evaluation presentation device according to claim 15. (Appendix 20) the state prediction means predicts a state of an area on the map for which the state acquisition means cannot acquire the state of the environment, assuming that each possible state of the area is uniformly distributed; 20. The behavior evaluation presentation device according to claim 19, wherein the evaluation means evaluates the behavior of the moving object that has moved to the first area and the behavior of the moving object that has moved to the second area. (Appendix 21) the state prediction means predicts a state of an area on the map for which the state acquisition means cannot acquire the state of the environment, assuming that each possible state of the area is distributed with a predetermined probability; 20. The behavior evaluation presentation device according to claim 19, wherein the evaluation means evaluates the behavior of the moving object that has moved to the first area and the behavior of the moving object that has moved to the second area. [Explanation of symbols]

[0078] 1. Behavioral evaluation presentation system 100 Behavioral evaluation presentation device 101 Route Planning Department 102 State prediction unit 103 Evaluation Department 104 Presentation section 110 Map storage unit 150 control section 151 Status acquisition unit FD flood AG Agent

Claims

1. A route planning means for planning a first route along which a moving body moves to a first area on a map at a specific time point, and a second route along which the moving body moves to a second area on the map different from the first area at the specific time point; predicting an environmental state of the first area in accordance with an estimated elapsed time required for the moving object to move from the specific time point to the first time point when the moving object moves along the planned first route and arrives at the first area at a first time point; a state prediction means for predicting an environmental state of the second area in accordance with an estimated elapsed time required for the moving object to move from the specific time point to the second time point when the moving object moves along the planned second route and arrives at the second area at a second time point; an evaluation means for evaluating the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, and for evaluating the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area; a presentation means for presenting the evaluation in the first area and the evaluation in the second area by superimposing them on the map; A behavior evaluation presentation system comprising:

2. the state prediction means predicts a state of each area at a time when the moving object is estimated to reach each area along the first route on the way to the first area; predicting a state of each area along the second route at a time when the moving object is estimated to arrive at each area along the second route; the evaluation means evaluates the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area and the predicted states of each area up to the first area, and evaluates the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area and the predicted states of each area up to the second area; 2. The behavioral evaluation presentation system of claim 1, wherein the presentation means presents the evaluations in the first area and each area up to the first area, and the evaluations in the second area and each area up to the second area, superimposed on the map.

3. The behavior evaluation presentation system described in claim 1 or 2, wherein the evaluation means evaluates the behavior of the moving body that has moved to the first area based on the predicted environmental state of the first area, the predicted states of each area up to the first area, and the reward obtained by the moving body until it moves to the first area, and evaluates the behavior of the moving body that has moved to the second area based on the predicted environmental state of the second area, the predicted states of each area up to the second area, and the reward obtained by the moving body until it moves to the second area.

4. The behavior evaluation presentation system according to claim 1, wherein the evaluation means calculates an evaluation value of the behavior of the moving body in each area of ​​the route based on route information indicating the route of the moving body, the predicted environmental state, and a reference function indicating an indicator of the behavior of the moving body.

5. further comprising a state acquisition means for acquiring the position of the moving body and the state of the environment in which the moving body moves, the state prediction means predicts an environmental state of the first area in accordance with an estimated elapsed time required for the moving object to travel from the acquired position of the moving object to the first area; predicting an environmental state of the second area according to an estimated elapsed time required for the moving object to travel from the acquired position of the moving object to the second area; The behavior evaluation presentation system according to claim 1 .

6. the state prediction means predicts a state of an area on the map for which the state acquisition means cannot acquire the state of the environment, assuming that each possible state of the area is uniformly distributed; The behavior evaluation presentation system according to claim 5 , wherein the evaluation means evaluates the behavior of the mobile object that has moved to the first area and the behavior of the mobile object that has moved to the second area.

7. the state prediction means predicts a state of an area on the map for which the state acquisition means cannot acquire the state of the environment, assuming that each possible state of the area is distributed with a predetermined probability; The behavior evaluation presentation system according to claim 5 , wherein the evaluation means evaluates the behavior of the mobile object that has moved to the first area and the behavior of the mobile object that has moved to the second area.

8. planning a first route along which a moving object moves to a first area on a map and a second route along which the moving object moves to a second area on the map; predicting an environmental state of the first area according to an estimated elapsed time required for the moving object to move along the first route; predicting an environmental state of the second area according to an estimated elapsed time required for the moving object to move along the second route; Evaluating the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, and evaluating the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area; A behavior evaluation presentation method, wherein the evaluation in the first area and the evaluation in the second area are presented by being superimposed on the map.

9. A route planning means for planning a first route along which a moving body moves to a first area on a map at a specific time point, and a second route along which the moving body moves to a second area on the map different from the first area at the specific time point; predicting an environmental state of the first area in accordance with an estimated elapsed time required for the moving object to move from the specific time point to the first time point when the moving object moves along the planned first route and arrives at the first area at a first time point; a state prediction means for predicting an environmental state of the second area in accordance with an estimated elapsed time required for the moving object to move from the specific time point to the second time point when the moving object moves along the planned second route and arrives at the second area at a second time point; an evaluation means for evaluating the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, and for evaluating the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area; a presentation means for presenting the evaluation in the first area and the evaluation in the second area by superimposing them on the map; A behavior evaluation presentation device comprising:

10. A process for planning a first route along which a moving body at a specific time moves to a first area on a map, and a second route along which the moving body at the specific time moves to a second area on the map different from the first area; a process of predicting an environmental state of the first area in accordance with an estimated elapsed time required for the moving object to move from the specific time point to the first time point when the moving object moves along the planned first route and arrives at the first area at a first time point; a process of predicting an environmental state of the second area in accordance with an estimated elapsed time required for the moving object to move from the specific time point to the second time point when the moving object moves along the planned second route and arrives at the second area at a second time point; a process of evaluating the behavior of the moving object that has moved to the first area based on the predicted environmental state of the first area, and evaluating the behavior of the moving object that has moved to the second area based on the predicted environmental state of the second area; and presenting the evaluation in the first area and the evaluation in the second area by superimposing them on the map.

Citation Information

Patent Citations

  • Disaster prevention cooperation system

    JP2017215625A

  • Vehicle control device, risk map generation device, and program

    JP2019106049A

  • Collision prevention device, mobile body, and program

    JP2021144435A

  • Predictive action decision device and action decision method

    WO2004068399A1

  • Mobile body, control method, and program

    WO2020262189A1