Information processing device, information processing method, and learning device
The information processing device simplifies the construction of a reward system for reinforcement learning by using an abstract road map and processor to determine rewards for vehicle driving, addressing the challenge of processing diverse driving patterns with reduced computational resources.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2026-03-13
AI Technical Summary
Constructing a reward system for reinforcement learning of vehicle driving behavior using inverse reinforcement learning is challenging due to the need for vast amounts of data and significant computer resources to process diverse driving patterns in various environments, especially when considering complex rules and uncodifiable manners.
An information processing device and method that uses an abstract road map to represent simulation domains, distinguishing vehicles and objects, and a processor to determine rewards based on expert actions, updating rules to simplify processing with simpler resources.
Enables efficient processing of reinforcement learning for vehicle driving patterns with reduced computational requirements by abstracting simulation domains and representing vehicles and objects, facilitating easier determination of rewards.
Smart Images

Figure 0007829409000001 
Figure 0007829409000002 
Figure 0007829409000003
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus and an information processing method for constructing a reward system used in reinforcement learning according to the method of inverse reinforcement learning, and a learning apparatus for performing reinforcement learning using the reward system constructed in this way.
Background Art
[0002] Conventionally, a learning apparatus for performing reinforcement learning on the behavior (driving mode) of a vehicle has been known (see, for example, Patent Document 1). Generally, reinforcement learning learns "behavior that maximizes value" in a certain environment through trial and error. Specifically, an agent (controller of an acting body) determines its behavior in a certain environment based on a policy. That behavior affects the environment, and it is evaluated from the environment that has changed due to that behavior whether the behavior was good, and the evaluation result is given to the agent as a reward. Then, the policy is updated based on the evaluation result (reward). Thereafter, the determination of the agent's behavior based on the policy, the evaluation (reward) of that behavior in the environment affected by that behavior, and the update of the policy based on the evaluation result are sequentially repeated, and the policy is finally updated (learned) so that the obtained reward (evaluation) becomes maximum.
[0003] By the way, in the above-described reinforcement learning, generally, it is difficult to determine a reward, which is an evaluation of behavior based on a policy. As a method for constructing a system for determining this reward, a method of inverse reinforcement learning is known. In this inverse reinforcement learning, based on the behavior performed by an expert in a given environmental model, it is estimated how good each behavior is (reward). By quantitatively obtaining this goodness (as a reward), it is possible to search for behaviors similar to those of an expert (reinforcement learning).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
[0005] When constructing a reward system used for reinforcement learning of the driving behavior of a vehicle (target vehicle) using inverse reinforcement learning techniques, it is necessary to program the driving behavior (actions: for example, "stop on the left shoulder until the emergency vehicle passes, then return to the lane and drive") of the target vehicle in various environments where moving objects (regular cars, emergency vehicles, people, bicycles, motorcycles, etc.) move along the road according to various scenarios (for example, "an emergency vehicle approaches from behind before an intersection"). There are many rules that vehicles must follow (for example, rules stipulated in the Road Traffic Act), and furthermore, there are manners that are difficult to codify, so the driving behavior of a vehicle driven by an expert in the various environments mentioned above will be diverse.
[0006] In order to collect information representing the driving patterns of vehicles driven by skilled drivers, it is preferable to record data from cameras and sensors installed on the vehicles, from the perspective of collecting more realistic environmental information. However, as mentioned above, this, coupled with the need for information on the diverse driving patterns of vehicles driven by skilled drivers in various environments, necessitates recording a vast amount of data. Furthermore, processing such a large amount of information requires significant computer resources (CPU power, software, memory capacity, etc.), and efficient and concise processing is difficult.
[0007] This invention has been made in view of these circumstances, and provides an information processing device and an information processing method that can easily perform processing related to inverse reinforcement learning even with simpler computer resources.
[0008] Furthermore, the present invention provides a reinforcement learning device that uses a reward system constructed by such an information processing device (information processing method). [Means for solving the problem]
[0009] The information processing device according to the present invention is an information processing device that constructs a reward system used for reinforcement learning of the driving patterns of a target vehicle according to the inverse reinforcement learning method, and comprises an abstract road map that abstracts and represents a simulation domain including roads in a planar manner, and a first type that represents a target vehicle driving on the roads in the simulation domain and is placed within the roads in the abstract road map. Text information , and the second type of moving object that moves along the road in the simulation area is represented separately according to its attributes and placed within the road in the abstract road map. Text information A simulation image including this is displayed. Sue A simulation image generation unit generates information, and the second type of information is generated within the simulation image that shows the situation in which the moving object moves along the road in the simulation area according to the scenario. character An environmental information acquisition unit acquires environmental information representing the situation in which the moving object is moving, and the driving mode of the target vehicle, which is driven in such a way that a predetermined movement target is achieved under the situation in which the moving object is moving on the road in the simulation domain according to the scenario, is represented in the simulation image as the first type character The system comprises: an action information acquisition unit that acquires expert action information represented as the movement pattern of the target vehicle; a processor (NN) that takes environmental information and action information as input information and determines reward information, which represents the reward used for reinforcement learning regarding the driving pattern of the target vehicle, as output information according to processing rules; and a learning control unit that updates the processing rules in the processor according to a predetermined learning algorithm so that when the acquired environmental information and expert action information are used as input information, a certain reward information is obtained as output information. The system is configured such that the learning control unit repeatedly updates the processing rules in the processor while changing the environmental information and corresponding expert action information based on the scenario.
[0010] This configuration represents the state of the simulation domain (real world), including roads, by abstracting the simulation domain in a planar manner, and the first type of vehicle that travels on the roads in the simulation domain, which is placed within the roads on the abstract road map. character , and the second type of moving object that moves along the road in the simulation area is represented separately according to its attributes and placed within the road in the abstract road map. character This allows the state of a simulation image, including the elements mentioned above, to be represented.
[0011] Then, the situation in which the moving object moves along the road in the simulation domain (real world) according to the scenario is described in the simulation image as the second type character Environmental information representing the situation in which the object is moving is acquired, and the driving manner of the target vehicle, which is driven in such a way that a predetermined movement target is achieved, under the situation in which the object is moving on the road in the simulation domain (real world) according to the scenario, is shown in the simulation image as the first type character Expert behavior information, expressed as the movement pattern of the vehicle, is acquired. In a processor that takes environmental information and behavior information as input information and determines reward information, which represents the reward used for reinforcement learning about the driving pattern of the target vehicle, as output information according to processing rules, the processing rules are updated according to a predetermined learning algorithm so that when the acquired environmental information and expert behavior information are used as input information, a certain reward information is obtained as output information.
[0012] Then, the processing rules in the processor are repeatedly updated while changing the environmental information and corresponding expert behavior information based on the scenario. The processor obtained as a result of these repeated updates of the processing rules can be used as a reward system in reinforcement learning regarding the driving patterns of the vehicle (target vehicle).
[0013] Furthermore, the information processing device according to the present invention is an information processing device that constructs a reward system used for reinforcement learning of the driving patterns of a target vehicle according to the inverse reinforcement learning method, and comprises an abstract road map that abstracts and represents a simulation domain including roads in a planar manner, and a first type that represents a target vehicle driving on the roads in the simulation domain and is placed within the roads in the abstract road map. Text information , a second type of moving object that moves along the road in the simulation domain is represented separately according to its attributes and placed within the road in the abstract road map. Text information , and a third type of fixed object whose state changes over time in the simulation domain is represented in a way that distinguishes it according to its state and is placed within the abstract road map. Text information A simulation image including this is displayed. Sue A simulation image generation unit generates information, and the second type of information is generated within the simulation image as the moving object moves along the road in the simulation area according to the scenario and the state of the fixed object changes. character As it moves, the third type character An environmental information acquisition unit acquires environmental information that represents the situation as a changing state, and the driving mode of the target vehicle, which is driven in such a way that a predetermined movement target is achieved, under the conditions in which the moving body moves along the road in the simulation domain according to the scenario and the state of the fixed objects changes, is represented in the simulation image as the first type character The system comprises: an action information acquisition unit that acquires expert action information represented as the movement pattern of the target vehicle; a processor (NN) that takes environmental information and action information as input information and determines reward information, which represents the reward used for reinforcement learning about the driving pattern of the target vehicle, as output information according to a certain processing rule; and a learning control unit that updates the processing rule in the processor according to a predetermined learning algorithm so that when the acquired environmental information and expert action information are used as input information, a certain reward information is obtained as output information. The system is configured such that the learning control unit repeatedly updates the processing rule in the processor while changing the environmental information and corresponding expert action information based on the scenario.
[0014] With such a configuration, the situation of the simulation area (real world) including the road is represented by an abstract road map that abstracts and represents the simulation area in a planar manner, a first type arranged in the road in the abstract road map representing the target vehicle traveling on the road in the simulation area character and a second type arranged in the road in the abstract road map that distinguishes and represents moving objects moving on the road in the simulation area according to their attributes character and a third type arranged in the abstract road map that represents fixed objects whose states change over time in the simulation area so as to distinguish them according to their states character can be expressed as the state of a simulation image including them.
[0015] Then, in the simulation image, the situation where the moving object moves on the road in the simulation area (real world) according to the scenario and the state of the fixed object changes is represented as the situation where the second type character moves and the third type character changes, and environmental information is obtained. Based on the situation where the moving object moves on the road in the simulation area (real world) according to the scenario and the state of the fixed object changes, the driving mode of the target vehicle driven to achieve a predetermined movement target is represented as the movement mode of the first type character in the simulation image, and expert action information is obtained. In a processor that uses environmental information and action information as input information and determines reward information representing a reward used for reinforcement learning of the driving mode of the target vehicle as output information according to processing rules, when the obtained environmental information and the expert action information are used as input information, the processing rules are updated according to a predetermined learning algorithm so that a certain reward information can be obtained as output information.
[0016] Then, while changing the environmental information and the corresponding expert action information based on the scenario, the update of the processing rules in the processor is repeatedly executed. As a result of repeatedly updating the processing rules in this way, the obtained processor can be used as a reward system for reinforcement learning regarding the driving mode of a vehicle (target vehicle).
[0018] In the information processing apparatus according to the present invention, before the simulation image generation unit generates Reporting The simulation image represented by the information is formed on a grid composed of a plurality of cells, and the action information acquisition unit uses, as a movement unit, the cells of the grid in the simulation image to represent the first type character of movement mode, and can be configured to acquire the expert action information representing the movement mode.
[0019] With such a configuration, since the expert action information represents the first type character of movement mode in the simulation image as a movement with the cells of the grid as a unit, the movement of the target vehicle represented by the first type character in the road of the abstracted road map can be represented more simply by the expert action information.
[0020] In the information processing apparatus according to the present invention, the environmental information acquisition unit includes an environmental information generation unit that generates the environmental information representing, in the simulation image, the situation in which the moving body moves on the road in the simulation area as the situation in which the second type character moves, obtained from the information actually captured by actually photographing the simulation area where the moving body moves according to the scenario. The action information acquisition unit obtains, from the driving mode of the target vehicle actually driven so that the predetermined movement target is achieved by the driver in the situation where the moving body moves according to the scenario in the simulation area, the actual driving mode as the first type characterThe system can be configured to include an action information generation unit that generates expert action information expressed as a movement pattern.
[0021] With this configuration, information obtained by actually photographing the simulation area (real world) where the moving object moves according to the scenario is used to determine the situation of the moving object moving along the roads in the simulation area within the simulation image. character Environmental information representing the situation in which the object is moving is generated (acquired). Furthermore, the actual driving manner of the target vehicle, which is actually driven by the driver to achieve a predetermined movement target in the situation in which the moving object is moving in the simulation domain (real world) according to the scenario, is represented in the simulation image as the first type character Expert behavior information, expressed as a movement pattern, is generated (acquired).
[0022] In the information processing apparatus according to the present invention, the environmental information acquisition unit obtains information representing the video of the simulation area in which a moving object moves according to the scenario generated by the driving simulator of the vehicle as the target vehicle, and in the simulation image, the situation in which the moving object moves on the road in the simulation area is classified into the second type character The environmental information generation unit generates the environmental information that represents the situation in which the moving object is located, and the behavior information acquisition unit obtains the first type of behavior from the driving behavior of the vehicle in the video of the driving simulator that displays the video of the simulation area in which the moving object moves according to the scenario, and is operated by the driver so that the predetermined movement target is achieved. character The system can be configured to include an action information generation unit that generates expert action information expressed as a movement pattern.
[0023] With this configuration, the information representing the video of the simulation area in which a moving object moves according to a scenario generated by the driving simulator of the target vehicle, is used to determine the situation in which the moving object moves on the road in that simulation area within the simulation image, as described in the second type character Environmental information representing the situation in which the moving object is located is generated (acquired). Furthermore, from the driving manner of the vehicle in the video of the driving simulator that displays the video of the simulation area in which the moving object moves according to the scenario, the driving manner is determined from the driving operation of the vehicle in the video so that the predetermined movement target is achieved, and the first type of driving manner in the simulation image is determined. character Expert behavior information, expressed as a movement pattern, is generated (acquired).
[0024] In the information processing apparatus according to the present invention, the environmental information acquisition unit obtains information from information obtained by actually photographing the simulation area in which a moving object moves according to the scenario and the state of the fixed object changes, and in the simulation image the situation in which a moving object moves along the road in the simulation area and the state of the fixed object changes, and the second type character As it moves, the third type character The system includes an environmental information generation unit that generates the environmental information represented as a changing situation, and the behavior information acquisition unit obtains the actual driving manner of the target vehicle, which is actually driven by the driver to achieve the predetermined movement target, in a situation in the simulation domain where the moving object moves according to the scenario and the state of the fixed object changes, from the first type of driving manner in the simulation image. character The system can be configured to include an action information generation unit that generates expert action information expressed as a movement pattern.
[0025] With this configuration, information obtained by actually photographing a simulation area (real world) where a moving object moves according to a scenario and the state of fixed objects changes is used to represent the situation in the simulation image where the moving object moves along the road in that simulation area and the state of fixed objects changes. character As it moves, the third type character Environmental information is generated (acquired) that represents the situation in which the conditions change. Furthermore, in the simulation domain (real world), the moving object moves according to the scenario and the state of the fixed object changes, and the actual driving pattern of the target vehicle, which is actually driven by the driver to achieve a predetermined movement target, is represented in the simulation image as the first type character Expert behavior information, expressed as a movement pattern, is generated (acquired).
[0026] In the information processing apparatus according to the present invention, the environmental information acquisition unit obtains information representing the video of the simulation area in which a moving object moves and the state of the fixed object changes according to the scenario generated by the driving simulator of the vehicle as the target vehicle, and then uses the information representing the video of the simulation area in which a moving object moves and the state of the fixed object changes within the simulation image to obtain the second type character As it moves, the third type character The environmental information generation unit generates the environmental information that represents the situation as a changing state, and the behavior information acquisition unit obtains the first type of driving behavior in the simulation image from the driving behavior of the vehicle in the video, which is operated by the driver so that the predetermined movement target is achieved in the driving simulator that displays the video of the simulation area in which the moving object moves and the state of the fixed object changes according to the scenario, and the driving behavior is expressed in the simulation image. character The system can be configured to include an action information generation unit that generates expert action information expressed as a movement pattern.
[0027] With this configuration, information representing the video of the simulation area in which a moving object moves and the state of fixed objects changes according to the scenario generated by the driving simulator of the target vehicle, is used to determine the situation in the simulation image in which the moving object moves along the road in the simulation area and the state of fixed objects changes. character As it moves, the third type character Environmental information is generated (acquired) that represents the situation in which the state changes. Furthermore, the driving manner of the vehicle in the video of the driving simulator, which displays the video of the simulation area in which the moving object moves and the state of the fixed object changes according to the scenario, is driven by the driver so as to achieve the predetermined movement target, and that driving manner is represented as the first type in the simulation image. character Expert behavior information, expressed as a movement pattern, is generated (acquired).
[0028] The present invention relates to an information processing method for constructing a reward system used for reinforcement learning of the driving patterns of a target vehicle, in accordance with the inverse reinforcement learning method, comprising: an abstract road map that abstractly represents a simulation domain including roads in a planar manner; and a first type of reward system that represents a target vehicle driving on the roads in the simulation domain and is placed within the roads in the abstract road map. Text information , and the second type of moving object that moves along the road in the simulation area is represented separately according to its attributes and placed within the road in the abstract road map. Text information A simulation image including this is displayed. Sue A simulation image generation step that generates information, and the second type of the situation in which the moving object moves along the road in the simulation area according to the scenario within the simulation image. character An environmental information acquisition step to acquire environmental information representing the situation in which the moving object is moving, and the driving mode of the target vehicle, which is driven in such a way that a predetermined movement target is achieved, in the simulation image under the situation in which the moving object is moving on the road in the simulation domain according to the scenario, the first type characterThe system includes an action information acquisition step that acquires expert action information represented as the movement pattern of the target vehicle, and a learning control step that uses a processor (NN) that takes environmental information and action information as input information and determines reward information representing the reward used for reinforcement learning about the driving pattern of the target vehicle as output information according to a processing rule, and updates the processing rule in the processor according to a predetermined learning algorithm so that when the acquired environmental information and expert action information are used as input information, a certain reward information is obtained as output information, wherein the learning control step repeatedly updates the processing rule in the processor while changing the environmental information and corresponding expert action information based on the scenario.
[0029] Furthermore, the information processing method according to the present invention is an information processing method for constructing a reward system used for reinforcement learning of the driving patterns of a target vehicle in accordance with the inverse reinforcement learning method, comprising: an abstract road map that abstracts and represents a simulation domain including roads in a planar manner; and a first type that represents a target vehicle driving on the roads in the simulation domain and is placed within the roads in the abstract road map. Text information , a second type of moving object that moves along the road in the simulation domain is represented separately according to its attributes and placed within the road in the abstract road map. Text information , and a third type of fixed object whose state changes over time in the simulation domain is represented in a way that distinguishes it according to its state and is placed within the abstract road map. Text information A simulation image including this is displayed. Sue A simulation image generation step that generates information, and a second type of simulation image that shows the situation in which the moving object moves along the road in the simulation area according to the scenario and the state of the fixed object changes. character As it moves, the third type characterAn environmental information acquisition step to acquire environmental information that represents the situation as a changing state, and the driving mode of the target vehicle, which is driven in such a way that a predetermined movement target is achieved, under the conditions in which the moving body moves along the road in the simulation domain according to the scenario and the state of the fixed objects changes, in the simulation image, the first type character The system includes an action information acquisition step that acquires expert action information represented as the movement pattern of the target vehicle, and a learning control step that uses a processor (NN) that takes environmental information and action information as input information and determines reward information representing the reward used for reinforcement learning about the driving pattern of the target vehicle as output information according to a processing rule, and updates the processing rule in the processor according to a predetermined learning algorithm so that when the acquired environmental information and expert action information are used as input information, a certain reward information is obtained as output information, wherein the learning control step repeatedly updates the processing rule in the processor while changing the environmental information and corresponding expert action information based on the scenario.
[0030] The learning device according to the present invention is a learning device that performs reinforcement learning on the driving patterns of a target vehicle, and has as a reward device a processor obtained by repeatedly updating the processing rules in accordance with the predetermined learning algorithm in the information processing device described above (as described in claim 1), an abstract road map that abstractly represents a simulation area including roads in a planar manner, and a first type that represents a target vehicle driving on the roads in the simulation area and is placed within the roads in the abstract road map Text information , and the second type of moving object that moves along the road in the simulation area is represented separately according to its attributes and placed within the road in the abstract road map. Text information A simulation image including this is displayed. Sue A simulation image generation unit that generates information, and a first type character The second type moves within the roads of the abstract road map according to the scenario. character A simulation control unit controls the image information so that it moves within the roads of the abstract road map, and the second type according to the scenario. character In the simulation image of the state in which the abstract road map is moving within the roads, a predetermined first type character In order to achieve the movement objective, the first type character The mode of movement is determined based on the policy, and the first type is determined in that mode of movement. character A controller that instructs the simulation control unit to move relative to the abstract road map, and the first type determined by the controller according to a predetermined reinforcement learning algorithm. character Action information representing the mode of movement, and the aforementioned second type character The abstraction road The system includes a learning processing unit that updates the policy in the controller based on reward information, which is the output information of the rewarder and takes environmental information representing the state of moving within a road on a map as input information, and until the movement target is achieved, the system uses the first type of movement determined by the policy of the controller. character The configuration involves repeatedly issuing movement instructions and updating the policy based on reward information from the rewarder.
[0031] With this configuration, Type 2 will be executed according to the scenario. character In a simulation image of the state of moving within the roads of an abstract road map, Type 1 character However, the controller moves relative to the abstract road map in a manner determined by a policy so that a predetermined movement target is achieved. character Environmental information representing the state of movement within the roads of the abstract road map, and the first type determined by the controller character Reward information is obtained as output information of a rewarder, which takes as input information the behavioral information representing the movement pattern in the aforementioned simulation image (Type 1) character (Evaluation of the movement pattern). Then, the policy in the controller is updated based on the reward information according to a predetermined reinforcement learning algorithm. Thereafter, until the movement target is achieved, the first type of movement pattern determined based on the policy is performed. characterMovement instructions, acquisition of reward information from the rewarder, and updating of the policy in the controller based on the reward information are repeated. In this way, Type 2 according to the scenario within the road of the abstract road map character Type 1 under the mobile environment character The optimal strategy for achieving the movement objective is determined (learned).
[0032] Furthermore, the learning device according to the present invention is a learning device that performs reinforcement learning regarding the driving manner of a target vehicle, and the information processing device described above (in claim 2) is obtained by repeatedly updating the processing rules in accordance with the predetermined learning algorithm. process An abstract road map that has a reward device and represents the simulation domain including roads in a planar abstraction, and a first type that represents the target vehicle traveling on the roads in the simulation domain and is placed within the roads in the abstract road map. Text information , a moving body that moves along the road in the simulation area belongs to Sex The second type, which is distinguished and represented accordingly and placed within the roads in the aforementioned abstract road map. Text information , and the third type of fixed object (traffic light, convex mirror, etc.) whose state changes over time in the simulation domain is represented in a way that distinguishes it according to its state and is placed within the abstract road map. Text information A simulation image including this is displayed. Sue A simulation image generation unit that generates information, and a first type character The second type moves within the roads of the abstract road map according to the scenario. character It moves within the roads of the aforementioned abstract road map and is of the third type character The above changes simulation image Represents A simulation control unit that controls information, and the second type according to the scenario character The third type moves within the roads of the abstract road map and character In the simulation image of the state in which the abstract road map changes, a predetermined first type character In order to achieve the movement objective, the first type characterThe mode of movement is determined based on the policy, and the first type is determined in that mode of movement. character A controller that instructs the simulation control unit to move relative to the abstract road map, and the first type determined by the controller according to a predetermined reinforcement learning algorithm. character Action information representing the mode of movement, and the aforementioned second type character The aforementioned abstraction path alley Move within the road shown in the diagram, and the aforementioned third type character The system includes a learning processing unit that updates the policy in the controller based on reward information, which is the output information of the rewarder and takes environmental information representing the changing state in the abstract road map as input information, and until the movement target is achieved, the system updates the policy in the first type of movement determined by the policy of the controller. character The configuration involves repeatedly issuing movement instructions and updating the policy based on reward information from the rewarder.
[0033] With this configuration, Type 2 will be executed according to the scenario. character It moves within the roads of the abstract road map and is of the third type character In a simulation image of a state in which the first type changes, character However, the controller moves relative to the abstract road map in a manner determined by a policy so that a predetermined movement target is achieved. character The third type moves within the roads of the abstract road map and character Environmental information representing the state in which the state changes, and the first type determined by the controller character Reward information is obtained as output information of a rewarder, which takes as input information the behavioral information representing the movement pattern in the aforementioned simulation image (Type 1) character (Evaluation of the movement pattern). Then, the policy in the controller is updated based on the reward information according to a predetermined reinforcement learning algorithm. Thereafter, until the movement target is achieved, the first type of movement pattern determined based on the policy is performed. characterMovement instructions, acquisition of reward information from the rewarder, and updating of the policy in the controller based on the reward information are repeated. In this way, Type 2 is performed according to the scenario within the road of the abstract road map. character As it moves, the third type character Type 1 in a changing environment character The optimal strategy for achieving the movement objective is determined (learned).
[0035] In the learning device according to the present invention, before the simulation image generation unit generates Reporting The simulation image shown in the report is formed on a grid composed of multiple squares and is determined by the controller, the first type character The movement mode is represented as movement in units of the grid squares, and the simulation control unit performs the first type according to the instructions from the controller. character The abstract road map moves in units of the grid cells. simulation image Represents It is possible to control and configure information.
[0036] With this configuration, the generated simulation image will have a movement pattern determined based on the policy, and will be of type 1. character This moves in grid units relative to the abstract road map. This allows the first type representing the target vehicle to be identified. character This allows for a simpler representation of movement within roads on an abstract road map. [Effects of the Invention]
[0037] According to the information processing device and information processing method of the present invention, the simulation area including roads is represented by an abstracted road map that abstracts the area in a planar manner, a first-class symbol representing the target vehicle, and a simplified simulation image that distinguishes moving objects according to their attributes and includes a second-class symbol (corresponding to other vehicles, pedestrians, bicycles, emergency vehicles, etc.). As a result, processing related to inverse reinforcement learning can be easily performed even with simpler computer resources.
[0038] Furthermore, according to the learning device of the present invention, reinforcement learning becomes possible using a rewarder as a reward system constructed by the above-mentioned information processing device (information processing method), which can easily perform processing related to inverse reinforcement learning even with simpler computer resources. [Brief explanation of the drawing]
[0039] [Figure 1] Figure 1 is a block diagram showing the configuration of an information processing device according to an embodiment of the present invention. [Figure 2] Figure 2 shows a driving simulator in which the driver operates the vehicle. [Figure 3] Figure 3 shows a first example of a simulation image (formed on a grid) that is created on a grid. [Figure 4] Figure 4 shows a second example of a simulation image. [Figure 5] Figure 5 shows a third example of a simulation image. [Figure 6] Figure 6 shows a fourth example of a simulation image. [Figure 7] Figure 7 is a flowchart showing the processing procedure in an information processing device. [Figure 8] Figure 8 shows an example of the movement of the first type symbol (outlined pentagon) in a simulation image. [Figure 9] Figure 9 shows a fifth example of a simulation image. [Figure 10]Figure 10 shows a sixth example of a simulation image. [Figure 11] Figure 11 shows another example of a grid configuration formed in the background of a simulation image. [Figure 12] Figure 12 shows an example image of one layer in a simulation image of a multilayer structure. [Figure 13] Figure 13 shows an example of images of other layers in a simulation image of a multilayer structure. [Figure 14] Figure 14 shows an example of a two-layered simulation image created by superimposing the image shown in Figure 12 and the layered image shown in Figure 13. [Figure 15] Figure 15 is a block diagram showing a learning device according to an embodiment of the present invention. [Modes for carrying out the invention]
[0040] Embodiments of the present invention will be described below with reference to the drawings.
[0041] An information processing device according to one embodiment of the present invention is configured as shown in Figure 1.
[0042] In Figure 1, the information processing device 10 is composed of a computer system including various hardware and software, and has a main control unit 11, an image processing unit 12, a processor 13, and a learning control unit 14. The information processing device 10 with this configuration constructs a reward system (specifically, the processing rules in the processor 13 used as a rewarder in the learning device described later) used for reinforcement learning about the driving patterns of the target vehicle, according to the inverse reinforcement learning method.
[0043] In a simulated environment (real world), a vehicle (such as a regular vehicle, emergency vehicle, person, bicycle, or motorcycle) moves along a road according to a scenario. The storage device 200 stores information related to the driving behavior of the target vehicle (information from the simulation source image described later) that is driven in such a way that a predetermined movement objective corresponding to that scenario is achieved. Such information can be obtained, for example, by using a vehicle driving simulator 100, as shown in Figure 2.
[0044] In Figure 2, the scenario represents the outline of environmental changes for the target vehicle. The movement objective represents the target movement pattern of the target vehicle under the environmental changes according to the scenario (for example, "Move from the starting point to the target point, passing through intersections without colliding with other moving objects (cars, pedestrians, etc.) and without deviating from the road"). The above scenario and the movement objective corresponding to that scenario can be determined, for example, based on traffic laws (e.g., the Road Traffic Act) or traffic manners.
[0045] The driving simulator 100 (for example, CARLA® Simulator) displays a 3D image of a simulation area in which a moving object (a regular vehicle, an emergency vehicle, a person, a bicycle, a motorcycle, etc.) moves according to a certain scenario. In the driving simulator 100 in which such a 3D image is displayed, the driver 150 (a skilled driver (expert)) performs driving operations on the target vehicle (image) in the 3D image so as to achieve a predetermined movement target corresponding to the scenario. From the driving simulator 100 in which the target vehicle (image) is driven in this way, information representing the driving pattern of the target vehicle (image) (for example, the driving trajectory in the 3D image) as the target vehicle (image) is driven to achieve a predetermined movement target under the circumstances in which a moving object moves on the road in the simulation area according to the scenario is obtained as information of the simulation source image. This information of the simulation source image is then stored in the storage device 200 in a state associated with the scenario.
[0046] For example, in the driving simulator 100, under a scenario where "an emergency vehicle (e.g., an ambulance) approaches the target vehicle (image) from behind with its siren blaring," the driver 150 (expert) operates the target vehicle (image) to achieve the movement objective of "moving to the target location without deviating from the road (designated lane) in accordance with the rules of road traffic law that prioritize emergency vehicles." Information representing the driving pattern (e.g., driving trajectory) of the target vehicle (image) being operated within the 3D image of the simulation domain is then obtained as information from the source image of the simulation.
[0047] Returning to Figure 1, the image processing unit 12 (simulation image generation unit) in the information processing device 10, under the control of the main control unit 11, reads information of a simulation source image (3D image) corresponding to a specified scenario from the storage device 200, and generates image information representing the simulation image to be processed from the information of the simulation source image. This simulation image includes an abstract road map that abstracts the simulation area (real world) including roads in a planar manner, first-class symbols that represent target vehicles traveling on the roads in the simulation area and are placed within the roads in the abstract road map, and second-class symbols that distinguish and represent moving objects moving on the roads in the simulation area according to their attributes (normal vehicles, emergency vehicles, people, bicycles, motorcycles, etc.) and are placed within the roads in the abstract road map.
[0048] Figure 3 shows an example of a simulation image. This simulation image is formed on a grid 50 composed of multiple (n × m) squares. For example, in a grid 50 composed of 32 × 32 squares, if the actual vertical and horizontal distances represented by each square are set to 2m□, a simulation image covering a simulation area (real world) of 64m × 64m can be formed. The simulation image formed on the grid 50 described above includes an abstracted road map MAP that abstractly represents the simulation area in a planar manner, including road R1, which has two lanes of traffic on one side and four lanes of traffic on both sides, and road R2, which has a single lane of traffic, intersecting at intersection IS. In this abstracted road map MAP, two pedestrian crossings PC1 and PC2 cross road R1, and two pedestrian crossings PC3 and PC4 cross road R2, surrounding intersection IS.
[0049] The above simulation image includes a "white pentagon mark" as a Type 1 symbol placed within road R1 on the abstract road map MAP, representing a target vehicle traveling on the road in the simulation area. The above simulation image also includes a "black pentagon mark" representing a normal vehicle (other than the target vehicle), a "V-shaped figure mark" representing an emergency vehicle (ambulance, police vehicle, etc.), a "black trapezoid mark" representing a bicycle, and a "black triangle mark" representing a pedestrian, as Type 2 symbols placed within roads R1 and R2 on the abstract road map MAP, distinguishing moving objects on the road in the simulation area according to their attributes (normal vehicle, emergency vehicle, bicycle, pedestrian, etc.).
[0050] Furthermore, the simulation image can also be formed as shown in Figure 4, for example. The simulation image shown in Figure 4 includes the same abstracted road map MAP as the simulation image shown in Figure 3. In this simulation image, ASCII characters (English letters) are used instead of graphic marks ("white pentagon mark", "black pentagon mark", "'V' shaped mark", "black trapezoid mark", "black triangle mark") as Type 1 and Type 2 symbols. Specifically, "E" is used as a Type 1 symbol to represent the target vehicle and placed within road R1, and as Type 2 symbols, "C" represents a normal vehicle (other than the target vehicle), "A" represents an emergency vehicle, "B" represents a bicycle, and "W" represents a pedestrian.
[0051] The simulation image is not limited to those shown in Figure 3 or Figure 4, and can be defined in various ways depending on the environment. For example, the simulation image formed on the grid 50 as shown in Figure 5 includes an abstract road map MAP that abstractly represents the simulation area including a road R having four lanes in total, with two lanes on each side. This simulation image also includes a "white pentagon mark" as a Type 1 symbol placed in the shoulder lane of road R to represent a target vehicle, and a "black pentagon mark with a left black arrow" as a Type 2 symbol placed in the center lane of road R to represent a vehicle flashing its left turn signal. Furthermore, in this simulation image, as shown in Figure 6, the letter "E" can be used instead of the "white pentagon mark" as a Type 1 symbol, and the letter "L" can be used instead of the "black pentagon mark with a left black arrow" as a Type 2 symbol to represent a vehicle flashing its left turn signal.
[0052] As mentioned above, the source image for the simulation represents the driving behavior of a target vehicle (image) operated by a skilled driver 150 in a 3D image, under the condition that the moving object moves along the road in the simulation area according to the scenario, so as to achieve a predetermined movement target. The simulation image represented by the image information generated by the image processing unit 12 from the information of such source image for the simulation represents the movement behavior of a Type 1 symbol moving along the road in the abstracted road map MAP, under the condition that the Type 2 symbol moves along the road in the abstracted road map MAP according to the scenario, so as to achieve a predetermined movement target.
[0053] Note that the simulation images shown in Figures 3 and 4 are, for example, Three regular vehicles (marked with a "black pentagon" and "C") are traveling in the shoulder lane of Road R1. An emergency vehicle (marked with a white pentagon and the letter "A") is traveling in the lane on the center side of Road R1 towards Intersection IS. A pedestrian (marked with a black triangle and a "W") is moving across the crosswalk PC4 on road R2 at intersection IS. A bicycle (marked with a black trapezoid and labeled "B") is traveling in the vehicle lane on the shoulder side of road R1, moving away from intersection IS. Under the following scenario, where the second type of symbols (Figure 3: "black pentagon mark", "white pentagon mark", "black triangle mark", "black trapezoid mark", Figure 4: "C", "A", "W", "B") are moving on the abstract road map, "The vehicle in question (marked with a white pentagon and "E") passes through intersection IS without deviating from road R1, in accordance with the Road Traffic Act's rule that prioritizes emergency vehicles." This can represent the movement patterns of a Type 1 symbol (Figure 3: "white pentagon mark", Figure 4: "E") moving along road R1 on the abstract road map MAP in order to achieve the movement objective.
[0054] Furthermore, the simulation images shown in Figures 5 and 6 are, for example, "A vehicle traveling in the center lane of road R with its left turn signal flashing (marked with a black pentagon with a left arrow, "L") enters (cuts in) the shoulder lane and continues driving." Under the following scenario, the second type of symbol (Figure 5: "Black pentagon mark with left arrow", Figure 6: "L") is moving, "The target vehicle moves from the starting point to the target point without colliding with the vehicle cutting in with its left turn signal flashing (Figure 5: "Pentagon mark with black arrow", Figure 6: "L"), and without deviating from the shoulder lane, while maintaining a predetermined distance from the vehicle entering ahead (Figure 5: "Pentagon mark with black arrow", Figure 6: "L")." This allows us to represent the movement patterns of Type 1 symbols (Figure 5: "white pentagon mark", Figure 6: "E") moving along road R on the abstract road map MAP in order to achieve the movement objective.
[0055] Returning to Figure 1, the image processing unit 12 generates environmental information representing the situation in which the Type 2 symbol moves along the road in the abstracted road map MAP, and expert behavior information representing the movement pattern of the Type 1 symbol under the situation in which the Type 2 symbol moves (corresponding to the driving trajectory of the target vehicle operated by a skilled driver 150 (expert)). This environmental information and expert behavior information are provided to the processor 13.
[0056] The processor 13 obtains a reward value as an evaluation value for the behavior of an object under a certain environment. It takes environmental information and behavior information as input information and determines reward information representing the reward as output information according to processing rules. This processor 13 has a neural network (NN: composed of an input layer, an intermediate layer (hidden layer), and an output layer). The weight coefficients between each node of the neural network NN correspond to the processing rules. The environmental information and expert behavior information obtained from image information representing a simulation image (including an abstracted map MAP, a first-class symbol, and a second-class symbol) generated by the image processing unit 12 are provided to the processor 13 (the input layer of the neural network NN) as input information.
[0057] The learning control unit 14, under the control of the main control unit 11, updates the processing rules (weight coefficients between each node of the neural network NN) in the processor 13 according to a predetermined learning algorithm, so that when the environmental information and the expert action information are input, a predetermined reward information (for example, corresponding to the expected value of an action performed by an expert in a certain environment) is obtained as output information.
[0058] In the information processing device 10, processing is performed according to the procedure shown in Figure 7 under the control of the main control unit 11.
[0059] In Figure 7, information on the source image of the simulation (3D image information) corresponding to scenario i is acquired from the storage device 200 (S11). The acquired information on the source image of the simulation is converted (generated) by the image processing unit 12 into image information of a simulation image (including an abstracted road map, a first-class symbol, and a second-class symbol: see Figures 3 to 6) (S12: Simulation image generation unit, simulation image generation step). From the image information of the simulation image, environmental information is generated that represents the situation in which a second-class symbol moves along a road in the abstracted road map for a predetermined time (for example, 0.5 seconds) (S13: Environmental information acquisition unit, environmental information generation unit, environmental information acquisition step), and expert action information is generated that represents the movement pattern of the first-class symbol for the same time under the situation in which the second-class symbol is moving (S14: Action information acquisition unit, action information generation unit, action information acquisition step). The expert action information (movement pattern of the first-class symbol) can be represented as movement in units of the grid 50, as shown in Figure 8. Specifically, the movement pattern of the Type 1 symbol can be represented by specifying the position of the square on the grid 50 (the position of the Type 1 symbol) at a predetermined interval (for example, every 0.5 seconds). For example, if the actual vertical and horizontal distance represented by each square on the grid 50 is 2m□, the movement pattern of the Type 1 symbol corresponding to the straight-line driving pattern of the target vehicle at a speed of 60km / h is represented on the grid 50 by changing positions by 4 squares at a 0.5-second interval.
[0060] As described above, the generated environmental information and expert behavior information are provided to the processor 13 (the input layer of the neural network NN) (S15). Then, the processing rules (weight coefficients between each node) of the processor 13 (neural network NN) are updated according to the instructions of the learning control unit 14 so that a predetermined reward information is obtained as output information when the environmental information and expert behavior information are used as input information (S16: learning control step). Thereafter, the same processing as described above (S13~S16) is repeatedly executed for a predetermined amount of time for all of the image information of the simulation image corresponding to scenario i (NO in S17).
[0061] Then, when the same processing (S13-S16) is completed for all the image information of the simulation image corresponding to scenario i (YES in S17), the next scenario i+1 is specified (YES in S18, S19), and the information of the simulation source image corresponding to the next scenario is obtained from the storage device 20 (S11). Then, the information of that simulation source image (3D video) is converted into image information of a simulation image (including abstracted road map, first-class symbols, and second-class symbols) (S12), and the same processing as described above (S13-S16) is performed for all the image information of that simulation image. When the processing of all the simulation source images stored in the storage device 200 is completed (NO in S18), the processing for constructing the reward system used for reinforcement learning about the driving of the target vehicle is completed.
[0062] As a result of the processing performed by the information processing device 10 according to the procedure described above, the processing rules (weight coefficients between each node of the neural network NN) in the processor 13 are repeatedly updated according to a predetermined learning algorithm. As a result, the processor 13 takes environmental information (including an abstracted road map and a Type 2 symbol) and the movement pattern (behavioral information) of a Type 1 symbol (target vehicle) under the environment represented by that environmental information as input information, and outputs reward information (output information) that represents the reward. Therefore, the processor 13 is configured to give a reward (expected value: evaluation) for any movement pattern of a Type 1 symbol (target vehicle) under an environment, with the movement pattern of a Type 1 symbol (target vehicle) based on the driving operation of a skilled driver (expert) under a certain environment being considered as the ideal. Therefore, the processor 13 can be used as a reward system (reward device of the learning device described later) in reinforcement learning regarding the driving patterns of a vehicle (target vehicle).
[0063] According to the information processing device 10 described above, the situation of the simulation area including roads is represented by an abstracted road map MAP that abstracts the situation in a planar manner, a first-class symbol representing the target vehicle, and a simplified simulation image that distinguishes moving objects according to their attributes and includes a second-class symbol. As a result, it becomes possible to easily perform the processing related to inverse reinforcement learning for constructing a reward system even with simpler computer resources.
[0064] Incidentally, the appearance of a convex mirror within the field of view of a vehicle differs depending on whether or not moving objects such as vehicles or people are reflected in it. Furthermore, the appropriate driving operation of the vehicle may also differ depending on whether or not moving objects are reflected in the convex mirror. Also, traffic lights change their illumination status depending on the time of day, and the appropriate driving operation of the vehicle may differ depending on the illumination status of the traffic lights.
[0065] As mentioned above, the simulation not only involves moving objects (regular vehicles, emergency vehicles, people, bicycles, motorcycles, etc.) along roads in the simulation domain according to a scenario, but also incorporates information from a simulation source image (3D video) representing the driving behavior of a target vehicle (image) operated by a skilled driver 150 so as to achieve a predetermined movement target, under conditions where the state of the aforementioned fixed objects (traffic lights, convex mirrors, etc.) installed in the simulation domain changes.
[0066] In this case, the simulation image converted (generated) from the source image and subject to processing can be configured to include Type 3 symbols, which are placed in the abstract road map MAP and represent fixed objects (traffic lights, curve mirrors, etc.) whose state changes over time, in a way that distinguishes them according to their state. The simulation image can then represent the movement patterns of Type 1 symbols moving along the roads in the abstract road map MAP so that a predetermined movement target is achieved, under conditions where Type 2 symbols move along the roads in the abstract road map MAP according to a scenario, and Type 3 symbols change.
[0067] The source image for the simulation, which includes the image of the convex mirror, can be converted into a simulation image formed on a grid 50, for example, as shown in Figure 9.
[0068] Figure 9 shows an abstracted road map MAP that abstractly represents the simulation area including roads R1 and R2, each a single lane of traffic intersecting at intersection IS, in a planar manner. It includes a "white pentagon mark" as a Type 1 symbol placed within road R1 to represent the target vehicle, and a "black pentagon mark" as a Type 2 symbol placed within road R2 to represent a normal vehicle. The simulation image also includes a circular "curve mirror graphic mark" fixedly placed at a predetermined corner of intersection IS on the abstracted road map MAP as a Type 3 symbol to distinguish curve mirrors whose state changes over time. This "curve mirror graphic mark" is a graphic mark with a black pentagon placed inside a circle when a moving object is reflected, and a simple circle mark (not shown) when no moving object is reflected.
[0069] Furthermore, the same source image (3D video) can be converted into a simulation image as shown in Figure 10. In the simulation image shown in Figure 10, the letter "E" can be used instead of the "white pentagon mark" (see Figure 9) as the first type symbol, and the letter "C" can be used instead of the "black pentagon mark" (see Figure 9) as the second type symbol. In addition, the English string "MI" representing a state where a moving object is reflected in the convex mirror, and the English letter "M" (not shown) representing a state where a moving object is not reflected in the convex mirror can be used instead of the "curve mirror graphic mark" (see Figure 9) as the third type symbol.
[0070] The simulation images shown in Figures 9 and 10 include field of view lines (two dotted lines) extending from the Type 1 symbol ("white pentagon mark", "E"), which represent the field of view α from the target vehicle. These simulation images illustrate a situation where a vehicle traveling on road R1 towards intersection IS cannot directly see a car traveling on road R2 towards intersection IS, but can see the car traveling on road R2 towards intersection IS through the image in the convex mirror.
[0071] The simulation images shown in Figures 9 and 10 are, for example, In a scenario where "a normal vehicle traveling on road R2 ("black pentagon mark", "C") enters intersection IS from outside the field of view α of the target vehicle on road R1 ("white pentagon mark", "E"), while being reflected in a convex mirror (black pentagon mark inside a circular convex mirror shape mark), and passes through intersection IS," the Type 2 symbol (Figure 9: "black pentagon mark", Figure 10: "E") moves along road R2 on the abstract road map MAP, and the Type 3 symbol changes (Figure 9: "shape mark with black pentagon placed inside a circle" → "circle mark", Figure 10, "MI" → "M"), This can represent the movement pattern of a Type 1 symbol (Figure 9: "white pentagon mark", Figure 10: "E") moving along road R1 on the abstracted road map MAP so that the movement objective "the target vehicle (Figure 9: "white pentagon mark", Figure 10: "E") moves from the starting point along road R1, passes through intersection IS without colliding with a normal vehicle (Figure 9: "black pentagon mark", Figure 10: "E"), and without deviating from road R1" is achieved.
[0072] The information processing device 10 processes the simulation images described above (see Figures 9 and 10) according to the same procedure as described above (see Figure 7).
[0073] From the image information of the aforementioned simulation image, environmental information is generated that represents the situation in the abstracted road map MAP where a Type 2 symbol (Figure 9: "black pentagon mark", Figure 10: "C") moves along the road and a Type 3 symbol changes (Figure 9: "figure mark with a black pentagon placed inside a circle" → "circle mark", Figure 10, "MI" → "M"). Expert information is generated that represents the movement pattern of a Type 1 symbol (Figure 9: "white pentagon mark", Figure 10: "E") under these conditions where a Type 2 symbol moves and a Type 3 symbol changes. The generated environmental information and expert behavior information are provided to the processor 13 (input layer of the neural network NN). The processing rules (weight coefficients between each node) of the processor 13 (neural network NN) are updated so that a predetermined reward information is obtained as output information when the environmental information and expert behavior information are used as input information.
[0074] Then, the processing rules in the processor 13 are repeatedly updated while changing the environmental information and corresponding expert behavior information based on the above scenario. The processor 13 obtained as a result of these repeated updates of the processing rules can be used as a reinforcement learning reward system for the driving patterns of a vehicle (target vehicle) that takes into account situations in which a moving object is moving and the state of fixed objects such as traffic lights and convex mirrors changes.
[0075] In the embodiment described above, the first type symbol moved relative to a fixed abstract road map MAP. However, the image information representing the simulation image may be controlled so that the abstract road map MAP moves relative to a fixed first type symbol. In this case, the image information representing the simulation image is controlled so that the second type symbol moves according to the scenario relative to the moving abstract road map MAP, and the third type symbol remains stationary relative to the moving abstract road map MAP.
[0076] As described above, when moving an abstract road map MAP relative to a fixed Type 1 symbol, the grid 50 on which the moving abstract road map MAP (simulation image) is formed can be set such that, as shown in Figure 11, the actual distance (vertical and horizontal distance) represented by each cell in a region further away from the Type 1 symbol (e.g., "E") than the actual distance (vertical and horizontal distance) represented by each cell in a predetermined region S (e.g., a rectangular region) containing the fixed Type 1 symbol (e.g., "E") is greater than the actual distance (vertical and horizontal distance) represented by each cell in the region further away from the Type 1 symbol (e.g., "E") than the predetermined region S (region outside the predetermined region S). Specifically, in the case of the grid 50 shown in Figure 11, the assignment of actual distances to the cells corresponds to a logarithmic scale centered on the Type 1 mark "E".
[0077] When forming a simulation image on the grid 50 as described above, it becomes possible to construct an abstract road map of a wider area (simulation area) on the grid 50 while appropriately maintaining the resolution of the movement pattern of the Class 1 symbol ("E"), which is represented as movement in units of squares within a predetermined area S that includes the Class 1 symbol representing the target vehicle, in the simulation image. Therefore, it is possible to save memory size for image information representing the simulation image corresponding to a certain simulation area.
[0078] In addition, in grid 50, it is possible to set it so that in only one of the two directions, vertically or horizontally, for example, in the vertical direction, the actual distance represented by each cell corresponding to the area outside the predetermined area S (vertical actual distance) is greater than the actual distance represented by each cell corresponding to the predetermined area S containing the first type symbol (for example, "E"). In this case, the actual distance represented by each cell in the horizontal direction is set to a constant value.
[0079] Furthermore, the actual distance assigned to each square of Grid 50 can be arbitrarily assigned, and it is also possible to assign different actual distances vertically and horizontally. In addition, the size of Grid 50 can also be set arbitrarily. For example, Grid 50 can be composed of 64x64 squares, which is more than the 32x32 squares mentioned above.
[0080] The image information representing the simulation image converted from the source image (3D video) may have a multi-layer structure.
[0081] In this case, the simulation image can have a two-layer structure, for example, a map / infrastructure layer as shown in Figure 12, and an object layer as shown in Figure 13. The map / infrastructure layer (see Figure 12) consists of symbols representing road shapes and road infrastructure (lanes, signals). Specifically, -----·+: Lane information (abstract road map) s: Stop line (abstract road map) r: Traffic light (red) (Type 3 symbol) b: Traffic light (green) (Type 3 symbol) y: Traffic light (yellow) (Type 3 symbol) w: Waypoint (planned route of the vehicle) x: No entry allowed The aforementioned map and infrastructure layers are represented using symbols like the ones shown.
[0082] The object layer (see Figure 13) consists of symbols that distinguish and represent moving objects, walls, and buildings according to their attributes. Specifically, E: Applicable vehicles (Type 1 symbol) C: Standard vehicle (Type 2 symbol) R: Vehicle with right turn signal flashing (Type 2 symbol) L: Vehicle with left turn signal flashing (Type 2 symbol) H: Vehicle with flashing hazard lights (Type 2 designation) B: Motorcycle (Type 2 symbol) A: Emergency vehicle (Type 2 designation) W: Pedestrian (Type 2 symbol) X: Other three-dimensional objects The object layer is represented using symbols like these.
[0083] By combining the map / infrastructure layer (see Figure 12) and the object layer (see Figure 13) described above, a simulation image like the one shown in Figure 14 can be formed. In such a simulation image, according to a certain scenario, the third-class symbols ("b", "y", "r") representing traffic lights in the map / infrastructure layer (see Figure 12) change, and the second-class symbols ("C", "R", "L", "B", "A", "W") representing each moving object in the object layer (see Figure 13) move. As a result, the simulation image shown in Figure 14 changes according to the aforementioned scenario.
[0084] Furthermore, in the simulation image shown in Figure 14, the area outside the forward field of view of the Class 1 symbol ("E") can be obscured by shading to match the situation detected by the actual target vehicle.
[0085] In the embodiments described above, the first and second type symbols were graphic marks or English letters (ASCII characters), but are not limited to these. They can be other graphic marks, other characters (Kanji, Hiragana, Katakana), seals, symbols, trademarks, etc., as long as they fall under the concept of "symbol." Similarly, in the embodiments described above, the third type symbol was an English letter (ASCII character) or a graphic mark (see Figure 9), but is not limited to these. It can be other graphic marks, other characters, seals, symbols, trademarks, etc., as long as they fall under the concept of "symbol."
[0086] In the embodiments described above, the simulation image (including an abstracted road map, Type 1 symbols, and Type 2 symbols, or including a Type 3 symbol in addition to those) is generated (converted) from information of a simulation source image (3D video) obtained by a skilled driver 150 (expert) performing driving operations on a target vehicle (image) in the 3D video so as to achieve a predetermined movement target corresponding to the scenario, in a driving simulator 100 (see Figure 2) where a 3D video of a simulation area is displayed in which a moving object (normal vehicle, emergency vehicle, person, bicycle, motorcycle, etc.) moves according to a certain scenario, or in which the state of fixed objects such as traffic lights and curve mirrors changes in addition to the movement of the moving object. However, the embodiments are not limited to this. In a simulation domain (real world), simulation images (including an abstracted road map, Type 1 symbols, and Type 2 symbols, or including Type 3 symbols in addition to them) can be generated from real-world situations where moving objects (regular vehicles, emergency vehicles, people, bicycles, motorcycles, etc.) move according to a certain scenario, or real-world situations where the state of fixed objects such as traffic lights and convex mirrors changes in addition to the movement of moving objects, and from the driving manner of a real vehicle (target vehicle) driven by a skilled driver to achieve a movement target in such situations. Furthermore, image information representing a simulation image (see Figures 3-6, 9, and 10) including a Type 1 symbol moving along a road in the abstracted road map so that a predetermined movement target is achieved, under the circumstances where a Type 2 symbol moves along a road in the abstracted road map according to a certain scenario, or under the circumstances where a Type 3 symbol changes in addition to the movement of the Type 2 symbol.
[0087] Next, we will describe the reward system constructed as described above, specifically a learning device that uses processor 13 (neural network NN) as a rewarder.
[0088] The learning device according to the present invention is configured as shown in Figure 15.
[0089] In Figure 15, the learning device 20 is composed of a computer system including various hardware and software, and has a main control unit 21, a simulation control unit 22, an image processing unit 23, a controller 24, a learning control unit 25, and a rewarder 26. The learning device 20 with this configuration performs reinforcement learning about the driving patterns of the target vehicle.
[0090] Library 30 manages scenarios and their corresponding movement targets. The main control unit 21 controls the entire learning device 20 and provides the simulation control unit 22 with information representing scenarios from among the scenarios and movement targets managed in Library 30, and provides the controller 24 with information representing the movement targets corresponding to those scenarios. The image processing unit 23 generates image information representing a simulation image (see Figures 3-6, 9, 10, 12-14) that includes the abstracted road map MAP, Type 1 symbols (target vehicles), Type 2 symbols (normal vehicles, emergency vehicles, people, bicycles, motorcycles, etc.), or Type 3 symbols (fixed objects such as traffic lights and curve mirrors) in addition to the above. This simulation image is formed on a grid 50 composed of multiple (n × m) squares, similar to the case of the information processing device 10 described above. The simulation control unit 22 controls the image processing unit 23 in the simulation image so that the first type symbols move along the roads in the abstracted road map MAP in accordance with the scenario, or in a situation where the second type symbols move along the roads in the abstracted road map MAP in addition to the movement of the second type symbols, according to movement instructions from the controller 24, which will be described later, based on the movement target.
[0091] The controller 24 takes as input information the state in which the Type 2 symbol moves within the roads of the abstracted road map MAP according to the scenario in the simulation image generated by the image processing unit 23 (see, for example, Figure 3 or Figure 4), or the state in which the Type 3 symbol moves in addition to the movement of the Type 2 symbol (see, for example, Figure 9 or Figure 10), the behavior information representing the movement pattern of the Type 1 symbol (target vehicle), and the movement target from the main control unit 21, and determines the movement pattern of the Type 1 symbol (for example, the next position of the Type 1 symbol) based on a policy so that the movement target is achieved. Then, the controller 24 instructs the simulation control unit 22 to move within the abstracted road map in the determined movement pattern (movement instruction). Based on instructions from the controller 24, the simulation control unit 22 controls the image information generated by the image processing unit 23 so that the Type 1 symbol moves to the next indicated position in an environment where the Type 2 symbol moves within the roads of the abstracted road map MAP according to the scenario, or in an environment where the Type 3 symbol changes in addition to the movement of the Type 2 symbol. Here, the movement of the Type 1 symbol can be represented as movement in units of the grid 50, as shown in Figure 8.
[0092] The controller 24 described above has a neural network NN (composed of an input layer, an intermediate layer (hidden layer), and an output layer). The weight coefficients between each node of this neural network NN correspond to a "policy" for determining the movement pattern of the Type 1 symbol (for example, the next position). In the controller 24, environmental information representing the movement of the Type 2 symbol in the simulation image generated by the image processing unit 23, or the same environmental information representing changes in the Type 3 symbol in addition to the movement of the Type 2 symbol, the same behavioral information representing the movement pattern of the Type 1 symbol, and the movement target information provided by the main control unit 21 are input to the input layer of the neural network NN. Then, information determined from the information input to the input layer based on the weight coefficients (policy) between each node is output from the output layer of the neural network NN as information representing the new position (grid position) of the Type 1 symbol on the abstracted road map (on the grid 50). This new position is given to the simulation control unit 22 as a movement instruction for the Type 1 symbol.
[0093] The learning control unit 25 receives as input information the movement pattern of the first type symbol in a simulation image generated by the image processing unit 23 under the control of the simulation control unit 22 based on movement instructions from the controller 24, and environmental information representing the movement status of the second type symbol in the simulation image, or environmental information representing the change status of the third type symbol in addition to the movement of the second type symbol. It then evaluates the movement pattern of the first type symbol in the environment represented by the environmental information according to a predetermined reinforcement learning algorithm (e.g., DeepQ-Network, MuZero, etc.). Based on the evaluation result, the learning control unit 25 updates the weight coefficients between each node of the neural network NN in the controller 24 that determined the movement pattern of the first type symbol (corresponding to the "policy" used to determine the movement pattern of the first type symbol). Here, the evaluation of the movement pattern of the first type symbol is performed by the reward given to that movement pattern. Information related to this reward (reward information) is provided by the rewarder 26.
[0094] As mentioned above, the processor 13 (see Figure 1), which has learned (updated) processing rules (weight coefficients between each node of the neural network NN) using the inverse reinforcement learning method, is used as the rewarder 26. The rewarder 26 (processor 13) is configured to take as input information environmental information (for example, see Figure 3 or Figure 4) that represents the state in which a Type 2 symbol moves within the roads of the abstracted road map MAP according to the scenario in the simulation image generated by the image processing unit 23, or environmental information that represents the state in which a Type 3 symbol moves in addition to the movement of a Type 2 symbol (for example, see Figure 9 or Figure 10), and behavioral information that represents the movement pattern of a Type 1 symbol (target vehicle), and to determine as output information reward information that represents the reward (evaluation) for the movement pattern of the Type 1 symbol in the environment represented by the environmental information.
[0095] In this learning device 10, until the movement target of the Type 1 symbol is achieved, for example, until the movement target of the Type 1 symbol (target vehicle) is achieved, such as "moving from the starting point without colliding with a Type 2 symbol (normal vehicle, pedestrian, etc.) and without deviating from the road (designated traffic lane), passing through an intersection," the movement instructions of the Type 1 symbol, the evaluation of the movement pattern of the Type 1 symbol by the learning control unit 25 (rewarding), and the updating of the "policy" (weight coefficients between each node of the neural network NN) are repeatedly performed at a predetermined period (e.g., a 0.5-second period) based on the "policy" (weight coefficients between each node of the neural network NN), so as to maximize the total reward represented by the reward information obtained from the rewarder 26, the movement instructions of the Type 1 symbol by the controller 24 (weight coefficients between each node of the neural network NN) are repeatedly performed, the movement pattern of the Type 1 symbol is evaluated (rewarding) by the learning control unit 25, and the "policy" (weight coefficients between each node of the neural network NN) are updated based on the evaluation result (reward).
[0096] During this process, whenever a Type 1 symbol overlaps with (including contact with) a Type 2 symbol (normal vehicle, emergency vehicle, etc.) in the simulation image, or deviates from the road (designated traffic lane), the movement of the Type 1 and Type 2 symbols in the simulation image stops, or in addition to stopping, the Type 3 symbol returns to its initial state. Subsequently, according to the scenario, the movement of the Type 2 symbol, or the change of the Type 3 symbol from its initial state along with the movement of the Type 2 symbol, is resumed, and the movement of the Type 1 symbol from its starting point is also resumed. The image information representing the simulation image is controlled in this manner. When the movement target of the Type 1 symbol is achieved, the reinforcement learning process in the learning device 20 is terminated. As a result, the optimal "strategy" that enables the movement target of the Type 1 symbol (target vehicle) to be achieved in the environment represented by the simulation image that changes based on the scenario (a Type 2 symbol moves within the abstracted road map, or a Type 3 symbol changes in addition to the movement of the Type 2 symbol), i.e., the optimal weight coefficients between each node of the neural network in the controller 24, is determined.
[0097] The learning device described above enables reinforcement learning of the target vehicle by using the processor 13, which is constructed as a reward system by the information processing device 10 described above, as the rewarder 26. And, as with the information processing device 10 described above, the environment to be used for reinforcement learning is represented by an abstracted road map MAP that abstracts the simulation domain of roads in a planar manner, a first-class symbol representing the target vehicle, and a simple simulation image that includes a second-class symbol that distinguishes moving objects according to their attributes, or a simple simulation image that includes a third-class symbol that, in addition to these, represents fixed objects whose state changes over time in the simulation domain in a way that distinguishes them according to their state and is placed within the abstracted road map MAP. This makes it possible to update and determine a simple and efficient reinforcement learning "strategy" for the driving manner (movement manner) of the target vehicle (first-class symbol).
[0098] Although embodiments of the present invention have been described above, these embodiments and modifications of each part are presented as examples only and are not intended to limit the scope of the invention. These novel embodiments described above can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. [Industrial applicability]
[0099] As described above, the information processing device (information processing method) according to the present invention has the effect of easily performing processing related to inverse reinforcement learning even with simpler computer resources, and is useful as an information processing device for constructing a reward system used in reinforcement learning according to the inverse reinforcement learning method. [Explanation of symbols]
[0100] 10 Information Processing Devices 11 Main Control Unit 12 Image Processing Unit 13 Processing Units 14 Learning Control Unit 20 Learning device 21 Main Control Unit 22 Simulation Control Unit 23 Image Processing Unit 24 Controllers 25 Learning Control Unit 26 Reward device 50 grids 100 Driving Simulators 200 Storage device
Claims
1. An information processing device that constructs a reward system used for reinforcement learning about the driving patterns of a target vehicle according to the inverse reinforcement learning method, A simulation image generation unit generates information representing a simulation image that includes an abstract road map that abstractly represents the simulation area including roads in a planar manner, information of first-type characters placed within the roads in the abstract road map to represent target vehicles traveling on the roads in the simulation area, and information of second-type characters placed within the roads in the abstract road map to distinguish and represent moving objects moving on the roads in the simulation area according to their attributes, An environmental information acquisition unit acquires environmental information that represents the situation in which the moving object moves along the road in the simulation area according to the scenario as the situation in which the second type of character moves within the simulation image. An action information acquisition unit acquires expert action information that represents the driving manner of the target vehicle, which is driven in such a way as the movement manner of the first type of character in the simulation image, under the condition that the moving object moves along the road in the simulation domain according to the scenario, so as to achieve a predetermined movement target. A processor that takes environmental information and behavioral information as input information and determines, according to processing rules, reward information representing the reward used for reinforcement learning regarding the driving behavior of the target vehicle as output information, The system includes a learning control unit that updates the processing rules in the processor so that, when the acquired environmental information and the expert behavior information are used as input information according to a predetermined learning algorithm, a certain reward information is obtained as output information, An information processing device wherein the learning control unit repeatedly updates the processing rules in the processor while changing the environmental information and corresponding expert behavior information based on the aforementioned scenario.
2. An information processing device that constructs a reward system used for reinforcement learning about the driving patterns of a target vehicle according to the inverse reinforcement learning method, A simulation image generation unit generates information representing a simulation image that includes an abstract road map that abstractly represents a simulation area including roads in a planar manner, information of a first type of character placed within the roads in the abstract road map to represent target vehicles traveling on the roads in the simulation area, information of a second type of character placed within the roads in the abstract road map to distinguish and represent moving objects moving on the roads in the simulation area according to their attributes, and information of a third type of character placed within the abstract road map to distinguish and represent fixed objects in the simulation area whose state changes over time according to their state. An environmental information acquisition unit acquires environmental information that represents the situation in which the moving object moves along the road in the simulation area according to a scenario and the state of the fixed object changes as the second type of character moves and the third type of character changes within the simulation image. An action information acquisition unit acquires expert action information that represents the driving pattern of the target vehicle, which is driven in such a way that a predetermined movement target is achieved, as the movement pattern of the first type of character in the simulation image, under conditions in which the moving object moves along the road in the simulation area according to the scenario and the state of the fixed object changes, A processor that takes environmental information and behavioral information as input information and determines, according to processing rules, reward information representing the reward used for reinforcement learning regarding the driving behavior of the target vehicle as output information, The system includes a learning control unit that updates the processing rules in the processor so that, when the acquired environmental information and the expert behavior information are used as input information according to a predetermined learning algorithm, a certain reward information is obtained as output information, An information processing device wherein the learning control unit repeatedly updates the processing rules in the processor while changing the environmental information and corresponding expert behavior information based on the aforementioned scenario.
3. The simulation image represented by the information generated by the simulation image generation unit is formed on a grid composed of multiple squares, The information processing apparatus according to claim 1 or 2, wherein the behavior information acquisition unit acquires expert behavior information representing the movement patterns of the first type of character, which are represented in the simulation image as movement in units of the grid.
4. The environmental information acquisition unit includes an environmental information generation unit that generates environmental information in which the situation in which the moving object moves along the road in the simulation area is represented as the situation in which the second type of character moves within the simulation image, based on information obtained by actually photographing the simulation area in which the moving object moves according to the scenario. The information processing apparatus according to claim 1, wherein the behavior information acquisition unit includes a behavior information generation unit that generates expert behavior information, which represents the actual driving manner of the target vehicle as the movement manner of the first type of character in the simulation image, from the driving manner of the target vehicle actually driven by the driver so as to achieve the predetermined movement target in a situation in which the moving object moves in the simulation domain according to the scenario.
5. The environmental information acquisition unit includes an environmental information generation unit that generates environmental information from information representing the video of the simulation area in which a moving object moves according to the scenario generated by the driving simulator of the vehicle as the target vehicle, and expresses the situation in which the moving object moves on the road in the simulation area as the situation in which the second type of character moves within the simulation image. The information processing apparatus according to claim 1, wherein the behavior information acquisition unit includes a behavior information generation unit that generates expert behavior information, which represents the driving manner of the vehicle in the video of the driving simulator that displays the video of the simulation area in which a moving object moves according to the scenario, and which is operated by a driver to achieve the predetermined movement target, as the movement manner of the first type of character in the simulation image.
6. The environmental information acquisition unit includes an environmental information generation unit that generates environmental information from information obtained by actually photographing the simulation area in which a moving object moves and the state of the fixed object changes according to the scenario, and represents the situation in which a moving object moves on the road in the simulation area and the state of the fixed object changes as a situation in which the second type of character moves and the third type of character changes within the simulation image. The information processing apparatus according to claim 2, wherein the behavior information acquisition unit includes a behavior information generation unit that generates expert behavior information, which represents the actual driving manner of the target vehicle as the movement manner of the first type character in the simulation image, from the driving manner of the target vehicle actually driven by the driver so as to achieve the predetermined movement target in a situation in the simulation domain where the moving object moves according to the scenario and the state of the fixed object changes.
7. The environmental information acquisition unit includes an environmental information generation unit that generates environmental information from information representing the video of the simulation area in which a moving object moves and the state of the fixed object changes according to the scenario generated by the driving simulator of the vehicle as the target vehicle, and represents the situation in which a moving object moves on the road in the simulation area and the state of the fixed object changes as a situation in which the second type of character moves and the third type of character changes within the simulation image. The information processing apparatus according to claim 2, wherein the behavior information acquisition unit includes a behavior information generation unit that generates expert behavior information, which represents the driving manner of the vehicle in the video of the driving simulator that displays a video of the simulation area in which a moving object moves according to the scenario and the state of the fixed object changes, and is operated by a driver so that a predetermined movement target is achieved, and the driving manner is represented as the movement manner of the first type of character in the simulation image.
8. An information processing method for constructing a reward system used for reinforcement learning about the driving patterns of a target vehicle, according to the inverse reinforcement learning method, A simulation image generation step generates information representing a simulation image that includes an abstract road map that abstractly represents the simulation area including roads in a planar manner, information of first type characters placed within the roads in the abstract road map that represent target vehicles traveling on the roads in the simulation area, and information of second type characters placed within the roads in the abstract road map that distinguish and represent moving objects moving on the roads in the simulation area according to their attributes, An environmental information acquisition step involves acquiring environmental information that represents the situation in which the moving object moves along the road in the simulation area according to a scenario as the situation in which the second type of character moves within the simulation image. Action information acquisition step: Under the circumstances in which the moving object moves along the road in the simulation domain according to the scenario, an action information acquisition step is obtained in which expert action information is acquired in which the driving manner of the target vehicle, driven so as to achieve a predetermined movement target, is represented as the movement manner of the first type of character in the simulation image. The system includes a learning control step which uses a processor that takes environmental information and behavioral information as input information and determines, according to processing rules, reward information representing the reward used for reinforcement learning regarding the driving behavior of the target vehicle as output information, and updates the processing rules in the processor according to a predetermined learning algorithm so that when the acquired environmental information and the expert behavioral information are used as input information, a certain reward information is obtained as output information, An information processing method wherein the learning control step repeatedly updates the processing rules in the processor while changing the environmental information and corresponding expert behavior information based on the above scenario.
9. An information processing method for constructing a reward system used for reinforcement learning about the driving patterns of a target vehicle, according to the inverse reinforcement learning method, A simulation image generation step that generates information representing a simulation image including an abstract road map that abstractly represents a simulation area including roads in a planar manner, information of type 1 characters placed within the roads in the abstract road map that represent target vehicles traveling on the roads in the simulation area, information of type 2 characters placed within the roads in the abstract road map that distinguish and represent moving objects moving on the roads in the simulation area according to their attributes, and information of type 3 characters placed within the abstract road map that distinguish and represent fixed objects in the simulation area whose state changes over time according to their state, An environmental information acquisition step involves acquiring environmental information in which the situation in which the moving object moves along the road in the simulation area according to a scenario and the state of the fixed object changes is represented in the simulation image as a situation in which the second type of character moves and the third type of character changes. An action information acquisition step involves acquiring expert action information, which represents the driving pattern of the target vehicle, driven in such a way that a predetermined movement target is achieved, as the movement pattern of the first type of character in the simulation image, under conditions in which the moving object moves along the road in the simulation domain according to the scenario and the state of the fixed object changes; The system includes a learning control step which uses a processor that takes environmental information and behavioral information as input information and determines, according to processing rules, reward information representing the reward used for reinforcement learning regarding the driving behavior of the target vehicle as output information, and updates the processing rules in the processor according to a predetermined learning algorithm so that when the acquired environmental information and the expert behavioral information are used as input information, a certain reward information is obtained as output information, An information processing method wherein the learning control step repeatedly updates the processing rules in the processor while changing the environmental information and corresponding expert behavior information based on the above scenario.
10. A learning device that performs reinforcement learning regarding the driving patterns of a target vehicle, The information processing device according to claim 1 has a processor obtained by repeatedly updating the processing rules according to the predetermined learning algorithm as a rewarder, A simulation image generation unit generates information representing a simulation image that includes an abstract road map that abstractly represents the simulation area including roads in a planar manner, information of first-type characters placed within the roads in the abstract road map to represent target vehicles traveling on the roads in the simulation area, and information of second-type characters placed within the roads in the abstract road map to distinguish and represent moving objects moving on the roads in the simulation area according to their attributes, A simulation control unit controls information representing the simulation image such that a first type of character moves within the roads of the abstracted road map, and a second type of character moves within the roads of the abstracted road map according to a scenario. In the simulation image of the second type of character moving within the roads of the abstract road map according to the scenario, a controller determines the movement pattern of the first type of character based on a policy so that a predetermined movement target for the first type of character is achieved, and instructs the simulation control unit to move the first type of character relative to the abstract road map in that movement pattern. The system includes a learning processing unit that updates the policy in the controller based on reward information, which is the output information of the rewarder, with input information being behavior information representing the movement pattern of the first type character determined by the controller and environmental information representing the state in which the second type character moves within the roads of the abstracted road map, according to a predetermined reinforcement learning algorithm. A learning device that repeatedly issues movement instructions for the first type of character in a movement mode determined by the policy from the controller, and updates the policy based on reward information from the rewarder, until the movement target is achieved.
11. A learning device that performs reinforcement learning regarding the driving patterns of a target vehicle, The information processing device according to claim 2 has a rewarder which is obtained by repeatedly updating the processing rules according to the predetermined learning algorithm, A simulation image generation unit generates information representing a simulation image that includes an abstract road map that abstractly represents a simulation area including roads in a planar manner, information of first-type characters placed within the roads in the abstract road map to represent target vehicles traveling on the roads in the simulation area, second-type characters placed within the roads in the abstract road map to distinguish and represent moving objects moving on the roads in the simulation area according to their attributes, and third-type characters placed within the abstract road map to distinguish and represent fixed objects in the simulation area whose state changes over time according to their state. A simulation control unit controls information representing the simulation image such that a first type of character moves within the roads of the abstracted road map, and a second type of character moves within the roads of the abstracted road map and a third type of character changes according to a scenario. In the simulation image in which the second type of character moves within the roads of the abstract road map according to the scenario and the third type of character changes in the abstract road map, a controller determines the movement mode of the first type of character based on a policy so that a predetermined movement target of the first type of character is achieved, and instructs the simulation control unit to move the first type of character relative to the abstract road map in that movement mode. The system includes a learning processing unit that updates the policy in the controller based on reward information, which is the output information of the rewarder, with input information being behavior information representing the movement pattern of the first type character determined by the controller according to a predetermined reinforcement learning algorithm, and environmental information representing the state in which the second type character moves within the roads of the abstracted road map and the third type character changes in the abstracted road map. A learning device that repeatedly issues movement instructions for the first type of character in a movement mode determined by the policy from the controller, and updates the policy based on reward information from the rewarder, until the movement target is achieved.
12. The simulation image represented by the information generated by the simulation image generation unit is formed on a grid composed of multiple squares, The movement pattern of the first type of character determined by the controller is expressed as movement in units of the grid squares, The learning device according to claim 10 or 11, wherein the simulation control unit controls the information representing the simulation image so that the first type of character moves relative to the abstract road map in units of the grid, according to instructions from the controller.
Citation Information
Patent Citations
Traffic light display device
JP2015076016A
Learning device, learning method, and program
JP2020035222A
Mobile body control device, mobile body control method, and program
JP2021155006A
Prediction device, prediction method, program and vehicle control system
JP2021196632A
System and method for managing flexible control of vehicles by diverse agents in autonomous driving simulation
US20220032935A1