Simulator, learning device, and vehicle control device
The abstract road map simulator simplifies vehicle reinforcement learning by using grid-based representations of roads and objects, addressing computational challenges and enabling efficient learning.
Patent Information
- Application Number
- JP2021123317
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-28
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2041-07-28
AI Technical Summary
Existing simulators for vehicle reinforcement learning require significant computational resources to faithfully represent real-world environments, making it difficult to incorporate into a simple and efficient learning framework.
A simulator using an abstract road map that represents roads in a two-dimensional form, distinguishing vehicles and objects by attributes, and simulating their movements within a grid-based environment, allowing for efficient reinforcement learning.
Enables efficient reinforcement learning by simplifying the representation of vehicle driving behavior, reducing computational requirements while maintaining accuracy in simulating real-world scenarios.
Smart Images

Figure 0007746058000001 
Figure 0007746058000002 
Figure 0007746058000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a simulator used for reinforcement learning of vehicle driving behavior, a learning device using the simulator, and a vehicle driving control device that controls vehicle driving using a strategy obtained by reinforcement learning in the learning device. [Background technology]
[0002] Conventionally, a learning device that performs reinforcement learning on the behavior (driving behavior) of a vehicle is known (see, for example, Patent Document 1). Generally, reinforcement learning involves learning "behavior that maximizes value" in a certain environment through trial and error. Specifically, an agent (a controller of a behavioral entity) determines its behavior in a certain environment based on a policy. The behavior affects the environment, and the behavior is evaluated based on the changed environment influenced by the behavior to determine whether the behavior was good or not. The evaluation result is given to the agent as a reward. The policy is then updated based on the evaluation result (reward). Thereafter, the process of determining the agent's behavior based on the policy, evaluating (rewarding) the behavior in the environment influenced by the behavior, and updating the policy based on the evaluation result is repeated in sequence, and the policy is updated (learned) in sequence so as to maximize the final reward (evaluation).
[0003] In the above-described reinforcement learning, a simulator can be used as a means for providing the environment. For example, a simulator used in reinforcement learning about the driving behavior (behavior) of a vehicle is configured by a computer and provides a virtual space including roads in the area surrounding the driving (behavior) simulated vehicle, various moving objects (other vehicles, pedestrians, bicycles, emergency vehicles, motorcycles, etc.) moving on the road, traffic lights, convex mirrors, buildings, etc. Then, by causing the simulated vehicle to drive (behave) in the virtual space (environment) configured by the simulator, an optimal policy for determining the driving behavior (behavior) of the simulated vehicle is learned and determined. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2020-35222 Summary of the Invention [Problem to be solved by the invention]
[0005] The results (policies) of reinforcement learning on the driving behavior (behavior) of a vehicle are used for driving control of an actual vehicle. Therefore, it is preferable that the virtual space provided by the simulator used in the reinforcement learning reproduces the actual environment (roads, other vehicles, pedestrians, bicycles, emergency vehicles, etc.) as faithfully as possible.
[0006] However, simulators that provide virtual spaces that more faithfully reproduce the real environment must handle a wide range of information, such as the relative positions of detected objects (obstacles such as other vehicles, infrastructure such as white lines and traffic lights) relative to the target vehicle, as well as the attributes and status of detected objects in the virtual space (people, cars, trucks, blinking turn signals, emergency driving status, signal lights on), resulting in a large number of dimensions of information. This requires a large amount of computer resources (CPU power, software, memory capacity, etc.), making it difficult to incorporate into a simple and efficient reinforcement learning framework. This makes efficient and effective learning impossible.
[0007] The present invention has been made in view of the above circumstances, and provides a simulator that can incorporate vehicle driving behavior into a simple and efficient reinforcement learning framework.
[0008] The present invention also provides a learning device that uses such a simulator.
[0009] Furthermore, the present invention provides a vehicle driving control device that controls the driving of a vehicle using the learning results (policies) obtained by reinforcement learning in the learning device. [Means for solving the problem]
[0010] The simulator according to the present invention is a simulator used for reinforcement learning of the driving behavior of a target vehicle, and is configured to use an abstract road map that represents a simulation area including roads in a two-dimensional abstract form. Information a first type representing the target vehicle and positioned within a road in the abstract road map; Character information and a second type of vehicle that is arranged on the road in the abstract road map by distinguishing and representing the vehicle that moves on the road in the simulation area according to its attributes. Character information A simulation image including Love a simulation image generating unit that generates the first type of information; character The second type of the abstract road map is displayed according to a scenario while allowing movement relative to the abstract road map. character to move within the roads of the abstract road map. Represents a simulation image and a simulation control unit that controls the information.
[0011] This configuration allows the creation of an abstract road map, which represents the real simulation area including roads in a two-dimensional abstract form, and a first-class map representing the target vehicle. character and a second type in which moving objects moving on roads in the simulation area are distinguished and represented according to their attributes. character and the first type character The second type of the abstract road map is displayed according to a scenario while allowing movement relative to the abstract road map. character A simulated image of a vehicle moving within the roads of the abstract road map is provided.
[0012] In the simulator according to the present invention, the simulation image Information representing A third type of object whose state changes over time in the simulation area is represented so as to be distinguished according to its state and is fixedly arranged in the abstract road map. Character information the simulation control unit executes the third type of simulation according to a scenario. character As the above changes simulation image Represents You can control information.
[0013] This configuration allows the creation of an abstract road map, which represents the real simulation area including roads in a two-dimensional abstract form, and a first-class map representing the target vehicle. character a second type in which moving objects moving on roads in the simulation area are distinguished and represented according to their attributes; character and a third type in which an object whose state changes over time in the simulation domain is represented in a manner that distinguishes it from the other object's state. character and the first type character The second type of the abstract road map is displayed according to a scenario while allowing movement relative to the abstract road map. character moves within the road of the abstract road map and the third type character A simulation image is provided in which the
[0015] The learning device according to the present invention is a learning device that performs reinforcement learning on the driving mode of a target vehicle using the simulator (claim 1), and in the simulation image of a state in which the second type characters move within the road of the abstract road map according to the scenario, character In order to achieve the moving target of the first type, character The movement mode of the first type is determined based on the policy, and the determined movement mode is used. character a controller that instructs the simulation control unit of the simulator to move relative to the abstract road map; and a controller that is controlled by the simulation control unit according to a predetermined reinforcement learning algorithm. The aforementioned The second type of information is displayed according to the scenario in the simulation image. character The first type of vehicle moves within the road of the abstract road map. character and based on the evaluation result, the controller controls the first type character a learning processing unit that updates the strategy used to determine the movement mode of the first-class character, and character The system is configured to repeatedly evaluate the movement patterns of the robots and update the measures based on the evaluation results.
[0016] With this configuration, the simulator can perform the second type of simulation according to the scenario. character In the simulation image of the state in which the user moves on the road of the abstract road map, character The vehicle moves relative to the abstract road map in a movement manner determined based on the strategy so that a predetermined movement goal is achieved. character The first type in an environment where the user moves within the roads of an abstract road map character The movement pattern of the vehicle is evaluated, and based on the evaluation results, the first type character The strategy used to determine the movement mode of the first type robot is updated. Then, the first type robot is moved in the movement mode determined based on the strategy until the movement goal is achieved. character Movement instruction, the first type character The evaluation of the movement patterns of the vehicles and the updating of the measures based on the evaluation results are repeated. In this way, the second type of vehicle is developed according to the scenario on the roads of the abstract road map. character Type 1 in a moving environment character The optimal strategy for achieving the movement goal is determined (learned).
[0017] The learning device according to the present invention is a learning device that performs reinforcement learning on the driving behavior of a target vehicle using the simulator (claim 2), and character moves within the road of the abstract road map and the third type character In the simulation image of the state in which the state in which the abstract road map is changed, the predetermined first type character In order to achieve the moving target of the first type, character The movement mode of the first type is determined based on the policy, and the determined movement mode is used. charactera controller that instructs the simulation control unit of the simulator to move; and a controller that is controlled by the simulation control unit according to a predetermined reinforcement learning algorithm. The aforementioned The second type of information is generated according to the scenario in the simulation image represented by the information. character moves within the road of the abstract road map and the third type character The first type of the above-mentioned under the environment where the above-mentioned abstract road map is changed character and based on the evaluation result, the controller controls the first type character a learning processing unit that updates the strategy used to determine the movement mode of the first type robot, and the first type robot is moved in the movement mode determined based on the strategy by the controller until the movement target is achieved. character the movement instruction of the first type by the learning processing unit, character The system is configured to repeatedly evaluate the movement patterns of the robots and update the measures based on the evaluation results.
[0018] With this configuration, the simulator can perform the second type of simulation according to the scenario. character Moves within the road of the abstract road map and character In the simulation image of the state where the abstract road changes, character The vehicle moves relative to the abstract road map in a movement manner determined based on the strategy so that a predetermined movement goal is achieved. character Moves within the road of the abstract road map and character The first type in an environment where the abstract road map changes character The movement pattern of the vehicle is evaluated, and based on the evaluation results, the first type character The strategy used to determine the movement mode of the first type robot is updated. Then, the first type robot is moved in the movement mode determined based on the strategy until the movement goal is achieved. character Movement instruction, the first type character The evaluation of the movement mode of the vehicle and the update of the policy based on the evaluation result are repeated. In this way, the second type of vehicle is selected according to the scenario within the road of the abstract road map. characterAs the number of people moves, the third type character Type 1 in a changing environment character The optimal strategy for achieving the movement goal is determined (learned).
[0019] In the learning device according to the present invention, the simulation control unit in the simulator executes the first type of simulation in accordance with an instruction from the controller. character before representing the simulation image as moving relative to the fixed abstract road map. Memoirs The information can be configured to control the
[0020] With this configuration, in the simulation image generated by the simulator, character moves relative to a fixed abstract road map in a movement manner determined based on a strategy.
[0021] In the learning device according to the present invention, the simulation control unit in the simulator controls the first type of the simulation model fixed in accordance with an instruction from the controller. character The simulation image is displayed so that the abstract road map moves relative to the Information The information can be configured to control the
[0022] With this configuration, in a simulation image generated by the simulator, the first type of movement is determined based on the strategy. character The first fixed object is moved within the road of the abstract road map. character The abstract road map moves relative to the
[0023] In the learning device according to the present invention, The aforementioned The simulation image represented by the information is formed on a grid consisting of a plurality of squares, and the first type of information determined by the controller characterThe movement mode of the first type robot is expressed as a movement with the squares of the grid as units, and the simulation control unit of the simulator moves the first type robot in accordance with an instruction from the controller. character is moved in units of squares of the grid relative to the abstract road map. Represents a simulation image The information can be controlled and configured.
[0024] With this configuration, in a simulation image generated by the simulator, the first type of movement is determined based on the strategy. character moves to each grid cell on the abstract road map. This allows the first type of vehicle to be represented by the target vehicle. character The abstraction of the road map allows for a simpler representation of movement within the roads.
[0025] In the learning device according to the present invention, The aforementioned The simulation image represented by the information is formed on a grid consisting of a plurality of squares, and the first type of information determined by the controller character The movement mode of the first type robot is expressed as a movement with the squares of the grid as units, and the simulation control unit of the simulator moves the first type robot in accordance with an instruction from the controller. character The above-mentioned method is performed so as to move the grid squares in units of the abstract road map to which the above-mentioned method is fixed. Represents a simulation image The information can be controlled and configured.
[0026] With this configuration, in a simulation image generated by the simulator, the first type of movement is determined based on the strategy. character The target vehicle moves on the road of the abstract road map in units of grid squares. This makes it possible to more simply express the movement of the first-class symbol representing the target vehicle on the road of the abstract road map.
[0027] In the learning device according to the present invention, The aforementioned The simulation image represented by the information is formed on a grid consisting of a plurality of squares, and the first type of information determined by the controller character The movement mode of the first type is expressed as a movement with the squares of the grid as a unit, and the simulation control unit of the simulator is fixed in accordance with an instruction from the controller. character The abstract road map moves in units of squares of the grid. Represents a simulation image The information can be controlled and configured.
[0028] With this configuration, in a simulation image generated by the simulator, the first type of movement is determined based on the strategy. character The first fixed road map is moved in units of grid squares within the roads of the abstract road map. character The abstract road map moves relative to the target vehicle. character The abstraction of the road map allows for a simpler representation of movement within the roads.
[0029] In the learning device according to the present invention, the grid is character The actual distance represented by each square in a predetermined area including the first type is calculated from the predetermined area. character The actual distance represented by each square in the area farther away from the can be .
[0030] With this configuration, a first type of image representing the target vehicle is generated in the simulation image generated by the simulator. character The first type is expressed as a movement of a square unit in a predetermined area including character It is possible to construct an abstract road map of a wider range (simulation area) on the grid while maintaining the resolution of the movement pattern of the vehicle. Love This can reduce the memory size of the information.
[0031] A vehicle driving control device according to the present invention is a vehicle driving control device that is mounted on a vehicle and controls driving of the vehicle, and includes a map acquisition unit that acquires road map information that represents a road map of a predetermined surrounding area of the vehicle, an environment information acquisition unit that acquires surrounding environment information that represents the state of the environment including moving objects in the predetermined surrounding area, a vehicle state acquisition unit that acquires vehicle state information that represents the traveling position of the vehicle and the traveling state of the vehicle, and an abstract road map that abstracts the predetermined surrounding area in a two-dimensional manner based on the road map information, the surrounding environment information, and the vehicle state information. Information a first type representing the vehicle moving within a road in the abstract road map; Character information and a second type of vehicle that moves on the roads in the predetermined surrounding area and that distinguishes and represents the vehicle according to its attributes and moves within the roads in the abstract road map. Character information Abstract surrounding area image including Love a surrounding area image generating unit for generating information on the surrounding area image before the information is generated by the surrounding area image generating unit; Memoirs In the abstracted surrounding area image represented by the information, the first type is character and an actual driving control unit that executes driving control of the vehicle so that the vehicle moves in the driving mode determined by the vehicle driving mode determination unit.
[0032] With this configuration, road map information showing a road map of a predetermined surrounding area of the vehicle, surrounding environment information showing the state of the environment including the moving object in the predetermined surrounding area, and vehicle state information showing the traveling position of the vehicle and the traveling state of the vehicle are acquired. Based on these pieces of information, an abstract road map showing a planar abstraction of the predetermined surrounding area corresponding to the navigation image used in the learning device (claim 6), a first type of image showing the vehicle moving on the road in the abstract road map, and character and a second type of vehicle that moves on the roads in the predetermined surrounding area and that distinguishes and represents the vehicle according to its attributes and moves within the roads in the abstract road map. characterAbstract surrounding area image including Love Then, a first type of information is generated in the abstracted surrounding area image based on the strategy obtained by the learning device (claim 6). character The driving mode of the vehicle is determined so as to move the vehicle. The driving control of the vehicle is performed so as to achieve driving in the determined driving mode.
[0033] The abstract road map mentioned above, type 1 character and Type 2 character Based on the learning results (strategies) obtained by reinforcement learning using simplified simulation images including the above, vehicle driving control can be performed so that a driving pattern that is suited to the actual conditions of a specified surrounding area is obtained.
[0034] Furthermore, a vehicle driving control device according to the present invention is a vehicle driving control device that is mounted on a vehicle and controls driving of the vehicle, and includes a map acquisition unit that acquires road map information that represents a road map of a predetermined surrounding area of the vehicle, an environment information acquisition unit that acquires surrounding environment information that represents a state of an environment including moving bodies in the predetermined surrounding area and objects that are installed in the predetermined surrounding area and whose state changes over time, a vehicle state acquisition unit that acquires vehicle state information that represents the traveling position of the vehicle and the traveling state of the vehicle, and an abstract road map that abstracts the predetermined surrounding area in a two-dimensional manner based on the road map information, the surrounding environment information, and the vehicle state information. Information a first type representing the vehicle moving within a road in the abstract road map; Character information, previous A second type of vehicle that moves on the roads in the specified surrounding area is displayed according to its attributes and moves on the roads in the abstract road map. Character information and a third type of object that is installed in the predetermined surrounding area and changes over time, and that is displayed in a fixed manner in the abstract road map according to its state. Character information Abstract surrounding area image including Love a surrounding area image generating unit for generating information on the surrounding area image before the information is generated by the surrounding area image generating unit; MemoirsIn the abstracted surrounding area image represented by the information, the first type of character and an actual driving control unit that executes driving control of the vehicle so that the vehicle moves in the driving mode determined by the vehicle driving mode determination unit.
[0035] With this configuration, road map information showing a road map of a predetermined surrounding area of the vehicle, and surrounding environment information showing the state of the environment including moving objects in the predetermined surrounding area and objects arranged in the predetermined surrounding area whose state changes over time are acquired. Based on this information, the predetermined surrounding area is abstracted in a planar form, corresponding to the navigation image used in the learning device (claim 7), table an abstract road map; a first type of vehicle representing the vehicle moving within roads in the abstract road map; character a second type of vehicle that moves on the roads in the predetermined surrounding area and that is distinguished and represented according to its attributes and moves within the roads in the abstract road map; character and a third type of object that is arranged in the predetermined surrounding area and whose state changes over time, and that is fixedly arranged in the abstract road map by distinguishing the object according to its state. character Then, an abstracted surrounding area image including a first type of image is generated in the abstracted surrounding area image based on the strategy obtained by the learning device (claim 7). character The driving mode of the vehicle is determined so as to move the vehicle. The driving control of the vehicle is performed so as to achieve driving in the determined driving mode.
[0036] The abstract road map mentioned above, type 1 character , Type 2 character and Type 3 character Based on a strategy obtained by reinforcement learning using simplified simulation images including the above, vehicle driving control can be performed so as to obtain a driving pattern that is suited to the conditions of a specific actual surrounding area. [Effects of the Invention]
[0037] According to the simulator of the present invention, an abstract road map that abstractly represents a simulation area including roads in a planar form, a first type of map that represents a target vehicle, character (corresponding to the own vehicle), and the second type, which distinguishes and represents moving objects according to their attributes. character Since the vehicle's driving behavior is represented by simple simulation images that include other vehicles, pedestrians, bicycles, emergency vehicles, etc., it can be incorporated into a simple and efficient reinforcement learning framework.
[0038] Furthermore, the learning device according to the present invention makes it possible to carry out reinforcement learning of the running behavior of a vehicle using the simulator.
[0039] Furthermore, according to the vehicle driving control device of the present invention, it is possible to control the driving of a vehicle so as to obtain a driving pattern that is suited to the conditions of a specified actual surrounding area, based on the learning results (strategies) obtained by reinforcement learning using simplified simulation images. [Brief explanation of the drawings]
[0040] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a learning system including a learning device and a simulator according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the configuration of a simulator according to an embodiment of the present invention. [Figure 3] FIG. 3 is a block diagram showing the configuration of a learning device according to an embodiment of the present invention. [Figure 4] FIG. 4 is a diagram showing a first example of a simulation image formed on a grid. [Figure 5] FIG. 5 is a diagram showing a second example of a simulation image. [Figure 6] FIG. 6 is a diagram showing an example of the movement of the first symbol (open pentagon) in the simulation image. [Figure 7]FIG. 7 is a diagram showing a third example of a simulation image. [Figure 8] FIG. 8 is a diagram showing a fourth example of the simulation image. [Figure 9] FIG. 9 is a diagram showing a fifth example of a simulation image. [Figure 10] FIG. 10 is a diagram showing a sixth example of the simulation image. [Figure 11] FIG. 11 is a diagram showing another example of the configuration of the grid formed in the background of the simulation image. [Figure 12] FIG. 12 is a diagram showing an example of an image of a certain layer in a simulated image of a multilayer structure. [Figure 13] FIG. 13 is a diagram showing an example of an image of another layer in a simulation image of a multilayer structure. [Figure 14] FIG. 14 is a diagram showing an example of a simulation image of a two-layer structure in which the image of the image shown in FIG. 12 and the image of the layer shown in FIG. 13 are superimposed. [Figure 15] FIG. 15 is a block diagram showing the configuration of a vehicle driving control device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0042] A learning system including a learning device and a simulator according to an embodiment of the present invention is configured as shown in FIG.
[0043] In FIG. 1, the learning system 10 is configured as a computer system including various hardware and software components, and includes a control unit 11, a simulator 12, a learning device 13, and a library 14. The system performs reinforcement learning on the driving behavior of a vehicle (target vehicle). Under the control of the control unit 11, the simulator 12 generates simulation images that provide an environment for reinforcement learning, and controls image information representing the simulation images according to a scenario. Under the control of the control unit 11, the learning device 13 performs reinforcement learning processing on the driving behavior of the target vehicle in accordance with a predetermined reinforcement learning algorithm in the environment represented by the simulation images that change according to the scenario in the simulator 12. The library 14 stores information representing each of the scenarios and moving targets. The scenario represents a plot of changes in the simulation images that represent the environment of the target vehicle. Furthermore, the moving target represents a target moving behavior of the target vehicle (e.g., "move from a starting point to a target point without colliding with other moving objects (cars, pedestrians, etc.) and without departing from the road, passing through intersections") in the environment represented by the simulation images that change according to the scenario. In the library 14, various scenarios are managed so as to correspond to the moving targets of the target vehicle related to the simulation image that changes according to the scenario. The scenarios and the moving targets corresponding to the scenarios can be set based on the contents of traffic regulations (e.g., the Road Traffic Act), for example.
[0044] The simulator 12 is configured as shown in FIG.
[0045] In FIG. 2, the simulator 12 includes an image processor 12a, a simulation controller 12b, and a monitor 12c. The image processor 12a (simulation image generator) generates image information representing a simulation image. FIG. 4 shows an example of a simulation image. This simulation image is formed on a grid 100 consisting of multiple (n × m) squares. For example, in the grid 100 consisting of 32 × 32 squares, if the actual vertical and horizontal distances represented by each square are set to 2 m, a simulation image covering a simulation area (actual area) of 64 m × 64 m can be formed. The simulation image formed on the grid 100 includes an abstract road map MAP that represents, in a planar abstraction, a simulation area including a road R1 with two lanes on each side and four lanes on both sides, and a road R2 with a single lane, which intersect at an intersection IS. In this abstract road map MAP, two pedestrian crossings PC1 and PC2 (fixed objects) crossing road R1 and two pedestrian crossings PC3 and PC4 (fixed objects) crossing road R2 are depicted surrounding the intersection IS. The simulation image includes a "white pentagon" as a first-class symbol placed on road R1 to represent a target vehicle that is the subject of reinforcement learning. The simulation image also includes second-class symbols placed on roads R1 and R2 of the abstract road map MAP to distinguish moving objects moving on roads in the simulation domain according to their attributes (car, emergency vehicle, bicycle, pedestrian, etc.), including a "black pentagon" representing a car (other than the target vehicle), a "height" symbol representing an emergency vehicle (ambulance, police vehicle, etc.), a "black trapezoid" representing a bicycle, and a "black triangle" representing a pedestrian.
[0046] The simulation image can also be formed, for example, as shown in Fig. 5. The simulation image shown in Fig. 5 includes the same abstract road map MAP as the simulation image shown in Fig. 4. In this simulation image, ASCII characters (English letters) are used instead of graphic marks ("white pentagon mark," "black pentagon mark," "he character graphic mark," "black trapezoid mark," and "black triangle mark") as first-class symbols and second-class symbols. Specifically, "E" is used as the first-class symbol representing a target vehicle and placed on road R1, and "C" representing a car (other than the target vehicle), "A" representing an emergency vehicle, "B" representing a bicycle, and "W" representing a pedestrian are used as second-class symbols.
[0047] Returning to FIG. 2, the simulation control unit 12b of the simulator 12 controls the image information representing the simulation image so that second-class symbols (automobile: black pentagon mark "C", emergency vehicle: "he" shape mark "A", bicycle: black trapezoid mark "B", pedestrian: black triangle mark "W") move on roads R1 and R2 of the abstract road map MAP in accordance with a scenario provided from the control unit 11 (see FIG. 1), while allowing the first-class symbol representing the target vehicle ("white pentagon mark", "E") to move within the abstract road map MAP. The simulation images shown in FIGS. 4 and 5, which are represented by image information controlled in this way, may be generated, for example, in the following scenario: An emergency vehicle (marked with a "へ" shape and "A") traveling in the center lane of road R1 passes through intersection IS with its siren blaring. Vehicles entering the intersection (IS) (black pentagon mark, C) should move to the shoulder and stop, and begin moving once an emergency vehicle ('へ' shape mark, A) has passed. Vehicles already in the intersection (IS) (black pentagon mark, C) must continue moving and exit the intersection (IC), then stop on the shoulder of the road and resume moving once the emergency vehicle ('へ' shape mark, A) has passed. Bicycles moving in the opposite lane on the shoulder of the road (black trapezoid mark "B") and pedestrians moving on crosswalk PC4 (black triangle mark "W") should maintain their movement. According to the above, the second type symbol (cars: black pentagon mark "C", emergency vehicles: "he" shape mark "A", bicycles: black trapezoid mark "B", pedestrians: black triangle mark "W") will change and move accordingly.
[0048] The monitor 12c displays a simulation image (see FIG. 4 or 5) that changes based on image information generated by the image processing unit 12a and controlled by the simulation control unit 12b. Through the simulation image displayed on the monitor 12c, the operator can visually confirm the reinforcement learning environment and the movement of the first-class symbol (target vehicle) that is the learning target.
[0049] Next, the learning device 13 is configured as shown in FIG.
[0050] In FIG. 3, the learning device 13 includes a controller 13a and a learning processing unit 13b. The controller 13a determines the movement mode (driving mode) of the first-class symbol (target vehicle) based on a policy so that a movement target corresponding to a scenario is achieved in a simulation image of a state (environment) in which the second-class symbol moves on roads R1 and R2 of an abstract road map (see FIG. 4 or 5) according to the scenario. The controller 13a then instructs the simulation control unit 12b of the simulator 12 to move the first-class symbol on the abstract road map in the determined movement mode (movement instruction). In the simulator 12, the simulation control unit 12b (see FIG. 2) controls the image information generated by the image processing unit 12b based on instructions from the learning device 13 (controller 13a) so that the first-class symbol moves in an environment in which the second-class symbol moves on roads R1 and R2 of the abstract road map according to the scenario.
[0051] The movement of the first type symbol can be represented as a movement in units of squares on the grid 100, as shown in Fig. 6. Specifically, the movement of the first type symbol can be represented by specifying the position of a square on the grid 100 (the position of the first type symbol) at a predetermined cycle (for example, every 0.5 seconds). For example, if the actual vertical and horizontal distance represented by each square on the grid 100 is 2m square, the movement of the first type symbol corresponding to the straight-ahead driving of the target vehicle at a speed of 60km / h is represented by a position that changes every four squares on the grid 100 at a cycle of 0.5 seconds.
[0052] The controller 13a has a neural network (comprised of an input layer, an intermediate layer (hidden layer), and an output layer). The coefficients between each node of this neural network correspond to a "policy" for determining the movement mode of the first-type symbol. In the controller 13a, information on the state (environment) of the simulation image (see FIG. 4 or FIG. 5) that changes according to a scenario in the simulator 12, a moving target corresponding to the scenario, and the movement mode of the first-type symbol in the simulation image that changes according to the scenario (the position of a square on an abstract road map) is input to the input layer of the neural network. Then, information determined based on the coefficients between each node from the information input to the input layer is output from the output layer of the neural network as information representing the new position (square position) of the first-type symbol on the abstract road map (on the grid 100). This new position is given to the simulator 12 (simulation control unit 12b: see FIG. 2) as a movement instruction for the first-type symbol.
[0053] The learning processing unit 13b in the learning device 13 evaluates the movement pattern of the first type symbol in the simulation image (see FIG. 4 or FIG. 5) (evaluation of behavior) according to a predetermined reinforcement learning algorithm (e.g., DeepQ-Network, MuZero, etc.), and updates the coefficients between the nodes of the neural network in the controller 13a (corresponding to the "policy" used to determine the movement pattern of the first type symbol) based on the evaluation result. The evaluation of the movement pattern of the first type symbol is performed by a reward given for that movement pattern.
[0054] The degree to which traffic rules and manners are observed can be scored, and the number of points can be used as a reward. For example, the reward can be If the first-type symbol does not overlap with the second-type symbol at the destination position...+ point a If the first-type symbol overlaps with the second-type symbol at the destination position... -point a When the distance between the first and second symbols is less than a predetermined value... -b point When the first type symbol deviates from the road... - point c If the first-class sign stops at the stop line on the road...+d point When the first type symbol moves in a direction (straight to the side) that would never occur in a real vehicle... - point e If the first symbol achieves the movement target, +d point It can be set as such.
[0055] In the learning device 13, until the movement goal of the first type symbol is achieved, for example, until the first type symbol "moves from the starting point through an intersection to the target point without colliding with a second type symbol (car, pedestrian, etc.) and without deviating from the road (designated vehicle lane)," the controller 13a repeatedly issues movement instructions for the first type symbol based on a "strategy" (using a neural network) at a predetermined period (for example, 0.5 second period), the learning processing unit 13b evaluates (gives rewards to) the movement pattern of the first type symbol, and updates the "strategy" (coefficients between each node in the neural network) based on the evaluation results (number of points) so as to maximize the total amount of rewards (scores).
[0056] During this process, each time the first-type symbol overlaps (including contact with) a second-type symbol (such as a car or an ambulance) in the simulation image or deviates from the road (a designated vehicle lane), the movement of the first and second-type symbols in the simulation image is stopped, and then the second-type symbol resumes moving from the beginning according to the scenario, and the image information representing the simulation image is controlled so that the first-type symbol resumes moving from its starting point. Then, when the movement goal of the first-type symbol is achieved, the reinforcement learning process in the learning device 13 using the simulator 12 is terminated. As a result, a "strategy" is determined that enables the first-type symbol (target vehicle) to achieve its movement goal in the environment represented by the simulation image (the second-type symbol moving within an abstract road map) that changes based on the scenario, i.e., the coefficients between each node of the neural network of the controller 13a that will achieve this strategy are determined.
[0057] According to the simulator 12 described above, the environment to be subjected to reinforcement learning is represented by a simple simulation image including an abstract road map MAP, which represents a two-dimensional abstraction of the simulation domain including roads R1 and R2, a first-class symbol (a "white pentagon" or "E") representing a target vehicle, and a second-class symbol (a "black pentagon" or "C" for automobiles, a "height" shape or "A" for emergency vehicles, a "black trapezoid" or "B" for bicycles, and a "black triangle" or "W" for pedestrians) representing moving objects according to their attributes. Therefore, the driving behavior (movement behavior) of the target vehicle (first-class symbol) can be incorporated into a simple and efficient reinforcement learning framework. The learning device 13 can then use the simulator 12 to perform reinforcement learning (updating and determining the "policy") on the driving behavior of the vehicle (target vehicle).
[0058] The simulation image represented by the image information generated by the image processing unit 12a of the simulator 12 and controlled by the simulation control unit 12b is not limited to the one shown in Fig. 4 or 5, and various types can be determined depending on the environment to be subjected to reinforcement learning. For example, as shown in Fig. 7, the simulation image formed on the grid 100 includes an abstract road map MAP that represents, in a planar abstraction, a simulation area including a road R having two lanes on each side and four lanes on each side. The simulation image also includes a "white pentagon" as a first type symbol that represents a target vehicle and is placed in the shoulder-side vehicle lane of the road R, and a "black pentagon with a left black arrow" as a second type symbol that represents a vehicle with its left turn signal flashing and is placed in the center-side vehicle lane of the road R. In addition, in this simulation image, as shown in Figure 8, the English letter "E" can be used instead of the "white pentagon mark" as the first type symbol, and the English letter "L" can be used instead of the "black pentagon mark with a left black arrow" to represent a car with the same left turn indicator flashing.
[0059] The simulator 12 (image processing unit 12a, simulation control unit 12b) can control the image information representing the simulation image (see Figure 7 or Figure 8) so that the "black pentagon with left arrow" (or "L") as a second type symbol moves in accordance with a scenario such as, for example, "a vehicle ("black pentagon with left arrow" ("L")) traveling in the center lane of road R with its left turn indicator flashing, enters (cuts in) into the shoulder lane and continues traveling." Furthermore, a learning device 13 (controller 13a, learning processing unit 13b) using such a simulator 12 moves a first-type symbol (target vehicle: "white pentagon mark" or "E") in an environment represented by a simulation image (a second-type symbol moving within an abstract road map) that changes according to the scenario, and determines a "measure" that enables the first-type symbol to achieve a movement goal, for example, "to move from the starting point to the target point without colliding with a second-type symbol (automobile: "pentagon mark with black arrow," "L"), without leaving the shoulder vehicle lane, and while maintaining a predetermined distance from the second-type symbol (automobile: "pentagon mark with black arrow," "L") entering ahead," i.e., the coefficients between each node of the neural network (controller 13a), according to a predetermined reinforcement learning algorithm.
[0060] It is also possible to generate a simulation image such as that shown in Fig. 9. This simulation image includes an abstract road map MAP that represents, in a planar abstraction, a simulation area including roads R1 and R2, each of which is a single vehicle lane that intersects at an intersection IS, a "white pentagon mark" as a first type symbol that represents a target vehicle and is placed on road R1, and a "black pentagon mark" as a second type symbol that is placed on road R2.
[0061] Incidentally, the state of a convex mirror that is visible to a target vehicle differs depending on whether it reflects a moving object such as a vehicle (time) or not (time). Taking this into consideration, the simulation image includes a "convex mirror graphic mark" fixedly placed at a predetermined corner of an intersection IS on the abstract road map MAP as a third type of symbol that distinguishes the convex mirror, whose state changes over time, according to its state. This "convex mirror graphic mark" is a graphic mark consisting of a black pentagon within a circle (see Figure 9) when a moving object is reflected, and is a simple circle mark (not shown) when a moving object is not reflected.
[0062] In such a simulation image, the English letter "E" can be used in place of the "white pentagon mark" as the first type symbol, and the English letter "C" can be used in place of the "black pentagon mark" as the second type symbol, as shown in Fig. 10. Furthermore, the English letter string "MI" can be used in place of the "convex mirror graphic mark" as the third type symbol, indicating that a moving object is reflected in the convex mirror, and the English letter "M" (not shown) can be used in place of the "convex mirror graphic mark" as the third type symbol.
[0063] 9 and 10 include a field of view range line (two dotted lines) extending from a first-class symbol (a "white pentagon mark," "E") that represents the field of view range α from the target vehicle. This simulation image shows a situation in which a target vehicle heading toward intersection IS on road R1 cannot directly see vehicles heading toward the intersection on road R2, but can see vehicles heading toward the intersection on road R2 through an image in a convex mirror.
[0064] The simulator 12 (image processing unit 12a, simulation control unit 12b) can control the image information representing the simulation image (see Figure 9 or Figure 10) so that, for example, the "black pentagon mark" (or "C") as a second-type symbol moves and the "convex mirror graphic mark" as a third-type symbol changes (from "graphic mark with a black pentagon placed in a circle" to "circle mark," or from "MI" to "M") according to a scenario in which "a car ("black pentagon mark," "C") traveling on road R2 enters intersection IS while being reflected in the convex mirror from outside the field of view α of the target vehicle ("white pentagon mark," "E") on road R1, and passes through the intersection." Furthermore, a learning device 13 (controller 13a, learning processing unit 13b) using such a simulator 12 moves a first-type symbol (target vehicle: "white pentagon mark" or "E") in an environment represented by a simulation image that changes according to the scenario, and determines a "strategy" that enables the first-type symbol to achieve a movement goal, for example, "to move from the starting point to the target point, passing through intersection IS without colliding with a second-type symbol (automobile: "black pentagon mark" or "C") moving on road R2, and without deviating from road R1," i.e., the coefficients between each node of the neural network (controller 13a), according to a predetermined learning algorithm.
[0065] In the above-described embodiment, the first type symbols move relative to the fixed abstract road map MAP, but the image information representing the simulation image may be controlled so that the abstract road map MAP moves relative to the fixed first type symbols. In this case, the image information representing the simulation image is controlled so that the second type symbols move according to a scenario relative to the moving abstract road map MAP, and so that the third type symbols remain stationary relative to the moving abstract road map MAP.
[0066] As described above, when the abstract road map MAP is moved relative to a fixed first-class symbol, the grid 100 on which the moving abstract road map MAP (simulation image) is formed can be set so that the actual distances (actual vertical and horizontal distances) represented by each square in a region (region outside the predetermined region S) farther from the first-class symbol (e.g., "E") than the actual distances (actual vertical and horizontal distances) represented by each square in a predetermined region S (e.g., a rectangular region) containing the fixed first-class symbol (e.g., "E") are greater, as shown in Fig. 11. Specifically, in the case of the grid 100 shown in Fig. 11, the allocation of actual distances to the squares corresponds to a logarithmic scale with the first-class mark "E" at the center.
[0067] When a simulation image is formed on the grid 100 as described above, it is possible to appropriately maintain the resolution of the movement of the first type symbol ("E") representing the target vehicle in the simulation image, which is expressed as movement in units of squares in a predetermined area S including the first type symbol, and to construct an abstract road map of a wider range (simulation area) on the grid 100. This makes it possible to save the memory size of image information representing a simulation image corresponding to a certain simulation area.
[0068] Note that the grid 100 can be set so that in only one of the two directions, vertical and horizontal, for example, in the vertical direction, the actual distance (vertical actual distance) represented by each square corresponding to a predetermined area S including a first-class symbol (for example, "E") is greater than the actual distance (vertical actual distance) represented by each square corresponding to an area outside the predetermined area S. In this case, in the horizontal direction, the actual distance represented by each square is set to a constant value.
[0069] Furthermore, the actual distances assigned to each square of the grid 100 can be arbitrarily assigned, and different actual distances can be assigned vertically and horizontally. Furthermore, the size of the grid 100 can also be set arbitrarily. For example, the grid 100 can be configured with 64 x 64 squares, which is more than the 32 x 32 squares mentioned above.
[0070] The simulation image represented by the image information generated and controlled in the simulator 12 may have a multi-layer structure.
[0071] In this case, the simulation image can have a two-layer structure, for example, having a map and infrastructure layer shown in FIG. 12 and an object layer shown in FIG. 13. The map and infrastructure layer (see FIG. 12) is made up of symbols representing road shapes and road infrastructure (lanes, traffic lights). Specifically, -----·+: Lane information (abstract road map) s:Stop line (abstract road map) r: Traffic light (red) (third-class symbol) b: Traffic light (green) (third type symbol) y: Traffic light (yellow) (third-class symbol) w: Waypoint (planned route of the target vehicle) x: No entry allowed The map and infrastructure layer is represented using symbols such as:
[0072] The object layer (see Figure 13) consists of symbols that distinguish between moving objects, walls, and buildings according to their attributes. Specifically, E: Target vehicle (Type 1 symbol) C: Automobile (second type symbol) R: Right turn signal flashing vehicle (second type symbol) L: Vehicle with left turn signal flashing (second type symbol) H: Hazard lights flashing vehicle (type 2 symbol) B: Motorcycle (Type 2 symbol) A: Emergency vehicle (type 2 symbol) W: Pedestrian (second type symbol) X: Other three-dimensional objects The object layer is represented using symbols such as
[0073] By combining the map / infrastructure layer (see FIG. 12) and the object layer (see FIG. 13) described above, a simulation image such as that shown in FIG. 14 can be formed. The simulator 12 (image processing unit 12a, simulation control unit 12b) controls the image information of each layer according to a certain scenario, so as to change the third-class symbols ("b," "y," "r") representing traffic lights in the map / infrastructure layer and move the second-class symbols ("C," "R," "L," "B," "A," "W") representing each moving object in the object layer. As a result, the simulation image shown in FIG. 14 changes according to the scenario. Furthermore, a learning device 13 (controller 13a, learning processing unit 13b) using such a simulator 12 moves a first-type symbol ("E") in the object layer (see Figure 13) in an environment represented by a simulation image (see Figure 14) that changes according to the scenario, and determines a "measure" that enables the first-type symbol to achieve the movement target corresponding to the scenario, i.e., the coefficients between each node of the neural network (controller 13a), according to a predetermined learning algorithm.
[0074] In the simulation image shown in FIG. 14, the area outside the forward field of view of the first type symbol ("E") can be shaded to match the situation detected in an actual target vehicle.
[0075] In the above-described embodiment, the first type symbol and the second type symbol are each a graphic mark or an English letter (ASCII character), but are not limited thereto and may be any other graphic mark, other character (kanji, hiragana, katakana), seal, code, mark, etc., as long as they fall within the concept of "symbol."Furthermore, in the above-described embodiment, the third type symbol is also an English letter (ASCII character) or a graphic mark (see FIG. 9), but is not limited thereto and may be any other graphic mark, other character, seal, code, mark, etc., as long as they fall within the concept of "symbol."
[0076] Next, a vehicle driving control device that utilizes the "strategy" (neural network) obtained by the learning system 10 (simulator 12, learning device 13) as described above will be described.
[0077] A vehicle driving control device according to an embodiment of the present invention is configured as shown in FIG.
[0078] 15, the vehicle driving control device mounted on the vehicle includes an input control unit 51, an image processing unit 52 (surrounding area image generation unit), a controller 53 (vehicle driving mode determination unit), and a driving control unit 54 (actual driving control unit). The input control unit 51, image processing unit 52, controller 53, and driving control unit 54 are configured by a computer system including various hardware and software. The vehicle includes a map information storage unit 61 that stores map information provided from an external device or an in-vehicle navigation device, a GPS unit 62 that acquires position information of the vehicle, a vehicle sensor 63 that detects the vehicle state (vehicle speed, acceleration, etc.), a camera 64 that captures an image of a predetermined surrounding area of the vehicle and generates surrounding image information, a radar 65 that detects obstacles ahead of the vehicle, and a group of sensors 66 that detect other environmental objects relative to the vehicle.
[0079] The input control unit 51 acquires road map information representing a road map of a predetermined surrounding area of the vehicle from the map information storage unit 61 based on the vehicle's position information from the GPS unit 62 (map acquisition unit). The input control unit 51 also acquires information from the camera 64, radar 65, and sensor group 66 (surrounding images, obstacles ahead, other environmental objects) as surrounding environment information representing the environment including moving objects (other vehicles, pedestrians, etc.) and structures in the predetermined surrounding area of the vehicle (environment information acquisition unit). The input control unit 51 also acquires the vehicle's position information from the GPS unit 62 and information representing the vehicle's running state (vehicle speed, acceleration, etc.) from the vehicle sensor 63 as vehicle state information (vehicle state acquisition unit). The input control unit 51 provides the road map information, surrounding environment information, and vehicle state information acquired as described above to the image processing unit 52 at predetermined intervals (for example, every 0.5 seconds).
[0080] The image processing unit 52 generates image information representing an abstract surrounding area image of the same type as the simulation image (see FIGS. 4, 5, 7, 8, 10, and 14) generated by the simulator 12, based on the road map information, the surrounding environment information, and the vehicle state information provided by the input control unit 51. The abstract surrounding area image can be generated as follows.
[0081] An abstract road map (see abstract road map MAP shown in Figures 4, 5, etc.) is generated from the road map information, which represents the specified surrounding area in a two-dimensional abstract manner. Based on the vehicle status information (including location information), a first-class symbol (graphic mark, English letter, etc.) representing the vehicle is arranged on the road in the abstract road map. Furthermore, from the surrounding environment information, for example, using AI image recognition technology, image portions corresponding to moving objects (cars, pedestrians, bicycles, emergency vehicles, etc.) distinguished according to their attributes and image portions corresponding to objects installed in the specified surrounding area whose status changes over time (traffic lights, convex mirrors, etc.) are extracted. Then, the image portions corresponding to the moving objects are replaced with second-class symbols according to the attributes of the moving objects, and the second-class symbols are arranged on the road in the abstract road map based on the relative positional relationship of the moving objects with respect to the vehicle. Furthermore, image portions corresponding to the objects whose status changes over time are replaced with third-class symbols according to the status of the objects, and the third-class symbols are arranged on the abstract road map.
[0082] This generates image information representing an abstract surrounding area image including an abstract road map that abstracts the predetermined surrounding area in a planar manner, first-class symbols that represent the vehicles and move within the roads in the abstract road map, second-class symbols (other vehicles, pedestrians, bicycles, emergency vehicles) that represent moving objects on the roads in the predetermined surrounding area in a differentiated manner according to their attributes and that move within the roads in the abstract road map, and third-class symbols that represent objects (traffic lights, convex mirrors) that are installed in the predetermined surrounding area and change over time in a differentiated manner according to their status and that are fixedly arranged within the abstract road map. The abstract surrounding area image (image information) is updated every time new road map information, surrounding environment information, and vehicle status information are provided from the input control unit 51 (for example, every 0.5 seconds).
[0083] The controller 53 includes a neural network having characteristics (coefficients between each node) equivalent to those of the neural network of the learning device 13 that has undergone reinforcement learning processing in the learning system 10 (simulator 12, learning device 13). The coefficients between each node of this neural network are determined by the reinforcement learning processing. In the controller 53, the neural network receives image information (environmental information) representing the abstracted surrounding area image from the image processing unit 52 and position information representing the position on the road of the abstracted road map of a first-class symbol corresponding to the detected driving position (current position) of the vehicle, and outputs information representing the destination position (target position) of the first-class symbol according to the coefficients between each node ("strategy"). Then, based on the current position and destination position (target position) of the first-class symbol, the controller 53 calculates the target speed and driving direction of the actual vehicle corresponding to this first-class symbol as the driving behavior of the vehicle, and provides information relating to the driving behavior (target speed and driving direction) to the driving control unit 54.
[0084] The driving control unit 54 controls the driving of the vehicle so that the vehicle is in a driving mode (target speed, driving direction) provided by the controller 53, based on map information from the map information storage unit 61, position information of the vehicle from the GPS unit 62, information indicating the state of the vehicle (vehicle speed, acceleration, etc.) from the vehicle sensors 63, surrounding image information from the camera 64, information about obstacles ahead from the radar 64, and information about other environmental objects from the sensor group 66. Specifically, the steering mechanism 71, accelerator mechanism 72, and braking mechanism 73 of the vehicle are controlled so that the vehicle moves in the driving direction at the target speed (to be in the driving mode).
[0085] When the vehicle travels in the aforementioned driving mode, the input control unit 51 acquires new location information, road map information representing a road map of a predetermined surrounding area of the vehicle, surrounding environment information, and vehicle state information, and provides the road map information, surrounding environment information, and vehicle state information to the image processing unit 52. The image processing unit 52 then updates image information representing an abstract surrounding area image including an abstract road map, first-class symbols, second-class symbols, and third-class symbols based on the road map information, surrounding environment information, and vehicle state information. The controller 53 determines the driving mode of the vehicle based on the abstract surrounding area image represented by the updated image information, and the driving control unit 54 executes driving control (steering control, accelerator control, braking control) of the vehicle so that the vehicle travels in the aforementioned driving mode. By repeatedly executing such processing, so-called automatic driving control of the vehicle is performed.
[0086] According to the vehicle driving control device as described above, the vehicle driving can be controlled so as to obtain a driving pattern that is suited to the actual conditions of a specified surrounding area, based on the learning results (strategies) obtained by reinforcement learning using simplified simulation images.
[0087] In addition, for example, in the case of controlling vehicle driving on a highway, if it is not necessary to take into account objects whose states change over time (traffic lights, convex mirrors, etc.), the abstract surrounding area image to be generated does not need to include third-type symbols representing such objects whose states change over time (traffic lights, convex mirrors, etc.).
[0088] In the above-described embodiment, the "strategy" in the controller 13a of the learning device 13 and the controller 53 in the vehicle driving control device is the coefficient between each node of the neural network, but it may also be information adopted in other models, for example, a rule-based model or information adopted in a program-controlled model.
[0089] Although the embodiments of the present invention have been described above, the modifications of these embodiments and their respective parts are presented as examples and are not intended to limit the scope of the invention. These novel embodiments described above can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. [Industrial Applicability]
[0090] The simulator according to the present invention has the effect of being able to incorporate vehicle driving behavior into a simple and efficient reinforcement learning framework, and is useful as a simulator used for reinforcement learning of vehicle driving behavior. [Explanation of symbols]
[0091] 10 Learning System 11 Control section 12 Simulator 12a Image processing unit 12b Simulation control section 12c monitor 13 Learning Device 13a Controller 13b Learning processing unit 14 Libraries 51 Input control section 52 Image processing section 53 Controller 54 Operation control unit 61 Map information storage unit 62 GPS units 63 Vehicle Sensors 64 Camera 65 Radar 66 sensors 71 Steering mechanism 72 Accelerator mechanism 73 Braking mechanism 100 Grid
Claims
1. A simulator used for reinforcement learning about a driving mode of a target vehicle, a simulation image generating unit that generates information representing a simulation image including information on an abstract road map that abstractly represents a simulation area including roads in a planar manner, information on first-class characters that represent the target vehicle and are arranged on the roads in the abstract road map, and information on second-class characters that represent moving objects that move on the roads in the simulation area in a differentiated manner according to their attributes and are arranged on the roads in the abstract road map; a simulation control unit that controls information representing the simulation image so that the second type characters move within the roads of the abstract road map according to a scenario, while allowing the first type characters to move relative to the abstract road map.
2. The information representing the simulation image includes information on third-type characters that are fixedly arranged within the abstract road map and represent objects whose states change over time in the simulation area so as to distinguish them according to their states, 2. The simulator according to claim 1, wherein said simulation control unit controls the information representing said simulation image so that said third-type characters change according to a scenario.
3. The information representing the simulation image is information represented in a multi-layer structure, 3. The simulator according to claim 1, wherein the multi-layer structure includes a layer including the first type of characters.
4. The information representing the simulation image is information represented in a multi-layer structure, 3. The simulator according to claim 1, wherein the multi-layer structure includes a layer including the second type of characters.
5. The information representing the simulation image is information represented in a multi-layer structure, 3. The simulator according to claim 2, wherein the multi-layer structure includes a layer including the third type of characters that is different from the layer including the first type of characters and the layer including the second type of characters.
6. A learning device that performs reinforcement learning on a driving mode of a target vehicle using the simulator according to claim 1, a controller that determines a movement mode of the first type character based on a strategy so that a predetermined movement target of the first type character is achieved in the simulation image of the second type character moving within the road of the abstract road map in accordance with the scenario, and instructs the simulation control unit of the simulator to move the first type character relative to the abstract road map in the determined movement mode; a learning processing unit that evaluates, in accordance with a predetermined reinforcement learning algorithm, a movement pattern of the first type character in an environment in which the second type character moves within a road of the abstract road map in accordance with the scenario in the simulation image represented by the information controlled by the simulation control unit, and updates the strategy used by the controller to determine the movement pattern of the first type character based on the evaluation result; a learning device that repeats the following steps until the movement target is achieved: instructing the controller to move the first type character in a movement manner determined based on the strategy; evaluating the movement manner of the first type character by the learning processing unit; and updating the strategy based on the evaluation results.
7. A learning device that performs reinforcement learning on a driving mode of a target vehicle using the simulator according to claim 2, a controller that determines a movement mode of the first type character based on a strategy so that a predetermined movement target of the first type character is achieved in the simulation image in which the second type character moves within the road of the abstract road map in accordance with the scenario and the third type character changes within the abstract road map, and instructs the simulation control unit of the simulator to move the first type character in the determined movement mode; a learning processing unit that evaluates, in accordance with a predetermined reinforcement learning algorithm, the movement modes of the first type characters in an environment in which the second type characters move within the roads of the abstract road map and the third type characters change within the abstract road map in accordance with the scenario in the simulation image represented by the information controlled by the simulation control unit, and updates the strategy used by the controller to determine the movement modes of the first type characters based on the evaluation result; a learning device that repeats the following steps until the movement target is achieved: instructing the controller to move the first type character in a movement manner determined based on the strategy; evaluating the movement manner of the first type character by the learning processing unit; and updating the strategy based on the evaluation results.
8. 8. The learning device according to claim 6, wherein the simulation control unit in the simulator controls the information representing the simulation image so as to move the first type characters relative to the fixed abstract road map in accordance with instructions from the controller.
9. 8. The learning device according to claim 6, wherein the simulation control unit in the simulator controls the information representing the simulation image so as to move the abstract road map relative to the fixed first-class characters in accordance with instructions from the controller.
10. the simulation image represented by the information generated by the simulator is formed on a grid consisting of a plurality of squares; the movement manner of the first type character determined by the controller is expressed as a movement in units of squares of the grid, 8. The learning device according to claim 6, wherein the simulation control unit of the simulator controls the information representing the simulation image so that the first type characters move in units of grid squares relative to the abstract road map in accordance with instructions from the controller.
11. the simulation image represented by the information generated by the simulator is formed on a grid consisting of a plurality of squares; the movement manner of the first type character determined by the controller is expressed as a movement in units of squares of the grid, 9. The learning device according to claim 8, wherein the simulation control unit of the simulator controls the information representing the simulation image so as to move the first type characters fixed on the abstract road map in units of grid squares in accordance with instructions from the controller.
12. the simulation image represented by the information generated by the simulator is formed on a grid consisting of a plurality of squares; the movement manner of the first type character determined by the controller is expressed as a movement in units of squares of the grid, The learning device according to claim 9, wherein the simulation control unit of the simulator controls the information representing the simulation image so that the abstract road map moves in units of squares of the grid relative to the first type characters fixed in accordance with instructions from the controller.
13. 11. The learning device according to claim 10, wherein the grid is configured such that the actual distance represented by each square in a predetermined area including the first type character is greater than the actual distance represented by each square in an area farther from the first type character than the predetermined area.
14. A vehicle driving control device that is mounted on a vehicle and controls driving of the vehicle, a map acquisition unit that acquires road map information representing a road map of a predetermined surrounding area of the vehicle; an environmental information acquisition unit that acquires surrounding environmental information representing a state of an environment including the moving object in the predetermined surrounding area; a vehicle state acquisition unit that acquires vehicle state information that indicates the traveling position of the vehicle and the traveling state of the vehicle; a surrounding area image generation unit that generates, based on the road map information, the surrounding environment information, and the vehicle state information, information representing an abstract surrounding area image, including information on an abstract road map that abstractly represents the predetermined surrounding area in a planar manner, information on first-class characters that represent the vehicle and move within the roads in the abstract road map, and information on second-class characters that represent moving objects that move on the roads in the predetermined surrounding area in a differentiated manner according to their attributes; and a vehicle driving mode determination unit that determines a driving mode of the vehicle so as to move the first-class characters based on the strategy obtained by the learning device according to claim 6 in the abstract surrounding area image represented by the information generated by the surrounding area image generation unit; an actual driving control unit that executes driving control of the vehicle so that the driving mode determined by the vehicle driving mode determination unit is achieved.
15. A vehicle driving control device that is mounted on a vehicle and controls driving of the vehicle, a map acquisition unit that acquires road map information representing a road map of a predetermined surrounding area of the vehicle; an environmental information acquisition unit that acquires surrounding environmental information representing a state of an environment including moving objects in the predetermined surrounding area and objects installed in the predetermined surrounding area whose states change over time; a vehicle state acquisition unit that acquires vehicle state information that indicates the traveling position of the vehicle and the traveling state of the vehicle; a surrounding area image generating unit that generates, based on the road map information, the surrounding environment information, and the vehicle state information, information representing an abstract surrounding area image, including information on an abstract road map that abstractly represents the predetermined surrounding area in a planar manner, information on first-class characters that represent the vehicle and move within the roads in the abstract road map, information on second-class characters that represent moving objects that move on the roads in the predetermined surrounding area in a differentiated manner according to their attributes, and information on third-class characters that represent objects that are installed in the predetermined surrounding area and change over time in a differentiated manner according to their states and are fixedly arranged in the abstract road map; a vehicle driving mode determination unit that determines a driving mode of the vehicle so as to move the first-class characters in the abstract surrounding area image represented by the information generated by the surrounding area image generation unit, based on the strategy obtained by the learning device according to claim 7; an actual driving control unit that executes driving control of the vehicle so that the driving mode determined by the vehicle driving mode determination unit is achieved.
Citation Information
Patent Citations
Traffic flow simulator, environment analysis system, traffic flow simulating method and storage medium
JP2000132783A
Learning device, learning method, and program
JP2020035222A
Automatically generating training data for a lidar using simulated vehicles in virtual space
US20200074230A1
Systems and methods for navigating with safe distances
WO2020035728A2
Radar spatial estimation
WO2020069025A1