Vehicle driving control device, vehicle control management device, and vehicle control management method

By adopting a combination of reinforcement learning and rule algorithms in the vehicle driving control equipment, determining the driving mode of the vehicle and recording and transmitting differences for improvement, the problem of insufficient reliability of vehicle driving control in the prior art is solved, and higher driving control reliability and safety are achieved.

JP7672314B2Active Publication Date: 2025-05-07NTT DATA AUTOMOBILIGENCE RES CENT LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021148297
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-13
Publication Date
2025-05-07
Estimated Expiration
2041-09-13

AI Technical Summary

Technical Problem

There are shortcomings in existing vehicle driving control equipment in ensuring vehicle safety, especially in normal driving control, and the prior art is difficult to improve the reliability of driving control.

Method used

Two different algorithms are used to determine the driving mode of a vehicle: one based on reinforcement learning and the other based on rules. By comparing the results of these two algorithms, the most appropriate driving mode is determined and the differences are recorded and transmitted to external devices when necessary for improvement.

Benefits of technology

The reliability of vehicle driving control is improved, and the algorithm is continuously improved through multiple algorithm comparisons and differential recordings, which improves the stability and safety of driving control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672314000001
    Figure 0007672314000001
  • Figure 0007672314000002
    Figure 0007672314000002
  • Figure 0007672314000003
    Figure 0007672314000003
Patent Text Reader

Abstract

To provide a vehicle operation control device for enabling highly reliable operation control.SOLUTION: A vehicle operation control device includes: an environment information acquisition section 12 for acquiring peripheral environment information; a vehicle state acquisition section 11 for acquiring vehicle state information; a first travel mode determination section 13 for determining a travel mode which has to be selected by a vehicle in a travel state to be expressed by the vehicle state information in response to a first algorithm under an environment to be expressed by the peripheral environment information; a second travel state determination section 14 for determining a travel mode which has to be selected by the vehicle in the travel state to be expressed by the vehicle state information in response to a second algorithm under the environment to be expressed by the peripheral environment information; a determination section 15 for determining whether or not the first travel state is the same as the second travel state; a travel state determination section 15 for determining the travel state which has to be selected by the vehicle based on the determination result; and an operation control section 16 for performing the operation control of the vehicle so as to allow the vehicle to travel in the determined travel mode.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a vehicle driving control device that controls driving of a vehicle, a vehicle control management device that manages the driving control of a vehicle, and a vehicle control management method. [Background technology]

[0002] Conventionally, various vehicle driving control devices that perform driving control of a vehicle (including accelerator control, braking control, and steering control) have been proposed. For example, a device that performs driving control of a vehicle so that the driving mode is determined according to an algorithm (policy) based on reinforcement learning about the driving mode of the vehicle (see, for example, Patent Document 1), or a device that performs driving control of a vehicle so that the driving mode is determined according to an algorithm corresponding to a rule based on knowledge about the driving mode of the vehicle (traffic laws, etc.) (see, for example, Patent Document 2), etc. are known. Specifically, in such a vehicle driving control device, accelerator control, braking control, and steering control are performed so that the driving mode (represented by the driving direction from the current position, the destination position, the driving speed, the acceleration, etc.) is determined according to a certain algorithm from the environment including the road conditions (road configuration, the conditions of other vehicles, the conditions of pedestrians, the conditions of traffic lights, the conditions of installed traffic signs, etc.) in a predetermined surrounding area of ​​the vehicle and the driving state of the vehicle (driving position, driving speed, acceleration, etc.). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2020-35222 A [Patent Document 2] JP 2019-74885 A Summary of the Invention [Problem to be solved by the invention]

[0004] Such a vehicle driving control device must always ensure the safety of the vehicle. Usually, a fail-safe method is used to ensure safety when the vehicle is abnormal (e.g., when a breakdown occurs). However, this fail-safe method is only intended for when the vehicle is abnormal, and does not improve the reliability of normal driving control.

[0005] The present invention has been made in consideration of the above circumstances, and provides a vehicle driving control device that enables more reliable driving control.

[0006] The present invention also provides a vehicle control management device that can contribute to improving vehicle driving control. [Means for solving the problem]

[0007] A vehicle driving control device according to the present invention is a vehicle driving control device that controls driving of a vehicle, and includes an environmental information acquisition unit that acquires surrounding environment information representing an environment including road conditions in a predetermined surrounding area of ​​the vehicle, a vehicle state acquisition unit that acquires vehicle state information representing a driving state of the vehicle, a first driving mode determination unit that determines a driving mode that the vehicle should take in the driving state represented by the vehicle state information acquired by the vehicle state acquisition unit according to a first algorithm under the environment represented by the surrounding environment information acquired by the environmental information acquisition unit, a second driving mode determination unit that determines a driving mode that the vehicle should take in the driving state represented by the acquired vehicle state information under the environment represented by the acquired surrounding environment information according to a second algorithm different from the first algorithm, a determination unit that determines whether the first driving mode obtained by the first driving mode determination unit and the second driving mode obtained by the second driving mode determination unit are the same, a driving mode determination unit that determines the driving mode that the vehicle should take based on the determination result by the determination unit, and a driving control unit that controls the driving of the vehicle so that the vehicle runs in the determined driving mode. a transmission unit that, when the determination unit determines that the first driving mode and the second driving mode are not the same, transmits information about the first driving mode and the second driving mode that are not the same to an external device as log information, together with the surrounding environment information and the vehicle state information that are the basis for determining the first driving mode and the second driving mode, respectively; The configuration has the following.

[0008] According to this configuration, when the surrounding environment information and the vehicle state information are acquired, the driving mode that the vehicle should take in the driving state represented by the vehicle state information is determined by a first algorithm in an environment including the road conditions in a predetermined surrounding area represented by the surrounding environment information, and the driving mode that the vehicle should take in the driving state represented by the vehicle state information is determined according to a second algorithm different from the first algorithm in an environment including the road conditions in a predetermined surrounding area represented by the surrounding environment information. Then, it is determined whether the first driving mode obtained by the determination according to the first algorithm and the second driving mode obtained by the determination according to the second algorithm are the same, and the driving mode that the vehicle should take is determined based on the determination result. The driving control of the vehicle is performed so that the vehicle runs in the determined driving mode. If it is determined that the first driving mode obtained by the judgment according to the first algorithm is not the same as the second driving mode obtained by the judgment according to a second algorithm different from the first algorithm, information on the different first and second driving modes is transmitted as log information to an external device together with surrounding environment information (representing the environment including road information in a predetermined surrounding area) and vehicle state information (representing the vehicle's driving state) on which the judgment of the first and second driving modes is based. The external device can receive the log information and can further use the log information for improving the first algorithm (for example, the specifications of reinforcement learning).

[0009] In the vehicle driving control device according to the present invention, the first algorithm may be an algorithm based on reinforcement learning regarding the driving behavior of the vehicle.

[0010] According to such a configuration, it is determined whether a first driving mode obtained by a determination according to an algorithm (e.g., a policy) based on reinforcement learning is the same as a second driving mode obtained by a determination according to a second algorithm. Then, based on the result of the determination, a driving mode to be adopted by the vehicle is determined, and the vehicle is controlled so that the vehicle drives in the determined driving mode.

[0011] In the vehicle driving control device according to the present invention, the second algorithm may be an algorithm corresponding to a rule based on knowledge about the driving behavior of the vehicle.

[0012] According to this configuration, it is determined whether a first driving mode obtained by a determination according to a first algorithm is the same as a second driving mode obtained by a determination according to an algorithm corresponding to a rule based on knowledge about the driving mode of the vehicle. Then, based on the result of the determination, a driving mode to be adopted by the vehicle is determined, and the vehicle is controlled so that the vehicle drives in the determined driving mode.

[0013] In the vehicle driving control device of the present invention, the driving mode determination unit may be configured to determine the second driving mode as the driving mode that the vehicle should adopt when the judgment unit determines that the first driving mode and the second driving mode are not the same.

[0014] According to this configuration, when it is determined that a first driving behavior obtained by a judgment according to a first algorithm and a second driving behavior obtained by a judgment according to a second algorithm different from the first algorithm are not the same, the judgment method according to the second algorithm is given priority, and the second driving behavior is determined as the driving behavior that the vehicle should adopt.

[0015] In the vehicle driving control device of the present invention, the driving mode determination unit may be configured to determine the first driving mode as the driving mode that the vehicle should adopt when the judgment unit determines that the first driving mode and the second driving mode are the same.

[0016] According to this configuration, when it is determined that a first driving mode obtained by a determination according to a first algorithm and a second driving mode obtained by a determination according to a second algorithm different from the first algorithm are the same, the first driving mode is determined as the driving mode that the vehicle should take. This makes it possible to ensure the reliability of the first driving mode by the method of determining the vehicle's driving mode according to the second algorithm.

[0019] The vehicle control management device of the present invention can be configured to have a receiving unit that receives the log information transmitted from the transmitting unit in the vehicle driving control device, and a log management unit that stores and manages the log information received by the receiving unit, and to provide the log information stored and managed by the log management unit for use in improving at least the first algorithm of the first algorithm and the second algorithm.

[0020] According to this configuration, log information transmitted from a vehicle driving control device mounted on the vehicle (information about the first driving mode and the second driving mode, which are different from each other, along with the surrounding environment information (environment of a specified surrounding area) and the vehicle state information (driving state of the vehicle) that are the basis for judging the first driving mode and the second driving mode, respectively) is stored and managed, and the stored and managed log information is used for improving at least the first algorithm of the first algorithm and the second algorithm. As a result, it becomes possible to sequentially improve at least the first algorithm.

[0021] In the vehicle control management device of the present invention, the first algorithm is an algorithm based on reinforcement learning about the vehicle's driving behavior, and the second algorithm is an algorithm corresponding to a rule based on knowledge about the vehicle's driving behavior, and the log information stored and managed in the log management unit can be configured to be used for improving at least the specifications of the reinforcement learning.

[0022] With this configuration, the stored and managed log information (the surrounding environment information (environment of a specified surrounding area) and the vehicle state information (vehicle driving state) that are the basis for determining each of the first driving state and the second driving state, and information about the first driving state and the second driving state that are not the same are provided for the work of improving at least the specifications of the reinforcement learning. As a result, it becomes possible to sequentially improve the specifications of the reinforcement learning for the vehicle's driving state.

[0023] The vehicle control management method of the present invention includes a receiving step of receiving the log information transmitted from the transmitting unit in the vehicle driving control device, a log management step of storing and managing the log information received in the receiving step, and an improvement step of performing improvement work on at least the first algorithm of the first algorithm and the second algorithm using the log information stored and managed in the log management step.

[0024] In the vehicle control management method of the present invention, the first algorithm is an algorithm based on reinforcement learning about the vehicle's driving behavior, the second algorithm is an algorithm corresponding to a rule based on knowledge about the vehicle's driving behavior, and the improvement step can be configured to perform improvement work on the specifications of the reinforcement learning based on the log information. Effect of the Invention

[0025] According to the vehicle control device of the present invention, it is determined whether a first driving mode and a second driving mode obtained by judgment according to two different algorithms for a common environment (surrounding environment information) and a common vehicle driving state are the same, and the driving mode that the vehicle should adopt is determined based on the judgment result, thereby enabling more reliable driving control.

[0026] In addition, according to the vehicle control management device of the present invention, log information transmitted from the vehicle driving control device mounted on the vehicle (information about the first driving mode and the second driving mode, which are different from each other, together with the surrounding environment information (environment of a specified surrounding area) and the vehicle state information (driving state of the vehicle) that form the basis for determining each of the first driving mode and the second driving mode) is used to improve at least the first algorithm of the first algorithm and the second algorithm, thereby contributing to improvement of vehicle driving control. [Brief description of the drawings]

[0027] [Figure 1]FIG. 1 is a diagram conceptually showing a system to which a vehicle driving control device according to an embodiment of the present invention and a vehicle control management device according to an embodiment of the present invention are applied. [Diagram 2] FIG. 2 is a block diagram showing a configuration of a vehicle driving control device according to an embodiment of the present invention. [Diagram 3] FIG. 3 is a block diagram showing a configuration of a vehicle control management device according to an embodiment of the present invention. [Figure 4] FIG. 4 is a flowchart showing a processing procedure in the control unit of the vehicle driving control device according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0028] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0029] A vehicle driving control device according to an embodiment of the present invention and a system including a vehicle control management device according to an embodiment of the present invention are configured as shown in FIG.

[0030] 1, each vehicle 100 is equipped with a vehicle driving control device 10 that performs driving control (including accelerator control, braking control, and steering control) of the vehicle 100. The vehicle driving control device 10 can transmit log information (described later) to a management server 50 (vehicle control management device) on the Internet NW via a mobile communication network. The management server 50 receives the log information transmitted from the vehicle driving control device 10 installed in each vehicle 100, and stores and manages the log information.

[0031] The vehicle driving control device 10 is configured as shown in FIG.

[0032] 2, a vehicle driving control device 10 mounted on a vehicle 100 includes a driving state detection module 11, an environment detection module 12, a reinforcement learning policy decision module 13, a rule-based decision module 14, a control unit 15, a driving control unit 16, and a transmission unit 17 (transmission unit). The driving state detection module 11, the environment detection module 12, the reinforcement learning policy decision module 13, the rule-based decision module 14, the control unit 15, and the driving control unit 16 are generally configured by a computer system including various hardware and software. The transmission unit 17 can transmit information to devices on the Internet NW via a mobile communication network under the control of the control unit 15.

[0033] The vehicle 100 is equipped with a map information storage unit 21 that stores map information provided from an external device or an in-vehicle navigation device, a GPS unit 22 that outputs a detection signal according to the position of the vehicle 100, a vehicle sensor 23 that outputs a detection signal according to the vehicle state (vehicle speed, acceleration, etc.), a camera 24 that captures an image of a specified area around the vehicle 100 to generate surrounding image information, a radar 25 that outputs a signal according to an obstacle in front of the vehicle, and a group of sensors 26 that detect other environmental objects relative to the vehicle 100.

[0034] In the vehicle driving control device 10, the driving state detection module 11 (vehicle state acquisition unit) generates position information representing the driving position (current position) of the vehicle 100 based on a detection signal from the GPS unit 22, generates vehicle speed information representing the speed and acceleration of the vehicle 100 based on a detection signal from the vehicle sensor 23, and acquires the position information and vehicle speed information as vehicle state information representing the driving state of the vehicle 100. The position information (current position) may be information representing latitude and longitude, information representing a position on a road map obtained from the map information storage unit 21, or information representing a position in another coordinate system. The driving state detection module 11 provides the acquired vehicle state information to each of the reinforcement learning policy judgment module 13 and the rule-based judgment module 14.

[0035] The environment detection module 12 (environment information acquisition unit) generates surrounding environment information representing the environment including road conditions (road configuration, other vehicle conditions, pedestrian conditions, traffic light conditions, installed traffic signs, etc.) in a predetermined surrounding area of ​​the vehicle 100 based on road map information from the map storage unit 21, detection signals from the GPS unit 22, surrounding image information from the camera 24, information on forward obstacles from the radar 25, and detection signals from the sensor group 26. The environment detection module 12 provides the acquired surrounding environment information to each of the reinforcement learning policy decision module 13 and the rule-based decision module.

[0036] The reinforcement learning policy determination module 13 (first driving behavior determination unit) determines the driving behavior that the vehicle 100 should take in the driving state (including the current position, speed, and acceleration) represented by the vehicle state information provided by the driving state detection module 11 under the environment (including road conditions such as road configuration, other vehicle conditions, pedestrian conditions, traffic light conditions, and installed traffic signs) represented by the surrounding environment information provided by the environment detection module 12, according to a policy (for example, a neural network function: a first algorithm) already obtained by reinforcement learning. As a result of the determination, for example, the position of the vehicle 100 after a predetermined time (for example, 0.5 seconds), the speed and acceleration when moving from the current position (detected driving position) to the position after the predetermined time, and the like are obtained as the first driving behavior.

[0037] In addition, the rule-based judgment module 14 (second driving mode judgment unit) judges the driving mode that the vehicle 100 should take in the driving state (including the current position, speed, and acceleration) represented by the vehicle state information provided from the driving state detection module 11 in the environment (including road conditions such as road configuration, other vehicle conditions, pedestrian conditions, traffic light conditions, and installed traffic signs) represented by the surrounding environment information provided from the environment detection module 12, according to an algorithm (rule-based: second algorithm) corresponding to rules based on knowledge (including, for example, traffic rules defined in laws and regulations such as the Road Traffic Act). As a result of the judgment, similar to the case of the reinforcement learning policy judgment module 13, for example, the position of the vehicle 100 after a predetermined time (for example, 0.5 seconds), and the speed and acceleration when moving from the current position (detected driving position) to the position after the predetermined time are obtained as the second driving mode.

[0038] The control unit 15 judges whether the first driving mode obtained by the reinforcement learning policy judgment module 13 is the same as the second driving mode obtained by the rule-based judgment module 14 based on a predetermined criterion (judgment unit), and determines the driving mode to be taken by the vehicle 100 based on the judgment result (driving state determination unit). The control unit 15 provides the determined driving mode to the driving control unit 16. The driving control unit 16 controls the driving of the vehicle 100, specifically, controls the steering mechanism 31, the accelerator mechanism 32, and the braking mechanism 33 mounted on the vehicle 100 so that the vehicle 100 drives in the provided (determined) driving mode. Furthermore, the control unit 15 controls the transmission unit 17 to transmit log information, which will be described later, to the management server 50. The transmission unit 17 transmits the log information to the management server 50 via the mobile communication network and the Internet NW under the control of the control unit 15.

[0039] The management server 50 is configured as shown in FIG.

[0040] 3, the management server 50 has a processing unit 51, a receiving unit 52 (receiving unit), and a storage unit 53. The receiving unit 52 receives the log information transmitted from the transmitting unit 17 of the vehicle driving control device 10 mounted on each vehicle 100 via a mobile communication network and the Internet NW. The processing unit 51 stores the log information received by the receiving unit 52 in the storage unit 53. The log information stored in the storage unit 53 is stored and managed under the control of the processing unit 51 (log management unit). The processing unit 51 can provide the log information stored and managed in the storage unit 53 to an analysis device 60.

[0041] The control unit 15 of the vehicle driving control device 10 described above executes processing, for example, according to the procedure shown in FIG.

[0042] 4, when the control unit 15 acquires a first driving mode determined by the reinforcement learning policy judgment module 13 as the driving mode that the vehicle 100 should take and a second driving mode determined by the rule-based judgment module 14 as the driving mode that the vehicle 100 should take (S11), it judges whether they are the same or not based on a predetermined criterion (S12). For example, when the driving mode is expressed as the position of the vehicle 100 after a predetermined time (e.g., 0.5 seconds), and the speed and acceleration when moving from the current position to that position, it is possible to judge not only the completely same position but also a predetermined range of the position as the same position, and not only the completely same speed and the completely same acceleration but also a predetermined range of the speed and a predetermined range of the acceleration as the same speed and the same acceleration, respectively.

[0043] When it is determined that the first driving mode and the second driving mode are the same (YES in S12), the control unit 15 determines the first driving mode obtained by the determination in the reinforcement learning policy determination module 13 as the driving mode to be taken by the vehicle 100, and provides the first driving mode to the driving control unit 16 (S13). The control unit 15 repeatedly executes the above-mentioned processing (S11 to S13) at a predetermined period (corresponding to the respective determination periods of the reinforcement learning policy determination module 13 and the rule-based determination module 14: for example, 0.5 seconds) until a predetermined end condition (for example, an operation to stop driving, activation of a fail-safe operation, etc.) is satisfied (S16). As a result, in the vehicle 100, the driving control unit 16 controls the steering mechanism 31, the accelerator mechanism 32, and the brake mechanism 33 so that the vehicle 100 becomes the driving mode (first driving mode) obtained by the determination in the reinforcement learning policy determination module 13 under the environment represented by the surrounding environment information acquired by the environment detection module 12 (driving control).

[0044] In the course of such processing, if it is determined that the first driving mode obtained by the determination in the reinforcement learning determination module 13 and the second driving mode obtained by the determination in the rule-based determination module 14 are not the same (NO in S12), the control unit 15 determines (mediates) the second driving mode obtained by the determination in the rule-based determination module 14 as the driving mode that the vehicle 100 should take, and provides the second driving mode to the driving control unit 16 (S14). Then, the control unit 15 further causes the transmission unit 17 (transmitter) to transmit information on the first driving mode and the second driving mode that are not the same as log information together with surrounding environment information (including road configuration, other vehicle status, pedestrian status, traffic light status, installed traffic signs, etc.) and vehicle state information (including current position, speed, and acceleration) that are the basis for determining each of the first driving mode and the second driving mode that are determined to be not the same (S15). The transmission unit 17 transmits the log information to the management server 50 via a mobile communication network and the Internet NW under the control of the control unit 15. The information about the first driving mode and the second driving mode, which are not the same and are included in the log information, may be information that represents the first driving mode and the second driving mode themselves, or may be other information obtained by processing the information that represents the first driving mode and the second driving mode.

[0045] While it is determined that the first driving mode obtained by the determination in the reinforcement learning policy determination module 13 and the second driving mode obtained by the determination in the rule-based determination module 14 are the same, the control unit 15 provides the first driving mode obtained by the determination in the reinforcement learning policy determination module 13 to the driving control unit 16 as the driving mode to be taken by the vehicle 100. Then, every time it is determined that the first driving mode and the second driving mode are not the same, the control unit 15 provides the second driving mode obtained by the determination in the rule-based determination module 14 to the driving control unit 16 as the driving mode to be taken by the vehicle 100. As a result, usually, the driving control of the vehicle 100 is performed so that the first driving mode obtained by the determination in the reinforcement learning policy determination module 13 is obtained. Then, every time it is determined that the first driving mode and the second driving mode are not the same, the driving control of the vehicle 100 is performed so that the second driving mode obtained by the determination in the rule-based determination module 14 is obtained.

[0046] When a predetermined end condition is satisfied during the above process (YES in S16), the control unit 15 ends the process related to the driving control of the vehicle 100.

[0047] According to the vehicle driving control device 10 described above, in normal times when it is determined that the first driving mode and the second driving mode are the same, driving control of the vehicle 100 is performed so that the vehicle 100 becomes the first driving mode obtained by the judgment of the reinforcement learning policy judgment module 13 under the environment represented by the surrounding environment information acquired by the environment detection module 12. As a result, the reliability of the first driving mode obtained by the judgment of the reinforcement learning policy judgment module 13 is guaranteed by the judgment of the rule-based judgment module 14 that judges the driving mode based on rules based on certain knowledge (traffic rules, etc.). Therefore, more reliable driving control of the vehicle 100 is possible.

[0048] On the other hand, when it is determined that the first driving mode and the second driving mode are not the same, that is, when it is determined that the driving mode (first driving mode) obtained by a judgment according to an algorithm (policy) based on a certain experience by a reinforcement learning technique and the driving mode (second driving mode) obtained by a judgment (rule-based judgment) according to an algorithm corresponding to a knowledge-based rule are not the same in a common environment (surrounding environment information) and a common driving state (vehicle state information), the driving control of the vehicle 100 is performed so that the second driving mode obtained by the judgment in the rule-based judgment module 14 according to the algorithm corresponding to the knowledge-based rule is obtained. By judging the driving mode according to an algorithm corresponding to more and more detailed rules, more reliable driving control of the vehicle 100 can be realized.

[0049] In the vehicle driving control device 10, as described above, when it is determined that a driving behavior (first driving behavior) obtained by a judgment according to a certain experience-based algorithm (policy) using a reinforcement learning technique is not the same as a driving behavior (second driving behavior) obtained by a judgment according to an algorithm corresponding to a knowledge-based rule (rule-based judgment) in a common environment (ambient environment information) and common driving state (vehicle state information), as shown in FIG. 4 (NO in S12), the transmitting unit 17 transmits information about the different first driving behavior and second driving behavior as log information to the management server 50 on the Internet NW, together with the ambient environment information representing the common environment and the vehicle state information representing the common driving state.

[0050] In the management server 50, when the receiving unit 52 receives the log information transmitted from the vehicle driving control device 10 (transmitting unit 17) of the vehicle 100 (receiving step), the processing unit 51 stores the received log information, i.e., information about the fact that different first driving information and second driving states have been obtained by judgment according to different algorithms (algorithms based on reinforcement learning, algorithms corresponding to knowledge-based rules) for a common environment (surrounding environment information) and a common vehicle driving state (vehicle state information), in the memory unit 53. Then, such log information is stored and managed in the memory unit 53 under the control of the processing unit 51 (log management step).

[0051] Various log information stored and managed in the storage unit 53 can be used to improve the algorithm based on reinforcement learning used in the reinforcement learning policy decision module 13, specifically, to improve the specifications of reinforcement learning for obtaining a policy (algorithm). For example, when the log information is provided from the management server 50 (processing unit 51) to the analysis device 60, the analysis device 60 analyzes information on the different first and second driving modes, together with surrounding environment information and vehicle state information, according to a predetermined procedure. Then, the analysis result can be used to improve the specifications of reinforcement learning, for example, a scenario for changing the environment in a virtual space realized by a simulator used for reinforcement learning (improvement step).

[0052] The log information can also be used to improve an algorithm corresponding to a knowledge-based rule used in the rule-based decision module 14. In this case, the algorithm can be improved to correspond to more detailed rules according to the situation based on the log information (information about the surrounding environment information, vehicle state information, and information about the different first and second driving modes).

[0053] According to the above-mentioned management server 50 (vehicle control management device), the log information transmitted from the vehicle driving control device 10 mounted on the vehicle 100 (information about the first driving mode and the second driving mode, which are not the same, along with the surrounding environment information (environment of a specified surrounding area) and the vehicle state information (vehicle driving state) that form the basis for determining each of the first driving mode and the second driving mode) can be used to improve at least one of the algorithm (policy) based on reinforcement learning used in the reinforcement learning policy judgment module 13 and the algorithm corresponding to the knowledge-based rules used in the rule-based judgment module 14, thereby contributing to improvement of the driving control of the vehicle 100 by the vehicle driving control device 10.

[0054] In the above-described embodiment, the first algorithm is an algorithm (policy) based on reinforcement learning about the vehicle's driving behavior, and the second algorithm is an algorithm corresponding to a rule based on knowledge about the vehicle's driving behavior, but is not limited thereto. The first algorithm and the second algorithm for determining the driving behavior to be taken by the vehicle 100 in the driving state represented by the vehicle state information under the environment represented by the surrounding environment information are not particularly limited as long as they are different.

[0055] Also, when it is determined that the first driving mode and the second driving mode are the same, the first driving mode is determined as the driving mode that the vehicle 100 should take, and when it is determined that the first driving mode and the second driving mode are not the same, the second driving mode is determined as the driving mode that the vehicle 100 should take, but this is not limited to this. When it is determined that the first driving mode and the second driving mode are the same, a new driving mode that takes into account the first driving mode and the second driving mode at a predetermined ratio, respectively, can be determined as the driving mode that the vehicle 100 should take. Also, when it is determined that the first driving mode and the second driving mode are not the same, a new driving mode that takes into account the first driving mode and the second driving mode at a ratio different from the ratio can be determined as the driving mode that the vehicle 100 should take.

[0056] Although the embodiment of the present invention has been described above, the embodiment and the modified examples of each part are presented as examples and are not intended to limit the scope of the invention. These novel embodiments described above can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the gist of the invention. [Industrial Applicability]

[0057] The vehicle driving control device according to the present invention has the effect of enabling more reliable driving control, and is useful as a vehicle driving control device that controls driving of a vehicle. Also, the vehicle control management device and vehicle control management method according to the present invention have the effect of contributing to improvement of vehicle driving control, and are useful as a vehicle control management device and vehicle control management method that manage driving control in a vehicle. [Explanation of symbols]

[0058] 10 Vehicle driving control device 11. Driving condition detection module 12 Environmental Detection Module 13 Reinforcement learning policy decision module 14 Rule-based decision module 15 Control section 16 Operation control unit 17 Transmitting unit 21 Map information storage section 22 GPS unit 23 Vehicle Sensors 24 Camera 25 Radar 26 Sensor Group 31 Steering mechanism 32 Accelerator mechanism 33 Braking mechanism 50 Management Server 51 Processing section 52 Receiving unit 53 Storage section 60 Analyzer 100 vehicles

Claims

1. A vehicle driving control device that controls driving of a vehicle, an environmental information acquisition unit that acquires surrounding environmental information representing an environment including a road condition in a predetermined surrounding area of ​​the vehicle; a vehicle state acquisition unit that acquires vehicle state information representing a traveling state of the vehicle; a first driving mode determination unit that determines a driving mode that the vehicle should take in the driving state represented by the vehicle state information acquired by the vehicle state acquisition unit under the environment represented by the surrounding environment information acquired by the environmental information acquisition unit in accordance with a first algorithm; a second driving mode determination unit that determines a driving mode that the vehicle should take in the driving state represented by the acquired vehicle state information under the environment represented by the acquired surrounding environment information in accordance with a second algorithm different from the first algorithm; a determination unit that determines whether or not a first driving mode obtained by the first driving mode determination unit and a second driving mode obtained by the second driving mode determination unit are the same; a driving mode determination unit that determines a driving mode to be taken by the vehicle based on a result of the determination by the determination unit; A driving control unit that controls driving of the vehicle so that the vehicle runs in the determined driving mode; a transmission unit that, when the determination unit determines that the first driving mode and the second driving mode are not the same, transmits information about the first driving mode and the second driving mode that are not the same to an external device as log information, along with the surrounding environment information and the vehicle state information that form the basis for determining the first driving mode and the second driving mode, respectively.

2. The vehicle driving control device according to claim 1 , wherein the first algorithm is an algorithm based on reinforcement learning regarding a driving behavior of the vehicle.

3. 3. The vehicle driving control device according to claim 1, wherein the second algorithm corresponds to a rule based on knowledge about a driving manner of the vehicle.

4. 4. The vehicle driving control device according to claim 1, wherein the driving mode determination unit determines the second driving mode as the driving mode to be adopted by the vehicle when the determination unit determines that the first driving mode and the second driving mode are not the same.

5. 5. The vehicle driving control device according to claim 1, wherein the driving mode determination unit determines the first driving mode as the driving mode to be adopted by the vehicle when the determination unit determines that the first driving mode and the second driving mode are the same.

6. a receiving unit for receiving the log information transmitted from the transmitting unit in the vehicle driving control device according to claim 1; a log management unit that stores and manages the log information received by the receiving unit, The log information stored and managed by the log management unit is used to improve at least the first algorithm of the first algorithm and the second algorithm.

7. The first algorithm is an algorithm based on reinforcement learning about a driving behavior of a vehicle, the second algorithm corresponds to rules based on knowledge of vehicle driving behavior; 7. The vehicle control management device according to claim 6, wherein the log information stored and managed by the log management unit is used for at least improving the specifications of the reinforcement learning.

8. a receiving step of receiving the log information transmitted from the transmitting unit in the vehicle driving control device according to claim 1; a log management step of storing and managing the log information received in the receiving step; A vehicle control management method comprising: an improvement step of performing improvement work on at least the first algorithm of the first algorithm and the second algorithm, using the log information stored and managed in the log management step.

9. The first algorithm is an algorithm based on reinforcement learning about a driving behavior of a vehicle, the second algorithm corresponds to rules based on knowledge of vehicle driving behavior; The improvement step includes improving the specifications of the reinforcement learning based on the log information.

9. The vehicle control management method according to claim 8.

Citation Information

Patent Citations

  • Automatic operation device

    JP2018024286A

  • Operation simulator of automatic driving vehicle, operation confirmation method of automatic driving vehicle, control device of automatic driving vehicle and method for controlling automatic driving vehicle

    JP2019074885A

  • Learning device, learning method, and program

    JP2020035222A

  • Automobile arithmetic system

    JP2020142770A

  • Deep learning-based autonomous vehicle control device, system including the same, and method thereof

    US20180275657A1