Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for behavior-based adaptive cruise control

By using reinforcement learning and dynamic neural networks, the behavior of the target vehicle is quantified, a driver model is established, and adaptive cruise control is adjusted. This solves the problem of unconsidered target vehicle behavior and enables personalized following distance adjustment and safety control.

CN114763139BActive Publication Date: 2025-11-21GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111541940.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-01-14
Filing Date
2021-12-16
Publication Date
2025-11-21
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

Existing adaptive cruise control systems fail to effectively consider the behavior of the target vehicle, resulting in following distances that do not match driver preferences, and personalized adjustments are difficult to achieve on resource-constrained embedded controllers.

Method used

Reinforcement learning (RL) combined with dynamic neural networks (DNN) is used to quantify and predict the behavior of target vehicles, establish a driver behavior model, adjust the adaptive cruise control of the master vehicle, optimize the following strategy through a reward function, and achieve low-cost learning on a resource-constrained ECU.

Benefits of technology

It enables dynamic adjustment of the following distance based on driver preferences and target vehicle behavior, improving the performance of adaptive cruise control, maintaining safety margins, and providing personalized lane following customization on resource-constrained controllers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114763139B_ABST
    Figure CN114763139B_ABST
Patent Text Reader

Abstract

In various embodiments, methods, systems, and vehicle devices are provided. A method for implementing adaptive cruise control (ACC) established by reinforcement learning (RL) includes: executing, by a processor, adaptive cruise control to receive a set of vehicle inputs regarding an operating environment and a current operation of a host vehicle; identifying, by the processor, a target vehicle operating in the host vehicle environment and quantifying a set of target vehicle parameters derived from the sensed inputs regarding the target vehicle; modeling, by the processor, state estimates of the host vehicle and the target vehicle by generating a set of speed and torque calculation values regarding each vehicle; generating, by the processor, a result set from at least one reward function based on one or more modeled state estimates of the host vehicle and the target vehicle; and processing the result set with driver behavior data contained in the RL to associate one or more control actions with the driver behavior data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to vehicles, and more specifically to methods, systems, and apparatus for evaluating driver behavior and detecting target vehicle behavior to train an intelligent model for adaptive cruise control functions, the intelligent model being correlated with driver style in vehicle operation. Background Technology

[0002] In recent years, significant progress has been made in the autonomous and semi-autonomous driving features of vehicles, such as Super Cruise (a hands-free semi-autonomous driving assistance feature that uses high-definition maps and sensors to observe the road to help the vehicle accelerate and decelerate) and LKA (Lane Keeping Assist, a semi-autonomous driving feature that assists steering to keep the vehicle centered in its lane). Vehicles can still be improved in many ways.

[0003] Adaptive Cruise Control (ACC) allows a vehicle to automatically adjust its speed based on driver preferences to maintain a preset distance from the vehicle ahead. With currently available conventional cruise control systems, the driver can manually adjust the clearance distance to the target vehicle and the speed of the lead vehicle. However, in semi-autonomous driving, the distance behind the target vehicle may not be suitable for the driver's preferences. When executing acceleration and deceleration requests in Adaptive Cruise Control (ACC), the behavior of the target vehicle is currently not taken into account.

[0004] The goal is to understand the operating environment of the primary vehicle by combining the target vehicle information, and to modify command requests to improve ACC performance.

[0005] The goal is to customize ACC to tailor the target following distance based on real-time, historical, and online driver-vehicle interactions, while still maintaining an appropriate safety margin.

[0006] The goal is to classify and learn the driving behavior of target vehicles based on different driving scenarios (e.g., surrounding targets), road geometry, and target vehicle dynamics.

[0007] The goal is to build a knowledge base for the primary vehicle based at least on online and historical information, driving area, target type, road category, and relative lane position, and based on interactions with target vehicles that follow driver behavior based on performance preferences.

[0008] The goal is to adjust the following cruise control distance based on a real-time or stored knowledge base, tailored to individual drivers or categories.

[0009] The goal is to implement low-cost learning and classification algorithms for driver recognition on resource-constrained embedded controllers, and to provide customer-centric lane-following customization without adding extra hardware.

[0010] Furthermore, other desirable features and characteristics of this disclosure will become apparent from the following detailed description and appended claims, taken in conjunction with the accompanying drawings and the foregoing technical and background information. Summary of the Invention

[0011] In at least one exemplary embodiment, a method for implementing adaptive cruise control using reinforcement learning (RL) is provided. The method includes: a processor performing adaptive cruise control (ACC) to receive a set of vehicle inputs regarding the operating environment of a primary vehicle and its current operation; the processor identifying a target vehicle operating in the primary vehicle environment and quantifying a set of target vehicle parameters derived from the sensed inputs; the processor modeling state estimates of the primary and target vehicles by generating a set of calculated speed and torque values ​​for each vehicle; the processor generating a result set from at least one reward function based on one or more modeled state estimates of the primary and target vehicles; and processing the result set with driver behavior data contained in the RL to associate one or more control actions with the driver behavior data.

[0012] In at least one embodiment, the method includes applying at least one control action by a processor associated with driver behavior data of the RL to adjust at least one operation of the adaptive cruise control of the primary vehicle.

[0013] In at least one embodiment, the method includes adjusting at least one control action associated with driver behavior data of the RL by the processor based on a control safety check.

[0014] In at least one embodiment, the method includes updating data of a learning matrix by a processor based on a result set generated from at least one reward function to create a profile of driver behavior.

[0015] In at least one embodiment, the method includes a processor calculating a reward function using a set of parameters, including acceleration and speed estimates of the primary and target vehicles, and speed and torque calculations.

[0016] In at least one embodiment, the method includes adjusting one or more distances between a master vehicle and a target vehicle by a processor based on learned driver behavior contained in data in a learning matrix.

[0017] In at least one embodiment, the method includes a control safety check that includes the speed difference between the safe speed of the target vehicle and the speed estimate of the master vehicle.

[0018] In another exemplary embodiment, a system is provided. The system includes: a set of inputs obtained by a processor, the set of inputs including a set of vehicle inputs having one or more measurement inputs of a primary vehicle operation and sensed inputs regarding the operating environment of the primary vehicle, the primary vehicle being used to perform control operations of an adaptive cruise control (ACC) system established by reinforcement learning (RL) contained in the primary vehicle; the vehicle ACC system being guided by a driver behavior prediction model established by the RL, the driver behavior prediction model learning driver expectations online, and also using a dynamic neural network (DNN) to process the set of vehicle inputs to adjust control operations based on historical data; the processor being configured to identify a target vehicle operating in the primary vehicle environment to quantify a set of target vehicle parameters sensed from the processor regarding the target vehicle, the processor being configured to model state estimates of the primary vehicle and the target vehicle based on a set of speed and torque calculations for each vehicle;

[0019] In at least one exemplary embodiment, the processor is configured to generate a result set from at least one reward function based on one or more state estimates of the primary vehicle and the target vehicle; and the processor is configured to process the result set using driver behavior data established by the RL to associate one or more control actions with the driver behavior data. In a similar embodiment, historical data may be used in the DNN to associate one or more control actions with the driver behavior data.

[0020] In at least one exemplary embodiment, the processor is configured to apply at least one control action associated with driver behavior data established by the RL to adjust at least one control action of the ACC system of the master vehicle.

[0021] In at least one exemplary embodiment, the processor is configured to adjust at least one control action associated with driver behavior data established by the RL based on a control safety check.

[0022] In at least one exemplary embodiment, the processor is configured to adjust at least one control action associated with driver behavior data established by the RL based on a control safety check.

[0023] In at least one exemplary embodiment, the processor is configured to use a set of parameters to calculate a reward function, the set of parameters including the acceleration and speed estimates of the master vehicle and the target vehicle, and the speed and torque calculation.

[0024] In at least one exemplary embodiment, the processor is configured to adjust one or more distances between the master vehicle and the target vehicle based on learned driver behavior contained in the data of the learning matrix.

[0025] In at least one exemplary embodiment, the processor is configured to perform a control safety check, which includes checking the speed difference between a safe speed and speed estimates of the target vehicle and the host vehicle.

[0026] In yet another exemplary embodiment, a vehicle device is provided. The vehicle device includes a vehicle controller comprising a processor, wherein the processor is built using reinforcement learning (RL) and configured to: perform adaptive cruise control by receiving a set of vehicle inputs regarding the operating environment of a primary vehicle and current operation; identify a target vehicle operating in the primary vehicle environment by the processor and quantify a set of target vehicle parameters derived from the sensed inputs; model state estimates of the primary vehicle and the target vehicle by the processor through generating a set of speed and torque calculations for each vehicle; generate a result set from at least one reward function based on one or more modeled state estimates of the primary vehicle and the target vehicle; and associate the result set with driver behavior data built using RL to associate one or more control actions with the driver behavior data.

[0027] In at least one exemplary embodiment, the vehicle device includes a processor configured to: apply at least one control action associated with driver behavior data established by the RL to adjust at least one operation of the adaptive cruise control of the primary vehicle.

[0028] In at least one exemplary embodiment, the vehicle equipment includes a processor configured to adjust at least one control action associated with driver behavior data established by the RL based on a control safety check.

[0029] In at least one exemplary embodiment, the vehicle device includes a processor configured to update data of a learning matrix based on a result set generated from at least one reward function to create a profile of driver behavior and quantify driver expectations.

[0030] In at least one exemplary embodiment, the vehicle device includes a processor configured to calculate a reward function using a set of parameters, including acceleration and speed estimates of the primary vehicle and the target vehicle, and speed and torque calculations.

[0031] In at least one exemplary embodiment, the vehicle device includes a processor configured to adjust one or more distances between a master vehicle and a target vehicle based on learned driver behavior contained in data for solving a learning matrix; learning online using a proposed RL or learning offline using a developed DNN. Attached Figure Description

[0032] Exemplary embodiments will now be described in conjunction with the following accompanying drawings, in which the same reference numerals denote the same elements, and in the drawings:

[0033] Figure 1 This is a functional block diagram illustrating an autonomous or semi-autonomous vehicle with a control system according to an exemplary embodiment, the control system controlling vehicle actions based on predicting driver behavior using a neural network in the vehicle control system.

[0034] Figure 2 This is a diagram illustrating an adaptive cruise control system according to various embodiments, which can be implemented using a neural network to predict driver behavior of the vehicle control system.

[0035] Figure 3 It shows according to Figure 1-2 The diagram shows components of an adaptive cruise control system according to various embodiments, which, according to various embodiments, can be implemented using a neural network to predict driver behavior of the vehicle control system.

[0036] Figure 4 It is illustrated that various embodiments are used for Figure 1-3 An example diagram of the reward function in the control method of the adaptive cruise control system is shown.

[0037] Figure 5 This illustrates various embodiments. Figure 1-3 Example diagrams illustrating the potential benefits of the exemplary use of the adaptive cruise control system; and

[0038] Figure 6 These are exemplary flowcharts according to various embodiments, illustrating in Figure 1-3 The steps used in the adaptive cruise control system are shown. Detailed Implementation

[0039] The following detailed description is merely exemplary in nature and is not intended to limit application and use. Furthermore, it is not intended to be bound by any express or implied theory presented in the foregoing technical field, background art, summary of the invention, or the following detailed description. As used herein, the term "module" refers to any hardware, software, firmware, electronic control components, processing logic, and / or processor device, employed individually or in any combination, including but not limited to: application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), electronic circuits, processors (shared, dedicated, or grouped), and memory executing one or more software or firmware programs, combinational logic circuitry, and / or other suitable components providing the described functionality.

[0040] Embodiments of this disclosure can be described herein based on functional and / or logical block components and various processing steps. It should be understood that such block components can be implemented by any number of hardware, software, and / or firmware components configured to perform specified functions. For example, embodiments of this disclosure can employ various integrated circuit components, such as memory elements, digital signal processing elements, logic elements, lookup tables, etc., which can perform various functions under the control of one or more microprocessors or other control devices. Furthermore, those skilled in the art will understand that embodiments of this disclosure can be practiced in conjunction with any number of systems, and the systems described herein are merely exemplary embodiments of this disclosure.

[0041] For the sake of brevity, conventional techniques related to signal processing, data transmission, signaling, control, machine learning, image analysis, and other functional aspects of the system (as well as the various operating components of the system) are not described in detail herein. Furthermore, the connecting lines shown in the various figures included herein are intended to represent exemplary functional relationships and / or physical connections between various elements. It should be noted that many alternative or additional functional relationships or physical connections may exist in the embodiments of this disclosure.

[0042] In an ACC system, the time interval to the target vehicle can be set in increments of 1 to 2.5 seconds. For example, if the target vehicle accelerates, the primary vehicle accelerates, but only to its maximum limit. If another vehicle appears in front of the target vehicle, ACC will automatically lock onto the new target vehicle and take a short time to identify it. When making these decisions, ACC is unaware of driver preferences regarding different geographical locations, settings, etc. ACC can monitor vehicle movement and collect data on driver preferences.

[0043] This disclosure provides methods, systems, and apparatus that implement intelligent systems and methods that mathematically quantify the behavior of a target vehicle and incorporate that behavior into an adaptive control design that provides what is desired by the adaptive cruise control characteristics.

[0044] Furthermore, this disclosure provides methods, systems, and apparatus for implementing an online method that incorporates additional intelligent features to enhance performance and desired following distance in an interactive manner.

[0045] In addition, methods, systems, and devices are provided that can classify the attributes of target vehicles and locations into adaptive cruise control.

[0046] refer to Figure 1According to various embodiments, the control system 100 is associated with the vehicle 10 (also referred to herein as the "main vehicle"). Typically, the control system (or simply the "system") 100 provides control over various actions of the vehicle 10 (e.g., torque control), which is established by reinforcement learning (RL), which is a DNN-type model or can be stored in a DNN-type model, and controls the operation in response to data from vehicle inputs, for example, as combined below. Figure 2-6 More detailed description.

[0047] In various exemplary embodiments, System 100 is capable of providing an ACC behavior prediction model that learns a driver's preference for following different target vehicles at varying distances. System 100 includes a method for classifying driver preferences based on driving scenarios (e.g., traffic signs, stop-and-go traffic, city driving, etc.). System 100 is capable of building a knowledge base for target following performance preferences by utilizing online and historical driver and environmental information. System 100 is capable of adjusting lane following performance using online driver-vehicle interactions and can adjust lane following control for individual drivers or categories based on real-time or stored knowledge. System 100 can also adjust lane following performance preferences on the vehicle based on driver ID and can provide a low-cost learning method with a driver identification and classification algorithm that can be executed on resource-constrained ECUs (which can be deployed on top of existing SuperCruise / LC lateral control algorithms). System 100 provides lane following customization while maintaining a safety margin.

[0048] In various exemplary embodiments, system 100 provides a process for using algorithms in the embedded controller software of system 100 of the master vehicle 10 to control torque and speed, thereby allowing a DNN to be used for ACC behavior prediction models. System 100 enables the learning of a driver's preference for following distances to different target vehicles (e.g., target vehicles) and classifies the driver's preferences based on driving scenarios; for example, traffic signs, stop-and-go traffic, city driving, etc. Using a Q-matrix, system 100 can build a knowledge base for target vehicle follow-through performance preferences by leveraging online and historical driver and environmental information.

[0049] like Figure 1 As shown, vehicle 10 typically includes a chassis 12, a body 14, front wheels 16, and rear wheels 18. The body 14 is mounted on the chassis 12 and substantially encloses the components of vehicle 10. The body 14 and chassis 12 may together form a frame. Each of wheels 16-18 is rotatably coupled to the chassis 12 near a corresponding corner of the body 14. In various embodiments, wheels 16, 18 include wheel assemblies that also include correspondingly associated tires.

[0050] In various embodiments, vehicle 10 is autonomous or semi-autonomous, and control system 100 and / or components thereof are integrated into vehicle 10. Vehicle 10 is, for example, a vehicle automatically controlled to transport passengers from one location to another. Vehicle 10 is depicted as a passenger car in the illustrated embodiment, but it should be understood that any other vehicle may be used, including motorcycles, trucks, sports utility vehicles (SUVs), recreational vehicles (RVs), boats, aircraft, etc.

[0051] As shown in the figure, vehicle 10 typically includes a propulsion system 20, a transmission system 22, a steering system 24, a braking system 26, a filter purification system 31, one or more user input devices 27, a sensor system 28, an actuator system 30, at least one data storage device 32, at least one controller 34, and a communication system 36. In various embodiments, the propulsion system 20 may include an internal combustion engine, an electric motor such as a traction motor, and / or a fuel cell propulsion system. The transmission system 22 is configured to transmit power from the propulsion system 20 to wheels 16 and 18 according to a selectable speed ratio. According to various embodiments, the transmission system 22 may include a stepped transmission, a continuously variable transmission (CVT), or other suitable transmission.

[0052] Braking system 26 is configured to provide braking torque to wheels 16 and 18. In various embodiments, braking system 26 may include friction brakes, brake-by-wire brakes, regenerative braking systems such as motors, and / or other suitable braking systems.

[0053] Steering system 24 affects the position of wheels 16 and / or 18. Although depicted as including a steering wheel for illustrative purposes, in some embodiments contemplated within the scope of this disclosure, steering system 24 may not include a steering wheel.

[0054] The controller 34 includes at least one processor 44 (and a neural network 33) and a computer-readable storage device or medium 46. As described above, in various embodiments, the controller 34 (e.g., its processor 44) pre-provides the steering control system 84 with data regarding the anticipated future path of the vehicle 10, including anticipated future steering commands, for controlling steering for a limited period of time in the event that communication with the steering control system 84 becomes unavailable. Furthermore, in various embodiments, the controller 34 communicates via a communication system 36, further described below, such as via a communication bus and / or a transmitter (…). Figure 1 (Not shown in the image) provides communication to the steering control systems 84 and 34.

[0055] In various embodiments, the controller 34 includes at least one processor 44 and a computer-readable storage device or medium 46. The processor 44 can be any custom or commercially available processor, central processing unit (CPU), graphics processing unit (GPU), auxiliary processor among several processors associated with the controller 34, semiconductor-based microprocessor (in the form of a microchip or chipset), any combination thereof, or any device generally used for executing instructions. The computer-readable storage device or medium 46 can include volatile and non-volatile storage, such as read-only memory, random access memory, and keep-alive memory (KAM). KAM is persistent or non-volatile memory that can be used to store multiple neural networks and various operational variables when the processor 44 is powered off. The computer-readable storage device or medium 46 can be implemented using any of a variety of known storage devices, such as PROMs (programmable read-only memory), EPROMs (electrically programmable PROMs), EEPROMs (electrically erasable PROMs), flash memory, or any other electrical, magnetic, optical, or combined storage device capable of storing data, some of which represent executable instructions used by the controller 34 when controlling the vehicle 10.

[0056] The instructions may include one or more separate programs, each comprising an ordered list of executable instructions for implementing logical functions. When executed by processor 44, these instructions receive and process signals from sensor system 28, execute logic, calculations, methods, and / or algorithms for automatically controlling components of vehicle 10, and generate control signals transmitted to actuator system 30 to automatically control components of vehicle 10 based on logic, calculations, methods, and / or algorithms. Although in Figure 1 Only one controller 34 is shown, but embodiments of vehicle 10 may include any number of controllers 34 that communicate via any suitable communication medium or combination of communication media and cooperate to process sensor signals, execute logic, calculations, methods and / or algorithms, and generate control signals to automatically control the features of vehicle 10.

[0057] like Figure 1 As shown, in addition to the steering system 24 and controller 34 described above, vehicle 10 typically includes a chassis 12, a body 14, front wheels 16, and rear wheels 18. The body 14 is mounted on the chassis 12 and substantially encloses the components of vehicle 10. The body 14 and chassis 12 may together form a frame. Each of wheels 16-18 is rotatably coupled to the chassis 12 near a corresponding corner of the body 14. In various embodiments, wheels 16, 18 include wheel assemblies that also include correspondingly associated tires.

[0058] In various embodiments, vehicle 10 is an autonomous vehicle, and control system 100 and / or components thereof are integrated into vehicle 10. Vehicle 10 is, for example, a vehicle that is automatically controlled to transport passengers from one location to another. Vehicle 10 is depicted as a passenger car in the illustrated embodiment, but it should be understood that any other vehicle may be used, including motorcycles, trucks, sports utility vehicles (SUVs), recreational vehicles (RVs), boats, aircraft, etc.

[0059] As shown in the figure, vehicle 10 typically also includes a propulsion system 20, a transmission system 22, a braking system 26, one or more user input devices 27, a sensor system 28, an actuator system 30, at least one data storage device 32, and a communication system 36. In various embodiments, the propulsion system 20 may include an internal combustion engine, an electric motor such as a traction motor, and / or a fuel cell propulsion system. The transmission system 22 is configured to transmit power from the propulsion system 20 to wheels 16 and 18 according to a selectable speed ratio. According to various embodiments, the transmission system 22 may include a stepped transmission, a continuously variable transmission (CVT), or other suitable transmission.

[0060] Braking system 26 is configured to provide braking torque to wheels 16 and 18. In various embodiments, braking system 26 may include friction brakes, brake-by-wire brakes, regenerative braking systems such as motors, and / or other suitable braking systems.

[0061] Steering system 24 affects the position of wheels 16 and / or 18. Although depicted as including a steering wheel for illustrative purposes, in some embodiments contemplated within the scope of this disclosure, steering system 24 may not include a steering wheel.

[0062] The controller 34 includes a vehicle controller that is directly affected by the output of the neural network 33 model. In an exemplary embodiment, feedforward operation may be applied to a regulation factor, which is the continuous output of the neural network 33 model, to generate a control action or other similar action for a desired torque (e.g., in the case of a continuous neural network 33 model, the continuous APC / SPARK prediction is the output).

[0063] In various embodiments, one or more user input devices 27 receive input from one or more passengers (and driver 11) of vehicle 10. In various embodiments, the input includes the desired destination of vehicle 10. In some embodiments, one or more input devices 27 include an interactive touchscreen in vehicle 10. In some embodiments, one or more input devices 27 include speakers for receiving audio information from passengers. In some other embodiments, one or more input devices 27 may include one or more other types of devices and / or user devices that can be coupled to passengers (e.g., smartphones and / or other electronic devices).

[0064] The sensor system 28 includes one or more sensors 40a-40n that sense observable conditions of the external and / or internal environment of the vehicle 10. Sensors 40a-40n include, but are not limited to, radar, lidar, global positioning system, optical camera, thermal camera, ultrasonic sensor, inertial measurement unit, and / or other sensors.

[0065] The actuator system 30 includes one or more actuators 42a-42n that control one or more vehicle features, such as, but not limited to, a filter purification system 31, an intake system 38, a propulsion system 20, a transmission system 22, a steering system 24, and a braking system 26. In various embodiments, the vehicle 10 may also include Figure 1 Interior and / or exterior vehicle features not shown, such as various doors, trunk and cabin features, such as air, music, lighting, touch screen display components (e.g., components used in conjunction with a navigation system), etc.

[0066] Data storage device 32 stores data for automatically controlling vehicle 10, including storing DNN data built by RL for predicting driver behavior in vehicle control. In various embodiments, data storage device 32 stores machine learning models of the DNN and other data models built by RL. Models built by RL can occur within DNN behavior prediction models or models built by RL (see...). Figure 2 The DNN prediction model (210) or RL prediction model is used. In an exemplary embodiment, instead of training the DNN separately, a set of learned functions is used to implement the DNN behavior prediction model (i.e., the DNN prediction model). In various embodiments, the neural network 33 (i.e., the DNN behavior prediction model) may be built by RL or trained by a remote system through a supervised learning method and transmitted or provided in the vehicle 10 (wirelessly and / or wired) and stored in the data storage device 32. The DNN behavior prediction model may also be trained by supervised or unsupervised learning based on input vehicle data of the master vehicle operation and / or sensed data about the master vehicle operating environment.

[0067] Data storage device 32 is not limited to control data, as other data can also be stored in it. For example, route information, i.e., a set of road segments (geographically associated with one or more defined maps), can be stored in data storage device 32, defining the route a user can take from a starting location (e.g., the user's current location) to a destination location. It should be understood that data storage device 32 can be part of controller 34, separate from controller 34, or part of controller 34 and a separate system.

[0068] The controller 34 implements a logic model built by RL or for a DNN based on a DNN behavioral model that has been trained with a set of values, including at least one processor 44 and a computer-readable storage device or medium 46. The processor 44 can be any custom or commercially available processor, central processing unit (CPU), graphics processing unit (GPU), auxiliary processor among several processors associated with the controller 34, semiconductor-based microprocessor (in the form of a microchip or chipset), any combination thereof, or any device typically used for executing instructions. The computer-readable storage device or medium 46 can include, for example, volatile and non-volatile memory in read-only memory, random access memory, and keep-alive memory (KAM). KAM is persistent or non-volatile memory that can be used to store various operational variables when the processor 44 is powered off. The computer-readable storage device or medium 46 may be implemented using any of a variety of known storage devices, such as PROMs (programmable read-only memory), EPROMs (electrically programmable PROMs), EEPROMs (electrically erasable PROMs), flash memory, or any other electrical, magnetic, optical, or combined storage device capable of storing data, some of which represent executable instructions used by the controller 34 when controlling the vehicle 10.

[0069] The instructions may include one or more separate programs, each comprising an ordered list of executable instructions for implementing logical functions. When executed by processor 44, these instructions receive and process signals from sensor system 28, execute logic, calculations, methods, and / or algorithms for automatically controlling components of vehicle 10, and generate control signals transmitted to actuator system 30 to automatically control components of vehicle 10 based on logic, calculations, methods, and / or algorithms. Although in Figure 1 Only one controller 34 is shown, but embodiments of vehicle 10 may include any number of controllers 34 that communicate via any suitable communication medium or combination of communication media and cooperate to process sensor signals, execute logic, calculations, methods and / or algorithms, and generate control signals to automatically control the features of vehicle 10.

[0070] Communication system 36 is configured to wirelessly transmit information to and from other entities 48, such as, but not limited to, other vehicles (“V2V” communication), infrastructure (“V2I” communication), remote transportation systems and / or user equipment (see reference). Figure 2(Described in more detail). In an exemplary embodiment, communication system 36 is a wireless communication system configured to communicate using the IEEE 802.11 standard or via a wireless local area network (WLAN) using cellular data communication. However, additional or alternative communication methods, such as dedicated short-range communication (DSRC) channels, are also considered to be within the scope of this disclosure. A DSRC channel refers to a one-way or two-way short-to-medium-range wireless communication channel specifically designed for automotive use and for a corresponding set of protocols and standards.

[0071] In various embodiments, the communication system 36 is used for communication between controllers 34, including data related to the expected future path of vehicle 10, including expected future steering commands. Furthermore, in various embodiments, the communication system 36 may facilitate communication between the steering control system 84 and / or other systems and / or devices.

[0072] In some embodiments, the communication system 36 is also configured for communication between the sensor system 28, input device 27, actuator system 30, one or more controllers (e.g., controller 34), and / or more other systems and / or devices. For example, the communication system 36 may include any combination of a controller area network (CAN) bus and / or direct wiring between the sensor system 28, actuator system 30, one or more controllers 34, and / or one or more other systems and / or devices. In various embodiments, the communication system 36 may include one or more transceivers for communication with one or more devices and / or systems of the vehicle 10, passenger devices (e.g., [missing information]), and other systems and / or devices. Figure 2 The user device 54) communicates with one or more remote information sources (e.g., GPS data, traffic information, weather information, etc.).

[0073] Now for reference Figure 2 , Figure 2 This diagram illustrates an adaptive cruise control (ACC) system 200 according to various embodiments, which can be implemented using a dynamic neural network to predict driver behavior in the vehicle control system. System 200 includes setting inputs from vehicle sensors 205 (such as...). Figure 1As described above, it includes inputs from sources such as radar, cameras, lidar, etc., as well as input 210 from the environment (the vehicle's surrounding environment) and vehicle instrument data. Input 210 is received by a state estimation module 220 constituting an online learning module 215. Input 210 is received by a driver adaptation module 225 constituting an online learning module 215. The online learning module 215 generates online corrections for the driver expectation adjustment module 230. ACC control commands 255 are received by the driver expectation adjustment module 230. The driver expectation adjustment module 230 is associated with driver control by associating with a driver control module 235 and stores relevant information and new settings in a neural network 240. The driver expectation adjustment module 230 performs a safety check via an ACC control safety check 245.

[0074] In one exemplary embodiment, an ACC control safety check is implemented for online adaptation. The quantization function of the safety check is defined as follows: V safety =d -2 +(v ACC -v est ) 2 Only when v ACC =v est When d→∞, V safety =0; otherwise V safety >0. The expectation is to improve ACC control and allow V safety →0. In other cases, V safety >>0 indicates that the implemented control needs adjustment or has malfunctioned.

[0075] To calculate V safety To avoid destabilizing or reducing the performance of the main vehicle's ACC control, the same or similar ACC torque can be applied to the vehicle dynamics model, and v can be calculated. Aprx =f(τ) Acc ,v est Then the vehicle dynamics can be examined against the speed-based safety measurement model, as follows: V safety,Aprx =(d+Δ(v) Aprx )) -2 +(v ACC -v est ) 2 Therefore, safety can be checked if the following conditions are met, indicating that the commands of the ACC adaptive control method are considered safe and can be improved (the safety check can also indicate V). safety,Aprx (Non-increasing) Aprx ≤0.

[0076] During safety inspections, the goal is generally to improve ACC control and ensure V safety→0. Offline data is also sent to the driver expectation adjustment module 230 via the offline data module 250.

[0077] Figure 3 It shows according to Figure 1-2 The diagram shows components of an adaptive cruise control system according to various embodiments, which, according to various embodiments, can be implemented using a neural network to predict driver behavior in the vehicle control system. Figure 3 In the process, the target vehicle and the main vehicle information, vehicle dynamic operation parameter information are received and quantified at the quantization module 310, including the steps of quantifying the target vehicle parameters and collecting and analyzing information based on forward perception to generate steady-state information about the target vehicle.

[0078] Quantitative and steady-state information are analyzed and quantified by the data learning module 305 and then sent to the state estimation module 325. The state estimation module 325 implements the function. To determine the state of the target and the main vehicle. State estimation is based on v ACC Determined speed, estimated speed v est Driver torque τ Driver and ACC torque τ ACC ACC torque Time vector and time vector of detected target vehicle brake light The parameters and other parameters. In addition, the reward function R... ij Between the main vehicle and the target vehicle Between the control and the calculation of the reward function R. ij At that time, the target vehicle and the status of the host vehicle are identified and received by inputting 345.

[0079] In one exemplary embodiment, a set of reward functions R (R1 to R4) used to determine the reward function are represented as follows: R1 = (τ ACC -τ Driver ) -2 , R3=d 2 and R4 = (v ACC -v est ) -2 To R n .

[0080] This estimation model allows the ACC learning reward function to be optimally correlated with the driver's vehicle handling style and driver characteristics. It can learn and save driver profiles for each "driver ID".

[0081] After checking policy action 350, Q-learning module 340 can update Q matrix Q(i,j)=αR ij+(1-α)Q(i,j). ACC control safety check 355 checks the speed of the main vehicle. After the security check, appropriate control actions can be applied 360 degrees.

[0082] Figure 4 It is illustrated that various embodiments are used for Figure 1-3 The diagram shows an example of the reward function in the control method of the adaptive cruise control system. Figure 4 In the example diagram, there are a primary vehicle 405 and a target vehicle 410. The target vehicle 410 has a set of reward functions, which are calculated based on driver feedback R1, instruction R2, distance R3, and ACC performance R4. i It is a reward function associated with each environment or state.

[0083] Figure 5 This illustrates various embodiments. Figure 1-3 An example diagram illustrating the potential benefits of an exemplary use of an adaptive cruise control system. Figure 5 The graph illustrates an exemplary embodiment of the potential benefits to the primary vehicle, where the adaptive ACC system is activated to react to the brake light status in the target vehicle, thereby reducing the axle torque of the primary vehicle, which results in or facilitates a smoother brake light response. The deceleration rate of the target vehicle is correspondingly more gradual.

[0084] Figure 6 These are exemplary flowcharts according to various embodiments, illustrating in Figure 1-3 The steps used in the adaptive cruise control system are shown.

[0085] exist Figure 6 In Task 610, flowchart 600 receives a set of vehicle inputs regarding the operating environment and current operation of the primary vehicle by performing adaptive cruise control to identify road geometry and environmental factors. Next, in Task 620, the driver's expected behavior to react to various states of the target vehicle is learned by identifying the target vehicle operating in the primary vehicle's environment and quantifying a set of target vehicle parameters derived from the sensed inputs. Furthermore, by modeling the state estimates of the primary and target vehicles, a set of speed and torque calculations is generated for each vehicle, and a result set is produced from at least one reward function based on one or more modeled state estimates of the primary and target vehicles. This result set is then processed with driver behavior data contained in the DNN to associate one or more control actions with the driver behavior data.

[0086] In an exemplary embodiment, for the target vehicle, the following distances are quantified to model the driver's driving patterns in the master vehicle. At task 630, based on the result set generated from the reward function, it is expected that a behavior matrix will be built with updated data to create a profile of the driver's behavior. At task 640, the ACC system is implemented to adjust subsequent distances by providing the desired acceleration and deceleration rates. The ACC system executes the control actions required for the master vehicle operation by calculating the reward function using a set of parameters for calculating and estimating the acceleration and velocity of the master and target vehicles.

[0087] In various exemplary embodiments, the implemented driver behavior prediction model logic can be created in offline training derived from a supervised or unsupervised learning process and can be implemented using other neural networks. For example, other neural networks may include trained convolutional neural networks (CNNs) and / or recurrent neural networks (RNNs), where similar approaches can be applied and used for vehicle control operations. Furthermore, alternative embodiments are conceivable, according to various embodiments, which include a neural network consisting of multiple (i.e., 3-layer) convolutional neural networks (CNNs) and also having dense layers (i.e., 2 dense layers) that have been trained offline and enable [interaction with] [other methods]. Figure 1 The system shown controls the operation of the ACC in a coordinated manner.

[0088] A dynamic neural network is used to inform the ACC controller of torque and speed characteristics and is configured as a pre-trained neural network. Therefore, in some embodiments, the torque prediction system process is configured only in operating mode. For example, in various embodiments, the dynamic neural network is trained during training mode before being used or provided in the vehicle (or other vehicles). Once the dynamic neural network is trained, it can be used in operating mode in the vehicle (e.g., Figure 1 The implementation is carried out in vehicles 10), where the vehicles are operated autonomously, semi-autonomously, or manually.

[0089] In various alternative exemplary embodiments, it should be understood that the neural network may also be implemented in both training and operating modes within the vehicle, and trained during initial operation in conjunction with time delays or similar methods for torque control prediction. Furthermore, in various embodiments, the vehicle may operate only in operating modes having neural networks that have been trained via training modes of the same vehicle and / or other vehicles.

[0090] As briefly mentioned, the various modules and systems described above can be implemented as one or more machine learning models that undergo supervised, unsupervised, semi-supervised, or reinforcement learning. Such models can be trained to perform tasks such as classification (e.g., binary or multi-class classification), regression, clustering, dimensionality reduction, and / or similar tasks. Examples of such models include, but are not limited to, artificial neural networks (ANNs) (such as recurrent neural networks (RNNs) and convolutional neural networks (CNNs)), decision tree models (such as classification and regression trees (CART)), ensemble learning models (such as boosting, bootstrap aggregation, gradient boosting machines, and random forests), Bayesian network models (such as Naive Bayes), principal component analysis (PCA), support vector machines (SVMs), clustering models (such as K-nearest neighbors, K-means, expectation maximization, hierarchical clustering, etc.), and linear discriminant analysis models.

[0091] It should be understood that Figure 1-6 The process can include any number of additional or alternative tasks. Figure 1-6 The tasks shown do not need to be performed in the order illustrated, and Figure 1-6 This process can be integrated into a more comprehensive process or process with additional functions not described in detail here. Furthermore, Figure 1-6 One or more tasks shown can be performed from Figure 1-6 The process shown in the embodiments is omitted, as long as the intended overall functionality remains intact.

[0092] The foregoing detailed description is merely illustrative in nature and is not intended to limit the embodiments of the subject matter or the application and use of such embodiments. As used herein, the word "exemplary" means "serving as an example, instance, or illustration." Any implementation described herein as exemplary is not necessarily to be construed as being more preferred or advantageous than other implementations. Furthermore, no one is intended to be bound by any express or implied theory presented in the foregoing technical field, background, or detailed description.

[0093] While at least one exemplary embodiment has been presented in the foregoing detailed description, it should be understood that numerous variations exist. It should also be understood that the one or more exemplary embodiments are merely examples and are not intended to limit the scope, applicability, or configuration of this disclosure in any way. Rather, the foregoing detailed description will provide those skilled in the art with a convenient roadmap for implementing one or more exemplary embodiments.

[0094] It should be understood that various changes can be made to the function and arrangement of the elements without departing from the scope of this disclosure as set forth in the appended claims and their legal equivalents.

Claims

1. A method for implementing adaptive cruise control (ACC) established through reinforcement learning (RL), comprising: The processor executes adaptive cruise control by receiving a set of vehicle inputs about the operating environment of the primary vehicle and the current operation. The processor identifies the target vehicle operating in the main vehicle environment and quantifies a set of target vehicle parameters obtained from the sensed input. The processor models the state estimates of the master vehicle and the target vehicle by generating a set of speed and torque calculations for each vehicle and by utilizing a first time vector of ACC torque and a second time vector of the target vehicle's brake light detection. The processor generates a result set from at least one reward function based on one or more modeled state estimates of the master vehicle and the target vehicle. and The result set is processed using driver behavior data established by RL to associate one or more control actions with the driver behavior data. The generation of this result set is performed based on multiple reward functions, including: A first reward function, which is related to driver feedback and is based on the difference between driver torque and ACC torque; A second reward function, which is related to the indication and is based on a first time vector of ACC torque and a second time vector of the target vehicle's brake light detection; The third reward function is distance-related and based on the distance between the master vehicle and the target vehicle; and The fourth reward function is related to ACC performance and is based on ACC speed and estimated speed.

2. The method according to claim 1, further comprising: The processor applies at least one control action associated with the driver behavior data established via RL to adjust at least one operation of the adaptive cruise control of the master vehicle.

3. The method according to claim 2, further comprising: The processor adjusts the at least one control action associated with the driver behavior data established by the RL based on the control safety check.

4. The method according to claim 3, further comprising: The processor updates the data of the learning matrix based on the result set generated from at least one reward function to create a profile of driver behavior.

5. The method according to claim 4, further comprising: The processor calculates the reward function using a set of parameters, including the acceleration and speed estimates of the master vehicle and the target vehicle, as well as the speed and torque calculations.

6. The method according to claim 5, further comprising: The processor adjusts one or more distances between the master vehicle and the target vehicle based on learned driver behavior contained in the learning matrix.

7. The method according to claim 6, wherein, The control safety check includes the speed difference between the safe speed and the speed estimate of the target vehicle and the host vehicle.

8. A system comprising: A set of inputs obtained by the processor, including a set of vehicle inputs with one or more measurement inputs of the master vehicle operation and sensed inputs about the operating environment of the master vehicle, is used to perform control operations of an adaptive cruise control (ACC) system built by reinforcement learning (RL) and included in the master vehicle. The vehicle's ACC system is guided by a driver behavior prediction model implemented by a processor through RL. The driver behavior prediction model learns the driver's expectations online and uses a neural network (NN) to process the set of vehicle inputs to adjust the control operation. The processor is configured to identify a target vehicle operating in the host vehicle environment to quantify a set of target vehicle parameters obtained from sensed inputs about the target vehicle. The processor is configured to model the state estimates of the master vehicle and the target vehicle based on a set of speed and torque calculations for each vehicle and a second time vector by utilizing a first time vector of ACC torque and a second time vector of the target vehicle's brake light detection. The processor is configured to generate a result set from at least one reward function based on one or more state estimates of the master vehicle and the target vehicle; and The processor is configured to process the result set using driver behavior data established by RL, in order to associate one or more control actions with the driver behavior data. The processor is further configured to generate a result set based on a plurality of reward functions, the plurality of reward functions including: A first reward function, which is related to driver feedback and is based on the difference between driver torque and ACC torque; A second reward function, which is related to the indication and is based on a first time vector of ACC torque and a second time vector of the target vehicle's brake light detection; The third reward function is distance-related and based on the distance between the master vehicle and the target vehicle; and The fourth reward function is related to ACC performance and is based on ACC speed and estimated speed.

9. The system of claim 8, further comprising the processor being configured to: Apply at least one control action associated with driver behavior data established by RL to adjust at least one control action of the ACC system of the master vehicle; Adjust the at least one control action associated with the driver behavior data established by the RL based on the control safety check; and Adjust at least one control action associated with driver behavior data of the NN based on control safety checks.

10. The system of claim 9, further comprising the processor being configured to: The reward function is calculated using a set of parameters, including the acceleration and speed estimates of the master vehicle and the target vehicle, as well as the speed and torque calculations. One or more distances between the master vehicle and the target vehicle are adjusted based on driver behavior learned from data contained in the learning matrix. in, The control safety check includes the speed difference between the safe speed and the speed estimate of the target vehicle and the host vehicle.

Citation Information

Patent Citations

  • Self-adaptive cruise control system considering driving behaviors and control method thereof

    CN112109708A

  • Adaptive longitudinal control using reinforcement learning

    US20190367025A1