Dynamic threshold control method and multi-threshold cooperative control method for autonomous vehicles
By constructing a dynamic threshold control method for autonomous vehicles, and utilizing multi-threshold collaborative control and reinforcement learning, the problem that fixed thresholds in autonomous driving systems are difficult to adapt to complex environments is solved. This achieves comprehensive optimization of safety, comfort, and robustness, and improves the accuracy of takeover decisions and the adaptive capability of the strategy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-27
- Publication Date
- 2026-04-03
AI Technical Summary
Existing autonomous driving systems use a fixed single threshold in their takeover trigger mechanism, which is difficult to adapt to complex and dynamic driver states and environmental risks, leading to excessive takeover, false alarms, and unstable threshold adjustments between different operating conditions.
A dynamic threshold control method for autonomous vehicles is adopted. By using onboard cameras, driver monitoring systems and vehicle sensors to obtain driver distraction level scores, vehicle operating status and environmental risk field intensity, a state space is constructed. Through a multi-threshold collaborative control mechanism and reinforcement learning-based strategy training, the takeover-related thresholds are adaptively adjusted in multiple scenarios and operating conditions.
It achieves human-machine collaborative control under the same risk assessment framework, improves the accuracy and smoothness of takeover triggering, enhances safety, comfort and system robustness, reduces false alarms and false alarms, reduces engineering parameter tuning workload, and improves the generalization ability and maintainability of the strategy.
Smart Images

Figure CN121573011B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of road vehicle control and relates to the takeover of autonomous vehicles. Specifically, it relates to a dynamic threshold control method and a multi-threshold collaborative control method for autonomous vehicles. Background Technology
[0002] With the rapid development of autonomous driving technology, the safety and reliability of vehicles operating in complex traffic environments has become one of the core research issues. Most existing autonomous driving systems employ fixed takeover trigger thresholds, meaning that takeover is triggered when the driver's attention level or environmental risk exceeds a certain static threshold. However, this single static threshold mechanism has significant limitations: on the one hand, drivers' attention states and reaction abilities exhibit significant individual differences and time-varying characteristics, making it difficult for static thresholds to fully characterize their dynamic changes; on the other hand, road environments possess high uncertainty and randomness, and a single threshold often fails to achieve a good balance between safety and driving comfort.
[0003] Reinforcement learning, as a data-driven adaptive decision-making method, has been widely applied in recent years in fields such as autonomous driving trajectory planning, risk assessment, and human-machine collaborative control. Reinforcement learning can continuously optimize strategies through interaction with the environment. However, existing research mostly focuses on optimizing single risk indicators, lacking multi-dimensional collaborative modeling of driver state, vehicle dynamics, and environmental risks, and still has shortcomings in the smoothness and stability of threshold adjustment. Therefore, there is an urgent need for a dynamic takeover threshold control method based on reinforcement learning that can improve safety while also considering driving comfort and system robustness. Summary of the Invention
[0004] Given that existing autonomous driving systems generally use a fixed single threshold in their takeover trigger mechanisms, which is difficult to adapt to complex and dynamic driver states and environmental risks, resulting in problems such as excessive takeover, false alarms, and unstable threshold adjustments across different operating conditions, this invention proposes a dynamic threshold control method for autonomous vehicles. This method utilizes onboard cameras, driver monitoring systems, and vehicle sensors to acquire driver distraction level scores, vehicle operating states, and environmental risk field strengths, and constructs a state space under a unified dimension. Under the same risk assessment framework, two dynamic boundaries, a warning threshold and a takeover threshold, are introduced. Through a multi-threshold collaborative control mechanism and reinforcement learning-based strategy training, adaptive adjustment of takeover-related thresholds under multiple scenarios and operating conditions is achieved.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] A dynamic threshold control method for autonomous vehicles, comprising the following steps:
[0007] Step 1. Collect vehicle operation data in real time and construct a comprehensive state space. ;in, This indicates the driver's level of distraction. Indicates the vehicle's operating status. Indicates the intensity of the environmental risk field;
[0008] Step 2. Set two types of dynamic threshold boundaries: warning threshold and takeover threshold.
[0009] Step 3. Train and optimize the dynamic threshold adjustment strategy based on reinforcement learning methods, using the comprehensive state space as input, the warning threshold and takeover threshold as the action space, and setting the reward function: ;
[0010] in, , , , , These are the weighting coefficients for different evaluation indicators. For collision cost terms, To take over the frequency term, This is an invalid warning item. For threshold smoothness, This is an exploratory item;
[0011] Collision cost term at time t The expression is:
[0012] ;
[0013] in, For collision indication function, Intensity of environmental risk fields; The safety threshold for environmental risk fields; These are the weighting coefficients;
[0014] Threshold smoothness term at time t The expression is:
[0015] ;
[0016] in, This represents the change in the warning threshold between the current time and the previous time. The current warning threshold is set at this moment. The warning threshold is the threshold set at the previous moment. = - This represents the change in the takeover threshold between the current time and the previous time. The current takeover threshold, The threshold value for takeover at the previous moment;
[0017] Exploratory items The expression is: When the number of accesses to the current state-action pair in the experience memory is less than a preset threshold, let =1, otherwise 0;
[0018] Step 4. Based on the current integrated state space, use the trained policy network to output the warning threshold and takeover threshold.
[0019] As a preferred embodiment of the present invention, step 1 employs a method combining kinetic and potential energy fields to construct the driving risk field, and finally calculates the intensity of the environmental risk field. The vehicle's operating status includes vehicle speed, steering wheel angle, and longitudinal acceleration.
[0020] As a preferred embodiment of the present invention, step 3 introduces [the following] during the training process. A strategy or entropy regularization mechanism ensures that the agent achieves a balance between exploration and exploitation between known optimal actions and unknown potential actions.
[0021] As a preferred embodiment of the present invention, a coupling constraint is set during the dynamic adjustment of the warning threshold and the takeover threshold: the warning threshold Always less than the takeover threshold Smoothing constraint: The rate of change of the warning threshold and the takeover threshold is limited by the maximum adjustment rate; Redundancy constraint: When the risk exceeds the set threshold, if the warning threshold... With takeover threshold If the spacing is lower than the set spacing, the threshold spacing will be forcibly increased.
[0022] As a preferred embodiment of the present invention, the kinetic energy field The expression is:
[0023] ;
[0024] In the formula, Let be the relative speed between the target vehicle j and the vehicle itself. The sign function is used to ensure that a positive risk gain is generated only when the target vehicle approaches the vehicle. The angle between the directions of the vehicle and the target vehicle. , These are model constants. Let be the relative distance between the vehicle and the j-th target vehicle. For the equivalent mass of the vehicle, , The actual physical mass of the vehicle. The current speed of the vehicle. This is the adjustment coefficient.
[0025] As a preferred embodiment of the present invention, the potential energy field The expression is:
[0026] ;
[0027] In the formula, The distance between the vehicle and the static obstacle; These are the perpendicular distances from the center of the vehicle to the dashed lane line and the solid lane line, respectively. For the corresponding risk factors, and set ; This refers to the lateral distance from the vehicle to the physical boundary of the road. As a boundary risk factor, To avoid smoothing constants with a denominator of zero.
[0028] The present invention also provides a multi-threshold cooperative control method for autonomous vehicles, the method comprising the following steps:
[0029] Step A. Based on the current driving environment, vehicle operating status, and driver status, the aforementioned dynamic threshold control method for autonomous vehicles is used to output dynamically changing warning thresholds and takeover thresholds;
[0030] Step B. Based on driver distraction level score Vehicle operating status and environmental risk field intensity Calculate the comprehensive risk index Furthermore, multi-threshold coordinated control is implemented based on the relationship between comprehensive risk indicators and early warning thresholds and takeover thresholds;
[0031] When comprehensive risk indicators Exceeding the warning threshold When a warning is triggered, a comprehensive risk indicator will be used; if a warning is issued, the risk indicators will be adjusted accordingly. Falling back to the warning threshold The following will only involve continuous risk monitoring; if the driver fails to respond promptly after a warning, status monitoring will continue, and risk information will be continuously accumulated. The overall risk indicators will be considered when... Exceeding the takeover threshold If necessary, a mandatory takeover will be immediately triggered, with the autonomous driving system taking over longitudinal and / or lateral control of the vehicle to ensure the vehicle remains in a safe state.
[0032] Step C. After the autonomous driving system takes over the vehicle, it continuously monitors the driver's distraction level score. Vehicle operating status and the intensity of environmental risk fields And calculate the comprehensive risk index. ;when If the temperature remains below the warning threshold for an extended period of time. When setting the ratio, the warning threshold is gradually increased based on the trained policy network. and takeover threshold This restores the value to the range corresponding to normal cruise conditions, thereby achieving dynamic recovery of the threshold; based on this, when the following conditions are met... < When the vehicle is in a stable operating state, the system prompts the driver to take back control of the vehicle via voice or interface, and smoothly switches longitudinal and / or lateral control of the vehicle to the driver after the driver issues a confirmation operation.
[0033] As a preferred embodiment of the present invention, a comprehensive risk index The expression is:
[0034]
[0035] in, , , ≥0 and satisfy , For vehicle operating state vector The calculated vehicle dynamic risk index, function It is used to reduce and aggregate multi-dimensional vehicle states and output a single scalar risk metric.
[0036] The present invention also provides an electronic device, comprising: one or more processors and a memory; wherein the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the above-described multi-threshold cooperative control method for autonomous vehicles.
[0037] The present invention also provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described multi-threshold cooperative control method for autonomous vehicles.
[0038] Advantages and beneficial effects of the present invention:
[0039] (1) Human-machine collaborative control mechanism extended from a single static threshold to a dual dynamic threshold.
[0040] Existing autonomous driving takeover strategies often employ a single, fixed takeover threshold, failing to distinguish between "early warning" and "forced takeover," leading to rigid transitions that either lack warnings for extended periods or abruptly force takeover. This invention introduces two dynamic boundaries—an early warning threshold and a takeover threshold—within the same risk assessment framework. By constraining the minimum distance between these two thresholds and ensuring that the early warning precedes the takeover, a human-machine collaborative control mechanism with a buffer zone is constructed. This allows drivers to prepare for takeover in advance, significantly reducing the abruptness and tension of the takeover process.
[0041] (2) Multi-source information collaborative modeling improves the accuracy of takeover triggering.
[0042] This invention simultaneously utilizes distraction level scores output by the driver monitoring system, vehicle operating status (vehicle speed, acceleration, etc.), and environmental risk field intensity calculated based on driving risk field theory. It maps all of this information to a unified state space, considering both the driver's internal state and vehicle dynamics and external environmental risks when making threshold adjustments and takeover decisions. Compared to takeover strategies based solely on single physical quantities such as fixed safety distance or TTC (Time To Collision), this invention can more precisely characterize the actual takeover risk, reducing false alarms and missed alarms.
[0043] (3) Introduce a driving risk field model to achieve adaptive adjustment of threshold according to the scenario.
[0044] This invention employs a driving risk field model combining kinetic and potential energy fields to quantify surrounding traffic flow, static obstacles, and road structure. The intensity of the environmental risk field directly influences the height of the warning and takeover thresholds, allowing the thresholds to adaptively change with factors such as traffic density, relative speed, and lane changes. Compared to schemes using fixed thresholds or simple rule adjustments, this invention can dynamically expand or contract the safety margin in complex scenarios such as highways, urban roads, and ramp merging, improving the sensitivity and rationality of takeover decisions in response to changing scenarios.
[0045] (4) Design of a multi-objective reward function for safety, comfort and robustness.
[0046] This invention explicitly introduces a collision cost term, a takeover frequency term, an invalid warning term, a threshold smoothness term, and an exploration term into the reward function of reinforcement learning. It jointly constrains multiple indicators, including "whether a collision occurs," "whether takeovers are too frequent," "whether there are a large number of invalid warnings," "whether the threshold fluctuates drastically," and "whether the strategy has sufficient exploration capabilities." The resulting takeover threshold adjustment strategy not only aims to "avoid collisions as much as possible" but also controls unnecessary takeovers and excessive warnings, ensuring that the threshold changes smoothly over time, achieving a comprehensive optimization of safety, comfort, and system robustness.
[0047] (5) Threshold smoothness and rate of change constraints suppress threshold jitter and improve long-term experience.
[0048] Traditional reinforcement learning-based strategies are prone to drastic policy fluctuations in practical deployments, leading to frequent jumps in the takeover threshold, which drivers perceive as "system emotional instability." This invention addresses this by introducing a threshold smoothness penalty term and a maximum adjustment rate. Constraints limit the magnitude of threshold changes between adjacent time points. When the threshold changes too quickly, a negative reward is given, thereby effectively suppressing threshold jitter and ensuring that the threshold changes slowly and continuously over a long period of time, significantly improving the long-term user experience for drivers.
[0049] (6) Supports multi-scenario training and automatic transfer, without relying on repeated manual parameter tuning.
[0050] This invention introduces various typical operating conditions during the training phase, including highways, urban roads, congested sections, and ramp merging. By explicitly considering the intensity of the environmental risk field and scene labels in the reward function and state space, the takeover threshold control strategy obtained through reinforcement learning possesses multi-scene adaptive capabilities. During system deployment, switching between different scenes eliminates the need for manual resetting of different takeover thresholds and warning logic, reducing engineering parameter tuning workload and improving the strategy's generalization ability and maintainability on real roads.
[0051] (7) Use simulation platforms for strategy training to reduce the risks and costs of real road testing.
[0052] This invention preferably employs a co-simulation platform such as Carsim–Prescan–Simulink to construct virtual traffic scenarios. Through reinforcement learning training on a large number of simulation samples, a relatively mature takeover threshold control strategy is obtained, which is then deployed to real vehicles for small-scale verification. Compared to directly relying on repeated threshold adjustments during real-vehicle testing, this invention can significantly reduce collision risks and testing costs in the early development stages, and shorten the system development cycle.
[0053] (8) A lightweight decision-making structure that can be deployed online, which facilitates engineering implementation and continuous upgrades.
[0054] The takeover threshold control structure of this invention adopts an "offline training + online inference" approach. Complex reinforcement learning training and risk field parameter calibration are completed offline. During online operation, it only needs to output adjustment actions for the warning threshold and takeover threshold based on the current state vector through a trained policy network or policy table, resulting in low computational and storage resource consumption on the vehicle computing platform. Furthermore, the policy model of this invention supports online updates after retraining with incremental data, facilitating continuous optimization of the takeover strategy after the vehicle is put into operation.
[0055] (9) The structure is clear and has good interpretability, making it easy to connect with current regulations and safety standards.
[0056] This invention constructs a complete causal chain from "distraction level score – vehicle status – environmental risk field intensity – early warning threshold – takeover threshold – takeover decision". Each module has a clear physical meaning and risk meaning, which makes it easy to explain the system's takeover logic to regulatory authorities and users, and is conducive to subsequent docking and certification with autonomous driving-related regulations and industry standards.
[0057] (10) It has good scalability and can be extended to different levels and brands of autonomous driving systems.
[0058] The state space construction method, dual-threshold collaborative control structure, and threshold adjustment approach based on reinforcement learning of this invention are universal. The state vector and reward weight can be adapted to different autonomous driving levels (such as L2, L3, etc.) and sensor configurations of different car manufacturers, and can be transferred to other autonomous driving platforms, showing good scalability and industrial application prospects. Attached Figure Description
[0059] Figure 1 The flowchart of the dynamic threshold control method for autonomous vehicles provided by the present invention is shown below.
[0060] Figure 2 Flowchart for constructing state-space features for this invention;
[0061] Figure 3 The flowchart of multi-threshold collaborative control for autonomous vehicles provided by the present invention. Detailed Implementation
[0062] To enable those skilled in the art to better understand the technical solutions and advantages of the present invention, the present application will be described in detail below with reference to the accompanying drawings, but this is not intended to limit the scope of protection of the present invention.
[0063] Example 1:
[0064] like Figure 1 , Figure 2 As shown, this embodiment provides a method for dynamic takeover threshold control of autonomous vehicles, which includes the following steps:
[0065] Step 1. Data Acquisition and State Space Construction (see...) Figure 2 ):
[0066] Step 1.1. Driver Distraction Level Assessment Acquisition:
[0067] The system collects driver facial expressions, hand gestures, and interactions with in-vehicle devices via in-vehicle cameras and a driver monitoring system. A multi-model fusion recognition method based on ResNet50, InceptionV3, and Xception is employed, combined with a Convolutional Block Attention (CBAM) module and a Compression and Activation (SE) module, to accurately identify driver distraction behaviors. Subsequently, a risk mapping mechanism is used to convert the identified distraction behaviors into quantified risk scores, forming a driver distraction level rating. , as input to the state space.
[0068] It should be noted that in this embodiment, the driver distraction level score can also be obtained in other ways. This application does not limit the specific method of obtaining the driver distraction level score. Those skilled in the art can refer to any of the methods in the prior art to obtain the driver distraction level score.
[0069] Step 1.2. Environmental Risk Field Intensity Acquisition:
[0070] Based on the theory of driving risk field, traffic environmental risk is decomposed into two dimensions: kinetic energy field and potential energy field. The kinetic energy field is used to characterize the dynamic risk of the vehicle, taking into account vehicle speed, acceleration and vehicle kinematic state; the potential energy field is used to characterize the road environmental risk, taking into account road geometry, obstacle distribution and spatial layout of surrounding traffic participants.
[0071] This invention employs a method combining kinetic and potential energy fields to construct a driving risk field, and finally calculates the intensity of the environmental risk field. It is used to characterize the uncertainty and potential danger of the external environment.
[0072] Specifically, kinetic energy field The expression is:
[0073]
[0074] In the formula, Let be the relative speed between the target vehicle j and the vehicle itself. The sign function is used to ensure that a positive risk gain is generated only when the target vehicle approaches the vehicle. The angle between the directions of the vehicle and the target vehicle. , These are model constants. Let be the relative distance between the vehicle and the j-th target vehicle. The equivalent mass of the vehicle is determined by its real-time speed. The relevant calculation formula is as follows:
[0075]
[0076] In the formula, The actual physical mass of the vehicle. The current speed of the vehicle. This is an adjustment factor. This formula reflects the physical characteristic that the higher the vehicle speed, the greater the inertia, and the more severe the risk impact in a potential collision.
[0077] Potential energy field The expression is:
[0078]
[0079] In the formula, The distance between a vehicle and a static obstacle is represented by an exponential function, which characterizes the characteristic that the risk increases sharply as the distance decreases. These are the perpendicular distances from the center of the vehicle to the dashed lane line and the solid lane line, respectively. For the corresponding risk factors, and set This is to demonstrate that the risk of crossing a solid line is higher than that of crossing a dotted line; This refers to the lateral distance from the vehicle to the physical boundary (curb) of the road. As a boundary risk factor, the strong constraint of the road edge is characterized by an inverse proportional function. To avoid smoothing constants with a denominator of zero.
[0080] In this embodiment, the undetermined constant coefficients in the above formula (such as...) , , , The model can be optimized and calibrated using genetic algorithms based on real driving data or simulation data to ensure that the risk field model can accurately reflect the risk distribution patterns in the real traffic environment.
[0081] Environmental risk field intensity The relationship between the kinetic energy field and the potential energy field is:
[0082]
[0083] Step 1.3. Vehicle Operating Status Acquisition:
[0084] Vehicle operating parameters, including vehicle speed, steering wheel angle, and longitudinal acceleration, are collected in real time via the vehicle's CAN bus and onboard sensors. The collected data is normalized according to their respective physical value ranges to obtain the vehicle's dynamic state input. .
[0085] Step 1.4. Construction of the state space:
[0086] By fusing the above three types of inputs, a comprehensive state space is constructed:
[0087]
[0088] in, This indicates the driver's level of distraction. Indicates the vehicle's operating status. This represents the intensity of the environmental risk field. This state space serves as the input basis for subsequent dynamic threshold adjustment and reinforcement learning training.
[0089] Step 2. Setting dynamic threshold boundaries:
[0090] In this embodiment, two types of dynamic threshold boundaries are set based on preset safety and comfort goals:
[0091] Warning thresholds are used to trigger early warnings when the driver is distracted or when environmental risks are at a moderate level, in order to remind the driver to regain attention;
[0092] Takeover threshold: This threshold is used to trigger mandatory takeover when the driver is distracted or the environmental risk exceeds a high-risk level, in order to ensure vehicle safety.
[0093] In this embodiment, both the warning threshold and the takeover threshold are dynamic thresholds, meaning that they are dynamically adjusted based on the driver's status and environmental risks. The specific rules are as follows:
[0094] (1) Threshold adjustment based on driver status:
[0095] When the driver's distraction level is rated An increase in elevation indicates that the driver's attention is distracted or their cognitive load is increased. In this case: decrease. Early warnings are triggered to ensure drivers have sufficient reaction time; reducing... To expedite the triggering of takeover, so as to avoid the driver being unable to regain control in time;
[0096] When the driver's distraction level is rated As the vehicle descends, it indicates that the driver has regained focus or improved reaction time. At this point: gradually increase... Reduce false alarms or unnecessary warnings; gradually improve This reduces excessive control and improves driving comfort.
[0097] (2) Threshold adjustment based on environmental risk:
[0098] When the intensity of the environmental risk field An increase in elevation indicates increased complexity of the road or traffic situation and a greater potential danger. In this case: Decrease. This enables earlier risk warnings and reduces... This ensures the system can take over promptly in high-risk environments;
[0099] when When the value decreases, it indicates that the road environment is safer. At this time: increase To avoid frequent and ineffective warnings; to improve This reduces unnecessary takeover operations.
[0100] In addition, to prevent drastic fluctuations in the dynamic threshold within a short period of time, a smoothness constraint is set for threshold adjustment. By limiting the rate of change of the threshold, the continuity and predictability of the warning and takeover process are ensured, thereby improving the driver's trust in and acceptability of the system's behavior.
[0101] Step 3. Reward Function Design:
[0102] Step 3.1. Overall Construction of the Reward Function:
[0103] To achieve a comprehensive optimization of safety, comfort, and system robustness, the reward function designed in this invention is as follows:
[0104]
[0105] in, , , , , These are the weighting coefficients for different evaluation indicators. For collision cost terms, To take over the frequency term, This is an invalid warning item. For threshold smoothness, This is an exploratory item.
[0106] In this embodiment, the collision cost term is addressed. When a vehicle poses a potential collision risk due to an unreasonable threshold setting, the system imposes a significant penalty. Weighting coefficient. Set to the maximum value to ensure that the risk of collision is minimized.
[0107] Specifically, collision cost item Environmental risk field intensity The duration exceeding the safety threshold, combined with the collision probability, can be defined as follows:
[0108]
[0109] In the formula, For collision indication function, The intensity of the environmental risk field is calculated based on the driving risk field model; The safety threshold for environmental risk fields; This is the weighting coefficient, with a value ranging from 0.5 to 5;
[0110] In this embodiment, the collision indication function Let be a binary variable, either 0 or 1. At discrete time step t, when a collision event is detected between the vehicle and any target traffic unit (target vehicle), let... In all time steps where no collision occurs, . Take the larger value. This represents a significant negative contribution to the total reward; even if no collision occurs, but when consistently higher At the same time, the second term will continue to increase, thereby achieving the effect of "combining the duration of the environmental risk field intensity exceeding the safety threshold with the collision probability", so that the agent actively stays away from high-risk areas during the training process.
[0111] In this embodiment, the takeover frequency item Used to suppress excessive takeover behavior, minimizing unnecessary takeovers while ensuring system safety. Weighting coefficient A reasonable range for adjusting the take-off frequency.
[0112] Specifically, the frequency of takeover It can be defined at the time step level:
[0113]
[0114] in, This is the takeover indicator function, which takes the value 1 if manual takeover or system-mandated takeover is triggered at the current time, and 0 otherwise. Thus, in the reward for this time step, That is, a fixed penalty is imposed on each takeover. The more takeovers occur (the more frequently takeovers are triggered), the greater the cumulative penalty (the reward function is significantly reduced), so as to encourage the agent to learn to reduce unnecessary takeovers.
[0115] Frequent warnings that do not lead to a risk event can undermine the driver's trust in the system; therefore, an invalid warning term is introduced into the reward function. And set penalties for the number of invalid warnings:
[0116]
[0117] Specifically, when the system issues an early warning at time t, the intensity of the environmental risk field at that time... Still below the safety threshold At that time, it was determined to be an invalid warning, and ordered Otherwise, it is 0.
[0118] Therefore, when an invalid warning occurs This will reduce rewards, encouraging strategies to trigger alerts only when there is a truly high risk. Weighting coefficient. Constraints on the effectiveness of corresponding early warnings.
[0119] In this embodiment, to avoid drastic fluctuations in the threshold within a short period of time, a threshold smoothness term is introduced into the reward function. The expression is:
[0120]
[0121] in, This represents the change in the warning threshold between the current time and the previous time. The current warning threshold is set at this moment. The warning threshold is the threshold set at the previous moment. = - This represents the change in the takeover threshold between the current time and the previous time. The current takeover threshold, The threshold value for takeover at the previous moment.
[0122] The greater the threshold change, The larger, The smaller, thus It has a stronger negative impact on the total reward, achieving the goal of drastically changing the penalty threshold.
[0123] Weighting coefficient This is used to balance the threshold adjustment rate. The reward value increases when the threshold adjustment remains stable within the allowed rate range.
[0124] In this embodiment, to maintain sufficient exploratory activity of the reinforcement learning agent during training and avoid getting trapped in local optima, an exploratory reward is introduced into the reward function. The expression is:
[0125]
[0126] Among them, when the current state-action pair When the number of accesses to the experience memory is less than a preset threshold, let =1, otherwise 0; thus, when the agent accesses a state-action pair that has not yet been fully explored, This will positively contribute to the total reward, encouraging the exploration of new threshold adjustment strategies under safe conditions, and improving the model's generalization ability across multiple scenarios. Weight coefficients Maintaining a balance between exploration and utilization.
[0127] Step 4. Reinforcement learning-based policy training:
[0128] Step 4.1. Agent Modeling:
[0129] This invention models the adjustment process of warning thresholds and takeover thresholds as a Markov decision process. It takes a high-dimensional state vector composed of "driver distraction level score - vehicle operating state - environmental risk field intensity" as input, and the action space is composed of threshold adjustment operations such as raising, lowering or keeping the warning thresholds and takeover thresholds unchanged. It constructs a comprehensive reward function that includes multiple indicators such as collision cost, takeover frequency, invalid warning, threshold smoothness and exploratory nature. It uses reinforcement learning to train and optimize the threshold adjustment strategy, thereby ensuring driving safety while taking into account driving comfort and system robustness, and realizing dynamic optimization and multi-scenario adaptation of the takeover triggering mechanism.
[0130] Specifically, state space Driver distraction level rating Vehicle operating status and environmental risk field intensity constitute.
[0131] Action space Including warning thresholds and takeover threshold The dynamic adjustment operation specifically involves raising, lowering, or maintaining the threshold.
[0132] reward function The comprehensive reward function defined in step 3 is used to evaluate the results of each threshold adjustment.
[0133] Step 4.2. Interaction Process:
[0134] A virtual traffic scenario (introducing various typical operating conditions such as highways, urban roads, congested sections, and ramp merging) is constructed using a joint simulation platform including Carsim–Prescan–Simulink. The agent continuously interacts with the environment in this virtual traffic environment: at time t, the agent adjusts its state based on its current state. Select Action That is, adjusting the threshold and The environment triggers corresponding warnings or takeovers based on the new thresholds and reports the new status. The reward function calculates the reward corresponding to this action. The agent updates its strategy accordingly.
[0135] Step 4.3. Strategy Update:
[0136] Training is performed using reinforcement learning methods:
[0137] Policy Network: Input State Space Output the probability of the action or the Q value;
[0138] Update mechanism: Based on the temporal difference learning method, utilizing state transition The strategy was iteratively optimized;
[0139] Convergence objective: Maximize the cumulative expected reward to obtain the optimal threshold adjustment strategy.
[0140] Specifically, the reinforcement learning model employs a value function-based reinforcement learning method based on Deep Q Network (DQN). It constructs a model using the current state vector... As input, output the action value corresponding to the candidate action. The Q-network is used, and its parameters are updated based on the time difference learning approach to maximize the cumulative expected reward. The Q-value update can be expressed as:
[0141]
[0142] in, The learning rate is preferably set within the range of 0.0001 to 0.01. The discount factor is preferably in the range of 0.90 to 0.99. In the state Next action Instant rewards received To maximize the expected return in the next state. This represents the next state. Below, candidate actions are selected from all possible actions. During training, mechanisms such as experience replay and target network can be combined to improve training stability and convergence performance.
[0143] Step 4.4. Exploring and Utilizing Balance:
[0144] Introduced during training The strategy or entropy regularization mechanism ensures that the agent achieves a balance between exploration and utilization between known optimal actions and unknown potential actions, avoiding getting trapped in local optima.
[0145] Step 4.5. Training Termination Conditions:
[0146] The training process terminates when one of the following conditions is met:
[0147] The cumulative rewards have reached the preset threshold;
[0148] The threshold adjustment performed stably across multiple scenarios, with a significant reduction in the number of takeover attempts and collision risks.
[0149] The policy network parameters converged, and the update magnitude was lower than the preset threshold.
[0150] Step 5. Based on the current integrated state vector, use the trained policy network to output the warning threshold and takeover threshold.
[0151] Furthermore, in this embodiment, to ensure system stability, the present invention sets the following constraints during the dynamic adjustment of the threshold:
[0152] Coupling constraints: Always less than This is to ensure that early warnings occur before takeover, thus avoiding system logic conflicts;
[0153] Smoothing constraint: The rate of change of the threshold is limited by the maximum adjustment rate. This is to prevent the threshold from fluctuating significantly in a short period of time;
[0154] Redundancy constraints: in comprehensive risk indicators Approaching or exceeding At that time, if - If so, then the threshold interval will be forcibly increased. The preset minimum threshold interval is used to constrain the warning threshold. With takeover threshold The minimum distance between them is determined to ensure that a sufficient reaction buffer is always maintained between them.
[0155] Example 2:
[0156] like Figure 3 As shown in the figure, this embodiment also provides a multi-threshold cooperative control method for autonomous vehicles, which includes the following steps:
[0157] Step A. Based on the current driving environment, vehicle operating status, and driver status, output dynamically changing warning thresholds and takeover thresholds using the dynamic threshold control method for autonomous vehicles described in Example 1;
[0158] Step B. Based on driver distraction level score Vehicle operating status and the intensity of environmental risk fields Calculate the comprehensive risk index Furthermore, multi-threshold coordinated control is implemented based on the relationship between comprehensive risk indicators and early warning thresholds and takeover thresholds;
[0159] When comprehensive risk indicators Approaching or exceeding the warning threshold When the system triggers a warning, including audible, visual, or tactile feedback, it reminds the driver to regain attention and adjust driving actions if necessary. If, after the warning, a comprehensive risk assessment is conducted... Falling back to the warning threshold In the following scenarios, the system will only continuously monitor risks; if the driver does not respond promptly after a warning, status monitoring will continue, and risk information will be continuously accumulated. When the overall risk indicators are considered... Exceeding the takeover threshold When this occurs, the system immediately triggers a mandatory takeover, with the autonomous driving system taking over longitudinal and / or lateral control of the vehicle. The takeover actions include automatic braking, steering correction, or speed adjustment to ensure the vehicle remains in a safe state.
[0160] Specifically, in this embodiment, the comprehensive risk index is constructed in the following manner. The expression is:
[0161]
[0162] in, , , ≥0 and satisfy , For vehicle operating state vector The calculated vehicle dynamic risk index, function It is used to reduce and aggregate multi-dimensional vehicle states and output a single scalar risk metric.
[0163] In this embodiment, the system records the current state data at the same time as the takeover is triggered, providing feedback samples for subsequent reinforcement learning training and parameter optimization.
[0164] Step C. After the autonomous driving system takes over the vehicle, it continuously monitors the driver's distraction level score. Vehicle operating status and the intensity of environmental risk fields And calculate the comprehensive risk index. ;when If the temperature remains below the warning threshold for an extended period of time. A certain proportion (e.g., 0.6%) When this happens, the warning threshold is gradually increased based on the trained policy network. and takeover threshold This restores the value to the range corresponding to normal cruise conditions, thereby achieving dynamic recovery of the threshold. Based on this, when the following conditions are met... < When the vehicle is in stable condition, the system will prompt the driver via voice or interface that he / she can take back control of the vehicle. After the driver issues a confirmation, the longitudinal and / or lateral control of the vehicle will be smoothly transferred to the driver.
[0165] In this embodiment, the entire takeover process adopts a gradual adjustment to avoid sudden changes in the threshold and ensure a smooth transition in human-machine collaboration.
[0166] The present invention also provides an electronic device, comprising: one or more processors and a memory; wherein the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the above-described multi-threshold cooperative control method for autonomous vehicles.
[0167] The present invention also provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described multi-threshold cooperative control method for autonomous vehicles.
[0168] Those skilled in the art will understand that all or part of the functions of the various methods / modules in the above embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the above functions can be implemented by executing the program with a computer. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be implemented.
[0169] In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the programs can also be stored in storage media such as servers, other computers, disks, optical discs, flash drives, or portable hard drives. They can be downloaded or copied to the memory of the local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be implemented.
[0170] The above-described specific examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art can make various simple deductions, modifications, or substitutions based on the principles of this invention. Therefore, the scope of protection of this invention should be determined by the scope of the claims.
Claims
1. A dynamic threshold control method for autonomous vehicles, characterized in that, The method includes the following steps: Step 1. Collect vehicle operation data in real time and construct a comprehensive state space. ;in, This indicates the driver's level of distraction. Indicates the vehicle's operating status. Indicates the intensity of the environmental risk field; Step 2. Set two types of dynamic threshold boundaries: warning threshold and takeover threshold. Step 3. Train and optimize the dynamic threshold adjustment strategy based on reinforcement learning methods, using the comprehensive state space as input, the warning threshold and takeover threshold as the action space, and setting the reward function: ; in, , , , , These are the weighting coefficients for different evaluation indicators. For collision cost terms, To take over the frequency term, This is an invalid warning item. For threshold smoothness, This is an exploratory item; Collision cost term at time t The expression is: ; in, For collision indication function, Intensity of environmental risk fields; The safety threshold for environmental risk fields; These are the weighting coefficients; Threshold smoothness term at time t The expression is: ; in, This represents the change in the warning threshold between the current time and the previous time. The current warning threshold is set as follows. The warning threshold is the threshold set at the previous moment. This represents the change in the takeover threshold between the current time and the previous time. The current takeover threshold, The threshold value for takeover at the previous moment; Exploratory items The expression is: When the number of accesses to the current state-action pair in the experience memory is less than a preset threshold, let =1, otherwise 0; Step 4. Based on the current integrated state space, use the trained policy network to output the warning threshold and takeover threshold.
2. The dynamic threshold control method for autonomous vehicles according to claim 1, characterized in that, In step 1, a driving risk field is constructed by combining kinetic and potential energy fields, and the intensity of the environmental risk field is finally calculated. The vehicle's operating status includes vehicle speed, steering wheel angle, and longitudinal acceleration.
3. The dynamic threshold control method for autonomous vehicles according to claim 1, characterized in that, Step 3 introduces during training A strategy or entropy regularization mechanism ensures that the agent achieves a balance between exploration and exploitation between known optimal actions and unknown potential actions.
4. The dynamic threshold control method for autonomous vehicles according to claim 1, characterized in that, Coupled constraints are set during the dynamic adjustment of the early warning threshold and the takeover threshold: Early warning threshold Always less than the takeover threshold Smoothing constraint: The rate of change of the warning threshold and the takeover threshold is limited by the maximum adjustment rate; Redundancy constraint: When the comprehensive risk index exceeds the set threshold, if the warning threshold... With takeover threshold If the interval is lower than the set interval, the threshold interval will be forcibly increased.
5. The dynamic threshold control method for autonomous vehicles according to claim 2, characterized in that, Kinetic field The expression is: ; In the formula, Let be the relative speed between the target vehicle j and the vehicle itself. The sign function is used to ensure that a positive risk gain is generated only when the target vehicle approaches the vehicle. The angle between the directions of the vehicle and the target vehicle. , These are model constants. Let be the relative distance between the vehicle and the j-th target vehicle. For the equivalent mass of the vehicle, , The actual physical mass of the vehicle. The current speed of the vehicle. This is the adjustment coefficient.
6. The dynamic threshold control method for autonomous vehicles according to claim 2, characterized in that, Potential energy field The expression is: ; In the formula, The distance between the vehicle and the static obstacle; These are the perpendicular distances from the center of the vehicle to the dashed lane line and the solid lane line, respectively. For the corresponding risk factors, and set ; This refers to the lateral distance from the vehicle to the physical boundary of the road. As a boundary risk factor, To avoid smoothing constants with a denominator of zero.
7. A multi-threshold cooperative control method for autonomous vehicles, characterized in that, The method includes the following steps: Step A. Based on the current driving environment, vehicle operating status, and driver status, the dynamic threshold control method for autonomous vehicles described in any one of claims 1 to 6 is used to output dynamically changing warning thresholds and takeover thresholds; Step B. Based on driver distraction level score Vehicle operating status and environmental risk field intensity Calculate the comprehensive risk index Furthermore, multi-threshold coordinated control is implemented based on the relationship between comprehensive risk indicators and early warning thresholds and takeover thresholds; When comprehensive risk indicators Exceeding the warning threshold When a warning is triggered, a comprehensive risk indicator will be used; if a warning is issued, the risk indicators will be adjusted accordingly. Falling back to the warning threshold The following will only involve continuous risk monitoring; if the driver fails to respond promptly after a warning, status monitoring will continue, and risk information will be continuously accumulated. The overall risk indicators will be considered when... Exceeding the takeover threshold If necessary, a mandatory takeover will be immediately triggered, with the autonomous driving system taking over longitudinal and / or lateral control of the vehicle to ensure the vehicle remains in a safe state. Step C. After the autonomous driving system takes over the vehicle, it continuously monitors the driver's distraction level score. Vehicle operating status and environmental risk field intensity And calculate the comprehensive risk index. ;when If the temperature remains below the warning threshold for an extended period of time. When setting the ratio, the warning threshold is gradually increased based on the trained policy network. and takeover threshold This restores the value to the range corresponding to normal cruise conditions, thereby achieving dynamic recovery of the threshold; when the condition is met... < When the vehicle is in a stable operating state, the system prompts the driver to take back control of the vehicle via voice or interface, and smoothly switches longitudinal and / or lateral control of the vehicle to the driver after the driver issues a confirmation operation.
8. The multi-threshold cooperative control method for autonomous vehicles according to claim 7, characterized in that, Comprehensive risk indicators The expression is: ; in, , , ≥0 and satisfy , For vehicle operating state vector The calculated vehicle dynamic risk index, function It is used to reduce and aggregate multi-dimensional vehicle states and output a single scalar risk metric.
Citation Information
Patent Citations
Driving takeover condition identification method and system based on multivariate situation awareness
CN121062759A
Lane departure early warning method and system and vehicle
CN121084418A