Intelligent vehicle human-machine shared collision avoidance control system and method under critical working condition
By combining the sensing system with offline learning, the driver's collision avoidance intention and ability are quantified, control weights are dynamically allocated, and machine operations are generated. This solves the human-machine conflict and safety issues in traditional human-machine shared collision avoidance technology, and achieves safe collision avoidance control under critical conditions.
Patent Information
- Application Number
- CN202510146478.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-02-10
AI Technical Summary
In critical operating conditions, traditional human-machine shared collision avoidance technology may cause driver resistance or misoperation. Existing methods lack theoretically based machine intervention standards, leading to human-machine conflicts and reduced safety.
The vehicle status and environmental information are collected through the sensing system, and the set of unreachable collision avoidance states is approximately solved by combining offline learning and large-scale data. The driver's collision avoidance intention and ability are quantified, the human-machine control weight is dynamically allocated, and machine operations are generated based on the reachability-inspired reinforcement learning algorithm to ensure the theoretical and safety of collision avoidance decisions.
It has achieved theoretical assessment of collision risks under critical conditions, reduced human-machine conflicts, improved vehicle driving safety and stability, ensured the driver's sense of control, and enhanced emergency response capabilities.
Smart Images

Figure CN119928843B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of active automobile safety, and in particular to a human-machine shared collision avoidance control system and method for intelligent vehicles under critical operating conditions. Background Art
[0002] In high-risk driving scenarios, complex operational requirements, tight reaction times, and reduced decision-making ability caused by cognitive overload make timely collision avoidance a major challenge. In such situations, intelligent vehicles need to be able to assist drivers in implementing quick and precise evasive maneuvers to address potential safety risks.
[0003] Traditional human-machine shared collision avoidance technologies typically generate collision avoidance trajectories through path replanning and trajectory tracking, and have proven to be highly effective in ensuring vehicle safety. The underlying principle is to achieve higher-precision obstacle avoidance control through collaborative operation between the intelligent system and the driver. However, these approaches often overlook the driver's real-time intentions in critical situations. When the vehicle needs to perform rapid or extreme maneuvers, forcing the system-generated trajectory to adhere to can trigger driver resistance or misoperation, increasing the risk of new collisions.
[0004] To reduce human-machine conflicts, driver-centric shared control methods have gradually developed. These methods introduce the concept of dynamic safety range and model predictive control technology to provide necessary intervention while respecting the driver's intentions. However, treating environmental constraints as soft constraints may reduce overall safety in some scenarios. In addition, collision avoidance control methods based on reinforcement learning have attracted attention in recent years. These methods avoid the limitations of traditional path planning by directly generating collision avoidance actions and can achieve robust collision avoidance control in complex dynamic environments. These methods still rely mainly on distance-based risk assessment and intervene near manually set thresholds. Within this framework, the theoretical basis for effective system intervention has not yet been established, that is, it is impossible to determine when machine intervention is truly necessary.
[0005] Although significant progress has been made in the field of human-robot shared control collision avoidance, including trajectory tracking-based frameworks, human-centered flexible constraint methods, and recent non-explicit collision avoidance path strategies based on reinforcement learning, these technical solutions still face several challenges. Traditional trajectory tracking strategies typically require the driver to follow a preset path, which can cause human-robot conflicts in dynamic or emergency operation scenarios and reduce driver acceptance. Attempts to improve cooperation by softening obstacle constraints may lead to reduced safety, and risk assessment methods based on distance or artificial potential fields (APF) lack rigorous theoretical guarantees regarding the necessity of intervention. Furthermore, current research lacks consideration of the driver's own collision avoidance intentions.
[0006] In response to the above problems, there is an urgent need to establish a human-machine co-driving control method and system in collision avoidance scenarios with theoretical machine intervention triggering standards and minimizing human-machine conflicts. Summary of the Invention
[0007] The present application aims to solve one of the technical problems in the related art at least to a certain extent.
[0008] To this end, the first purpose of this application is to propose a human-machine shared collision avoidance control method for intelligent vehicles under critical working conditions, so as to theoretically evaluate the collision risk of the vehicle status and the collision avoidance intention of the driver's actions, and minimize human-machine conflicts while ensuring that no collision occurs.
[0009] The second purpose of this application is to propose an intelligent vehicle human-machine shared collision avoidance control system under critical working conditions.
[0010] The third objective of this application is to provide an electronic device.
[0011] The fourth object of this application is to provide a computer-readable storage medium.
[0012] A fifth object of this application is to provide a computer program product.
[0013] To achieve the above objectives, the first embodiment of the present application proposes a human-machine shared collision avoidance control method for an intelligent vehicle under critical conditions, comprising:
[0014] Collect and transmit information about the vehicle's current state and surrounding environment through a sensor system, which includes a camera, millimeter-wave radar, an integrated inertial navigation unit (IMU), and related networking facilities;
[0015] The set of unreachable states for collision avoidance is approximately solved using offline learning and large-scale vehicle data. The Hamilton-Jacobi reachability value function and action-value function are updated offline iteratively using a large amount of pre-collected real-vehicle collision avoidance data to obtain the reachability value function and action-value function.
[0016] quantifying the driver's collision avoidance capability and collision avoidance intention based on the accessibility value function and the action-value function, and dynamically allocating human-machine control weights according to the quantified collision avoidance capability and collision avoidance intention;
[0017] Generate machine actions based on a reachability-inspired reinforcement learning algorithm that incorporates the driver's collision avoidance intention, collision avoidance capability, and reachability information into the state space;
[0018] The final operation to be executed is generated by combining the human-machine control weight, the driver operation and the machine operation, and the final operation to be executed is sent to each actuator of the vehicle in real time through the communication system to execute the collision avoidance operation.
[0019] Optionally, use offline learning and large-scale vehicle data to approximate the set of unreachable states for collision avoidance. Then, perform offline iterative updates on the Hamilton-Jacobi reachability value function and action-value function using a large amount of pre-collected real-vehicle collision avoidance data to obtain the reachability value function and action-value function, including:
[0020] The obstacle is represented as an elliptical envelope or a T-shaped collision surface, where the parameters of the target obstacle are recorded as [X0, Y0, a, b], which represent the center coordinates of the obstacle and the elliptical shape parameters respectively;
[0021] The target obstacle is represented as a safety constraint state set Y. The vehicle state in this set is considered to be a dangerous state where a collision has occurred, which is defined as:
[0022]
[0023] in, Represents the state and kinematic information of the vehicle, X, Y are the global position coordinates of the vehicle, is the yaw angle, v x ,v y are the longitudinal velocity and the lateral velocity, respectively, and r represents the yaw rate;
[0024] Substituting the vehicle state and driver operation in a large amount of pre-collected real-vehicle collision avoidance data into the reachability value function network and action value function network corresponding to the safety constraint state set T, and calculating the sample reachability value function value and sample action value function value;
[0025] The sample reachability value function and sample action value are updated offline iteratively to further optimize the accuracy of collision avoidance decision. After the iteration, the updated reachability value function V is obtained. h Network and action-value function Q h network and deploy it in smart vehicle advanced driver assistance systems.
[0026] Optionally, based on the accessibility value function and the action-value function, quantifying the driver's collision avoidance capability and collision avoidance intention, and dynamically allocating human-machine control weights according to the quantified collision avoidance capability and collision avoidance intention, including:
[0027] Substitute the current vehicle state, obstacle envelope and driver operation collected by the sensor system into the updated accessibility value function V corresponding to the constraint state set T h Network and action-value function Q h Network, get the reachability value V h (x) and action-value Q h (x,u d );
[0028] Through the accessibility value V h (x) represents the optimal collision avoidance distance in the current state and calculates the driver's collision avoidance ability CAA, which is calculated using the following formula:
[0029]
[0030] Among them, α CAA is the sensitivity adjustment parameter, C CAA is the offset constant;
[0031] By the driver's current operation u d The corresponding action-value Q h (x,u d ) The reachability value V corresponding to the best action h (x), calculate the driver's collision avoidance intention CAI, which is calculated using the following formula:
[0032]
[0033] According to the collision avoidance capability CAA and the collision avoidance intention CAI, the human-machine control weight γ is dynamically allocated. The control weight is calculated by the following formula:
[0034] γ=max(γ min ,(1-s CAI )(1-s CAA ))
[0035] Among them, s CAI With s CAA is the normalization function, which is defined by the following formulas:
[0036]
[0037] Among them, k CAI1 and k CAA1 Control s separately CAI and s CAA Sensitivity to changes in CAI and CAA, k CAI2 and k CAA2 Definitions CAI and s CAA Reaching a threshold of 0.5 represents the balance point between collision avoidance intention and capability.
[0038] Optionally, machine actions are generated based on a reachability-inspired reinforcement learning algorithm, where the reinforcement learning algorithm incorporates the driver's collision avoidance intention, collision avoidance capability, and reachability information into the state space, including:
[0039] The driver's collision avoidance intention, collision avoidance ability, accessibility information and other vehicle dynamic states are explicitly integrated into the state space of reinforcement learning, which includes: the vehicle's dynamic state x, the accessibility value function V h (x), the driver's collision avoidance ability CAA, the driver's collision avoidance intention CAI, the human-machine control weight γ, and the control operation u of the driver and the machine d and u m , where the action space u m Contains only the front wheel steering angle δ mf ;
[0040] In the state space, combining the dynamic and kinematic states of the vehicle, a driver's collision avoidance maneuvers during training are simulated by constructing a driver action generation model. The driver action generation model simulates the driver's collision avoidance decisions by dynamically adjusting the relative relationship between the driver's input and the obstacle.
[0041] A reinforcement learning reward function is constructed based on the driver's collision avoidance intention. The reward function is designed with the goal of reducing human-machine conflict and maintaining the original task performance. The reward function uses the obstacle envelope as a constraint to guide the generation of machine control actions during the reinforcement learning process.
[0042] During the reinforcement learning training process, by maximizing the above reward function, the machine prioritizes safety while reducing human-machine conflicts and generates the final control action.
[0043] Optionally, in the state space, a driver action generation model is constructed in combination with the vehicle's dynamic and kinematic states to simulate the driver's collision avoidance operations during training. The driver action generation model dynamically adjusts the relative relationship between the driver's input and the obstacle to simulate the driver's collision avoidance decision, including:
[0044] The current vehicle state, obstacle envelope, and driver’s control input are substituted into the driver action generation model. By calculating the relative position and speed between the vehicle and the obstacle, the driver’s steering operation is calculated and adjusted so that the driver’s preview direction is aligned with the tangent of the obstacle boundary, thereby ensuring that the vehicle can avoid the obstacle. The preview angle θ c Calculated according to the following formula:
[0045]
[0046] Among them, (X trg ,Y trg ) represents the tangent point on the ellipse, which lies on the boundary of the ellipse defined as follows:
[0047]
[0048] Where λ is a scaling factor that is proportional to the accessibility value V h (x) and the collision avoidance intention CAI adjustment ellipse size, defined as:
[0049]
[0050] Among them, k λ is a hyperparameter.
[0051] Optionally, a reinforcement learning reward function is constructed based on the driver's collision avoidance intention. The reward function is designed based on the goal of reducing human-machine conflict and maintaining original task performance. The reward function uses the obstacle envelope as a constraint to guide the generation of machine control actions during the reinforcement learning process, including:
[0052] The cooperation between the driver and the machine is optimized by reinforcement learning reward function to reduce human-machine conflict. The reward function includes a safety reward R sf and collaboration reward R co ;
[0053] The safety reward R sf It is calculated based on the distance between the vehicle state and the set of unreachable states for collision avoidance, and is defined as:
[0054]
[0055] Among them, d0 is the scaling parameter, k sf1 To emphasize the negative weighting factor of the importance of being far away from the boundary of intelligent vehicle advanced driver assistance system, V h (x)>k sf2 States with high penalties result in significant penalties to ensure that the agent learns to avoid unsafe states;
[0056] The collaboration reward R co Optimize human-machine collaboration by quantifying the deviation between machine and driver behavior, defined as:
[0057] R co =-k co ·γ·(u M -u D ) 2
[0058] Among them, the penalty term (u M -u D ) 2 is the deviation between the machine and the driver’s behavior, k co is a scaling factor, and the weighting factor γ reflects the necessity of machine intervention based on the driver’s collision avoidance ability and intention.
[0059] Optionally, the combining of the human-machine control weight, the driver operation, and the machine operation to generate a final operation to be executed, and sending the final operation to be executed to each actuator of the vehicle in real time via a communication system to execute a collision avoidance operation, including:
[0060] According to the human-machine control weight γ, machine operation u m and driver operation d , generate the final operation to be executed u f , the formula is:
[0061] u f =γ·u m +(1-γ)·u d
[0062] The final operation to be performed u is transmitted through the communication system f Send real-time information to various vehicle actuators, including the active differential steering system, drive system, and braking system, to perform corresponding collision avoidance maneuvers;
[0063] During the execution process, the vehicle's status is monitored in real time through continuous sensor feedback, and the final operation is dynamically adjusted according to the relative position of the vehicle and the obstacle to ensure that the vehicle successfully avoids the obstacle and returns to its normal driving trajectory.
[0064] To achieve the above objectives, the second embodiment of the present application proposes an intelligent vehicle human-machine shared collision avoidance control system under critical working conditions, including:
[0065] A data acquisition module, which is used to collect and transmit information about the vehicle's current state and surrounding environment through a sensor system, including a camera, millimeter-wave radar, an integrated inertial navigation unit (IMU), and related networking facilities;
[0066] The reachability analysis module uses offline learning and large-scale vehicle data to approximate the set of unreachable states for collision avoidance. It also uses a large amount of pre-collected real-vehicle collision avoidance data to iteratively update the Hamilton-Jacobi reachability value function and action-value function offline to obtain the reachability value function and action-value function.
[0067] a driver intention and capability evaluation module, configured to quantify the driver's collision avoidance capability and intention based on the accessibility value function and the action-value function, and dynamically assign human-machine control weights based on the quantified collision avoidance capability and intention;
[0068] a reinforcement learning control module for generating machine actions based on a reachability-inspired reinforcement learning algorithm that incorporates the driver's collision avoidance intent, collision avoidance capability, and reachability information into a state space;
[0069] The control decision module is used to combine the human-machine control weights, driver operations and machine operations to generate a final operation to be executed, and send the final operation to be executed to each actuator of the vehicle in real time through the communication system to perform the collision avoidance operation.
[0070] To achieve the above-mentioned purpose, a third embodiment of the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;
[0071] The memory stores computer-executable instructions;
[0072] The processor executes the computer-executable instructions stored in the memory to implement the method as described in any one of the first aspects above.
[0073] To achieve the above-mentioned purpose, the fourth embodiment of the present application proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method as described in any one of the above-mentioned first aspects.
[0074] To achieve the above-mentioned purpose, the fifth embodiment of the present application proposes a computer program product, including a computer program, which, when executed by a processor, implements the method as described in any one of the above-mentioned first aspects.
[0075] The technical solutions provided by the embodiments of this application bring at least the following beneficial effects:
[0076] (1) The Hamilton-Jacobi reachability method is used to theoretically evaluate collision risk. By calculating the reachability value, the system can make timely intervention decisions when danger is approaching. This method enables the machine to determine whether intervention is necessary based on the real-time collision risk, effectively avoiding collisions and reducing human-machine conflicts, ensuring the safety of vehicles in complex environments.
[0077] (2) The system uses a human-machine shared control system that can precisely control the vehicle's trajectory and achieve smoother yaw angle adjustment. During obstacle avoidance, the system's real-time adjustment capabilities can effectively respond to sudden obstacles or complex working conditions, thereby improving the vehicle's stability and safety, especially significantly improving its response capabilities in emergency situations. Whether in emergency situations or during daily driving, the vehicle can maintain a higher level of safety.
[0078] (3) By dynamically quantifying the driver's collision avoidance capabilities and intentions, and combining the driver's real-time actions with the system's calculations, the system intelligently allocates human-machine control weights. At critical moments, the system effectively takes over vehicle control while respecting the driver's actions at other times, maintaining the driver's sense of control. This intelligent mechanism not only optimizes the driving experience but also balances human-machine collaboration in complex driving environments, improving driving safety and comfort.
[0079] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0081] Figure 1 A flowchart of a human-machine shared collision avoidance control method for an intelligent vehicle under critical conditions provided by an embodiment of the present application;
[0082] Figure 2 A schematic diagram of the architecture of a human-machine shared collision avoidance control method for an intelligent vehicle under critical conditions provided by an embodiment of the present application;
[0083] Figure 3 A schematic diagram illustrating the impact of the driver's collision avoidance ability and collision avoidance intention on the machine control weight provided in an embodiment of the present application;
[0084] Figure 4 A comparison diagram of the deployment effects of the present application and other existing technical solutions provided in the embodiments of the present application;
[0085] Figure 5 This is a schematic structural diagram of an intelligent vehicle human-machine shared collision avoidance control system under critical conditions provided by an embodiment of the present application. DETAILED DESCRIPTION
[0086] The following describes in detail embodiments of the present application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0087] To address the shortcomings of existing technologies, the present invention provides a method for shared collision avoidance control between humans and machines in intelligent vehicles under critical operating conditions. This method utilizes large-scale data to approximate the set of unreachable collision avoidance states. The proximity of the vehicle's state to the set of unreachable collision avoidance states serves as a theoretical basis for assessing collision risk, thereby determining the necessity of machine intervention. This method can theoretically assess the collision risk of the vehicle's state and the driver's collision avoidance intent, minimizing human-machine conflict while ensuring that no collision occurs.
[0088] Reference Figure 1 and Figure 2 , the method comprises the following steps:
[0089] Step S1: Collect and transmit the vehicle's current status and surrounding environment information through the sensing system, which includes a camera, millimeter-wave radar, integrated inertial navigation unit (IMU), and related networking facilities.
[0090] In the embodiment of the present application, this step is performed by Figure 2 Module 10 is executed. Figure 2 The sensing system includes cameras, millimeter-wave radar, integrated inertial navigation unit (IMU), and related networking equipment. These sensors can acquire and transmit dynamic state data such as the vehicle's position, speed, acceleration, and the relative position and size of obstacles in the surrounding environment in real time. This data is crucial for subsequent collision avoidance decisions.
[0091] Considering that a large amount of historical collision avoidance data is required when solving the collision avoidance unreachable state set, the sensing system should have strong data integration, collection, and storage capabilities to ensure that the necessary collision avoidance data can be stored and quickly retrieved to support subsequent analysis and calculation tasks.
[0092] Furthermore, in the embodiments of the present application, the collected data will be transmitted via a communication network to the autonomous driving system and the advanced driver assistance system (ADAS) of the intelligent vehicle. Specifically, Ethernet and CAN FD (Controller Area Network Flexible Data) protocols are used for data fusion transmission, ensuring timely data delivery in a high-speed, reliable environment. This efficient data transmission method ensures the sharing of real-time information, thereby providing the necessary data support for subsequent collision avoidance operations.
[0093] Through this data collection and transmission method, the embodiment of the present application realizes real-time and accurate vehicle status monitoring and surrounding environment perception, providing strong support for the collision avoidance decision-making of intelligent vehicles in complex traffic scenarios.
[0094] In step S2, offline learning and large-scale vehicle data are used to approximate the set of unreachable states for collision avoidance, and the Hamilton-Jacobi reachability value function and action-value function are updated offline iteratively using a large amount of pre-collected real vehicle collision avoidance data to obtain the reachability value function and action-value function.
[0095] In the embodiment of the present application, module 20 obtains obstacle information and vehicle kinematics and dynamics information from module 10, and substitutes the relevant vehicle and environment information into the pre-trained neural network to obtain the value function V of the current state. h (x), and then combined with the driver's current operation to get Q h (x,u d ).
[0096] Specifically, in an embodiment of the present application, the obstacle is represented as an elliptical envelope or a T-shaped collision surface, where the parameters of the target obstacle are recorded as [X0, Y0, a, b], which represent the center coordinates and elliptical shape parameters of the obstacle, respectively. These parameters are used to describe the geometric characteristics of the obstacle to ensure that the obstacle is represented more accurately in the system.
[0097] In addition, in this embodiment of the present application, the target obstacle is represented as a safety constraint state set T. The vehicle state within this set is considered to be a dangerous state where a collision has occurred, which is defined as:
[0098]
[0099] in, Represents the state and kinematic information of the vehicle, X, Y are the global position coordinates of the vehicle, is the yaw angle, v x ,v y are the longitudinal velocity and lateral velocity respectively, and r represents the yaw rate.
[0100] In addition, in the solution process, the embodiment of the present application uses the approximation technology to convert the Hamilton-Jacobi reachability value function V h and the action-value function Q h The solution process is simplified and updated offline through a large amount of real vehicle collision avoidance data collected in advance.
[0101] The specific training process is as follows: First, the vehicle state and driver operation in a large amount of pre-collected real-vehicle collision avoidance data are substituted into the reachability value function network and action value function network corresponding to the safety constraint state set T, and the sample reachability value function value and sample action value function value are calculated. These sample data provide a basis for subsequent optimization.
[0102] Next, the reachability value function and action-value function of these samples are updated offline iteratively to optimize the accuracy of collision avoidance decisions. Through continuous updates, this application can achieve more accurate collision avoidance decisions in various driving scenarios.
[0103] After the iteration, the updated reachability value function V is obtained h Network and action-value function Q h network and deploy it in the intelligent vehicle advanced driver assistance system (ADAS) for subsequent application processes.
[0104] Through this process, the embodiment of the present application effectively improves the human-machine shared control system's ability to assess collision avoidance risks and the accuracy of collision avoidance decisions, providing strong support for the human-machine shared control system, ensuring that safety decisions can be made in a timely manner in complex environments, avoiding collisions and improving vehicle driving safety.
[0105] In one possible embodiment, module 10 detects a sudden accident ahead at a vehicle speed of 50 km / h and transmits relevant information to module 20 for driving state identification during the collision avoidance process. This information is also sent to module 32 for temporary storage. In this embodiment, the target obstacle can be described as an elliptical envelope with a long diameter of 8 meters and a short diameter of 6 meters. To ensure a human-centric approach, this embodiment implements shared control of the steering system, while the driver has full control over the braking and driving systems. To further specify the operating conditions, this embodiment assumes that the driver wishes to avoid collisions at the original speed.
[0106] In step S3, based on the reachability value function and the action-value function, the driver's collision avoidance capability and collision avoidance intention are quantified, and the human-machine control weight is dynamically allocated according to the quantified collision avoidance capability and collision avoidance intention.
[0107] Step S3 is a key step in the embodiment of the present application. It aims to optimize the collaboration between the driver and the machine by quantifying the driver's collision avoidance ability and collision avoidance intention, and dynamically allocating human-machine control weights based on the quantification results, thereby ensuring that the best collision avoidance decisions are made in complex environments.
[0108] The specific steps include:
[0109] First, the current vehicle state, obstacle envelope and driver operation collected by the sensor system are substituted into the updated accessibility value function V corresponding to the constraint state set T. h Network and action-value function Q h network, and then calculate the reachability value V in the current state h (x) and action-value Q h (x,u d ), these values provide the necessary basis for subsequent collision avoidance capability and collision avoidance intention calculations.
[0110] Next, through the reachability value V h (x) represents the optimal collision avoidance distance in the current state and further calculates the driver's collision avoidance ability CAA. The collision avoidance ability is calculated using the following formula:
[0111]
[0112] Among them, α CAA C is the sensitivity adjustment parameter used to control the sensitivity of collision avoidance to changes in accessibility values; CAA is an offset constant used to adjust the baseline value of collision avoidance capability. This formula can effectively quantify the driver's collision avoidance capability in the current state, thus providing a basis for subsequent control weight allocation.
[0113] Based on the calculation of the collision avoidance capability, the embodiment of the present application continues to calculate the collision avoidance capability according to the current control operation u of the driver.d The corresponding action-value Q h (x,u d ) The reachability value V corresponding to the best action h (x), and further calculate the driver's collision avoidance intention CAI. The collision avoidance intention is calculated using the following formula:
[0114]
[0115] The CAI (Collision Avoidance Intention) reflects the driver's intention to avoid collisions, with a value ranging from 0 to 1, where a value closer to 1 indicates a stronger intention. This indicator can be used to quantify the driver's motivation to avoid collisions in the current driving state.
[0116] Furthermore, this embodiment of the application proposes a method for assigning weights γ to the human-machine controller that considers the driver's collision avoidance ability (CAA) and collision avoidance intent (CAI). This method should ensure that the weights fluctuate between 0 and 1 and are inversely proportional to the collision avoidance ability and collision avoidance intent. This method should be decoupled from the generation of machine actions to ensure the interpretability and transparency of the final collision avoidance maneuver.
[0117] Specifically, module 20 can update the collision avoidance capability CAA and the collision avoidance intention CAI in real time and transmit them to module 30. The human-machine weight allocation mechanism in module 30 calculates the control weight γ of the machine control according to the following method:
[0118] γ=max(γ min ,(1-s CAI )(1-s CAA ))
[0119] Among them, γ min The minimum control weight ensures that machine intervention has minimal impact in all situations. This formula intelligently balances the control efforts of the driver and the machine, ensuring that the machine can effectively take over control at critical moments and reduce the risk of collision.
[0120] In addition, to ensure that the calculation of the control weight is more accurate, the embodiment of the present application also uses a normalization function s CAI With s CAA To adjust the sensitivity to collision avoidance intention and collision avoidance ability. The specific formula is as follows:
[0121]
[0122] Among them, k CAI1 and k CAA1 Control s separately CAI and s CAA Sensitivity to changes in CAI and CAA, k CAI2 and k CAA2 DefinitionsCAI and s CAA Reaching a threshold of 0.5 represents the balance point between collision avoidance intention and capability.
[0123] Figure 3 This is a schematic diagram showing how the driver's collision avoidance ability and intention affect the vehicle's control weights, as provided in an embodiment of the present application. Through these calculations, the embodiment of the present application can dynamically adjust the balance of human-machine control in real time, enabling effective driver-machine coordination in dangerous situations, ensuring the vehicle can successfully avoid obstacles and reduce the risk of collision.
[0124] Step S4: generating machine operations based on a reachability-inspired reinforcement learning algorithm, where the reinforcement learning algorithm combines the driver's collision avoidance intention, collision avoidance capability, and reachability information into the state space.
[0125] The embodiment of the present application also proposes a method for generating shared driving machine actions based on reachability-inspired reinforcement learning. First, the reachability value function, driving ability, and driving intention are explicitly integrated into the state space of reinforcement learning (in addition, vehicle dynamics and movement habit states should also be included). Secondly, a driver action generation model is constructed in combination with the collision avoidance intention CAI to simulate and generate driver collision avoidance operations during training. Finally, based on reducing human-machine conflicts and maintaining the original task performance, a reward function is designed and the obstacle envelope is used as a constraint. The online reinforcement learning method is used to solve the machine action u m .
[0126] Specifically, step S4 further includes:
[0127] Step S41: The driver's collision avoidance intention, collision avoidance ability, accessibility information and other vehicle dynamic states are explicitly integrated into the state space of reinforcement learning. The state space includes: the vehicle's dynamic state x, the accessibility value function V h (x), the driver's collision avoidance ability CAA, the driver's collision avoidance intention CAI, the human-machine control weight γ, and the control operation u of the driver and the machine d and u m , where the action space u m Contains only the front wheel steering angle δ mf .
[0128] In the embodiment of the present application, the state space can be expressed as: s = [x, V h (x),CAA,CAI,γ,u d ,u m ], the state space includes several key components.
[0129] Specifically, first, the dynamic state x of the vehicle can fully describe the kinematic state of the vehicle at the current moment, providing a basis for subsequent collision avoidance decisions.
[0130] Secondly, the reachability value function V h (x) can measure the possibility of collision avoidance in the current vehicle state, reflecting whether the vehicle is close to the critical state of collision in the current state and predicting its long-term safety in this state.
[0131] The driver's collision avoidance ability (CAA) and collision avoidance intention (CAI) are also incorporated into the state space, quantifying the driver's ability and intention to avoid collisions in the current situation, respectively. The driver's collision avoidance ability is calculated using a reachability value function, reflecting whether the driver can take timely collision avoidance actions. The collision avoidance intention is quantified using the action value of the driver's current operation and state, reflecting the driver's willingness to take measures to avoid collisions in the current situation. This allows the reinforcement learning agent to evaluate and supplement the driver's behavior.
[0132] The human-machine control weight γ determines the balance between machine and driver control. Through this weight, the application can dynamically adjust the degree of machine intervention based on the driver's collision avoidance ability and intention, ensuring that the machine can take over control in a timely manner to avoid a collision when necessary.
[0133] It should be emphasized that the control operation of the machine mf Only includes the front wheel steering angle δ mf , because in this application, the machine mainly avoids obstacles by controlling steering, without involving other control operations (such as braking or acceleration).
[0134] By integrating this information into the state space of reinforcement learning, the present application can comprehensively assess the current driving situation and respond appropriately. Moreover, based on these inputs, embodiments of the present application can more accurately predict the driver's intentions and determine whether machine intervention is needed, thereby achieving more intelligent collision avoidance control. Overall, this approach improves the collaboration between the driver and the machine, ensuring that optimal collision avoidance decisions can be made in a variety of complex environments.
[0135] Step S42, in the state space, combining the dynamic and kinematic states of the vehicle, by constructing a driver action generation model to simulate the driver's collision avoidance operation during the training process. The driver action generation model simulates the driver's collision avoidance decision by dynamically adjusting the relative relationship between the driver input and the obstacle.
[0136] In the embodiment of the present application, module 30 uses a reinforcement learning algorithm to generate machine operation u m .
[0137] It should be noted that the reinforcement learning algorithm requires an accurate environment interaction to train the policy network. For this process, the embodiment of the present application provides a driver action generation model, such as Figure 4As shown in , the model is mainly divided into two stages: (1) obstacle avoidance stage, in which the driver steers around the obstacle; (2) recovery stage, in which the driver readjusts to the original trajectory after avoiding the obstacle. In the first stage, the driver's steering behavior can be represented as maneuvering around an elliptical envelope slightly larger than the obstacle, as shown in Figure 4 shown in the upper part of .
[0138] During the calculation process, the present embodiment incorporates the current vehicle state, obstacle envelope, and driver control input into the driver action generation model. By calculating the relative position and velocity between the vehicle and the obstacle, the driver's steering operation is calculated and adjusted. The goal of this adjustment is to align the driver's preview direction with the tangent of the obstacle boundary, thereby ensuring that the vehicle can successfully avoid the obstacle and prevent a collision.
[0139] In order to achieve this goal, the embodiment of the present application sets the driver's preview angle θ c It is defined as the angular difference between the vehicle's current position and the obstacle, and is calculated as follows:
[0140]
[0141] Among them, (X trg ,Y trg ) represents the tangent point on the ellipse, and the current position of the vehicle is (X, Y). Through this formula, the embodiment of the present application can calculate in real time the steering angle that the driver should take to ensure that the vehicle's driving direction is aligned with the boundary tangent of the obstacle, thereby avoiding collision.
[0142] To accurately describe the obstacle, this application represents the target obstacle as an ellipse, whose boundary is described by the following equation:
[0143]
[0144] Where λ is a scaling factor that is proportional to the accessibility value V h (x) and the collision avoidance intention CAI adjustment ellipse size, defined as:
[0145]
[0146] Among them, k λ is a hyperparameter that controls the sensitivity of the scaling factor. Through this scaling factor, the size of the obstacle will be dynamically adjusted based on the current collision avoidance intention and the vehicle's accessibility information, ensuring more accurate collision avoidance decisions.
[0147] The recovery process can be regarded as tracking the original trajectory, which has been well studied in the public, and will not be described in detail in the embodiments of this application.
[0148] This process ensures that the intelligent vehicle's advanced driver assistance system accurately simulates collision avoidance behavior when the driver encounters an obstacle. By simulating the driver's collision avoidance decisions, the intelligent vehicle's advanced driver assistance system not only reflects how the driver maneuvers the vehicle to avoid obstacles, but also continuously optimizes the model during training, providing reliable support for collision avoidance decisions during actual driving.
[0149] In step S43, a reinforcement learning reward function is constructed based on the driver's collision avoidance intention. The reward function is designed based on the goal of reducing human-machine conflict and maintaining the original task performance. The reward function uses the obstacle envelope as a constraint condition to guide the generation of machine control actions during the reinforcement learning process.
[0150] In this embodiment of the present application, the reward design for the reinforcement learning algorithm used to generate machine actions comprehensively considers collision avoidance safety and driver cooperation. The reward function is designed based on two primary goals: reducing human-machine conflict and maintaining performance of the original task. The reward function plays a key role in the reinforcement learning process, guiding the generation of machine control actions and ensuring that the driver and machine work together to avoid collisions.
[0151] First, the reward function considers two main aspects: safety reward and collaboration reward.
[0152] Safety Reward R sf The goal is to ensure that the vehicle remains within a safe state range and avoids entering an unreachable state set, which can effectively prevent collisions. The safety reward is calculated using the following formula:
[0153]
[0154] Among them, V h (x) represents the accessibility value of the vehicle’s current state, reflecting whether the vehicle is in a safe state; d0 in the formula is a scaling parameter that adjusts the sensitivity of the reward function; k sf1 A negative weighting factor is used to emphasize the importance of staying away from the boundary of the intelligent vehicle advanced driver assistance system, which is used to emphasize the importance of staying away from unsafe states. h (x)>k sf2 , that is, if the vehicle state is in the collision avoidance unreachable state set, a larger penalty will be incurred to ensure that the reinforcement learning model can guide the vehicle to avoid entering unsafe areas.
[0155] Secondly, the collaboration reward R co The goal is to optimize the collaboration between the machine and the driver and reduce conflicts caused by excessive or insufficient intervention. The collaboration reward is calculated by quantifying the deviation between the machine and driver's behavior. The specific formula is as follows:
[0156] Rco =-k co ·γ·(u M -u D ) 2
[0157] Among them, the penalty term (u M -u D ) 2 is the deviation between the machine and the driver’s behavior. The penalty term is used to guide the machine control to be consistent with the driver’s behavior as much as possible to avoid excessive intervention; co is a scaling factor used to adjust the weight of the collaborative reward, and the weighting factor γ reflects the necessity of machine intervention based on the driver's collision avoidance ability and intention. When the driver's collision avoidance ability and intention are weak, the value of γ is large, indicating a strong need for machine intervention, which encourages the machine to intervene more. Conversely, when the driver's ability is strong, the machine intervenes less.
[0158] Step S44: During the reinforcement learning training process, by maximizing the above-mentioned reward function, the machine prioritizes safety while reducing human-machine conflicts, and generates the final control action.
[0159] In the reinforcement learning process of the embodiments of this application, the goal of training is to enable the machine to make the best control decisions in the actual environment. To this end, the machine will learn based on the previously defined safety rewards and collaboration rewards, optimizing the interaction between the driver and the machine by maximizing the total reward. Specifically, the machine will continuously adjust the control strategy to reduce the deviation from the driver's operation, ensuring that collisions can be effectively avoided during human-machine collaboration, while maintaining the driver's comfort and avoiding excessive intervention.
[0160] In this way, the control weights between the machine and the driver can be dynamically adjusted during the reinforcement learning process, allowing the machine to quickly take over control when necessary, ensuring that the vehicle can successfully avoid obstacles in potentially dangerous situations. Conversely, when the driver's collision avoidance ability is strong, the machine reduces intervention, maximizing respect for the driver's operational intent. This dynamic adjustment process reduces human-machine conflict while ensuring that safety is the highest priority in every decision.
[0161] Ultimately, after reinforcement learning training, the machine is able to generate final control actions that are tailored to the current environment and context, enabling more intelligent, precise, and safe collision avoidance maneuvers. This process significantly improves the system's performance in complex driving scenarios, enabling intelligent vehicles to maintain driving safety in various traffic conditions and optimizing the driving experience.
[0162] In practical applications, according to the principles and demonstrations provided in the above embodiments, a policy network that can be deployed in an ADAS system can be trained to generate the machine steering action u in the process of human-machine co-driving in critical scenarios. m
[0163] In step S5, the final operation to be executed is generated by combining the human-machine control weight, the driver's operation and the machine operation. The final operation to be executed is sent to each actuator of the vehicle in real time through the communication system to execute the collision avoidance operation.
[0164] In an embodiment of the present application, based on the aforementioned calculation results, combined with the human-machine control weights, the driver's operations, and the machine's operations, the final operation to be executed is generated, and the operation is transmitted to each actuator of the vehicle in real time through the communication system to ensure that the corresponding collision avoidance operation is performed.
[0165] Specifically, first, according to the human-machine control weight γ, machine operation u m and driver operation d , generate the final operation to be executed u f , the formula is:
[0166] u f =γ·u m +(1-γ)·u d
[0167] Through this formula, the degree of machine intervention can be dynamically adjusted according to the weight ratio in the current situation to ensure a collaborative balance between the machine and the driver.
[0168] Once the final pending operation u f Once the calculated maneuver is complete, the intelligent vehicle's advanced driver assistance system (ADAS) transmits this maneuver in real time to the vehicle's various actuators via the communication system. These actuators, including the active differential steering system, drive system, and braking system, execute corresponding collision avoidance maneuvers based on the received maneuver signals. For example, when encountering an obstacle, the ADAS might instruct the steering system to make necessary adjustments or instruct the braking system to slow down to avoid the obstacle.
[0169] Furthermore, during execution, the embodiments of the present application continuously monitor sensor feedback, acquiring real-time vehicle status information. Using this real-time feedback, the system continuously assesses the relative position of the vehicle and the obstacle, dynamically adjusting the final action based on real-time changes to ensure the vehicle successfully avoids the obstacle and returns to its normal trajectory. This adjustment ensures the vehicle can successfully avoid the obstacle and return to its normal trajectory during the collision avoidance process, avoiding secondary risks caused by excessive intervention or errors.
[0170] In the test of the embodiment of the present application, Figure 4The results shown reflect the outstanding performance of the embodiment of the present application in the human-machine shared control system. Figure 4 The first sub-graph shows that during the obstacle avoidance process, the human-machine shared control system using the present application is less affected by obstacle interference than the baseline driving mode, and within the same 6 seconds, the vehicle travels a distance of up to 4.5 meters, indicating that the present application can effectively improve the vehicle's driving efficiency and maintain good original driving task performance without sacrificing driving safety.
[0171] Figure 4 The second sub-figure shows the effect of yaw angle adjustment. The yaw angle adjustment method used in this application is smoother, avoiding the overshoot that can occur in traditional human-controlled modes. This means that when encountering obstacles, the vehicle can adjust its driving trajectory more stably, reducing the driver's operational burden and improving driving comfort and safety.
[0172] Figure 4 The third sub-figure shows the evolution of the reachability value during the obstacle avoidance process. Using our approach, the reachability value steadily increases and stabilizes at a low peak value as the vehicle approaches the obstacle (approximately at t = 2 seconds), demonstrating that the shared human-machine control system is able to proactively respond to obstacles and ensure safety. Compared to the baseline approach, our approach significantly reduces response latency and avoids high peak values, demonstrating a more accurate and timely collision avoidance response.
[0173] In summary, the above test results prove that the application of this application in the human-machine shared control system can improve the obstacle avoidance performance while ensuring the safety of the vehicle is better than the traditional driving mode.
[0174] In order to implement the above embodiments, the present application also proposes an intelligent vehicle human-machine shared collision avoidance control system under critical working conditions. Figure 5 The present application embodiment provides a schematic diagram of a structure of an intelligent vehicle human-machine shared collision avoidance control system under critical working conditions. Figure 5 The system includes:
[0175] The data acquisition module 100 collects and transmits information about the vehicle's current state and surrounding environment through a sensor system that includes a camera, millimeter-wave radar, integrated inertial navigation unit (IMU), and related networking facilities.
[0176] The reachability analysis module 200 uses offline learning and large-scale vehicle data to approximate the set of collision avoidance unreachable states, and performs offline iterative updates on the Hamilton-Jacobi reachability value function and action-value function using a large amount of pre-collected real vehicle collision avoidance data to obtain the reachability value function and action-value function;
[0177] The driver intention and capability evaluation module 300 quantifies the driver's collision avoidance capability and intention based on the reachability value function and the action-value function, and dynamically allocates human-machine control weights based on the quantified collision avoidance capability and intention;
[0178] The reinforcement learning control module 400 generates machine operations based on a reachability-inspired reinforcement learning algorithm that incorporates the driver's collision avoidance intention, collision avoidance capability, and reachability information into the state space;
[0179] The control decision module 500 combines the human-machine control weight, the driver's operation and the machine operation to generate the final operation to be executed, and sends the final operation to be executed to each actuator of the vehicle in real time through the communication system to perform the collision avoidance operation.
[0180] Regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0181] In order to implement the above embodiments, the present application also proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided by the above embodiments.
[0182] In order to implement the above embodiments, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided by the above embodiments.
[0183] In order to implement the above embodiments, the present application also proposes a computer program product, including a computer program, which implements the methods provided by the above embodiments when executed by a processor.
[0184] The collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this application are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0185] It is important to note that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold beyond these legitimate uses. Furthermore, such collection / sharing should be conducted only after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes the relevant user information before using the feature. Furthermore, any necessary steps must be taken to safeguard and secure access to such personal information and ensure that others with access to personal information comply with its privacy policy and procedures.
[0186] This application contemplates providing implementations that allow users to selectively block the use or access of personal information data. Specifically, this disclosure contemplates providing hardware and / or software to prevent or block access to such personal information data. Risks can be minimized by limiting data collection and deleting data once it is no longer needed. Furthermore, where applicable, such personal information can be de-identified to protect user privacy.
[0187] In the descriptions of the foregoing embodiments, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are mutually inconsistent.
[0188] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0189] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0190] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0191] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0192] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0193] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0194] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
[0195] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved. This is not limited herein.
[0196] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A human-machine shared collision avoidance control method for intelligent vehicles under critical working conditions, characterized in that: include: Collect and transmit information about the vehicle's current state and surrounding environment through a sensor system, which includes a camera, millimeter-wave radar, an integrated inertial navigation unit (IMU), and related networking facilities; The set of unreachable states for collision avoidance is approximately solved using offline learning and large-scale vehicle data. The Hamilton-Jacobi reachability value function and action-value function are updated offline iteratively using a large amount of pre-collected real-vehicle collision avoidance data to obtain the reachability value function and action-value function. quantifying the driver's collision avoidance capability and collision avoidance intention based on the accessibility value function and the action-value function, and dynamically allocating human-machine control weights according to the quantified collision avoidance capability and collision avoidance intention; Generate machine actions based on a reachability-inspired reinforcement learning algorithm that incorporates the driver's collision avoidance intention, collision avoidance capability, and reachability information into the state space; Combining the human-machine control weights, the driver's operation, and the machine's operation, a final operation to be executed is generated, and the final operation to be executed is sent in real time to each actuator of the vehicle via a communication system to execute a collision avoidance operation; Based on the accessibility value function and the action-value function, the driver's collision avoidance capability and collision avoidance intention are quantified, and according to the quantified collision avoidance capability and collision avoidance intention, human-machine control weights are dynamically allocated, including: Substitute the current vehicle state, obstacle envelope and driver operation collected by the sensor system into the constraint state set The corresponding updated reachability value function Networks and Action-Value Functions Network, get the reachability value and action-value ;in, Represents the state and kinematic information of the vehicle, is the global position coordinate of the vehicle, is the yaw angle, are the longitudinal velocity and the lateral velocity, represents the yaw rate; By reachability value Characterize the optimal collision avoidance distance of the current state and calculate the driver's collision avoidance ability , the collision avoidance capability is calculated by the following formula: in, is the sensitivity adjustment parameter, is the offset constant; By the driver's current operation Corresponding action-value The reachability value corresponding to the best action , calculate the driver's collision avoidance intention , the collision avoidance intention is calculated by the following formula: According to the collision avoidance capability and collision avoidance intention , dynamically allocate human-machine control weights , the control weight is calculated by the following formula: in, and is the normalization function, which is defined by the following formulas: in, and Separate control and right and sensitivity to change, and definition and Reaching a threshold of 0.5 represents the balance point between collision avoidance intention and capability.
2. The method according to claim 1, characterized in that The set of unreachable states for collision avoidance is approximately solved using offline learning and large-scale vehicle data. The Hamilton-Jacobi reachability value function and action-value function are updated offline iteratively using a large amount of pre-collected real-vehicle collision avoidance data to obtain the reachability value function and action-value function, including: The obstacle is represented as an elliptical envelope or a T-shaped collision surface, where the parameters of the target obstacle are recorded as , respectively represent the center coordinates and ellipse shape parameters of the obstacle; The target obstacle is represented as a set of safety constraint states , the vehicle state in this set is considered to be a dangerous state where a collision has occurred, which is defined as: in, Represents the state and kinematic information of the vehicle, is the global position coordinate of the vehicle, is the yaw angle, are the longitudinal velocity and the lateral velocity, represents the yaw rate; Substitute the vehicle state and driver operation from a large amount of pre-collected real vehicle collision avoidance data into the safety constraint state set The corresponding reachability value function network and action value function network are used to calculate the sample reachability value function value and the sample action value function value; The sample reachability value function and sample action value are updated offline iteratively to further optimize the accuracy of collision avoidance decision-making. After the iteration is completed, the updated reachability value function is obtained. Networks and Action-Value Functions network and deploy it in smart vehicle advanced driver assistance systems.
3. The method according to claim 2, characterized in that Generate machine actions based on a reachability-inspired reinforcement learning algorithm that incorporates the driver's collision avoidance intention, collision avoidance capability, and reachability information into the state space, including: The driver's collision avoidance intention, collision avoidance ability, accessibility information and other vehicle dynamic states are explicitly integrated into the state space of reinforcement learning, which includes: the dynamic state of the vehicle , reachability value function , the driver's collision avoidance ability , the driver's collision avoidance intention , human-machine control weight and the control operations of the driver and the machine and , where the action space Contains only the front wheel steering angle ; In the state space, combining the dynamic and kinematic states of the vehicle, a driver's collision avoidance maneuvers during training are simulated by constructing a driver action generation model. The driver action generation model simulates the driver's collision avoidance decisions by dynamically adjusting the relative relationship between the driver's input and the obstacle. A reinforcement learning reward function is constructed based on the driver's collision avoidance intention. The reward function is designed with the goal of reducing human-machine conflict and maintaining the original task performance. The reward function uses the obstacle envelope as a constraint to guide the generation of machine control actions during the reinforcement learning process. During the reinforcement learning training process, by maximizing the above reward function, the machine prioritizes safety while reducing human-machine conflicts and generates the final control action.
4. The method according to claim 3, characterized in that In the state space, the driver's collision avoidance maneuvers during training are simulated by building a driver action generation model in combination with the vehicle's dynamic and kinematic states. The driver action generation model simulates the driver's collision avoidance decisions by dynamically adjusting the relative relationship between the driver's input and the obstacle, including: The current vehicle state, obstacle envelope, and driver's control operation input are substituted into the driver action generation model. By calculating the relative position and speed between the vehicle and the obstacle, the driver's steering operation is calculated and adjusted so that the driver's preview direction is aligned with the tangent of the obstacle boundary, thereby ensuring that the vehicle can avoid the obstacle and the preview angle Calculated according to the following formula: in, represents the tangent point on the ellipse, located on the boundary of the ellipse defined by: in, is a scaling factor based on the reachability value and collision avoidance intention Adjust the size of the ellipse, defined as: in, is a hyperparameter.
5. The method according to claim 4, characterized in that The reinforcement learning reward function is constructed based on the driver's collision avoidance intention. The reward function is designed based on the goal of reducing human-machine conflict and maintaining the original task performance. The reward function uses the obstacle envelope as a constraint to guide the generation of machine control actions during the reinforcement learning process, including: Optimize the cooperation between the driver and the machine through reinforcement learning reward function to reduce human-machine conflicts. The reward function includes safety reward. and collaboration rewards ; The safety reward It is calculated based on the distance between the vehicle state and the set of unreachable states for collision avoidance, and is defined as: in, is the scaling parameter, To emphasize the negative weighting factor of the importance of being far away from the boundary of intelligent vehicle advanced driver assistance system, States with high penalties result in significant penalties to ensure that the agent learns to avoid unsafe states; The collaboration reward Optimize human-machine collaboration by quantifying the deviation between machine and driver behavior, defined as: Among them, the penalty is the deviation between the machine and the driver's behavior, is the scaling factor, weight factor Reflects the necessity of machine intervention based on the driver's collision avoidance ability and intention.
6. The method according to claim 5, characterized in that The method combines the human-machine control weight, the driver operation, and the machine operation to generate a final operation to be executed, and sends the final operation to be executed to each actuator of the vehicle in real time through the communication system to execute the collision avoidance operation, including: According to the human-machine control weight , machine operation and driver operation , generating the final operations to be executed , the formula is: The final operation to be performed is transmitted through the communication system Send real-time information to various vehicle actuators, including the active differential steering system, drive system, and braking system, to perform corresponding collision avoidance maneuvers; During the execution process, the vehicle's status is monitored in real time through continuous sensor feedback, and the final operation is dynamically adjusted according to the relative position of the vehicle and the obstacle to ensure that the vehicle successfully avoids the obstacle and returns to its normal driving trajectory.
7. An intelligent vehicle human-machine shared collision avoidance control system under critical working conditions, characterized by: include: A data acquisition module, which is used to collect and transmit information about the vehicle's current state and surrounding environment through a sensor system, including a camera, millimeter-wave radar, an integrated inertial navigation unit (IMU), and related networking facilities; The reachability analysis module uses offline learning and large-scale vehicle data to approximate the set of unreachable states for collision avoidance. It also uses a large amount of pre-collected real-vehicle collision avoidance data to iteratively update the Hamilton-Jacobi reachability value function and action-value function offline to obtain the reachability value function and action-value function. a driver intention and capability evaluation module, configured to quantify the driver's collision avoidance capability and intention based on the accessibility value function and the action-value function, and dynamically assign human-machine control weights based on the quantified collision avoidance capability and intention; a reinforcement learning control module for generating machine actions based on a reachability-inspired reinforcement learning algorithm that incorporates the driver's collision avoidance intent, collision avoidance capability, and reachability information into a state space; A control decision module, configured to combine the human-machine control weights, the driver's operation, and the machine's operation to generate a final operation to be executed, and transmit the final operation to be executed to each actuator of the vehicle in real time via a communication system to execute a collision avoidance operation; Based on the accessibility value function and the action-value function, the driver's collision avoidance capability and collision avoidance intention are quantified, and according to the quantified collision avoidance capability and collision avoidance intention, human-machine control weights are dynamically allocated, including: Substitute the current vehicle state, obstacle envelope and driver operation collected by the sensor system into the constraint state set The corresponding updated reachability value function Networks and Action-Value Functions Network, get the reachability value and action-value ;in, Represents the state and kinematic information of the vehicle, is the global position coordinate of the vehicle, is the yaw angle, are the longitudinal velocity and the lateral velocity, represents the yaw rate; By reachability value Characterize the optimal collision avoidance distance of the current state and calculate the driver's collision avoidance ability , the collision avoidance capability is calculated by the following formula: in, is the sensitivity adjustment parameter, is the offset constant; By the driver's current operation Corresponding action-value The reachability value corresponding to the best action , calculate the driver's collision avoidance intention , the collision avoidance intention is calculated by the following formula: According to the collision avoidance capability and collision avoidance intention , dynamically allocate human-machine control weights , the control weight is calculated by the following formula: in, and is the normalization function, which is defined by the following formulas: in, and Separate control and right and sensitivity to change, and definition and Reaching a threshold of 0.5 represents the balance point between collision avoidance intention and capability.
8. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.
Citation Information
Patent Citations
Intervention type sharing control method and device for autonomous vehicle in forward collision avoidance scene
CN115923845A
Automatic control method and system for the virtual confinement of a land vehicle within a track
WO2024003674A1