Intelligent vehicle man-machine sharing collision avoidance control system and method under critical working condition

By adopting a human-machine shared collision avoidance control method in intelligent vehicles, combining sensing systems, offline learning and reinforcement learning algorithms, the driver's collision avoidance ability and intention are quantified, and the control weight is dynamically allocated, which solves the problem of difficult-to-determined human-machine conflict and the necessity of machine intervention in traditional technologies, and achieves higher driving safety and collaboration effects.

CN119928843AActive Publication Date: 2025-05-06TSINGHUA UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510146478.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-06
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

In critical operating conditions, traditional human-machine sharing anti-collision technology may cause drivers to resist or misoperate, increase new collision risks, and the risk assessment method based on distance or artificial potential fields lacks strict theoretical guarantees, and the necessity of machine intervention cannot be determined.

Method used

A method of sharing collision avoidance control for intelligent vehicles under critical working conditions is proposed. Vehicle status and environmental information is collected through the sensing system, offline learning and large-scale vehicle data approximation solve the unreachable state set of collision avoidance, quantify the driver's collision avoidance ability and intention, dynamically allocate the human-machine control weight, and generate machine operations based on accessibility-inspired reinforcement learning algorithms.

Benefits of technology

It achieves the minimization of human-machine conflicts while ensuring no collisions, improves the driving safety of vehicles in complex environments, improves the cooperation between drivers and machines, and ensures that the optimal collision avoidance decisions are made in multiple complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119928843A_ABST
    Figure CN119928843A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent vehicle man-machine sharing collision avoidance control system and method under critical working conditions, and the method comprises the steps: collecting and transmitting the current state and surrounding environment information of a vehicle through a sensing system; off-line learning and large-scale vehicle data are used for approximately solving a collision avoidance inaccessible state set, and a Hamiltonian-Jacobian accessibility value function and an action-value function are subjected to off-line iteration updating through a large amount of real vehicle collision avoidance data collected in advance; based on the accessibility value function and the action-value function, the collision avoidance capability and the collision avoidance intention of the driver are quantified, and the man-machine control weight is dynamically distributed; and generating machine operation based on a reachability heuristic reinforcement learning algorithm, generating final operation to be executed in combination with the man-machine control weight, the driver operation and the machine operation, and executing collision avoidance operation. According to the method, the collision risk of the vehicle state and the collision avoidance intention of the driver action can be evaluated theoretically, and man-machine conflicts are reduced to the greatest extent on the premise of ensuring that no collision occurs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of active automobile safety, and in particular to a human-machine shared collision avoidance control system and method for an intelligent vehicle under critical conditions. Background Art

[0002] In high-risk driving scenarios, complex operational requirements, tight reaction time, and reduced decision-making ability caused by cognitive overload make timely collision avoidance a major challenge. In this case, intelligent vehicles need to be able to assist drivers in achieving fast and precise evasive maneuvers to deal with possible safety risks.

[0003] Traditional human-machine shared collision avoidance technology usually generates collision avoidance trajectories through path replanning and trajectory tracking, which has been proven to be effective in ensuring vehicle safety. Its basic principle is to achieve higher-precision obstacle avoidance control through collaborative operation between the intelligent system and the driver. However, these methods often ignore the driver's real-time intentions in emergency scenarios. When the vehicle needs to perform fast or extreme operations, forcing the system to follow the trajectory generated by the system may cause resistance or misoperation from the driver, thereby increasing the risk of new collisions.

[0004] In order to reduce human-machine conflicts, driver-centric shared control methods have gradually developed. These methods provide necessary intervention while respecting the driver's intentions by introducing the concept of dynamic safety range and model predictive control technology. However, treating environmental constraints as soft constraints may reduce overall safety in some scenarios. In addition, collision avoidance control methods based on reinforcement learning have received attention in recent years. These methods avoid the limitations of traditional path planning by directly generating collision avoidance actions and are able to achieve robust collision avoidance control in complex dynamic environments. These methods still rely mainly on distance-based risk assessment and intervene near manually set thresholds. Under this framework, the theoretical basis for effective system intervention has not yet been established, that is, it is impossible to determine when machine intervention is truly necessary.

[0005] At present, although research has made significant progress in the field of human-machine shared control collision avoidance, including trajectory tracking-based frameworks, human-centered flexible constraint methods, and recent non-explicit collision avoidance path strategies based on reinforcement learning. However, these technical solutions still face some challenges. Traditional trajectory tracking strategies usually require drivers to follow preset paths, which may cause human-machine conflicts in dynamic or emergency operation scenarios and reduce driver acceptance. Attempts to improve cooperation by softening obstacle constraints may lead to reduced safety, and risk assessment methods based on distance or artificial potential field (APF) lack strict theoretical guarantees on the necessity of intervention. In addition, current research lacks consideration of the driver's own collision avoidance intention.

[0006] In response to the above problems, there is an urgent need to establish a human-machine co-driving control method and system in collision avoidance scenarios with theoretical machine intervention triggering standards to minimize human-machine conflicts. Summary of the invention

[0007] The present application aims to solve one of the technical problems in the related art at least to some extent.

[0008] To this end, the first purpose of this application is to propose a human-machine shared collision avoidance control method for intelligent vehicles under critical conditions, so as to theoretically evaluate the collision risk of the vehicle state and the collision avoidance intention of the driver's actions, and minimize the human-machine conflict while ensuring that no collision occurs.

[0009] The second objective of the present application is to propose an intelligent vehicle human-machine shared collision avoidance control system under critical conditions.

[0010] The third objective of the present application is to provide an electronic device.

[0011] A fourth objective of the present application is to provide a computer-readable storage medium.

[0012] A fifth object of the present application is to provide a computer program product.

[0013] To achieve the above-mentioned purpose, the first embodiment of the present application proposes a human-machine shared collision avoidance control method for an intelligent vehicle under critical conditions, comprising:

[0014] Collect and transmit the vehicle's current state and surrounding environment information through a sensor system, wherein the sensor system includes a camera, a millimeter-wave radar, an integrated inertial navigation unit (IMU), and related networking facilities;

[0015] Use offline learning and large-scale vehicle data to approximate the set of unreachable states for collision avoidance, and use a large amount of real vehicle collision avoidance data collected in advance to perform offline iterative updates on the Hamilton-Jacobi reachability value function and action-value function to obtain the reachability value function and action-value function.

[0016] quantifying the driver's collision avoidance capability and collision avoidance intention based on the accessibility value function and the action-value function, and dynamically allocating human-machine control weights according to the quantified collision avoidance capability and collision avoidance intention;

[0017] Generate machine actions based on a reachability-inspired reinforcement learning algorithm that incorporates driver collision avoidance intent, collision avoidance capability, and reachability information into a state space;

[0018] The final operation to be executed is generated by combining the human-machine control weight, the driver operation and the machine operation, and the final operation to be executed is sent to each actuator of the vehicle in real time through the communication system to execute the collision avoidance operation.

[0019] Optionally, offline learning and large-scale vehicle data are used to approximately solve the set of collision avoidance unreachable states, and the Hamilton-Jacobi reachability value function and the action-value function are updated offline iteratively through a large amount of pre-collected real vehicle collision avoidance data to obtain the reachability value function and the action-value function, including:

[0020] The obstacle is represented as an elliptical envelope or a T-shaped collision surface, where the parameters of the target obstacle are recorded as [X0, Y0, a, b], which represent the center coordinates of the obstacle and the elliptical shape parameters respectively;

[0021] The target obstacle is represented as a safety constraint state set Y, and the vehicle state in this set is considered to be a dangerous state where a collision has occurred, which is defined as:

[0022]

[0023] in, Represents the state and kinematic information of the vehicle. X, Y are the global position coordinates of the vehicle. is the yaw angle, v x ,v y are the longitudinal velocity and the lateral velocity respectively, and r represents the yaw rate;

[0024] Substituting the vehicle state and driver operation in a large amount of real vehicle collision avoidance data collected in advance into the reachability value function network and action value function network corresponding to the safety constraint state set T, and calculating the sample reachability value function value and the sample action value function value;

[0025] The sample reachability value function and the sample action value are updated offline iteratively to further optimize the accuracy of the collision avoidance decision. After the iteration, the updated reachability value function V is obtained. h Network and action-value function Q h network and deploy it in smart vehicle advanced driver assistance systems.

[0026] Optionally, based on the accessibility value function and the action-value function, the collision avoidance capability and collision avoidance intention of the driver are quantified, and according to the quantified collision avoidance capability and collision avoidance intention, a human-machine control weight is dynamically allocated, including:

[0027] Substitute the current vehicle state, obstacle envelope and driver operation collected by the sensor system into the updated accessibility value function V corresponding to the constraint state set T h Network and action-value function Q h Network, get the reachability value V h (x) and the action-value Q h (x,u d );

[0028] Through the accessibility value V h (x) represents the optimal collision avoidance distance in the current state, and calculates the driver's collision avoidance ability CAA, which is calculated by the following formula:

[0029]

[0030] Among them, α CAA is the sensitivity adjustment parameter, C CAA is the offset constant;

[0031] By the driver's current operation u d The corresponding action-value Q h (x,u d ) The reachability value V corresponding to the best action h (x), calculate the driver's collision avoidance intention CAI, which is calculated by the following formula:

[0032]

[0033] According to the collision avoidance capability CAA and the collision avoidance intention CAI, the human-machine control weight γ is dynamically allocated, and the control weight is calculated by the following formula:

[0034] γ=max(γ min ,(1-s CAI )(1-s CAA ))

[0035] Among them, s CAI With s CAA is the normalization function, which is defined by the following formulas:

[0036]

[0037] Among them, k CAI1 and k CAA1 Separate control CAI and CAA Sensitivity to changes in CAI and CAA, k CAI2 and k CAA2 Definitions CAI and CAA Reaching a threshold of 0.5 represents the balance point between collision avoidance intention and capability.

[0038] Optionally, the machine operation is generated based on a reachability-inspired reinforcement learning algorithm, wherein the reinforcement learning algorithm combines the driver's collision avoidance intention, collision avoidance capability, and reachability information into the state space, including:

[0039] The driver's collision avoidance intention, collision avoidance ability, accessibility information and other vehicle dynamic states are explicitly integrated into the state space of reinforcement learning, which includes: the vehicle's dynamic state x, accessibility value function V h (x), the driver’s collision avoidance ability CAA, the driver’s collision avoidance intention CAI, the human-machine control weight γ, and the control operation u of the driver and the machine d and u m , where the action space u m Only the front wheel steering angle δ mf ;

[0040] In the state space, combining the dynamic and kinematic states of the vehicle, a driver's collision avoidance operation during training is simulated by constructing a driver action generation model, wherein the driver action generation model simulates the driver's collision avoidance decision by dynamically adjusting the relative relationship between the driver input and the obstacle;

[0041] Constructing a reinforcement learning reward function based on the driver's collision avoidance intention, wherein the reward function is designed based on the goal of reducing human-machine conflict and maintaining the original task performance, and the reward function combines the obstacle envelope as a constraint to guide the generation of machine control actions during the reinforcement learning process;

[0042] During the reinforcement learning training process, by maximizing the above reward function, the machine can prioritize safety while reducing human-machine conflicts and generate the final control action.

[0043] Optionally, in the state space, in combination with the dynamic and kinematic states of the vehicle, a driver action generation model is constructed to simulate the driver's collision avoidance operation during the training process, and the driver action generation model dynamically adjusts the relative relationship between the driver input and the obstacle to simulate the driver's collision avoidance decision, including:

[0044] The current vehicle state, obstacle envelope, and driver’s control operation input are substituted into the driver action generation model. By calculating the relative position and speed between the vehicle and the obstacle, the driver’s steering operation is calculated and adjusted so that the driver’s preview direction is aligned with the tangent of the obstacle boundary, thereby ensuring that the vehicle can avoid the obstacle. The preview angle θ c Calculated according to the following formula:

[0045]

[0046] Among them, (X trg ,Y trg ) represents the tangent point on the ellipse, located on the boundary of the ellipse defined as follows:

[0047]

[0048] Where λ is a scaling factor based on the accessibility value V h (x) and the collision avoidance intention CAI adjustment ellipse size, defined as:

[0049]

[0050] Among them, k λ is a hyperparameter.

[0051] Optionally, the reinforcement learning reward function is constructed according to the driver's collision avoidance intention, the reward function is designed based on the goal of reducing human-machine conflict and maintaining the original task performance, and the reward function combines the obstacle envelope as a constraint condition to guide the generation of machine control actions during the reinforcement learning process, including:

[0052] The cooperation between the driver and the machine is optimized by reinforcement learning reward function to reduce human-machine conflict. The reward function includes a safety reward R sf and collaboration reward R co ;

[0053] The safety reward R sf It is calculated based on the distance between the vehicle state and the set of unreachable states for collision avoidance, and is defined as:

[0054]

[0055] Among them, d0 is the scaling parameter, k sf1 To emphasize the negative weighting factor of the importance of being far away from the boundary of the intelligent vehicle advanced driver assistance system, V h (x)>k sf2 States that result in significant penalties ensure that the agent learns to avoid unsafe states;

[0056] The collaboration reward R co Optimize human-machine collaboration by quantifying the deviation between machine and driver behavior, defined as:

[0057] R co =-k co ·γ·(u M -u D ) 2

[0058] Among them, the penalty term (u M -u D ) 2 is the deviation between the machine and the driver’s behavior, k co is a scaling factor, and the weighting factor γ reflects the necessity of machine intervention based on the driver’s collision avoidance ability and intention.

[0059] Optionally, the combining the human-machine control weight, the driver operation and the machine operation to generate a final operation to be executed, and sending the final operation to be executed to each actuator of the vehicle in real time through a communication system to execute a collision avoidance operation, including:

[0060] According to the human-machine control weight γ, machine operation u m And driver operation d , generate the final operation to be executed u f , the formula is:

[0061] u f =γ·u m +(1-γ)·u d

[0062] The final operation to be performed u is transmitted through the communication system f Sent in real time to the vehicle's actuators, including the active differential steering system, drive system and braking system, to perform corresponding collision avoidance operations;

[0063] During the execution process, the vehicle's status is monitored in real time through continuous sensor feedback, and the final operation is dynamically adjusted according to the relative position of the vehicle and the obstacle to ensure that the vehicle successfully avoids the obstacle and returns to its normal driving trajectory.

[0064] To achieve the above-mentioned purpose, the second embodiment of the present application proposes an intelligent vehicle human-machine shared collision avoidance control system under critical working conditions, including:

[0065] A data acquisition module, which is used to collect and transmit information about the vehicle's current state and surrounding environment through a sensor system, wherein the sensor system includes a camera, a millimeter-wave radar, an integrated inertial navigation unit (IMU), and related networking facilities;

[0066] The reachability analysis module is used to use offline learning and large-scale vehicle data to approximately solve the set of unreachable states for collision avoidance, and to perform offline iterative updates on the Hamilton-Jacobi reachability value function and action-value function through a large amount of pre-collected real vehicle collision avoidance data to obtain the reachability value function and action-value function;

[0067] A driver intention and capability evaluation module, used to quantify the driver's collision avoidance capability and collision avoidance intention based on the accessibility value function and the action-value function, and dynamically allocate a human-machine control weight according to the quantified collision avoidance capability and collision avoidance intention;

[0068] A reinforcement learning control module for generating machine operations based on a reachability-inspired reinforcement learning algorithm that incorporates the driver's collision avoidance intent, collision avoidance capability, and reachability information into a state space;

[0069] The control decision module is used to combine the human-machine control weight, the driver operation and the machine operation to generate a final operation to be executed, and send the final operation to be executed to each actuator of the vehicle in real time through the communication system to execute the collision avoidance operation.

[0070] To achieve the above-mentioned purpose, the third aspect of the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0071] The memory stores computer-executable instructions;

[0072] The processor executes the computer-executable instructions stored in the memory to implement the method as described in any one of the first aspects above.

[0073] To achieve the above-mentioned purpose, the fourth aspect embodiment of the present application proposes a computer-readable storage medium, in which computer execution instructions are stored. When the computer execution instructions are executed by a processor, they are used to implement the method as described in any one of the above-mentioned first aspects.

[0074] To achieve the above-mentioned purpose, the fifth aspect of the present application proposes a computer program product, including a computer program, which, when executed by a processor, implements the method as described in any one of the above-mentioned first aspects.

[0075] The technical solution provided by the embodiments of the present application brings at least the following beneficial effects:

[0076] (1) The Hamilton-Jacobi reachability method is used to theoretically evaluate the collision risk. By calculating the reachability value, the system can make timely intervention decisions when danger is approaching. This method enables the machine to determine whether intervention is needed based on the real-time collision risk, effectively avoiding collisions and reducing human-machine conflicts, ensuring the safety of vehicles in complex environments.

[0077] (2) The system uses a human-machine shared control system that can accurately control the vehicle trajectory and achieve smoother yaw angle adjustment. During the obstacle avoidance process, the system's real-time adjustment capability can effectively respond to sudden obstacles or complex working conditions, thereby improving the stability and safety of the vehicle, especially the response capability in emergency situations. Whether in emergency situations or daily driving, the vehicle can maintain a higher level of safety.

[0078] (3) By dynamically quantifying the driver's collision avoidance ability and intention, and combining the driver's real-time operation with the system's calculation, the human-machine control weight is intelligently allocated. At critical moments, the system can effectively take over vehicle control, while respecting the driver's operation at other times and maintaining the driver's sense of control. This intelligent mechanism not only optimizes the driving experience, but also balances human-machine collaboration in complex driving environments, improving driving safety and comfort.

[0079] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0081] Figure 1 A schematic diagram of a flow chart of a human-machine shared collision avoidance control method for an intelligent vehicle under critical conditions provided by an embodiment of the present application;

[0082] Figure 2 A schematic diagram of the architecture of a human-machine shared collision avoidance control method for an intelligent vehicle under critical conditions provided by an embodiment of the present application;

[0083] Figure 3 A schematic diagram showing the influence of the driver's collision avoidance ability and collision avoidance intention on the machine control weight provided in an embodiment of the present application;

[0084] Figure 4 A comparison diagram of the deployment effects of the present application and other prior art solutions provided in the embodiments of the present application;

[0085] Figure 5 A schematic diagram of the structure of a human-machine shared collision avoidance control system for an intelligent vehicle under critical conditions provided in an embodiment of the present application. DETAILED DESCRIPTION

[0086] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0087] In view of the shortcomings of the prior art, the embodiment of the present application provides a human-machine shared collision avoidance control method for intelligent vehicles under critical working conditions, which uses large-scale data to approximate the set of unreachable collision avoidance states, and uses the degree of proximity between the vehicle state and the set of unreachable collision avoidance states as a theoretical basis for evaluating the collision risk, thereby judging the necessity of machine intervention. This method can theoretically evaluate the collision risk of the vehicle state and the collision avoidance intention of the driver's actions, and minimize human-machine conflicts while ensuring that no collision occurs.

[0088] Reference Figure 1 and Figure 2 , the method comprises the following steps:

[0089] Step S1, collecting and transmitting the current state of the vehicle and the surrounding environment information through the sensor system, the sensor system includes a camera, a millimeter wave radar, a combined inertial navigation IMU and related networking facilities.

[0090] In the embodiment of the present application, this step is performed by Figure 2 Module 10 is used to execute. Figure 2 The sensor system includes cameras, millimeter-wave radars, combined inertial navigation IMUs and related networking facilities. These sensors can acquire and transmit the vehicle's position information, speed, acceleration and other dynamic state data in real time, as well as the relative position and size of obstacles in the surrounding environment. These data are crucial for subsequent collision avoidance decisions.

[0091] Considering that a large amount of historical collision avoidance data is required when solving the collision avoidance unreachable state set, the sensing system should have strong data integration, collection and storage capabilities to ensure that the necessary collision avoidance data can be stored and quickly retrieved to support subsequent analysis and calculation tasks.

[0092] In addition, in the embodiment of the present application, the collected data will be transmitted to the advanced driver assistance system (ADAS) of the autonomous driving system and the intelligent vehicle through the communication network. Specifically, Ethernet and CAN FD (Controller Area Network Flexible Data Transmission) protocols are used for data fusion transmission to ensure that data is delivered in a high-speed and reliable environment in a timely manner. This efficient data transmission method can ensure the sharing of real-time information, thereby providing the necessary data support for subsequent collision avoidance operations.

[0093] Through this data collection and transmission method, the embodiment of the present application realizes real-time and accurate vehicle status monitoring and surrounding environment perception, providing strong support for the collision avoidance decision-making of intelligent vehicles in complex traffic situations.

[0094] Step S2, using offline learning and large-scale vehicle data to approximately solve the set of unreachable states for collision avoidance, and performing offline iterative updates on the Hamilton-Jacobi reachability value function and the action-value function through a large amount of real vehicle collision avoidance data collected in advance, so as to obtain the reachability value function and the action-value function.

[0095] In the embodiment of the present application, module 20 obtains obstacle information and vehicle kinematics and dynamics information from module 10, and substitutes relevant vehicle and environment information into the pre-trained neural network to obtain the value function V of the current state. h (x), and combined with the driver's current operation, we can get Q h (x,u d ).

[0096] Specifically, in an embodiment of the present application, the obstacle is represented as an elliptical envelope or a T-shaped collision surface, wherein the parameters of the target obstacle are recorded as [X0, Y0, a, b], which respectively represent the center coordinates and elliptical shape parameters of the obstacle, and these parameters are used to describe the geometric characteristics of the obstacle to ensure that the obstacle is represented more accurately in the system.

[0097] In addition, in the embodiment of the present application, the target obstacle is represented as a safety constraint state set T, and the vehicle state in the set is regarded as a dangerous state where a collision has occurred, which is defined as:

[0098]

[0099] in, Represents the state and kinematic information of the vehicle. X, Y are the global position coordinates of the vehicle. is the yaw angle, v x ,v y are the longitudinal velocity and the lateral velocity respectively, and r represents the yaw rate.

[0100] In addition, in the solution process, the embodiment of the present application uses the approximation technology to convert the Hamilton-Jacobi reachability value function V h and the action-value function Q h The solution process is simplified and updated offline through a large amount of real vehicle collision avoidance data collected in advance.

[0101] The specific training process is as follows: First, the vehicle state and driver operation in a large amount of real-vehicle collision avoidance data collected in advance are substituted into the reachability value function network and action value function network corresponding to the safety constraint state set T, and the sample reachability value function value and sample action value function value are calculated. These sample data provide a basis for subsequent optimization.

[0102] Next, the reachability value function and action-value function of these samples are updated offline iteratively to optimize the accuracy of collision avoidance decisions. Through continuous updates, the present application can achieve more accurate collision avoidance decisions in a variety of driving scenarios.

[0103] After the iteration, the updated reachability value function V is obtained. h Network and action-value function Q h network and deploy it in the intelligent vehicle advanced driver assistance system (ADAS) for subsequent application processes.

[0104] Through this process, the embodiment of the present application effectively improves the human-machine shared control system's ability to assess collision avoidance risks and the accuracy of collision avoidance decisions, and provides strong support for the human-machine shared control system, ensuring that safety decisions can be made in a timely manner in complex environments, avoiding collisions and improving vehicle driving safety.

[0105] In a possible embodiment, module 10 detects a sudden accident ahead when the vehicle speed is 50 km / h, and transmits the relevant information to module 20 for driving state recognition during the collision avoidance process, and sends the information to module 32 for temporary storage. The target obstacle in the embodiment can be described as an elliptical envelope with a long diameter of 8m and a short diameter of 6m. In order to ensure the human-centric strategy, in this embodiment, shared control is performed on the steering system, and the driving and braking systems are fully controlled by the driver. For further specific working conditions, this embodiment assumes that the driver wants to avoid collisions at the original speed.

[0106] Step S3, based on the accessibility value function and the action-value function, quantify the driver's collision avoidance ability and collision avoidance intention, and dynamically allocate human-machine control weights according to the quantified collision avoidance ability and collision avoidance intention.

[0107] Step S3 is a key step in the embodiment of the present application, which aims to optimize the collaboration between the driver and the machine by quantifying the driver's collision avoidance ability and collision avoidance intention and dynamically allocating human-machine control weights according to the quantification results, thereby ensuring the best collision avoidance decision in a complex environment.

[0108] The specific steps include:

[0109] First, substitute the current vehicle state, obstacle envelope and driver operation collected by the sensor system into the updated accessibility value function V corresponding to the constraint state set T h Network and action-value function Q h network, and then calculate the reachability value V in the current state h (x) and the action-value Q h (x,u d ), these values ​​provide the necessary basis for the subsequent calculation of collision avoidance capability and collision avoidance intention.

[0110] Next, through the reachability value V h (x) represents the optimal collision avoidance distance of the current state, and further calculates the driver's collision avoidance ability CAA. The collision avoidance ability is calculated by the following formula:

[0111]

[0112] Among them, α CAA C is the sensitivity adjustment parameter, which is used to control the sensitivity of collision avoidance to changes in accessibility values; CAA is an offset constant used to adjust the baseline value of collision avoidance capability. This formula can effectively quantify the driver's collision avoidance capability in the current state, thus providing a basis for subsequent control weight allocation.

[0113] Based on the calculation of the collision avoidance capability, the embodiment of the present application continues to calculate the collision avoidance capability according to the current control operation u of the driver.d The corresponding action-value Q h (x,u d ) The reachability value V corresponding to the best action h (x), and further calculate the driver's collision avoidance intention CAI. The collision avoidance intention is calculated by the following formula:

[0114]

[0115] Among them, the collision avoidance intention CAI reflects the degree of the driver's intention to avoid collision, and the value range is from 0 to 1. The closer it is to 1, the stronger the driver's collision avoidance intention is. Through this indicator, the driver's motivation to avoid collision in the current driving state can be quantified.

[0116] In addition, the embodiment of the present application also proposes a method for allocating weights γ of a human-machine controller that takes into account the driver's collision avoidance capability CAA and collision avoidance intention CAI. The method should allow the weight to float between 0 and 1 and be inversely proportional to the collision avoidance capability and collision avoidance intention. The method should be decoupled from the machine action generation to ensure the interpretability and transparency of the final collision avoidance operation.

[0117] Specifically, module 20 can update the driving collision avoidance capability CAA and the driving collision avoidance intention CAI in real time and transmit them to module 30. The human-machine weight allocation mechanism in module 30 calculates the control weight γ of the machine control according to the following method:

[0118] γ=max(γ min ,(1-s CAI )(1-s CAA ))

[0119] Among them, γ min is the minimum control weight, ensuring that machine intervention has the minimum effect in any situation. Through this formula, the control effects of the driver and the machine can be intelligently balanced to ensure that the machine can effectively take over control at critical moments and reduce the risk of collision.

[0120] In addition, to ensure that the calculation of the control weight is more accurate, the embodiment of the present application also uses a normalization function s CAI With s CAA To adjust the sensitivity to collision avoidance intention and collision avoidance ability. The specific formula is as follows:

[0121]

[0122] Among them, k CAI1 and k CAA1 Separate control CAI and CAA Sensitivity to changes in CAI and CAA, k CAI2 and k CAA2 DefinitionsCAI and CAA Reaching a threshold of 0.5 represents the balance point between collision avoidance intention and capability.

[0123] Figure 3 This is a schematic diagram of the impact of the driver's collision avoidance ability and collision avoidance intention on the machine control weight provided by the embodiment of the present application. Through these calculations, the embodiment of the present application can dynamically adjust the balance of human-machine control in real time, so that the driver and the machine can effectively cooperate in dangerous situations, ensuring that the vehicle can smoothly avoid obstacles and reduce the risk of collision.

[0124] Step S4, generating machine operations based on a reachability-inspired reinforcement learning algorithm, where the reinforcement learning algorithm combines the driver's collision avoidance intention, collision avoidance capability, and reachability information into the state space.

[0125] The embodiment of the present application also proposes a method for generating shared driving machine actions based on reachability-inspired reinforcement learning. First, the reachability value function, driving ability, and driving intention are explicitly integrated into the state space of reinforcement learning (in addition, vehicle dynamics and movement habit states should also be included). Secondly, a driver action generation model is constructed in combination with the collision avoidance intention CAI to simulate and generate driver collision avoidance operations during training. Finally, a reward function is designed based on reducing human-machine conflicts and maintaining the original task performance, and the obstacle envelope is used as a constraint. The online reinforcement learning method is used to solve the machine action u m .

[0126] Specifically, step S4 further includes:

[0127] Step S41, explicitly integrate the driver's collision avoidance intention, collision avoidance ability, accessibility information and other vehicle dynamic states into the state space of reinforcement learning, which includes: the vehicle's dynamic state x, accessibility value function V h (x), the driver’s collision avoidance ability CAA, the driver’s collision avoidance intention CAI, the human-machine control weight γ, and the control operation u of the driver and the machine d and u m , where the action space u m Only the front wheel steering angle δ mf .

[0128] In the embodiment of the present application, the state space can be expressed as: s = [x, V h (x),CAA,CAI,γ,u d ,u m ], the state space includes several key components.

[0129] Specifically, first, the dynamic state x of the vehicle can fully describe the kinematic state of the vehicle at the current moment, providing a basis for subsequent collision avoidance decisions.

[0130] Secondly, the reachability value function V h (x) can measure the possibility of collision avoidance in the current vehicle state, reflect whether the vehicle is close to the critical state of collision in the current state, and predict its long-term safety in this state.

[0131] The driver's collision avoidance ability CAA and collision avoidance intention CAI are also included in the state space, which quantify the driver's ability and intention to avoid collisions in the current situation. The driver's collision avoidance ability is calculated through the accessibility value function, reflecting whether the driver can make timely collision avoidance actions, and the collision avoidance intention is quantified by the action value under the driver's current operation and state, reflecting whether the driver is willing to take measures to avoid collisions in the current situation, so that the reinforcement learning agent can evaluate and supplement the driver's behavior.

[0132] The human-machine control weight γ determines the balance between machine and driver control. Through this weight, the application can dynamically adjust the machine's intervention level according to the driver's collision avoidance ability and intention, ensuring that the machine can take over control in time to avoid collision when necessary.

[0133] It should be emphasized that the control operation of the machine mf Only the front wheel steering angle δ mf , because in this application, the machine mainly avoids obstacles by controlling steering, without involving other control operations (such as braking or acceleration).

[0134] By integrating this information into the state space of reinforcement learning, the present application can comprehensively evaluate the current driving situation and make appropriate responses. Moreover, based on these inputs, the embodiments of the present application can more accurately predict the driver's intentions and decide whether machine intervention is needed, thereby achieving more intelligent collision avoidance control. Overall, this method improves the collaboration between the driver and the machine, ensuring that the best collision avoidance decisions can be made in a variety of complex environments.

[0135] Step S42, in the state space, combining the dynamic and kinematic states of the vehicle, by constructing a driver action generation model to simulate the driver's collision avoidance operation during the training process. The driver action generation model simulates the driver's collision avoidance decision by dynamically adjusting the relative relationship between the driver input and the obstacle.

[0136] In the embodiment of the present application, module 30 uses a reinforcement learning algorithm to generate machine operation u m .

[0137] It should be noted that the reinforcement learning algorithm requires an accurate environment interaction to train the policy network. For this process, the embodiment of the present application provides a driver action generation model, such as Figure 4As shown in , the model is mainly divided into two stages: (1) obstacle avoidance stage, in which the driver steers around the obstacle; (2) recovery stage, in which the driver readjusts to the original trajectory after avoiding the obstacle. In the first stage, the driver's steering behavior can be represented as maneuvering around an elliptical envelope slightly larger than the obstacle, as shown in Figure 4 shown in the upper part of .

[0138] During calculation, the embodiment of the present application substitutes the current vehicle state, obstacle envelope and driver's control operation input into the driver action generation model, calculates and adjusts the driver's steering operation by calculating the relative position and speed between the vehicle and the obstacle. The goal of the adjustment is to align the driver's preview direction with the tangent of the obstacle boundary, thereby ensuring that the vehicle can smoothly avoid the obstacle and avoid collision.

[0139] In order to achieve this goal, the embodiment of the present application sets the driver's preview angle θ c It is defined as the angular difference between the vehicle's current position and the obstacle, and is calculated as follows:

[0140]

[0141] Among them, (X trg ,Y trg ) represents the tangent point on the ellipse, and the current position of the vehicle is (X, Y). Through this formula, the embodiment of the present application can calculate in real time the steering angle that the driver should take to ensure that the vehicle's driving direction is aligned with the boundary tangent of the obstacle, thereby avoiding collision.

[0142] In order to accurately describe the obstacle, this application represents the target obstacle as an ellipse, whose boundary is described by the following equation:

[0143]

[0144] Where λ is a scaling factor based on the accessibility value V h (x) and the collision avoidance intention CAI adjustment ellipse size, defined as:

[0145]

[0146] Among them, k λ is a hyperparameter that controls the sensitivity of the scaling factor. Through this scaling factor, the size of the obstacle will be dynamically adjusted according to the current collision avoidance intention and the vehicle's accessibility information, ensuring a more accurate collision avoidance decision.

[0147] The recovery process can be regarded as tracking the original trajectory, which has been well studied in the public, and the embodiments of the present application will not be described in detail here.

[0148] Through the above process, it can ensure that when the driver encounters an obstacle, the intelligent car advanced driver assistance system can simulate accurate collision avoidance behavior. By simulating the driver's collision avoidance decision, the intelligent car advanced driver assistance system can not only reflect how the driver controls the vehicle to avoid obstacles, but also continuously optimize the model during the training process, thereby providing reliable support for collision avoidance decisions in actual driving.

[0149] Step S43, constructing a reinforcement learning reward function according to the driver's collision avoidance intention. The reward function is designed based on the goal of reducing human-machine conflict and maintaining the original task performance. The reward function combines the obstacle envelope as a constraint condition to guide the generation of machine control actions during the reinforcement learning process.

[0150] In the embodiment of the present application, the reward design of the reinforcement learning algorithm used to generate machine operations comprehensively considers collision avoidance safety and driver cooperation, and the design of its reward function is based on two main goals: reducing human-machine conflicts and maintaining the performance of the original task. The reward function plays a key role in the reinforcement learning process, guiding the generation of machine control actions and ensuring that the driver and the machine can work together to avoid collisions.

[0151] First, the reward function considers two main aspects: safety reward and cooperation reward.

[0152] Safety Reward R sf Aims to ensure that the vehicle remains within a safe state range and avoids entering an unreachable state set, which can effectively prevent collisions. The safety reward is calculated using the following formula:

[0153]

[0154] Among them, V h (x) represents the accessibility value of the vehicle’s current state, reflecting whether the vehicle is in a safe state; d0 in the formula is a scaling parameter that adjusts the sensitivity of the reward function; k sf1 A negative weighting factor that emphasizes the importance of staying away from the boundary of the intelligent vehicle advanced driver assistance system is used to emphasize the importance of staying away from unsafe conditions. h (x)>k sf2 , that is, if the vehicle state is in the set of unreachable states for collision avoidance, a larger penalty will be incurred to ensure that the reinforcement learning model can guide the vehicle to avoid entering unsafe areas.

[0155] Secondly, the collaboration reward R co Aims to optimize the collaboration between the machine and the driver and reduce conflicts caused by too much or too little intervention. The collaboration reward is calculated by quantifying the deviation between the machine and the driver's behavior. The specific formula is as follows:

[0156] Rco =-k co ·γ·(u M -u D ) 2

[0157] Among them, the penalty term (u M -u D ) 2 is the deviation between the machine and the driver’s behavior. The penalty term is used to guide the machine control to be consistent with the driver’s behavior as much as possible to avoid excessive intervention; k co is a scaling factor used to adjust the weight of the collaborative reward, and the weight factor γ reflects the necessity of machine intervention based on the driver's collision avoidance ability and intention. When the driver's collision avoidance ability and intention are weak, the value of γ is large, indicating that the necessity of machine intervention is strong, thereby promoting more machine intervention; conversely, when the driver's ability is strong, the machine intervention will be relatively less.

[0158] Step S44, during the reinforcement learning training process, by maximizing the above-mentioned reward function, the machine prioritizes safety while reducing human-machine conflicts, and generates the final control action.

[0159] In the reinforcement learning process of the embodiment of the present application, the goal of training is to enable the machine to make the best control decisions in the actual environment. To this end, the machine will learn based on the previously defined safety rewards and collaboration rewards, and optimize the interaction between the driver and the machine by maximizing the total reward. Specifically, the machine will continuously adjust the control strategy to reduce the deviation from the driver's operation, ensuring that collisions can be effectively avoided during human-machine collaboration, while maintaining the driver's comfort and avoiding excessive intervention.

[0160] In this way, the control weight between the machine and the driver can be dynamically adjusted during the reinforcement learning process, so that the machine can quickly take over control when necessary, ensuring that the vehicle can smoothly avoid obstacles in potentially dangerous situations; and when the driver's collision avoidance ability is strong, the machine will reduce intervention and respect the driver's operating intention to the greatest extent. This dynamic adjustment process can reduce human-machine conflicts while ensuring that safety is the highest priority for every decision.

[0161] Ultimately, after reinforcement learning training, the machine can generate the final control action that conforms to the current environment and situation, thereby achieving more intelligent, precise and safe collision avoidance operations. This process greatly improves the performance of the system in complex driving situations, allowing intelligent vehicles to ensure driving safety in various traffic conditions and optimize the driving experience.

[0162] In practical applications, according to the principles and demonstrations provided in the above embodiments, a policy network that can be deployed in an ADAS system can be trained to generate a machine steering action u in a human-machine co-driving process in a critical scenario. m

[0163] Step S5, combining the human-machine control weight, the driver operation and the machine operation, generates a final operation to be executed, and sends the final operation to be executed to each actuator of the vehicle in real time through the communication system to execute the collision avoidance operation.

[0164] In an embodiment of the present application, based on the aforementioned calculation results, combined with the human-machine control weights, the driver's operation and the machine's operation, the final operation to be executed is generated, and the operation is transmitted in real time to each actuator of the vehicle through the communication system to ensure the execution of the corresponding collision avoidance operation.

[0165] Specifically, firstly, according to the human-machine control weight γ, machine operation u m And driver operation d , generate the final operation to be executed u f , the formula is:

[0166] u f =γ·u m +(1-γ)·u d

[0167] Through this formula, the degree of machine intervention can be dynamically adjusted according to the weight ratio in the current situation to ensure a collaborative balance between the machine and the driver.

[0168] Once the final pending operation u f Once the operation is calculated, the intelligent vehicle advanced driver assistance system will send the operation to the various actuators of the vehicle in real time through the communication system. The actuators include active differential steering systems, drive systems, and braking systems, which can perform corresponding collision avoidance actions based on the received operation signals. For example, when facing an obstacle, the intelligent vehicle advanced driver assistance system may instruct the steering system to make necessary adjustments, or instruct the braking system to slow down to avoid the obstacle.

[0169] In addition, during the execution process, the embodiment of the present application will continuously monitor sensor feedback, obtain vehicle status information in real time, and through these real-time feedback, continuously evaluate the relative position of the vehicle and the obstacle, and dynamically adjust the final operation according to the real-time changes to ensure that the vehicle can smoothly avoid the obstacle and return to the normal driving trajectory. This adjustment ensures that the vehicle can smoothly avoid obstacles and return to the normal driving trajectory during the collision avoidance process, avoiding secondary risks caused by excessive intervention or mistakes.

[0170] In the test of the embodiment of the present application, Figure 4The results shown reflect the outstanding performance of the embodiments of the present application in a human-machine shared control system. Figure 4 The first sub-figure shows that during the obstacle avoidance process, the human-machine shared control system using the present application is less affected by obstacle interference than the baseline driving mode, and within the same 6 seconds, the vehicle travels a distance of up to 4.5 meters, indicating that the present application can effectively improve the vehicle's driving efficiency and maintain good original driving task performance without sacrificing driving safety.

[0171] Figure 4 The second sub-figure shows the effect of yaw angle adjustment. The yaw angle adjustment under the present application method is smoother, avoiding the overshoot phenomenon that may occur in the traditional human control mode. This means that when encountering obstacles, the vehicle can adjust the driving trajectory more stably, thereby reducing the driver's operating burden and improving driving comfort and safety.

[0172] Figure 4 The third sub-graph shows the change in accessibility value during obstacle avoidance. Under the method of this application, the accessibility value of the vehicle steadily increases and stabilizes at a lower peak value when approaching an obstacle (approximately at t=2 seconds), which shows that the human-machine shared control system can respond to obstacles in advance and ensure safety. Compared with the baseline method, this application significantly reduces the response delay and avoids higher peaks, reflecting a more accurate and timely collision avoidance response.

[0173] In summary, the above test results prove that the application of the present invention in the human-machine shared control system can improve the obstacle avoidance performance while ensuring the safety of the vehicle is better than the traditional driving mode.

[0174] In order to implement the above embodiments, the present application also proposes an intelligent vehicle human-machine shared collision avoidance control system under critical conditions. Figure 5 The present application provides a schematic diagram of a structure of a human-machine shared collision avoidance control system for an intelligent vehicle under critical conditions. Figure 5 The system comprises:

[0175] The data acquisition module 100 collects and transmits the current state of the vehicle and the surrounding environment information through the sensor system, and the sensor system includes a camera, a millimeter wave radar, a combined inertial navigation IMU and related networking facilities;

[0176] The reachability analysis module 200 uses offline learning and large-scale vehicle data to approximately solve the collision avoidance unreachable state set, and performs offline iterative updates on the Hamilton-Jacobi reachability value function and the action-value function through a large amount of pre-collected real vehicle collision avoidance data to obtain the reachability value function and the action-value function;

[0177] The driver intention and capability evaluation module 300 quantifies the driver's collision avoidance capability and collision avoidance intention based on the accessibility value function and the action-value function, and dynamically allocates the human-machine control weight according to the quantified collision avoidance capability and collision avoidance intention;

[0178] The reinforcement learning control module 400 generates machine operations based on a reachability-inspired reinforcement learning algorithm, which combines the driver's collision avoidance intention, collision avoidance ability, and reachability information into the state space;

[0179] The control decision module 500 combines the human-machine control weight, the driver operation and the machine operation to generate the final operation to be executed, and sends the final operation to be executed to each actuator of the vehicle in real time through the communication system to perform the collision avoidance operation.

[0180] Regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0181] In order to implement the above embodiments, the present application also proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided by the above embodiments.

[0182] In order to implement the above embodiments, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided by the above embodiments.

[0183] In order to implement the above embodiments, the present application also proposes a computer program product, including a computer program, which implements the methods provided by the above embodiments when executed by a processor.

[0184] The collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this application are in compliance with relevant laws and regulations and do not violate public order and good morals.

[0185] It should be noted that personal information from users should be collected for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. In addition, such collection / sharing should be carried out after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign the agreement / authorization including authorization of relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others who have access to personal information data comply with its privacy policy and procedures.

[0186] The present application is expected to provide an implementation scheme for users to selectively block the use or access of personal information data. That is, the present disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by limiting data collection and deleting the data. In addition, when applicable, such personal information is de-identified to protect the privacy of the user.

[0187] In the description of the aforementioned embodiments, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0188] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0189] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0190] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.

[0191] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0192] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0193] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0194] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.

[0195] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this application can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution of this application can be achieved, and this document is not limited here.

[0196] The above specific implementations do not constitute a limitation on the protection scope of this application. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this application should be included in the protection scope of this application.

Claims

1. A human-machine shared collision avoidance control method for intelligent vehicles under critical conditions, characterized in that: include: Collect and transmit the vehicle's current state and surrounding environment information through a sensor system, wherein the sensor system includes a camera, a millimeter-wave radar, an integrated inertial navigation unit (IMU), and related networking facilities; Use offline learning and large-scale vehicle data to approximate the set of unreachable states for collision avoidance, and use a large amount of real vehicle collision avoidance data collected in advance to perform offline iterative updates on the Hamilton-Jacobi reachability value function and action-value function to obtain the reachability value function and action-value function. quantifying the driver's collision avoidance capability and collision avoidance intention based on the accessibility value function and the action-value function, and dynamically allocating human-machine control weights according to the quantified collision avoidance capability and collision avoidance intention; Generate machine actions based on a reachability-inspired reinforcement learning algorithm that incorporates driver collision avoidance intent, collision avoidance capability, and reachability information into a state space; The final operation to be executed is generated by combining the human-machine control weight, the driver operation and the machine operation, and the final operation to be executed is sent to each actuator of the vehicle in real time through the communication system to execute the collision avoidance operation.

2. The method according to claim 1, characterized in that The unreachable state set for collision avoidance is approximately solved using offline learning and large-scale vehicle data. The Hamilton-Jacobi reachability value function and action-value function are updated offline iteratively through a large amount of real vehicle collision avoidance data collected in advance to obtain the reachability value function and action-value function, including: The obstacle is represented as an elliptical envelope or a T-shaped collision surface, where the parameters of the target obstacle are recorded as [X0, Y0, a, b], which represent the center coordinates of the obstacle and the elliptical shape parameters respectively; The target obstacle is represented as a safety constraint state set T, and the vehicle state in this set is considered to be a dangerous state where a collision has occurred, which is defined as: in, Represents the state and kinematic information of the vehicle, X, T are the global position coordinates of the vehicle, is the yaw angle, v x , v y are the longitudinal velocity and the lateral velocity respectively, and r represents the yaw rate; Substituting the vehicle state and driver operation in a large amount of real vehicle collision avoidance data collected in advance into the reachability value function network and action value function network corresponding to the safety constraint state set T, and calculating the sample reachability value function value and the sample action value function value; The sample reachability value function and the sample action value are updated offline iteratively to further optimize the accuracy of the collision avoidance decision. After the iteration, the updated reachability value function V is obtained. h Network and action-value function Q h network and deploy it in smart vehicle advanced driver assistance systems.

3. The method according to claim 2, characterized in that Based on the accessibility value function and the action-value function, the collision avoidance capability and collision avoidance intention of the driver are quantified, and according to the quantified collision avoidance capability and collision avoidance intention, a human-machine control weight is dynamically allocated, including: Substitute the current vehicle state, obstacle envelope and driver operation collected by the sensor system into the updated accessibility value function V corresponding to the constraint state set T h Network and action-value function Q h Network, get the reachability value V h (x) and the action-value Q h (x,u d ); Through the accessibility value V h (x) represents the optimal collision avoidance distance in the current state, and calculates the driver's collision avoidance ability CAA, which is calculated by the following formula: Among them, α CAA is the sensitivity adjustment parameter, C CAA is the offset constant; By the driver's current operation u d The corresponding action-value Q h (x,u d ) The reachability value V corresponding to the best action h (x), calculate the driver's collision avoidance intention CAI, which is calculated by the following formula: According to the collision avoidance capability CAA and the collision avoidance intention CAI, the human-machine control weight γ is dynamically allocated, and the control weight is calculated by the following formula: γ=max(γ min ,(1-s CAI )(1-s CAA )) Among them, s CAI With s CAA is the normalization function, which is defined by the following formulas: Among them, k CAI1 and k CAA1 Separate control CAI and CAA Sensitivity to changes in CAI and CAA, k CAI2 and k CAA2 Definitions CAI and CAA Reaching a threshold of 0.5 represents the balance point between collision avoidance intention and capability.

4. The method according to claim 3, characterized in that Generating machine actions based on a reachability-inspired reinforcement learning algorithm that incorporates the driver's collision avoidance intention, collision avoidance capability, and reachability information into a state space, including: The driver's collision avoidance intention, collision avoidance ability, accessibility information and other vehicle dynamic states are explicitly integrated into the state space of reinforcement learning, which includes: the vehicle's dynamic state x, the accessibility value function V h (x), the driver’s collision avoidance ability CAA, the driver’s collision avoidance intention CAI, the human-machine control weight γ, and the control operation u of the driver and the machine d and u m , where the action space u m Only the front wheel steering angle δ mf ; In the state space, combining the dynamic and kinematic states of the vehicle, a driver's collision avoidance operation during training is simulated by constructing a driver action generation model, wherein the driver action generation model simulates the driver's collision avoidance decision by dynamically adjusting the relative relationship between the driver input and the obstacle; Constructing a reinforcement learning reward function based on the driver's collision avoidance intention, wherein the reward function is designed based on the goal of reducing human-machine conflict and maintaining the original task performance, and the reward function combines the obstacle envelope as a constraint to guide the generation of machine control actions during the reinforcement learning process; During the reinforcement learning training process, by maximizing the above reward function, the machine can prioritize safety while reducing human-machine conflicts and generate the final control action.

5. The method according to claim 4, characterized in that In the state space, the driver's collision avoidance operation during training is simulated by building a driver action generation model in combination with the dynamic and kinematic states of the vehicle. The driver action generation model dynamically adjusts the relative relationship between the driver input and the obstacle to simulate the driver's collision avoidance decision, including: The current vehicle state, obstacle envelope, and driver’s control operation input are substituted into the driver action generation model. By calculating the relative position and speed between the vehicle and the obstacle, the driver’s steering operation is calculated and adjusted so that the driver’s preview direction is aligned with the tangent of the obstacle boundary, thereby ensuring that the vehicle can avoid the obstacle. The preview angle θ c Calculated according to the following formula: Among them, (X trg , Y trg ) represents the tangent point on the ellipse, located on the boundary of the ellipse defined as follows: Where λ is a scaling factor based on the accessibility value V h (x) and the collision avoidance intention CAI adjustment ellipse size, defined as: Among them, k λ is a hyperparameter.

6. The method according to claim 5, characterized in that The reinforcement learning reward function is constructed according to the driver's collision avoidance intention. The reward function is designed based on the goal of reducing human-machine conflict and maintaining the original task performance. The reward function combines the obstacle envelope as a constraint condition to guide the generation of machine control actions during the reinforcement learning process, including: The cooperation between the driver and the machine is optimized by reinforcement learning reward function to reduce human-machine conflict. The reward function includes a safety reward R sf and collaboration reward R co ; The safety reward R sf It is calculated based on the distance between the vehicle state and the set of unreachable states for collision avoidance, and is defined as: Among them, d0 is the scaling parameter, k sf1 To emphasize the negative weighting factor of the importance of being far away from the boundary of the intelligent vehicle advanced driver assistance system, V h (x)>k sf2 States that result in significant penalties ensure that the agent learns to avoid unsafe states; The collaboration reward R co Optimize human-machine collaboration by quantifying the deviation between machine and driver behavior, defined as: R co =-k co ·γ·(in M -in D ) 2 Among them, the penalty term (u M -u D ) 2 is the deviation between the machine and the driver’s behavior, k co is a scaling factor, and the weighting factor γ reflects the necessity of machine intervention based on the driver’s collision avoidance ability and intention.

7. The method according to claim 6, characterized in that The method combines the human-machine control weight, the driver operation and the machine operation to generate a final operation to be executed, and sends the final operation to be executed to each actuator of the vehicle in real time through a communication system to execute a collision avoidance operation, including: According to the human-machine control weight γ, machine operation u m And driver operation d , generate the final operation to be executed u f , the formula is: in f =γ·u m +(1-γ)·u d The final operation to be performed u is transmitted through the communication system f Sent in real time to the vehicle's actuators, including the active differential steering system, drive system and braking system, to perform corresponding collision avoidance operations; During the execution process, the vehicle's status is monitored in real time through continuous sensor feedback, and the final operation is dynamically adjusted according to the relative position of the vehicle and the obstacle to ensure that the vehicle successfully avoids the obstacle and returns to its normal driving trajectory.

8. An intelligent vehicle human-machine shared collision avoidance control system under critical conditions, characterized in that: include: A data acquisition module, which is used to collect and transmit information about the vehicle's current state and surrounding environment through a sensor system, wherein the sensor system includes a camera, a millimeter-wave radar, an integrated inertial navigation unit (IMU), and related networking facilities; The reachability analysis module is used to use offline learning and large-scale vehicle data to approximately solve the set of unreachable states for collision avoidance, and to perform offline iterative updates on the Hamilton-Jacobi reachability value function and action-value function through a large amount of pre-collected real vehicle collision avoidance data to obtain the reachability value function and action-value function; A driver intention and capability evaluation module, used to quantify the driver's collision avoidance capability and collision avoidance intention based on the accessibility value function and the action-value function, and dynamically allocate a human-machine control weight according to the quantified collision avoidance capability and collision avoidance intention; A reinforcement learning control module for generating machine operations based on a reachability-inspired reinforcement learning algorithm that incorporates the driver's collision avoidance intent, collision avoidance capability, and reachability information into a state space; The control decision module is used to combine the human-machine control weight, the driver operation and the machine operation to generate a final operation to be executed, and send the final operation to be executed to each actuator of the vehicle in real time through the communication system to execute the collision avoidance operation.

9. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Intervention type sharing control method and device for autonomous vehicle in forward collision avoidance scene

    CN115923845A

  • Off-line reinforcement learning method for generating security policy and related component

    CN117494833A

  • Man-machine co-driving switching control method based on trajectory prediction

    CN117601857A

  • Driver permission allocation strategy based on fuzzy reasoning algorithm

    CN118833254A

  • Automatic control method and system for the virtual confinement of a land vehicle within a track

    WO2024003674A1