An intelligent vehicle safety decision-making method and system based on dynamic perception
By collecting vehicle dynamic characteristics and human preference data, a hierarchical decision-making structure and transfer learning mechanism are built, and the problem of adaptability of intelligent connected vehicles in different models and user preferences is solved, the quality of safety decisions and user acceptance is improved, and the deployment cost and intervention rate are reduced.
Patent Information
- Application Number
- CN202510453293.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing intelligent connected vehicle safety decision-making system is difficult to adapt to the characteristics of different models, and cannot accurately capture human complex preferences for safety decisions, resulting in low user acceptance and unnecessary frequent intervention.
The vehicle dynamic characteristic data is collected through sensors to generate vehicle model feature vectors, build a hierarchical decision structure (strategy layer, tactical layer, execution layer), combine implicit and explicit preference data to build short-term and long-term planning preference models, and use an experience migration network to migrate decision knowledge to the target vehicle model, generating the optimal trajectory and control actions.
It quickly adapts to the dynamic characteristics of different models, improves user acceptance and the safety performance of the system, reduces deployment costs and intervention rates, and enhances the personalization level and learning efficiency of the system.
Smart Images

Figure CN119953390B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent connected vehicles, and more specifically, to a vehicle type intelligent safety decision-making method and system based on dynamic perception. Background Art
[0002] With the development of artificial intelligence and autonomous driving technologies, the safety decision-making system of intelligent connected vehicles has become a research hotspot. The existing safety decision-making systems mainly have the following technical problems: It is difficult for traditional safety decision-making systems to autonomously learn the optimal safety strategy according to the characteristics of different vehicle types. It is difficult for traditional reinforcement learning systems to directly learn safety decision-making strategies that conform to human preferences from environmental feedback. Especially in scenarios where safety and comfort need to be balanced, relying solely on environmental reward signals cannot accurately capture the complex preferences of humans for safety decisions, resulting in the decision-making strategies generated by the system meeting the safety indicators but having low user acceptance and high intervention rates. The existing human feedback learning systems mainly focus on immediate feedback and ignore the preference learning for long-term planned trajectories. This leads to the situation that during long-term driving, although the local control actions meet the safety requirements, the overall trajectory planning does not conform to the driver's preferences, resulting in unnecessary frequent interventions and unsatisfactory experiences.
[0003] Therefore, there is a need for an intelligent safety decision-making method that can automatically adapt to the characteristics of different vehicle types and at the same time consider human short-term control preferences and long-term planning preferences, so as to improve the adaptability of the system in a multi-vehicle type environment and user acceptance. Summary of the Invention
[0004] The present invention provides a vehicle type intelligent safety decision-making method and system based on dynamic perception, which solves the problems of safety decision generalization for different vehicle type characteristics and human preference learning in related technologies, and improves the technical problems of the vehicle type adaptability of the system and user acceptance.
[0005] The present invention provides a vehicle type intelligent safety decision-making method based on dynamic perception, including the following steps:
[0006] Collect vehicle dynamic characteristic data through sensors to generate a vehicle type feature vector;
[0007] Based on the vehicle type feature vector, construct a hierarchical decision-making structure including a strategy layer, a tactic layer, and an execution layer, where the strategy layer is responsible for generating long-term planning goals, the tactic layer is responsible for generating mid-term maneuver decisions, the execution layer is responsible for generating short-term control actions, and the vehicle type feature vector is used as the common input parameter of each layer of neural network model to adjust the decision output of each layer to adapt to the dynamic characteristics of different vehicle types;
[0008] Collect human preference data including implicit preference data and explicit preference data;
[0009] Construct a short - term control preference model and a long - term planning preference model based on human preference data;
[0010] Use an experience transfer network to transfer the decision - making knowledge of the source vehicle model to the target vehicle model, generate multiple candidate trajectories, evaluate and select the optimal trajectory through the long - term planning preference model, and generate specific control actions based on the short - term control preference model.
[0011] In a preferred embodiment, the vehicle dynamic characteristic data includes the center - of - mass height, steering sensitivity, braking force distribution characteristics, and roll rate;
[0012] Process the data through a feature extraction algorithm to generate a vehicle model feature vector , where:
[0013] ;
[0014] Among them, represents the vehicle model feature vector, represents the set of vehicle model feature parameters, , , respectively represent the , , th vehicle model feature parameters, represents the total number of parameters, represents the vehicle model feature vector generation function, , , respectively represent the , , th elements in the vector, represents the vector dimension.
[0015] In a preferred embodiment, in the hierarchical decision - making structure:
[0016] The policy - layer neural network model generates long - term planning goals, satisfying:
[0017] ;
[0018] The tactical - layer neural network model generates mid - term maneuver decisions, satisfying:
[0019] ;
[0020] The execution - layer neural network model generates short - term control actions, satisfying:
[0021] ;
[0022] Among them, represents the long-term planning goal generated by the policy layer, represents the mid-term maneuver decision, represents the short-term control action, represents the current environmental state, represents the vehicle model feature vector, , and respectively represent the neural network models of the policy layer, the tactical layer and the execution layer.
[0023] In a preferred embodiment, the human preference data collection includes:
[0024] Implicit preference data collection: Recording driver operation intervention, steering wheel grip strength, pedal operation characteristics and eye movement tracking data to generate an implicit preference feature vector ;
[0025] Explicit preference data collection: Generating candidate trajectory pairs for the driver to select and constructing an explicit preference data set , satisfying:
[0026] ;
[0027] Among them, represents the explicit preference data set, represents the th trajectory pair, represents the corresponding preference label, represents the number of samples in the training data set.
[0028] In a preferred embodiment, the construction of the short-term control preference model includes:
[0029] Based on the implicit preference data, constructing a short-term control preference reward function , satisfying:
[0030] ;
[0031] Training the reward function parameters by maximizing the following optimization objective :
[0032] ;
[0033] Among them, represents the short-term control preference reward function, represents optimizing the reward parameter , represents a set of data sampled from the implicit preference data distribution , represents the environmental state, Represents a control action, Represents an implicit preference feature vector, Represents a vehicle model feature vector, Represents an action sample that the driver does not prefer, Represents the sigmoid function, Represents the state Take an unpreferred action The reward value of, Represents the logarithmic function.
[0034] In a preferred embodiment, the construction of the long-term planning preference model includes:
[0035] Based on the explicit preference data set, construct a long-term planning preference function , satisfying:
[0036] ;
[0037] Train the preference function parameters by maximizing the following likelihood function :
[0038] ;
[0039] Wherein, Represents the probability that trajectory is more preferred than trajectory given the vehicle model feature , and represent two candidate trajectories, Represents the sigmoid function, Represents optimizing the parameters of the preference function, Represents the preference function, Represents a trajectory feature extraction network, Represents the th preference label of the trajectory pair, Represents the logarithmic function, Represents the number of samples in the training data set.
[0040] In a preferred embodiment, the knowledge transfer includes:
[0041] Construct an experience transfer network , and transfer the decision-making knowledge of the source vehicle model to the target vehicle model, satisfying:
[0042] ;
[0043] Train the transfer network parameters by minimizing the following loss function :
[0044] ;
[0045] Among them, and respectively represent the decision model parameters of the source vehicle model and the target vehicle model, and respectively represent the feature vectors of the source vehicle model and the target vehicle model, represents the interaction data of the target vehicle model, represents the action value function of the target vehicle model, represents the experience transfer network, represents the parameters of the transfer network to be optimized, represents the state, action, reward, next state transition sample, represents the discount factor.
[0046] In a preferred embodiment, the decision generation includes:
[0047] Based on the current environmental state and the vehicle model feature vector, generate multiple candidate trajectories that satisfy the safety constraints;
[0048] Evaluate the priorities of each candidate trajectory through the long-term planning preference model, and select the optimal trajectory that satisfies:
[0049] ;
[0050] where represents the optimal trajectory with the highest evaluation score, , , respectively represent the , , th candidate trajectories, represents the total number of candidate trajectories, represents the preference evaluation score of the trajectory;
[0051] Based on the optimal trajectory and the short-term control preference model, generate a control action that satisfies:
[0052] ;
[0053] where, represents the control action with the highest reward value in the set of available actions, represents searching in the set of available actions ) to find the action that maximizes the short-term control preference reward function , represents the current environmental state, Represents a short-term control preference reward function.
[0054] In a preferred embodiment, the information interaction mechanism between the layers in the hierarchical decision-making structure includes:
[0055] The upper-layer decision serves as a constraint condition for the lower-layer decision. From the th layer to the th layer, the downward information flow satisfies:
[0056] ;
[0057] The execution result of the lower-layer decision serves as feedback information for the upper-layer decision. From the th layer to the th layer, the upward information flow satisfies:
[0058] ;
[0059] Among them, represents the downward information flow from the th layer to the th layer, represents the upward information flow from the th layer to the th layer, represents the decision result of the th layer, represents the execution feedback of the th layer, and respectively represent the downward and upward information processing functions, represents the vehicle model feature vector.
[0060] In a preferred embodiment, a vehicle model intelligent safety decision-making system based on dynamic perception is used to implement a vehicle model intelligent safety decision-making method based on dynamic perception, and is characterized by including:
[0061] A vehicle model feature dynamic perception module, configured to collect vehicle dynamic characteristic data and generate a vehicle model feature vector;
[0062] A hierarchical decision-making structure module, based on the vehicle model feature vector, constructs a hierarchical decision-making structure including a strategy layer, a tactic layer, and an execution layer, where the strategy layer is responsible for generating long-term planning goals, the tactic layer is responsible for generating medium-term maneuver decisions, the execution layer is responsible for generating short-term control actions, and the vehicle model feature vector serves as a common input parameter for the neural network models of each layer to adjust the decision output of each layer to adapt to the dynamic characteristics of different vehicle models;
[0063] A human preference data collection module, configured to collect human preference data including implicit preference data and explicit preference data;
[0064] A preference learning module for constructing a short-term control preference model and a long-term planning preference model based on the human preference data;
[0065] A transfer learning module for transferring the decision-making knowledge of the source vehicle model to the target vehicle model;
[0066] A decision generation module for generating multiple candidate trajectories, evaluating and selecting the optimal trajectory, and generating specific control actions.
[0067] The beneficial effects of the present invention are as follows:
[0068] Strong vehicle model adaptability: Through the vehicle model feature dynamic perception and transfer learning mechanism, the system can quickly adapt to the dynamic characteristics of different vehicle models. Experimental results show that for a new vehicle model, the adaptation time of the safety strategy is reduced by 85%, without the need to re-collect a large amount of data and retrain the model, significantly reducing the deployment cost and debugging workload.
[0069] High decision-making quality: The decision-making strategy integrating human preferences can take into account the riding comfort while ensuring safety. Experimental results show that compared with traditional methods, the success rate of the present invention in handling dangerous scenarios is increased by 31%, while reducing 35% of unnecessary interventions, improving the system reliability and robustness.
[0070] High user acceptance: Through the dual-time-scale preference learning mechanism, the system can capture the driver's preferences for both immediate control actions and long-term planning trajectories simultaneously, and the generated decisions are more consistent with human expectations. Experimental results show that in the human-machine collaboration scenario, the user acceptance is increased by 65%, significantly reducing the discomfort caused by the system intervention to the driver.
[0071] High personalization degree: The present invention can generate customized safety decision-making strategies according to different vehicle model characteristics and driver preferences, making the decisions more in line with the specific scenario requirements. The system can identify the personalized preferences of different drivers and optimize the decisions based on this, enhancing the adaptability of the system.
[0072] High learning efficiency: By adopting a hierarchical decision-making structure and a transfer learning method, the system can effectively learn from limited samples. Experimental results show that compared with traditional methods, the sample utilization efficiency of the present invention is increased by 3.2 times, significantly reducing the data demand and accelerating the system training and optimization process. Description of the Drawings
[0073] Figure 1 is a flowchart of a vehicle model intelligent safety decision-making method based on dynamic perception of the present invention;
[0074] Figure 2 is a detailed flowchart of generating a vehicle model feature vector by collecting vehicle dynamic characteristic data through sensors of the present invention;
[0075] Figure 3 is a detailed flowchart for constructing a hierarchical decision-making structure of the present invention;
[0076] Figure 4 is a detailed flowchart for collecting human preference data of the present invention;
[0077] Figure 5 is a detailed flowchart for constructing a short-term control preference model and a long-term planning preference model of the present invention;
[0078] Figure 6 is a detailed flowchart for generating multiple candidate trajectories and selecting optimal and specific control actions of the present invention. Detailed implementation manners
[0079] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.
[0080] In at least one embodiment of the present invention, a vehicle type intelligent safety decision-making method based on dynamic perception is disclosed. As Figures 1 to 6 shown, it includes the following steps:
[0081] Step 1: Collect vehicle dynamic characteristic data through sensors and generate a vehicle type feature vector;
[0082] Specifically, it includes the following steps:
[0083] Step 1.1: Sensor data collection;
[0084] Use various sensors mounted on the vehicle to collect vehicle dynamic characteristic data, including:
[0085] Body attitude data collected by an inertial measurement unit (IMU), such as the height of the center of mass, tilt angle, roll rate, etc.;
[0086] Braking force distribution characteristic data collected by wheel speed sensors;
[0087] Steering response characteristic data collected by steering system sensors;
[0088] Suspension system response characteristic data collected by acceleration sensors.
[0089] Step 1.2: Vehicle type feature extraction;
[0090] The collected sensor data is processed using an application feature extraction algorithm to extract a vehicle type feature parameter set :
[0091] ;
[0092] Among them, represents the vehicle type feature parameter set, 、 、 respectively represent the 、 、 rd, represents the total number of parameters.
[0093] Step 1.3, generation of vehicle type feature vectors;
[0094] Based on the extracted vehicle type feature parameter set , a vehicle type feature vector is constructed:
[0095] ;
[0096] Among them, represents the vehicle type feature vector, represents the vehicle type feature parameter set, represents the vehicle type feature vector generation function, 、 、 respectively represent the 、 、 th elements in the vector, represents the vector dimension.
[0097] Step 2, based on the vehicle type feature vector, construct a hierarchical decision-making structure including a strategy layer, a tactic layer, and an execution layer. The strategy layer is responsible for generating long-term planning goals, the tactic layer is responsible for generating medium-term maneuver decisions, and the execution layer is responsible for generating short-term control actions. The vehicle type feature vector is used as the common input parameter for each layer of the neural network model to adjust the decision output of each layer to adapt to the dynamic characteristics of different vehicle types;
[0098] Specifically, it includes the following steps:
[0099] Step 2.1, construction of the strategy layer;
[0100] A strategy layer neural network model is established, which is responsible for generating long-term planning goals:
[0101] ;
[0102] Among them, Represents the current environmental state, represents the vehicle model feature vector, represents the generated long-term planning goal, represents that the parameter is of the policy layer neural network model.
[0103] The policy layer neural network model is specifically implemented as a multi-layer perceptron structure, including:
[0104] Input layer: Receives the environmental state (including vehicle position, speed, surrounding obstacle information, etc.) and the vehicle model feature vector ;
[0105] Hidden layer: Contains 3 fully connected layers, with 256, 128, and 64 neurons in each layer respectively, and uses the ReLU activation function;
[0106] Output layer: Generates the long-term planning goal , including information such as the target position, desired speed, and desired driving trajectory.
[0107] In the highway lane-changing scenario, the policy layer neural network model receives the traffic flow information of the current lane and the target lane and the vehicle model feature vector, and outputs the overall planning goal for completing the lane-changing task, such as the target lane position, lane-changing completion time, and desired speed, etc.
[0108] Step 2.2, Tactical layer construction;
[0109] Establish a tactical layer neural network model , responsible for generating mid-term maneuver decisions:
[0110] ;
[0111] Among them, represents the generated mid-term maneuver decision, represents that the parameter is of the tactical layer neural network model, represents the current environmental state, represents the generated long-term planning goal, represents the vehicle model feature vector.
[0112] The tactical layer neural network model is implemented as a recurrent neural network structure combined with an attention mechanism, including:
[0113] Input processing unit: Fuses the features of the environmental state , long-term planning goal and vehicle model feature vector ;
[0114] Attention layer: Based on the vehicle type feature vector assigns different weights to different environmental factors, enabling the model to focus on features more relevant to a specific vehicle type;
[0115] Recurrent unit: Adopts long short-term memory (LSTM) units to process temporal information and remember past decisions;
[0116] Output layer: Generates mid-term maneuver decisions , including lane change timing, acceleration and deceleration strategies, etc.
[0117] In the scenario of turning on urban roads, the tactical layer neural network model receives the turning target provided by the strategy layer, combines the current intersection traffic conditions and the vehicle type feature vector to determine mid-term maneuver strategies such as turning timing, turning speed, and turning trajectory.
[0118] Step 2.3, Execution layer construction;
[0119] Build an execution layer neural network model , responsible for generating short-term control actions:
[0120] ;
[0121] Among them, represents the generated short-term control action, represents the execution layer neural network model with parameters , represents the current environmental state, represents the generated mid-term maneuver decision, represents the vehicle type feature vector.
[0122] The execution layer neural network model is implemented as a hybrid structure combining a feedforward neural network and a vehicle dynamics model, including:
[0123] Input layer: Receives the current state , mid-term maneuver decision and vehicle type feature vector ;
[0124] Hidden layer: Two fully connected layers, each with 128 and 64 neurons respectively, using the LeakyReLU activation function;
[0125] Vehicle dynamics constraint layer: Establishes constraint conditions based on the vehicle type feature vector to ensure that the generated control actions conform to the vehicle's physical characteristics;
[0126] Output layer: Generates specific control actions , including steering angle, acceleration, and braking force, etc.
[0127] In an emergency obstacle avoidance scenario, the execution layer neural network model receives the obstacle avoidance strategy provided by the tactical layer, combines the current vehicle state and the vehicle type feature vector, generates precise steering angle and braking force commands, and realizes safe obstacle avoidance while maintaining vehicle stability.
[0128] Step 2.4, the inter-layer information interaction mechanism;
[0129] Establish an information interaction mechanism between each decision-making layer. The upper-layer decision serves as a constraint condition for the lower-layer decision, and the execution result of the lower-layer decision serves as feedback information for the upper-layer decision:
[0130] ;
[0131] ;
[0132] Among them, represents the downward information flow from the th layer to the th layer, represents the decision result of the th layer, represents the upward information flow from the th layer to the th layer, represents the execution feedback of the th layer, and respectively represent the downward and upward information processing functions.
[0133] Step 3, collect human preference data including implicit preference data and explicit preference data;
[0134] Specifically, it includes the following steps:
[0135] Step 3.1, implicit preference data collection;
[0136] Build an implicit preference data collection module to record the driver's physiological and operation response data:
[0137] Driver operation intervention data: Record the driver's intervention behavior in the system decision-making, such as taking over control, correcting the steering wheel, etc.;
[0138] Steering wheel grip force data: Monitor the change of the driver's grip force through the pressure sensor on the steering wheel;
[0139] Pedal operation characteristic data: Record the operation force, speed and other characteristics of the accelerator pedal and the brake pedal;
[0140] Eye movement tracking data: Collect the driver's physiological response data such as the line of sight fixation point and pupil dilation.
[0141] Apply a processing algorithm to the collected implicit preference data to generate an implicit preference feature vector :
[0142] ;
[0143] Among them, represents the implicit preference feature vector, represents the original implicit preference data, represents the implicit preference feature extraction function, represents the vehicle model feature vector.
[0144] Step 3.2, Explicit preference data collection;
[0145] Build an explicit preference data collection module to regularly generate candidate trajectory pairs for the driver to select:
[0146] Use the current decision model to generate multiple candidate trajectories that meet safety constraints , among which, , , respectively represent the , , th candidate trajectory, represents the total number of candidate trajectories;
[0147] Randomly select two trajectories from the candidate trajectories and , to form a trajectory pair ;
[0148] Display the simulated execution effect of the trajectory pair to the driver through the in-vehicle human-machine interface;
[0149] Record the driver's preference selection to form preference data marked as or , indicating or .
[0150] Build an explicit preference data set :
[0151] ;
[0152] Among them, represents the explicit preference data set, represents the th trajectory pair, represents the corresponding preference label, represents , represents , Indicates the size of the dataset.
[0153] Step 4: Collect human preference data including implicit preference data and explicit preference data;
[0154] The specific steps are as follows:
[0155] Step 4.1: Construct a short-term control preference model;
[0156] Based on the implicit preference data , construct a short-term control preference reward function :
[0157] ;
[0158] Where represents the short-term control preference reward function, represents the environmental state, represents the control action, represents the implicit preference feature vector, represents the vehicle model feature vector, represents the parameters of the reward function.
[0159] The short-term control preference model is implemented as a two-tower neural network structure, specifically including:
[0160] State-action processing tower:
[0161] Input layer: Receive the environmental state and the control action ;
[0162] Feature extraction layer: Consists of 3 fully connected layers, with dimensions of 128, 64, and 32 layer by layer;
[0163] Fusion layer: Combine state and action features.
[0164] Preferred vehicle model processing tower:
[0165] Input layer: Receive the implicit preference feature vector and the vehicle model feature vector ;
[0166] Feature extraction layer: Consists of 2 fully connected layers, with dimensions of 64 and 32.
[0167] Cross-attention layer: Calculate the attention weights between the two towers and fuse the information of the two towers;
[0168] Output layer: Generate the preference reward value.
[0169] In the vehicle following scenario, the short-term control preference model receives states such as the current vehicle distance and relative speed, and control actions such as acceleration. Combining the driver's previous operation intervention mode and vehicle type characteristics, it evaluates the comfort preference score of the action and guides the system to generate a following strategy that better conforms to the user's habits.
[0170] Reward function parameters It is trained by maximizing the following optimization objective:
[0171] ;
[0172] Where, represents optimizing the reward parameter , represents a set of data sampled from the implicit preference data distribution , represents the environmental state, represents the control action, represents the implicit preference feature vector, represents the vehicle type feature vector, represents the action samples that the driver does not prefer, represents the sigmoid function, represents the state when taking the non-preferred action of the reward value, represents the logarithmic function.
[0173] Step 4.2, Construction of the long-term planning preference model;
[0174] Based on the explicit preference data set , construct the long-term planning preference function :
[0175] ;
[0176] Where, represents the probability that, given the vehicle type feature , the trajectory is more preferred than the trajectory , and represent two candidate trajectories, represents the vehicle type feature vector, represents the trajectory feature extraction network with parameter , represents the sigmoid function.
[0177] The long-term planning preference model is implemented as a comparison network structure based on Transformer, specifically including:
[0178] Trajectory Encoder:
[0179] Input layer: Receives a sequence of trajectories and , each trajectory containing a sequence of state-action pairs;
[0180] Position Encoding: Adds temporal position information;
[0181] Transformer Encoder: Consists of 4 self-attention layers to extract trajectory features.
[0182] Vehicle Model Condition Module:
[0183] Input layer: Receives a vehicle model feature vector ;
[0184] Mapping layer: Maps the vehicle model features to a conditional vector.
[0185] Comparison Module:
[0186] Cross-attention layer: Calculates the correlation between the two trajectory encodings and the vehicle model condition;
[0187] Fusion layer: Combines the trajectory features and the vehicle model condition information;
[0188] Output layer: Outputs a preference probability value.
[0189] In the scenario of passing through a complex intersection, the long-term planning preference model evaluates different trajectory plans for passing through the intersection (such as passing directly, decelerating and waiting, etc.), and selects the passing strategy that best meets the driver's long-term comfort expectations based on the previously collected driver preference data and vehicle model characteristics.
[0190] Preference function parameters Are trained by maximizing the following likelihood function:
[0191] ;
[0192] where denotes optimizing the parameters of the preference function , denotes the probability that trajectory is more preferred than trajectory given the vehicle model features , denotes the preference function, denotes the trajectory feature extraction network, denotes the th preference label of the trajectory pair, denotes the logarithmic function, denotes the number of samples in the training dataset.
[0193] Step 4.3, preference model integration;
[0194] Integrate the short-term control preference model and the long-term planning preference model into a unified preference evaluation system:
[0195] ;
[0196] Among them, represents the preference evaluation system, represents the complete trajectory, represents the reference trajectory, represents the discount factor, represents the vehicle model feature vector, represents the weight coefficient, which is used to balance the importance of the short-term control preference and the long-term planning preference, represents the short-term control preference model, represents the long-term planning preference model, represents the logarithmic function.
[0197] Step 5, use the experience transfer network to transfer the decision-making knowledge of the source vehicle model to the target vehicle model, generate multiple candidate trajectories, evaluate and select the optimal trajectory through the long-term planning preference model, and generate specific control actions based on the short-term control preference model;
[0198] The specific steps are as follows:
[0199] Step 5.1, construction of the transfer learning model;
[0200] Construct an experience transfer network , which is used to transfer the decision-making knowledge of the source vehicle model to the target vehicle model:
[0201] ;
[0202] Among them, and respectively represent the decision model parameters of the source vehicle model and the target vehicle model, and respectively represent the feature vectors of the source vehicle model and the target vehicle model, represents a transfer network with parameter .
[0203] The experience transfer network is implemented as a parameter adjustment network of a meta-learning architecture, specifically including:
[0204] Vehicle model feature comparison module:
[0205] Input layer: Receive the source vehicle model feature vector and the target vehicle model feature vector ;
[0206] Feature difference extraction layer: Calculate the similarity and difference features between vehicle models;
[0207] Output layer: Generate a feature mapping matrix .
[0208] Parameter mapping module:
[0209] Input layer: Receive the source vehicle model parameters ;
[0210] Parameter grouping layer: Group and process the parameters according to functions and levels;
[0211] Adjustment layer: Transform the source parameters according to the feature mapping matrix ;
[0212] Output layer: Generate the target vehicle model parameters .
[0213] In the new vehicle deployment scenario, assuming there is a safety decision-making model for sedans, and now it needs to be adapted to SUV models. The experience transfer network analyzes the feature differences (such as the centroid height, steering response, etc.) between the two vehicle models, automatically adjusts the decision-making model parameters to adapt to the dynamic characteristics of SUV models, so as to quickly generate a safety decision-making strategy suitable for SUV models without the need for a large amount of retraining.
[0214] Transfer network parameters Are trained by minimizing the following loss function:
[0215] ;
[0216] Where Represents the interaction data of the target vehicle model, Represents the action value function of the target vehicle model, Represents the experience transfer network, Represents the parameters of the transfer network Are optimized, Represents the state, action, reward, next state transition sample, Represents the discount factor.
[0217] Step 5.2, Candidate trajectory generation;
[0218] Based on the current environmental state And the vehicle model feature vector , Use a hierarchical decision-making structure to generate multiple candidate trajectories that meet safety constraints:
[0219] ;
[0220] Where 、 , respectively represent the , , th candidate trajectories, represents the total number of candidate trajectories, represents the trajectory generation function, represents the safety constraint condition, represents the vehicle type feature vector.
[0221] Step 5.3, Optimal trajectory selection;
[0222] Evaluate the priorities of each candidate trajectory based on the long-term planning preference model and select the optimal trajectory:
[0223] ;
[0224] Among them, represents the preference evaluation score of the trajectory , , , respectively represent the , , th candidate trajectories, represents the total number of candidate trajectories, represents the selected optimal trajectory, represents the vehicle type feature vector.
[0225] Step 5.4, Control action generation;
[0226] Based on the selected optimal trajectory and the short-term control preference model, generate specific control actions:
[0227] ;
[0228] Among them, represents the set of actionable actions that conform to the trajectory , represents the generated control action, represents selecting an action in the actionable action space such that the objective function achieves the maximum value.
[0229] Real application example of this embodiment:
[0230] The example involves safety decision adaptation for four different vehicle types, including sedans, SUVs, MPVs, and pickups. The application scenarios cover the following typical driving situations:
[0231] Highway lane change;
[0232] Emergency obstacle avoidance;
[0233] Vehicle following deceleration;
[0234] Ramp merging;
[0235] In these scenarios, the system needs to generate decision-making strategies that are both safe and comfortable based on vehicle type characteristics and driver preferences. The following will detail the implementation process and effect verification of the system in these scenarios.
[0236] Implementation example of dynamic perception of vehicle type characteristics:
[0237] Collect dynamic characteristic data for four vehicle types. The main sensor data collection results for each vehicle type are shown in Table 1:
[0238] Table 1: Main sensor data collection results for different vehicle types;
[0239]
[0240] Apply the vehicle type characteristic extraction algorithm to process the collected sensor data. The characteristic vectors after dimensionality reduction of the generated vehicle type characteristic vectors (simplified to 6-dimensional vectors) are shown in Table 2:
[0241] Table 2: Characteristic vectors of different vehicle types;
[0242]
[0243] These vehicle type characteristic vectors are used as input parameters for the hierarchical decision-making structure to adjust the decision-making strategy to adapt to different vehicle type characteristics.
[0244] Implementation example of human preference data collection:
[0245] Implicit preference data collection:
[0246] The system collected the operation response data of multiple drivers on different vehicle types during actual road tests. The intervention frequency and operation intensity of drivers on SUVs and pickups are significantly higher than those on sedans and MPVs, indicating that drivers have different control preferences and comfort expectations for different vehicle types.
[0247] Explicit preference data collection:
[0248] The system shows candidate trajectory pairs in different scenarios to the driver through the human-machine interface and records the driver's preference selections. The explicit preference data collection results in the highway lane change scenario are shown in Table 3:
[0249] Table 3: Explicit preference data collection samples (10 drivers) in the highway lane change scenario;
[0250]
[0251] Data shows that car and MPV drivers prefer medium-speed and smooth lane changes, while SUV and pickup truck drivers prefer slow and conservative lane changes, reflecting the influence of vehicle type characteristics on drivers' long-term trajectory preferences.
[0252] Implementation examples of transfer learning and decision generation:
[0253] In this embodiment, a benchmark safety decision-making model is first trained on cars, and then knowledge is transferred to other vehicle types using transfer learning. Transfer learning automatically adjusts the model parameters to adapt to the characteristics of different vehicle types. For example, for SUVs and pickup trucks, the steering gain is reduced, the yaw damping and steering lead are increased, and the comfort factor is reduced, making the decision-making more conservative and stable.
[0254] Using the transferred model, the system generates multiple candidate trajectories in an emergency obstacle avoidance scenario and selects the optimal trajectory based on the long-term planning preference model. The evaluation results of candidate trajectories for different vehicle types are shown in Table 4:
[0255] Table 4: Evaluation of candidate trajectories in emergency obstacle avoidance scenario (preference score, full score 10);
[0256]
[0257] The system selects different optimal trajectories for different vehicle types: cars select medium deceleration during mid-turn, SUVs and pickup trucks select early turn with small deceleration, and MPVs select medium deceleration during mid-turn. This reflects the system's adaptability to the dynamic characteristics of different vehicle types and consideration of drivers' long-term preferences.
[0258] Verification of technical effects:
[0259] The two most important technical effects of this embodiment are: strong vehicle type adaptability and high user acceptance.
[0260] These two effects are verified through experimental data below.
[0261] Verification of vehicle type adaptability:
[0262] The vehicle type adaptability is verified by comparing the differences in time and effect between the traditional method and this embodiment in adapting to new vehicle types. The comparison of adaptation time is shown in Table 5:
[0263] Table 5: Comparison of adaptation time for different vehicle types (hours);
[0264]
[0265] Data shows that this embodiment reduces the average safety policy adaptation time for new vehicle types by 85.2%, which is very close to the expected 85%.
[0266] User acceptance verification:
[0267] Verify the improvement effect of user acceptance by testing the satisfaction and intervention rate of drivers with the system's decisions. The user experience evaluation results in the human-machine collaborative driving scenario are shown in Table 6:
[0268] Table 6: Comparison of user experience evaluation results (20 test drivers);
[0269]
[0270] The data shows that the comprehensive index of user acceptance in the human-machine collaborative scenario of this embodiment has increased by 64.9%, which is basically consistent with the expected 65%, proving that this method can effectively improve the consistency between system decisions and human expectations.
[0271] The above describes the embodiments of the present invention, but these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.
Claims
1. An intelligent safety decision-making method for vehicle models based on dynamic perception, characterized in that, It includes the following steps: Collect vehicle dynamic characteristic data through sensors to generate a vehicle type feature vector; Based on the vehicle type feature vector, construct a hierarchical decision-making structure including a strategy layer, a tactic layer, and an execution layer. The strategy layer is responsible for generating long-term planning goals, the tactic layer is responsible for generating mid-term maneuver decisions, and the execution layer is responsible for generating short-term control actions. The vehicle type feature vector serves as the common input parameter for the neural network models of each layer to adjust the decision-making outputs of each layer to adapt to the dynamic characteristics of different vehicle types; Collect human preference data including implicit preference data and explicit preference data; Based on the human preference data, construct a short-term control preference model and a long-term planning preference model; Use an experience transfer network to transfer the decision-making knowledge of the source vehicle type to the target vehicle type, generate multiple candidate trajectories, evaluate and select the optimal trajectory through the long-term planning preference model, and generate specific control actions based on the short-term control preference model.
2. The intelligent safety decision-making method for vehicle models based on dynamic perception according to claim 1, characterized in that The vehicle dynamic characteristic data includes the center of mass height, steering sensitivity, braking force distribution characteristics, and roll rate; Process the data through a feature extraction algorithm to generate a vehicle model feature vector , where: ; Among them, represents the vehicle model feature vector, represents the vehicle model feature parameter set, , , respectively represent the , , th vehicle model feature parameter, represents the total number of parameters, represents the vehicle model feature vector generation function, , , respectively represent the , , th elements in the vector, represents the vector dimension.
3. The intelligent safety decision-making method for vehicle models based on dynamic perception according to claim 1 is characterized in that, In the hierarchical decision-making structure: Policy layer neural network model Generate long-term planning goals that satisfy: ; Tactical layer neural network model Generate mid-term maneuver decisions that satisfy: ; Execution layer neural network model Generate short-term control actions that satisfy: ; Among them, represents the long-term planning goal generated by the policy layer, represents the mid-term maneuver decision, represents the short-term control action, represents the current environmental state, represents the vehicle model feature vector, 、 and represent the neural network models of the policy layer, the tactical layer, and the execution layer respectively.
4. The intelligent vehicle safety decision-making method based on dynamic perception according to claim 1, characterized in that, The collection of the human preference data includes: Implicit preference data collection: Record driver operation intervention, steering wheel grip force, pedal operation characteristics, and eye-tracking data to generate an implicit preference feature vector ; Explicit preference data collection: Generate candidate trajectory pairs for drivers to select and construct an explicit preference dataset , satisfying: ; Among them, represents the explicit preference dataset, represents the th trajectory pair, represents the corresponding preference label, represents the number of samples in the training dataset.
5. The intelligent safety decision-making method for vehicle models based on dynamic perception according to claim 1, characterized in that, The construction of the short-term control preference model includes: Construct a short-term control preference reward function based on implicit preference data , satisfying: ; Train the reward function parameters by maximizing the following optimization objective : ; Among them, represents the short-term control preference reward function, represents optimizing the reward parameter for optimization, represents a set of data sampled from the implicit preference data distribution for sampling, represents the environmental state, represents the control action, represents the implicit preference feature vector, represents the vehicle model feature vector, represents the action samples that the driver does not prefer, represents the sigmoid function, represents the state when taking the non-preferred action for the reward value, represents the logarithmic function.
6. The intelligent safety decision-making method for vehicle models based on dynamic perception according to claim 1, wherein, The construction of the long-term planning preference model includes: Construct a long-term planning preference function based on the explicit preference dataset , satisfying: ; Train the preference function parameters by maximizing the following likelihood function :[[-]] ; Among them, represents the probability that, under the condition of a given vehicle model feature the trajectory is more preferred than the trajectory ; and represent two candidate trajectories, represents the sigmoid function, represents optimizing the parameters of the preference function, represents the preference function, represents the trajectory feature extraction network, represents the th preference label of the trajectory pair, represents the logarithmic function, represents the number of samples in the training dataset.
7. The intelligent safety decision-making method for vehicle models based on dynamic perception according to claim 1, characterized in that, The knowledge transfer includes: Construct an experience transfer network , transfer the decision-making knowledge of the source vehicle model to the target vehicle model, satisfying: ; Training the transfer network parameters by minimizing the following loss function : ; Among them, and respectively represent the decision model parameters of the source vehicle model and the target vehicle model, and respectively represent the feature vectors of the source vehicle model and the target vehicle model, represents the interaction data of the target vehicle model, represents the action value function of the target vehicle model, represents the experience transfer network, represents the parameters of the transfer network for optimization, represents the state, action, reward, next state transition sample, represents the discount factor.
8. A vehicle type intelligent safety decision-making method based on dynamic perception according to claim 1, characterized in that The decision generation includes: Based on the current environmental state and the vehicle type feature vector, generate multiple candidate trajectories that meet safety constraints; Evaluate the priorities of each candidate trajectory through the long-term planning preference model and select those that meet: ; The optimal trajectory, where represents the optimal trajectory with the highest evaluation score, , , respectively represent the -th, -th, -th candidate trajectory, represents the total number of candidate trajectories, represents the preference evaluation score of the trajectory; Based on the optimal trajectory and the short-term control preference model, generate those that meet: ; The control action, where, represents the control action with the highest reward value in the set of actionable actions, represents searching in the set of actionable actions ) to find the action that maximizes the short-term control preference reward function obtains the maximum value , represents the current environmental state, represents the short-term control preference reward function.
9. The intelligent safety decision-making method for vehicle models based on dynamic perception according to claim 1, characterized in that The information interaction mechanism between the layers in the hierarchical decision-making structure includes: The upper-level decision serves as a constraint condition for the lower-level decision. The downstream information flow from layer to layer satisfies the following: ; The execution result of the lower-level decision serves as the feedback information for the upper-level decision. The upstream information flow from layer to layer satisfies the following: ; Among them, represents the downlink information flow from layer to layer . represents the uplink information flow from layer to layer . represents the decision result of layer . represents the execution feedback of layer . and respectively represent the downlink and uplink information processing functions, represents the vehicle model feature vector.
10. An intelligent vehicle type safety decision-making system based on dynamic perception, which is used to implement an intelligent vehicle type safety decision-making method based on dynamic perception according to any one of claims 1 to 9, characterized in that, It includes: A vehicle type feature dynamic perception module, which is used to collect vehicle dynamic characteristic data and generate a vehicle type feature vector; A hierarchical decision-making structure module, which constructs a hierarchical decision-making structure including a strategy layer, a tactic layer, and an execution layer based on the vehicle type feature vector. The strategy layer is responsible for generating long-term planning goals, the tactic layer is responsible for generating mid-term maneuver decisions, and the execution layer is responsible for generating short-term control actions. The vehicle type feature vector serves as the common input parameter for the neural network models of each layer to adjust the decision-making outputs of each layer to adapt to the dynamic characteristics of different vehicle types; A human preference data collection module, which is used to collect human preference data including implicit preference data and explicit preference data; A preference learning module, which is used to construct a short-term control preference model and a long-term planning preference model based on the human preference data; A transfer learning module, which is used to transfer the decision-making knowledge of the source vehicle type to the target vehicle type; A decision generation module, which is used to generate multiple candidate trajectories, evaluate and select the optimal trajectory, and generate specific control actions.
Citation Information
Patent Citations
Systems and methods for predicting trajectories of multiple vehicles
US20230085296A1
Reinforcement learning for traffic simulation
US20250029489A1