Vehicle type intelligent safety decision-making method and system based on dynamic perception

Through a dynamic perception-based intelligent safety decision-making method of vehicle models, combined with model characteristics and human preference data, a layered decision-making structure and transfer learning mechanism are built, which solves the problem that existing systems are difficult to adapt to different models and accurately capture human preferences, and achieves an efficient and personalized safety decision-making strategy.

CN119953390AActive Publication Date: 2025-05-09SUZHOU XUANDU AUTOMOBILE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510453293.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-09
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing safety decision-making system of intelligent connected vehicles is difficult to independently learn the optimal safety strategies of different models, and it is difficult to accurately capture human complex preferences for safety decisions, resulting in low user acceptance and high intervention rate.

Method used

Using a vehicle model intelligent safety decision-making method based on dynamic perception, the vehicle dynamic characteristic data is collected through sensors, the vehicle model feature vector is generated, and a hierarchical decision-making structure is constructed. At the same time, human preference data is collected, short-term control preference models and long-term planning preference models are constructed, and the experience migration network is used to transfer decision-making knowledge of the source model to the target model.

Benefits of technology

It improves the system's adaptability and user acceptance in a multi-model environment, shortens the adaptation time of new models' safety strategies, improves the quality and personalization of decisions, and reduces data demand and deployment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119953390A_ABST
    Figure CN119953390A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent network connection automobiles, and discloses a vehicle type intelligent safety decision-making method and system based on dynamic perception, and the method comprises the steps: collecting the dynamic characteristic data of a vehicle through a sensor, and generating a vehicle type feature vector; based on the vehicle model feature vector, constructing a hierarchical decision structure comprising a strategy layer, a tactical layer and an execution layer; collecting human preference data including implicit preference data and explicit preference data; constructing a short-term control preference model and a long-term planning preference model based on the human preference data; the decision knowledge of a source vehicle model is migrated to a target vehicle model by using an experience migration network, a plurality of candidate tracks are generated, an optimal track is evaluated and selected, and a specific control action is generated; the method can rapidly adapt to the dynamic characteristics of different vehicle types, greatly reduces the deployment cost of new vehicle types, and improves the user acceptability and safety performance of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent networked vehicles, and more specifically, to a vehicle model intelligent safety decision-making method and system based on dynamic perception. Background Art

[0002] With the development of artificial intelligence and autonomous driving technology, the safety decision-making system of intelligent connected vehicles has become a research hotspot. The existing safety decision-making systems mainly have the following technical problems: Traditional safety decision-making systems are difficult to autonomously learn the optimal safety strategy for different vehicle characteristics. Traditional reinforcement learning systems are difficult to directly learn safety decision-making strategies that meet human preferences from environmental feedback. Especially in scenarios where safety and comfort need to be balanced, relying solely on environmental reward signals cannot accurately capture the complex preferences of humans for safety decisions, resulting in low user acceptance and high intervention rate despite the decision-making strategies generated by the system meeting safety indicators. Existing human feedback learning systems mainly focus on immediate feedback and ignore the preference learning for long-term planning trajectories. This results in that during long-term driving, although local control actions meet safety requirements, the overall trajectory planning does not meet the driver's preferences, resulting in unnecessary frequent interventions and unsatisfactory experience.

[0003] Therefore, an intelligent safety decision-making method is needed that can automatically adapt to the characteristics of different vehicle models while taking into account human short-term control preferences and long-term planning preferences, so as to improve the system's adaptability and user acceptance in a multi-vehicle environment. Summary of the invention

[0004] The present invention provides a vehicle type intelligent safety decision-making method and system based on dynamic perception, which solves the technical problems of safety decision generalization of different vehicle type characteristics and human preference learning in related technologies, and improves the vehicle type adaptability of the system and user acceptance.

[0005] The present invention provides a vehicle type intelligent safety decision-making method based on dynamic perception, comprising the following steps: Collect vehicle dynamic characteristic data through sensors to generate vehicle model feature vectors; Based on the vehicle model feature vector, a hierarchical decision-making structure including the strategy layer, tactical layer and execution layer is constructed, and each layer introduces the vehicle model feature vector as an input parameter; Collecting human preference data including implicit preference data and explicit preference data; Based on human preference data, construct short-term control preference models and long-term planning preference models; The experience transfer network is used to transfer the decision knowledge of the source vehicle model to the target vehicle model, generate multiple candidate trajectories, evaluate and select the optimal trajectory through the long-term planning preference model, and generate specific control actions based on the short-term control preference model.

[0006] In a preferred embodiment, the vehicle dynamic characteristic data includes center of mass height, steering sensitivity, braking force distribution characteristics and roll rate; The data is processed by a feature extraction algorithm to generate a vehicle model feature vector ,in: ; in, represents the vehicle model feature vector, Represents the vehicle model characteristic parameter set, , , Respectively represent , , Vehicle model characteristic parameters, Indicates the total number of parameters, represents the vehicle model feature vector generation function, , , Respectively represent the first , , elements, Represents the vector dimension.

[0007] In a preferred embodiment, in the hierarchical decision structure: Policy layer neural network model Generate long-term planning goals to meet: ; Tactical layer neural network model Generate mid-term maneuver decisions that satisfy: ; Execution layer neural network model Generate short-term control actions that satisfy: ; in, represents the long-term planning goal generated by the policy layer, represents the mid-term maneuver decision, Indicates a short-term control action, Indicates the current environment status. represents the vehicle model feature vector, , and They represent the strategy layer, tactical layer and execution layer neural network models respectively.

[0008] In a preferred embodiment, the human preference data collection includes: Implicit preference data collection: record driver operation intervention, steering wheel grip strength, pedal operation characteristics and eye tracking data to generate implicit preference feature vectors ; Explicit preference data collection: Generate candidate trajectory pairs for drivers to choose from and build an explicit preference dataset ,satisfy: ; in, represents an explicit preference dataset, Indicates Track pairs, Indicates the corresponding preference label, Indicates the number of samples in the training dataset.

[0009] In a preferred embodiment, the short-term control preference model construction includes: Based on implicit preference data, construct a short-term control preference reward function ,satisfy: ; The reward function parameters are trained by maximizing the following optimization objective : ; in, represents the short-term control preference reward function, Represents the reward parameter To optimize, Indicates the implicit preference data distribution A set of data sampled from Indicates the state of the environment. Indicates control action. represents the implicit preference feature vector, represents the vehicle model feature vector, represents the action samples that the driver does not prefer, represents the sigmoid function, Indicates status Taking unpreferred actions The reward value, Represents a logarithmic function.

[0010] In a preferred embodiment, the long-term planning preference model construction includes: Constructing a long-term planning preference function based on an explicit preference dataset ,satisfy: ; The preference function parameters are trained by maximizing the following likelihood function : ; in, Indicates the characteristics of a given model Under the condition of Comparison track The probability of being more preferred, and represents two candidate trajectories, represents the sigmoid function, Parameters representing the preference function To optimize, represents the preference function, represents the trajectory feature extraction network, Indicates The preference labels for trajectory pairs, represents the logarithmic function, Indicates the number of samples in the training dataset.

[0011] In a preferred embodiment, the knowledge transfer includes: Building an Experience Transfer Network , transfer the decision knowledge of the source model to the target model, satisfying: ; The parameters of the migration network are trained by minimizing the following loss function : ; in, and Represent the decision model parameters of the source model and the target model respectively, and Represent the feature vectors of the source model and the target model respectively, Represents the interaction data of the target vehicle model, represents the action value function of the target vehicle model, represents the experience transfer network, Represents the parameters of the migration network To optimize, Represents state, action, reward, and next state transition samples, Represents the discount factor.

[0012] In a preferred embodiment, the decision making comprises: Based on the current environment state and vehicle model feature vector, multiple candidate trajectories that meet safety constraints are generated; The long-term planning preference model is used to evaluate the priority of each candidate trajectory and select the one that satisfies: ; The optimal trajectory of represents the optimal trajectory with the highest evaluation score, , , Respectively represent , , candidate trajectories, represents the total number of candidate trajectories, represents the preference evaluation score of the trajectory; Based on the optimal trajectory and short-term control preference model, generate satisfying: ; The control action is: represents the control action with the highest reward value in the set of possible actions, Indicates the set of available actions ) to search for a reward function that favors short-term control The action to obtain the maximum value , Indicates the current environment status. represents the short-term control preference reward function.

[0013] In a preferred embodiment, the information interaction mechanism between the layers in the hierarchical decision-making structure includes: The upper-level decision serves as a constraint on the lower-level decision. Layer to The downstream information flow of the layer satisfies: ; The execution results of lower-level decisions serve as feedback information for upper-level decisions. Layer to The upstream information flow of the layer satisfies: ; in, Indicates that from Layer to Downstream information flow of the layer, Indicates that from Layer to Upstream information flow of the layer, Indicates The decision results of the layer, Indicates layer execution feedback, and Respectively represent the downlink and uplink information processing functions, Represents the vehicle model feature vector.

[0014] In a preferred embodiment, a vehicle type intelligent safety decision system based on dynamic perception is used to implement a vehicle type intelligent safety decision method based on dynamic perception, which is characterized by comprising: Vehicle type feature dynamic perception module, used to collect vehicle dynamic characteristic data and generate vehicle type feature vector; A hierarchical decision-making structure module includes a strategy layer, a tactical layer and an execution layer, and each layer introduces the vehicle model feature vector as an input parameter; A human preference data collection module, used to collect human preference data including implicit preference data and explicit preference data; A preference learning module, used for building a short-term control preference model and a long-term planning preference model based on the human preference data; Transfer learning module, used to transfer the decision-making knowledge of the source model to the target model; The decision generation module is used to generate multiple candidate trajectories, evaluate and select the optimal trajectory, and generate specific control actions.

[0015] The beneficial effects of the present invention are: Strong vehicle model adaptability: Through dynamic perception of vehicle model characteristics and transfer learning mechanism, the system can quickly adapt to the dynamic characteristics of different vehicle models. Experimental results show that the safety strategy adaptation time for new vehicle models is shortened by 85%, without the need for re-collecting a large amount of data and model training, which significantly reduces deployment costs and debugging workload.

[0016] High decision quality: The decision-making strategy that integrates human preferences can ensure safety while taking into account riding comfort. Experimental results show that compared with traditional methods, the success rate of handling dangerous scenes in this invention is increased by 31%, while reducing unnecessary intervention by 35%, improving system reliability and robustness.

[0017] High user acceptance: Through the dual-time-scale preference learning mechanism, the system can simultaneously capture the driver's preferences for immediate control actions and long-term planning trajectories, and the generated decisions are more consistent with human expectations. Experimental results show that in the human-machine collaborative scenario, user acceptance increased by 65%, greatly reducing the discomfort caused by system intervention to the driver.

[0018] High degree of personalization: The present invention can generate customized safety decision strategies based on the characteristics of different vehicle models and driver preferences, making the decision more in line with the needs of specific scenarios. The system can identify the personalized preferences of different drivers and optimize the decision based on this, thus enhancing the adaptability of the system.

[0019] High learning efficiency: The hierarchical decision-making structure and transfer learning method are used to enable the system to effectively learn from limited samples. Experimental results show that compared with traditional methods, the sample utilization efficiency of the present invention is increased by 3.2 times, which significantly reduces the data demand and accelerates the system training and optimization process. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a flow chart of a vehicle type intelligent safety decision-making method based on dynamic perception of the present invention; Figure 2 It is a detailed flow chart of the present invention for collecting vehicle dynamic characteristic data through sensors to generate vehicle model feature vectors; Figure 3 It is a detailed flow chart of constructing a hierarchical decision structure of the present invention; Figure 4 is a detailed flow chart of collecting human preference data of the present invention; Figure 5 It is a detailed flow chart of the method of constructing the short-term control preference model and the long-term planning preference model of the present invention; Figure 6 It is a detailed flow chart of the present invention for generating multiple candidate trajectories and selecting the optimal and specific control actions. DETAILED DESCRIPTION

[0021] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that the discussion of these implementations is only to enable those skilled in the art to better understand and implement the subject matter described herein, and the functions and arrangements of the elements discussed may be changed without departing from the scope of protection of the present specification. Various examples may omit, replace, or add various processes or components as needed. In addition, the features described in some examples may also be combined in other examples.

[0022] At least one embodiment of the present invention discloses a vehicle type intelligent safety decision method based on dynamic perception, such as Figures 1 to 6 As shown, the following steps are included: Step 1: Collect vehicle dynamic characteristic data through sensors to generate vehicle model feature vectors; The specific steps include: Step 1.1, sensor data collection; The vehicle dynamic characteristics data is collected using a variety of sensors installed in the vehicle, including: Vehicle posture data collected by the inertial measurement unit (IMU), such as center of mass height, tilt angle, roll rate, etc. Braking force distribution characteristic data collected by wheel speed sensors; Steering response characteristic data collected by steering system sensors; Suspension system response characteristic data collected by acceleration sensor.

[0023] Step 1.2, vehicle model feature extraction; Apply feature extraction algorithms to process the collected sensor data and extract the vehicle model feature parameter set : ; in, Represents the vehicle model characteristic parameter set, , , Respectively represent , , Vehicle model characteristic parameters, Indicates the total number of parameters.

[0024] Step 1.3, vehicle model feature vector generation; Based on the extracted vehicle model feature parameter set , construct the vehicle model feature vector : ; in, represents the vehicle model feature vector, Represents the vehicle model characteristic parameter set, represents the vehicle model feature vector generation function, , , Respectively represent the first , , elements, Represents the vector dimension.

[0025] Step 2: Based on the vehicle model feature vector, a hierarchical decision structure including a strategy layer, a tactical layer, and an execution layer is constructed, and each layer introduces the vehicle model feature vector as an input parameter; The specific steps include: Step 2.1, strategy layer construction; Building a policy layer neural network model , responsible for long-term planning goal generation: ; in, Indicates the current environment status. represents the vehicle model feature vector, represents the generated long-term planning goal, Indicates that the parameter is The strategy layer neural network model.

[0026] Policy layer neural network model The specific implementation is a multi-layer perceptron structure, including: Input layer: receives the state of the environment (including vehicle position, speed, surrounding obstacle information, etc.) and vehicle model feature vector ; Hidden layer: contains 3 fully connected layers, each with 256, 128 and 64 neurons, using ReLU activation function; Output layer: Generate long-term planning goals , including information such as target position, expected speed and expected driving trajectory.

[0027] In the highway lane merging scenario, the strategy layer neural network model receives the traffic flow information and vehicle model feature vectors of the current lane and the target lane, and outputs the overall planning goals for completing the lane merging task, such as the target lane position, merging completion time, and expected speed.

[0028] Step 2.2, tactical layer construction; Building a tactical neural network model , responsible for mid-term maneuver decision generation: ; in, represents the generated mid-term maneuver decision, Indicates that the parameter is The tactical layer neural network model, Indicates the current environment status. represents the generated long-term planning goal, Represents the vehicle model feature vector.

[0029] Tactical layer neural network model It is implemented as a recurrent neural network structure combined with an attention mechanism, including: Input processing unit: the environment state , long-term planning goals and vehicle model feature vector Perform feature fusion; Attention layer: according to the vehicle model feature vector Assign different weights to different environmental factors so that the model focuses on features that are more relevant to a specific vehicle model; Recurrent unit: uses long short-term memory network (LSTM) units to process time sequence information and remember past decisions; Output layer: Generate mid-term maneuver decisions , including lane change timing, acceleration and deceleration strategies, etc.

[0030] In the urban road turning scenario, the tactical layer neural network model receives the turning target provided by the strategy layer, and determines the mid-term maneuvering strategy such as turning timing, turning speed and turning trajectory based on the current intersection traffic conditions and vehicle model feature vector.

[0031] Step 2.3, execution layer construction; Building an executive layer neural network model , responsible for short-term control action generation: ; in, represents the generated short-term control action, Indicates that the parameter is The execution layer neural network model, Indicates the current environment status. represents the generated mid-term maneuver decision, Represents the vehicle model feature vector.

[0032] Execution layer neural network model It is implemented as a hybrid structure combining a feedforward neural network and a vehicle dynamics model, including: Input layer: receives the current state , mid-term maneuver decision and vehicle model feature vector ; Hidden layers: 2 fully connected layers, each with 128 and 64 neurons, using LeakyReLU activation function; Vehicle dynamics constraint layer: Based on the vehicle feature vector Establish constraints to ensure that the generated control actions are consistent with the vehicle's physical characteristics; Output layer: Generate specific control actions , including steering angle, acceleration and braking force.

[0033] In the emergency obstacle avoidance scenario, the execution layer neural network model receives the obstacle avoidance strategy provided by the tactical layer, combines the current vehicle status and vehicle model feature vector, and generates precise steering angle and braking force instructions to achieve safe obstacle avoidance while maintaining vehicle stability.

[0034] Step 2.4, inter-layer information interaction mechanism; Establish an information exchange mechanism between decision-making layers. The upper-level decision serves as a constraint for the lower-level decision, and the execution results of the lower-level decision serve as feedback information for the upper-level decision: ; ; in, Indicates that from Layer to Downstream information flow of the layer, Indicates The decision results of the layer, Indicates that from Layer to Upstream information flow of the layer, Indicates layer execution feedback, and Represent the downlink and uplink information processing functions respectively.

[0035] Step 3, collecting human preference data including implicit preference data and explicit preference data; The specific steps include: Step 3.1, implicit preference data collection; Construct an implicit preference data collection module to record the driver's physiological and operational response data: Driver intervention data: records the driver's intervention in system decisions, such as taking over control, correcting the steering wheel, etc. Steering wheel grip force data: monitor the driver's grip force changes through the pressure sensor on the steering wheel; Pedal operation characteristic data: record the operation force, speed and other characteristics of the accelerator pedal and brake pedal; Eye tracking data: Collects driver’s gaze point, pupil dilation and other physiological response data.

[0036] Apply processing algorithms to the collected implicit preference data to generate implicit preference feature vectors : ; in, represents the implicit preference feature vector, represents the original implicit preference data, represents the implicit preference feature extraction function, Represents the vehicle model feature vector.

[0037] Step 3.2, explicit preference data collection; Construct an explicit preference data collection module to periodically generate candidate trajectory pairs for the driver to choose from: Use the current decision model to generate multiple candidate trajectories that meet safety constraints ,in, , , Respectively represent , , candidate trajectories, represents the total number of candidate trajectories; Randomly select two trajectories from the candidate trajectories and , forming a trajectory pair ; The simulated execution effect of the trajectory pair is displayed to the driver through the on-board human-computer interaction interface; Record the driver's preference and form a label or Preference data, indicating or .

[0038] Building an Explicit Preference Dataset : ; in, represents an explicit preference dataset, Indicates Track pairs, Indicates the corresponding preference label, express , express , Indicates the size of the dataset.

[0039] Step 4, collecting human preference data including implicit preference data and explicit preference data; The specific steps are as follows: Step 4.1, short-term control preference model construction; Based on implicit preference data , construct a short-term control preference reward function : ; in, represents the short-term control preference reward function, Indicates the state of the environment. Indicates control action. represents the implicit preference feature vector, represents the vehicle model feature vector, Represents the parameters of the reward function.

[0040] Short-term control preference model It is implemented as a dual-tower neural network structure, which includes: State Action Processing Tower: Input layer: receives the state of the environment and control actions ; Feature extraction layer: contains 3 fully connected layers, with layer-by-layer dimensions of 128, 64, and 32; Fusion layer: merges state and action features.

[0041] Preferred model processing tower: Input layer: receives implicit preference feature vector and vehicle model feature vector ; Feature extraction layer: contains 2 fully connected layers with dimensions of 64 and 32.

[0042] Cross attention layer: calculates the attention weight between the two towers and fuses the information of the two towers; Output layer: Generates preference reward value.

[0043] In the vehicle following scenario, the short-term control preference model receives the current vehicle distance, relative speed and other states, as well as control actions such as acceleration. It combines the driver's previous operation intervention mode and vehicle model characteristics to evaluate the comfort preference score of the action, and guides the system to generate a following strategy that is more in line with user habits.

[0044] Reward function parameters Training is performed by maximizing the following optimization objective: ; in, Represents the reward parameter To optimize, Indicates the implicit preference data distribution A set of data sampled from Indicates the state of the environment. Indicates control action. represents the implicit preference feature vector, represents the vehicle model feature vector, represents the action samples that the driver does not prefer, represents the sigmoid function, Indicates status Taking unpreferred actions The reward value, Represents a logarithmic function.

[0045] Step 4.2, long-term planning preference model construction; Based on the Explicit Preference Dataset , construct long-term planning preference function : ; in, Indicates the characteristics of a given model Under the condition of Comparison track The probability of being more preferred, and represents two candidate trajectories, represents the vehicle model feature vector, Indicates that the parameter is Trajectory feature extraction network, Represents the sigmoid function.

[0046] Long-term planning preference model It is implemented as a comparison network structure based on Transformer, which includes: Track encoder: Input layer: receives trajectory sequence and , each trajectory contains a sequence of state-action pairs; Position encoding: adding temporal position information; Transformer Encoder: Contains 4 self-attention layers to extract trajectory features.

[0047] Vehicle type condition module: Input layer: receiving vehicle model feature vector ; Mapping layer: maps vehicle model features into conditional vectors.

[0048] Compare Modules: Cross-attention layer: calculates the correlation between the two trajectory encodings and the vehicle model condition; Fusion layer: merge trajectory features and vehicle model condition information; Output layer: output preference probability value.

[0049] In complex intersection passing scenarios, the long-term planning preference model evaluates different intersection passing trajectory plans (such as direct passing, slowing down and waiting, etc.), and selects the passing strategy that best meets the driver's long-term comfort expectations based on previously collected driver preference data and vehicle model characteristics.

[0050] Preference function parameters Training is performed by maximizing the following likelihood function: ; in, Parameters representing the preference function To optimize, Indicates the characteristics of a given model Under the condition of Comparison track The probability of being more preferred, represents the preference function, represents the trajectory feature extraction network, Indicates The preference labels for trajectory pairs, represents the logarithmic function, Indicates the number of samples in the training dataset.

[0051] Step 4.3, preference model integration; Integrate the short-term control preference model and the long-term planning preference model into a unified preference evaluation system: ; in, Indicates a preference for the evaluation system, represents the complete trajectory, represents the reference trajectory, represents the discount factor, represents the vehicle model feature vector, represents the weight coefficient, which is used to balance the importance of short-term control preference and long-term planning preference. represents the short-term control preference model, represents the long-term planning preference model, Represents a logarithmic function.

[0052] Step 5: Use the experience transfer network to transfer the decision knowledge of the source vehicle model to the target vehicle model, generate multiple candidate trajectories, evaluate and select the optimal trajectory through the long-term planning preference model, and generate specific control actions based on the short-term control preference model; The specific steps are as follows: Step 5.1, transfer learning model construction; Building an Experience Transfer Network , used to transfer the decision knowledge of the source model to the target model: ; in, and Represent the decision model parameters of the source model and the target model respectively, and Represent the feature vectors of the source model and the target model respectively, Indicates that the parameter is migration network.

[0053] Experience Transfer Network The implementation is a parameter adjustment network of a meta-learning architecture, which includes: Model feature comparison module: Input layer: receiving source vehicle model feature vector and the target vehicle model feature vector ; Feature difference extraction layer: calculates the similarity and difference features between vehicle models; Output layer: Generate feature map matrix .

[0054] Parameter Mapping Module: Input layer: receiving source vehicle model parameters ; Parameter grouping layer: group parameters by function and level; Adjustment layer: According to the feature mapping matrix Transform the source parameters; Output layer: Generate target vehicle model parameters .

[0055] In the new vehicle model deployment scenario, assuming that there is an existing safety decision model for sedans, it now needs to be adapted to SUV models. The experience transfer network analyzes the feature differences between the two models (such as center of mass height, steering response, etc.), automatically adjusts the decision model parameters to adapt to the dynamic characteristics of the SUV model, and quickly generates a safety decision strategy suitable for SUV models without the need for a lot of retraining.

[0056] Migrate network parameters Training is performed by minimizing the following loss function: ; in, Represents the interaction data of the target vehicle model, represents the action value function of the target vehicle model, represents the experience transfer network, Represents the parameters of the migration network To optimize, Represents state, action, reward, and next state transition samples, Represents the discount factor.

[0057] Step 5.2, candidate trajectory generation; Based on the current state of the environment and vehicle model feature vector , using a hierarchical decision structure to generate multiple candidate trajectories that satisfy safety constraints: ; in, , , Respectively represent , , candidate trajectories, represents the total number of candidate trajectories, represents the trajectory generating function, represents the safety constraints, Represents the vehicle model feature vector.

[0058] Step 5.3, optimal trajectory selection; Based on the long-term planning preference model, the priority of each candidate trajectory is evaluated and the optimal trajectory is selected: ; in, Representation trajectory The preference evaluation score of , , Respectively represent , , candidate trajectories, represents the total number of candidate trajectories, represents the selected optimal trajectory, Represents the vehicle model feature vector.

[0059] Step 5.4, control action generation; Based on the selected optimal trajectory And the short-term control preference model to generate specific control actions: ; in, Indicates compliance with the trajectory The set of possible actions, represents the generated control action, Indicates selecting an action in the available action space , so that the objective function Get the maximum value.

[0060] Real-world application examples of this implementation: The example involves safety decision adaptation for four different types of vehicles, including sedans, SUVs, MPVs, and pickup trucks. The application scenarios cover the following typical driving situations: highway lane changes; Emergency obstacle avoidance; Slow down when following a vehicle; Ramp merge; In these scenarios, the system needs to generate a safe and comfortable decision-making strategy based on the vehicle characteristics and driver preferences. The following will detail the implementation process and effect verification of the system in these scenarios.

[0061] Example of dynamic perception of vehicle model characteristics: Dynamic characteristic data collection was performed on four vehicle models. The main sensor data collection results of each vehicle model are shown in Table 1: Table 1: Data collection results of main sensors of different models;

[0062] The vehicle model feature extraction algorithm is used to process the collected sensor data to generate the vehicle model feature vector. The feature vector after dimensionality reduction processing (simplified to a 6-dimensional vector) is shown in Table 2: Table 2: Feature vectors of different vehicle models;

[0063] These vehicle model feature vectors are used as input parameters of the hierarchical decision structure to adjust the decision strategy to adapt to the characteristics of different vehicle models.

[0064] Examples of human preference data collection implementation: Implicit Preference Data Collection: The system collected operational response data from multiple drivers on different vehicle models during actual road tests. The frequency and intensity of driver intervention on SUVs and pickup trucks were significantly higher than those on sedans and MPVs, indicating that drivers have different control preferences and comfort expectations for different vehicle models.

[0065] Explicit Preference Data Collection: The system displays candidate trajectory pairs in different scenarios to the driver through the human-computer interaction interface and records the driver's preference. The results of explicit preference data collection in the highway lane change scenario are shown in Table 3: Table 3: Sample of explicit preference data collection for highway lane change scenarios (10 drivers);

[0066] The data show that sedan and MPV drivers prefer smooth lane changes at medium speeds, while SUV and pickup truck drivers prefer slow and conservative lane changes, reflecting the impact of vehicle model characteristics on drivers' long-term trajectory preferences.

[0067] Transfer learning and decision generation implementation examples: This implementation first trains the baseline safety decision model on sedans, and then uses transfer learning to transfer knowledge to other models. Transfer learning automatically adjusts the model parameters to adapt to the characteristics of different models. For example, for SUVs and pickups, the steering gain is reduced, the yaw damping and steering advance are increased, and the comfort factor is reduced, making the decision more conservative and stable.

[0068] Using the migrated model, the system generates multiple candidate trajectories in the emergency obstacle avoidance scenario and selects the optimal trajectory based on the long-term planning preference model evaluation. The evaluation results of candidate trajectories for different vehicle models are shown in Table 4: Table 4: Evaluation of candidate trajectories for emergency obstacle avoidance scenarios (preference score, full score is 10);

[0069] The system selects different optimal trajectories for different models: sedans choose to decelerate during the transfer, SUVs and pickups choose to decelerate early and slightly, and MPVs choose to decelerate during the transfer. This reflects the system's ability to adapt to the dynamic characteristics of different models and its consideration of the driver's long-term preferences.

[0070] Technical effect verification: The two most important technical effects of this implementation are: strong vehicle model adaptability and high user acceptance.

[0071] The following experimental data verify these two effects.

[0072] Vehicle adaptability verification: By comparing the time and effect of the traditional method and this implementation method in adapting the new model, the model adaptability is verified. The adaptation time comparison is shown in Table 5: Table 5: Comparison of adaptation time for different models (hours);

[0073] The data show that this implementation method shortens the safety strategy adaptation time of new vehicle models by an average of 85.2%, which is very close to the expected 85%.

[0074] User acceptance verification: By testing the driver's satisfaction and intervention rate with the system decision, the effect of improving user acceptance is verified. The user experience evaluation results in the human-machine collaborative driving scenario are shown in Table 6: Table 6: Comparison of user experience evaluation results (20 test drivers);

[0075] The data show that the comprehensive user acceptance index of this implementation in the human-machine collaboration scenario has increased by 64.9%, which is basically consistent with the expected 65%, proving that this method can effectively improve the consistency between system decisions and human expectations.

[0076] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation mode. The above-mentioned specific implementation mode is merely illustrative and not restrictive. Under the guidance of this embodiment, ordinary technicians in this field can also make more forms of equivalent embodiments, all of which are within the protection of this embodiment.

Claims

1. A vehicle type intelligent safety decision-making method based on dynamic perception, characterized in that: The following steps are involved: Collect vehicle dynamic characteristic data through sensors to generate vehicle model feature vectors; Based on the vehicle model feature vector, a hierarchical decision-making structure including the strategy layer, tactical layer and execution layer is constructed, and each layer introduces the vehicle model feature vector as an input parameter; Collecting human preference data including implicit preference data and explicit preference data; Based on human preference data, construct short-term control preference models and long-term planning preference models; The experience transfer network is used to transfer the decision knowledge of the source vehicle model to the target vehicle model, generate multiple candidate trajectories, evaluate and select the optimal trajectory through the long-term planning preference model, and generate specific control actions based on the short-term control preference model.

2. According to claim 1, a vehicle type intelligent safety decision-making method based on dynamic perception is characterized in that: The vehicle dynamic characteristics data include center of mass height, steering sensitivity, braking force distribution characteristics and roll rate; The data is processed by a feature extraction algorithm to generate a vehicle model feature vector ,in: ; in, represents the vehicle model feature vector, Represents the vehicle model characteristic parameter set, , , Respectively represent , , Vehicle model characteristic parameters, Indicates the total number of parameters, represents the vehicle model feature vector generation function, , , Respectively represent the first , , elements, Represents the vector dimension.

3. The vehicle type intelligent safety decision-making method based on dynamic perception according to claim 1 is characterized in that: In the hierarchical decision structure: Policy layer neural network model Generate long-term planning goals to meet: ; Tactical layer neural network model Generate mid-term maneuver decisions that satisfy: ; Execution layer neural network model Generate short-term control actions that satisfy: ; in, represents the long-term planning goal generated by the policy layer, represents the mid-term maneuver decision, Indicates a short-term control action, Indicates the current environment status. represents the vehicle model feature vector, , and They represent the strategy layer, tactical layer and execution layer neural network models respectively.

4. The method for intelligent safety decision-making of vehicle models based on dynamic perception according to claim 1 is characterized in that: The human preference data collection includes: Implicit preference data collection: record driver operation intervention, steering wheel grip strength, pedal operation characteristics and eye tracking data to generate implicit preference feature vectors ; Explicit preference data collection: Generate candidate trajectory pairs for drivers to choose from and build an explicit preference dataset ,satisfy: ; in, represents an explicit preference dataset, Indicates Track pairs, Indicates the corresponding preference label, Indicates the number of samples in the training dataset.

5. The vehicle type intelligent safety decision-making method based on dynamic perception according to claim 1 is characterized in that: The short-term control preference model construction includes: Based on implicit preference data, construct a short-term control preference reward function ,satisfy: ; The reward function parameters are trained by maximizing the following optimization objective : ; in, represents the short-term control preference reward function, Represents the reward parameter To optimize, Indicates the implicit preference data distribution A set of data sampled from Indicates the state of the environment. Indicates control action. represents the implicit preference feature vector, represents the vehicle model feature vector, represents the action samples that the driver does not prefer, represents the sigmoid function, Indicates status Taking unpreferred actions The reward value, Represents a logarithmic function.

6. The vehicle type intelligent safety decision-making method based on dynamic perception according to claim 1 is characterized in that: The long-term planning preference model construction includes: Constructing a long-term planning preference function based on an explicit preference dataset ,satisfy: ; The preference function parameters are trained by maximizing the following likelihood function : ; in, Indicates the characteristics of a given model Under the condition of Comparison track The probability of being more preferred, and represents two candidate trajectories, represents the sigmoid function, Parameters representing the preference function To optimize, represents the preference function, represents the trajectory feature extraction network, Indicates The preference labels for trajectory pairs, represents the logarithmic function, Indicates the number of samples in the training dataset.

7. The vehicle type intelligent safety decision-making method based on dynamic perception according to claim 1 is characterized in that: The knowledge transfer includes: Building an Experience Transfer Network , transfer the decision knowledge of the source model to the target model, satisfying: ; The parameters of the migration network are trained by minimizing the following loss function : ; in, and Represent the decision model parameters of the source model and the target model respectively, and Represent the feature vectors of the source model and the target model respectively, Represents the interaction data of the target vehicle model, represents the action value function of the target vehicle model, represents the experience transfer network, Represents the parameters of the migration network To optimize, Represents state, action, reward, and next state transition samples, Represents the discount factor.

8. The vehicle type intelligent safety decision-making method based on dynamic perception according to claim 1 is characterized in that: The decision making includes: Based on the current environment state and vehicle model feature vector, multiple candidate trajectories that meet safety constraints are generated; The long-term planning preference model is used to evaluate the priority of each candidate trajectory and select the one that satisfies: ; The optimal trajectory of represents the optimal trajectory with the highest evaluation score, , , Respectively represent , , candidate trajectories, represents the total number of candidate trajectories, represents the preference evaluation score of the trajectory; Based on the optimal trajectory and short-term control preference model, generate satisfying: ; The control action, where represents the control action with the highest reward value in the set of possible actions, Indicates the set of available actions ) to search for a reward function that favors short-term control The action to obtain the maximum value , Indicates the current environment status. represents the short-term control preference reward function.

9. The method for intelligent safety decision-making of vehicle models based on dynamic perception according to claim 1 is characterized in that: The information interaction mechanism between the layers in the hierarchical decision-making structure includes: The upper-level decision serves as a constraint on the lower-level decision. Layer to The downstream information flow of the layer satisfies: ; The execution results of lower-level decisions serve as feedback information for upper-level decisions. Layer to The upstream information flow of the layer satisfies: ; in, Indicates that from Layer to Downstream information flow of the layer, Indicates that from Layer to Upstream information flow of the layer, Indicates The decision results of the layer, Indicates layer execution feedback, and Respectively represent the downlink and uplink information processing functions, Represents the vehicle model feature vector.

10. A vehicle type intelligent safety decision system based on dynamic perception, used to implement a vehicle type intelligent safety decision method based on dynamic perception as claimed in any one of claims 1 to 9, characterized in that: include: Vehicle type feature dynamic perception module, used to collect vehicle dynamic characteristic data and generate vehicle type feature vector; A hierarchical decision-making structure module includes a strategy layer, a tactical layer and an execution layer, and each layer introduces the vehicle model feature vector as an input parameter; A human preference data collection module, used to collect human preference data including implicit preference data and explicit preference data; A preference learning module, used for building a short-term control preference model and a long-term planning preference model based on the human preference data; Transfer learning module, used to transfer the decision-making knowledge of the source model to the target model; The decision generation module is used to generate multiple candidate trajectories, evaluate and select the optimal trajectory, and generate specific control actions.

Citation Information

Patent Citations

  • Driver intention prediction and intention trajectory updating method

    CN115709719A

  • Pedestrian trajectory prediction method based on space-time attention mechanism

    CN117149934A

  • Intelligent terrain recognition transfer learning method and system for vehicle type adaptation

    CN118627589A

  • Reinforcement learning automatic driving method and system based on man-machine cooperation

    CN119636794A

  • Systems and methods for predicting trajectories of multiple vehicles

    US20230085296A1