A method and system for constructing a personalized driving style comfort field

CN122571797BActive Publication Date: 2026-09-18JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611017817.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-18
Estimated Expiration
2046-07-09

AI Technical Summary

Technical Problem

[0004]1.通用的驾驶风格分类无法反映驾驶员之间的个体差异,在边界模糊时驾驶员风格识别判断不清

Benefits of technology

[0069] Compared with the prior art, the beneficial effects of the present invention are: 1. By constructing an individualized comfort field based on the target driver's historical driving data, instead of a uniform fixed threshold, it can more accurately reflect the individual differences of the driver.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122571797B_ABST
    Figure CN122571797B_ABST
Patent Text Reader

Abstract

This invention belongs to the technical fields of intelligent vehicle control, driving behavior modeling, and vehicle dynamics risk assessment, and specifically provides a method and system for constructing a personalized driving style comfort field. It includes the following steps: S1: historical data collection, driver identification and grouping; S2: state and action feature extraction and coordinate unification; S3: construction of an individualized comfort field based on inverse reinforcement learning and two-dimensional occupancy density modulation; S4: comfort mapping based on physical prior constraints and data-driven modulation; S5: online comfort assessment and control application. This invention constructs an individualized comfort field based on the target driver's historical driving data, replacing a uniform fixed threshold. This more accurately reflects individual driver differences, improves the sample coverage and stability of individualized modeling, and combines data-driven modeling capabilities with physical interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent vehicle control, driving behavior modeling, and vehicle dynamics risk assessment, specifically to a method and system for constructing a personalized driving style comfort field. Background Technology

[0002] With the continuous development of intelligent vehicle control technology and personalized driving experience research, vehicle control systems are no longer satisfied with just general safety and stability. Instead, they aim to further match the driving habits and styles of different drivers while ensuring safety boundaries. Most existing vehicle risk assessment or stability constraint methods adopt uniform threshold boundaries, fixed friction circle limits, or general risk models, applying the same evaluation criteria to all drivers.

[0003] However, in actual driving, different drivers exhibit significant differences in longitudinal acceleration, lateral acceleration, yaw response, and steering maneuvering. For more aggressive drivers, their typical operating conditions may be closer to the vehicle's adhesion limits; for more conservative drivers, the same dynamic state may be subjectively perceived as high-risk. Using a uniform risk boundary for assessment will lead to the following problems:

[0004] 1. General driving style classifications cannot reflect individual differences among drivers, and driver style identification is unclear when the boundaries are blurred.

[0005] 2. Directly increasing the risk value based solely on sparse samples can easily lead to oversaturation of large-area working conditions within the friction circle, reducing the hierarchy and interpretability of the comfort field.

[0006] 3. Although pure black-box risk prediction methods have a certain fitting ability, they lack physical interpretability when combined with vehicle dynamics constraints, making them difficult to use directly for the design of control costs or constraint terms.

[0007] Therefore, there is a need for an individualized driving style comfort field construction and real-time evaluation method that takes into account individual driver differences, dynamic interpretability, and online real-time performance, in order to improve the personalized adaptation capability and safety decision-making capability of intelligent vehicle control. Summary of the Invention

[0008] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0009] To solve the above-mentioned technical problems, according to one aspect of the present invention, the present invention provides the following technical solution: a method for constructing a personalized driving style comfort field, comprising the following steps:

[0010] S1: Historical data collection, driver identification and grouping;

[0011] S2: State and action feature extraction and coordinate unification: Extract vehicle dynamics state vector and driving action vector from each driving record, and fill in missing fields with default values;

[0012] S3: Construction of an individualized comfort field based on inverse reinforcement learning and two-dimensional occupancy density modulation: The historical longitudinal and lateral accelerations of the target driver are mapped onto a two-dimensional longitudinal-lateral acceleration plane. The two-dimensional plane is divided into grids, the occupancy density of driving behavior in each grid is statistically analyzed, and a smoothed behavior distribution model is constructed. At the same time, physical features are extracted from the state-action samples of the target driver, and the reward weights are individualized through maximum entropy inverse reinforcement learning.

[0013] S4: Comfort mapping based on physical prior constraints and data-driven modulation: After obtaining the driver's individualized reward function and behavior density distribution, the normalized reward is mapped into an individualized comfort field by combining the vehicle friction circle or adhesion limit constraint.

[0014] S5: Online comfort assessment and control application: Obtain the current state of the vehicle, extract the corresponding longitudinal and lateral accelerations, project them into the constructed individualized comfort field, and obtain the current comfort value; the comfort value can be converted into a risk value and used as a cost item for trajectory planning, a controller weight adjustment item, a constraint boundary correction item, a stability protection trigger condition, or a human-machine interaction warning signal.

[0015] As a preferred embodiment of the personalized driving style comfort field construction method described in this invention, the specific method of S1 is as follows: obtain historical driving data files corresponding to multiple driving segments, extract the driver identifier from the file name, and group the historical driving data files according to the driver identifier; for a target driver, concatenate the multiple historical driving data files corresponding to him into a unified dataset, and the driver identifier is determined by the leading letter prefix in the file name.

[0016] In a preferred embodiment of the personalized driving style comfort field construction method described in this invention, the vehicle dynamics state vector in step S2 is represented as follows:

[0017]

[0018] in, Indicates longitudinal velocity. Indicates lateral velocity. Indicates longitudinal acceleration. Indicates lateral acceleration. Indicates yaw rate. Indicates the steering wheel angle. Indicates steering wheel torque;

[0019] The driving motion vector is represented as:

[0020] .

[0021] As a preferred embodiment of the personalized driving style comfort field construction method described in this invention, the two-dimensional occupancy density model in S3 is represented as follows:

[0022]

[0023] in, Indicates the first Number of samples in each grid cell and These represent the grid indices for the longitudinal and lateral acceleration directions, respectively. This represents the total number of historical samples. and They represent the first The longitudinal and lateral accelerations of the historical samples, Indicates the first Two-dimensional grid cells; the above two-dimensional occupancy density model is smoothed to obtain a smoothed density model; the two-dimensional density histogram is smoothed twice using one-dimensional separable convolution kernels executed along the vertical and horizontal directions respectively; the sampling density is obtained by bilinear interpolation or equivalent local interpolation for the density values ​​between grid nodes;

[0024] Constructing a reward function based on state-action features:

[0025]

[0026] in, For feature vectors, This is the individualized reward weight vector obtained through inverse reinforcement learning. The input state vector, This is the driving motion vector.

[0027] As a preferred embodiment of the personalized driving style comfort field construction method described in this invention, the specific method of S4 is as follows: defining the current dynamic utilization rate as:

[0028]

[0029] in, Indicates the current query point The corresponding kinetic utilization rate Indicates longitudinal acceleration. Indicates lateral acceleration. This refers to the road surface adhesion coefficient or equivalent friction coefficient. It is the acceleration due to gravity;

[0030] For any query point, first calculate the basic reward value according to the reward function, and then normalize it:

[0031]

[0032] in, For query point The corresponding original reward value, and These represent longitudinal acceleration and lateral acceleration, respectively. and These are the lower and upper bounds estimated based on the reward range of the training samples, respectively. This means that the result within the parentheses is restricted to the range of 0 to 1. The base comfort level corresponding to the normalized reward;

[0033] Original reward value Through individualized reward weight vector With the feature vector of the query point The inner product is calculated to obtain:

[0034]

[0035] The feature vector is 11-dimensional, with the following components in order: longitudinal vehicle speed, absolute value of lateral velocity, longitudinal acceleration energy, lateral acceleration energy, yaw rate energy, steering wheel angle energy, absolute value of steering wheel torque, friction circle utilization rate, over-limit quantity, over-limit penalty, and normalized friction coefficient; among which, For query point The corresponding feature vector, For the eigenvector of the th One portion, Individualized reward weight vector The Each component; all of the above features are derived from the input state vector. Physical constants are directly calculated analytically, without learnable parameters; all learnable information is concentrated in the individualized reward weight vector. It is obtained by offline training from driver trajectory data using the maximum entropy inverse reinforcement learning algorithm.

[0036] As a preferred embodiment of the personalized driving style comfort field construction method described in this invention, wherein the... and The estimation method is as follows: traverse all training samples, and for each sample... Substituting into the above inner product, the reward value is calculated. Take the minimum value of all samples. and maximum value And each of them extends outwards by a 5% margin; among which, The training sample number. and The first The longitudinal and lateral accelerations of the sample. For the first The original reward value corresponding to each sample:

[0037]

[0038] If the reward value distribution is extremely concentrated, it degenerates into a fixed interval. This is to prevent numerical anomalies caused by division by zero.

[0039] After normalization, then The operation truncates the result to To achieve basic comfort ;in, This means that the result within the parentheses is restricted to the range of 0 to 1. Reflecting the driver's relative preference for this acceleration combination under the current operating condition: operating conditions that frequently appear in the training data and are highly evaluated by the reward function. Conditions that are close to 1, but deviate from the usual driving range or are penalized by the reward function correspond to... Close to 0.

[0040] As a preferred embodiment of the personalized driving style comfort field construction method described in this invention, to avoid the sparse region collapsing directly to an extremely low value inside the friction circle, the sampling density is mapped to behavioral comfort through logarithmic compression:

[0041]

[0042] in, The behavioral comfort level is obtained by logarithmic compression of the sampling density. For the smoothed density model at the query point Sampling density at that location This represents the maximum density value of the smoothed density model.

[0043] To further reduce the numerical contrast between high and low density, contrast compression is applied to behavioral comfort:

[0044]

[0045] in, To compare compression indexes, The behavioral comfort level is obtained by logarithmic compression of the sampling density. For the behavioral comfort after comparison and compression;

[0046] When the current operating condition is within the boundary of the friction circle, construct an internal support base with utilization-related parameters:

[0047]

[0048] in, The internal support base is located within the boundary of the friction circle. For kinetic utilization rate, This indicates taking the smaller value within the parentheses; when the current operating condition is outside the friction circle boundary, the over-limit is defined as follows:

[0049]

[0050] in, For exceeding the limit, For kinetic utilization rate, This indicates that only the positive over-limit portion is taken when the utilization rate exceeds 1; and a small support term is constructed outside the boundary:

[0051]

[0052] in, For minor support outside the boundary, For exceeding the limit, It is an exponential function; from this, a unified density modulation term is obtained:

[0053]

[0054] in, According to and The resulting unified support base Indicates the area within the boundary of the friction circle. Indicates the area outside the friction circle boundary. For density modulation terms, To assess behavioral comfort after compression, an out-of-bounds smooth exponential decay term is constructed:

[0055]

[0056] in, The term represents the exponential decay term of the out-of-boundary smoothing. For exceeding the limit; the final individualized comfort value is expressed as a continuous function determined by the normalized reward, density modulation term, and out-of-bounds attenuation term:

[0057]

[0058] in, For query point Individualized comfort values ​​at the location, To normalize the base comfort level corresponding to the reward. For density modulation terms, For the attenuation term outside the boundary, To find the maximum median value on the grid, which is used for peak normalization of the final comfort field, This means that the result within the parentheses is restricted to the range of 0 to 1;

[0059] Thus, the range of values ​​is obtained. The individualized driving style comfort field within the system, where a higher comfort value indicates that the current operating condition is more in line with the driver's usual driving style and is closer to or even exceeds the vehicle's traction limits; a lower comfort value indicates that the current operating condition deviates more from the driver's usual driving style, or is closer to or even exceeds the vehicle's traction limits; corresponding risk values... Represented as:

[0060]

[0061] in, For query point Risk value at the location.

[0062] As a preferred embodiment of the personalized driving style comfort field construction method described in this invention, in step S5, the current comfort value is obtained through grid lookup, bilinear interpolation, or equivalent local mapping; an inverse reinforcement learning reward function with physical prior constraints is used, combined with a two-dimensional occupancy density model and friction circle boundary constraints to construct a continuous comfort field.

[0063] A personalized driving style comfort field construction system includes:

[0064] Historical driving data acquisition module: used to acquire the historical driving data of the target driver, the historical driving data including at least one or more parameters among vehicle longitudinal speed, lateral speed, longitudinal acceleration, lateral acceleration, yaw rate, steering wheel angle and steering wheel torque;

[0065] Data preprocessing and driver grouping module: This module is used to perform field alignment, numerical completion, coordinate system unification, and driver grouping on the raw driving data to obtain the state sequence and action sequence corresponding to the target driver.

[0066] Individualized comfort field construction module: used to map the target driver's historical driving data into the vehicle dynamics constraint space, and construct an individualized comfort field based on the driver's historical behavior distribution;

[0067] Online risk assessment module: used to determine the corresponding real-time comfort value or comfort level based on the current state's position in the individualized comfort level when the vehicle is running online;

[0068] Control application interface module: used to provide the real-time comfort value or comfort level to the vehicle trajectory planning module, lateral control module, longitudinal control module, stability control module or safety intervention module.

[0069] Compared with the prior art, the beneficial effects of the present invention are: 1. By constructing an individualized comfort field based on the target driver's historical driving data, instead of a uniform fixed threshold, it can more accurately reflect the individual differences of the driver.

[0070] 2. By automatically grouping and stitching multiple historical data segments according to driver identification, the sample coverage and stability of individualized modeling can be improved.

[0071] 3. By modeling within the longitudinal-lateral acceleration dynamic space and introducing friction circle utilization constraints, reward learning, and behavior density modulation, it combines data-driven modeling capabilities with physical interpretability.

[0072] 4. By applying logarithmic compression and contrastive compression to the behavior density, and combining the inner support base of the friction circle, smooth attenuation outside the boundary, and reward normalization, the problem of premature numerical collapse or large-area saturation in sparse areas inside the friction circle can be reduced, making the comfort level smoother and more reasonable.

[0073] 5. By using offline modeling and online querying, real-time comfort values ​​can be quickly output during the online operation phase, making it suitable for deployment in vehicle controllers or decision-making systems.

[0074] 6. The comfort field output is a continuous comfort value rather than a single category label, which is convenient for direct use in trajectory planning, chassis control, stability management and safety intervention. Attached Figure Description

[0075] To more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and detailed embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0076] Figure 1This is a schematic diagram of the distribution of the target driver's historical driving data on the longitudinal acceleration-lateral acceleration plane in an embodiment of the present invention;

[0077] Figure 2 This is a schematic diagram of the three-dimensional surface (top view) of the individualized driving style comfort field in an embodiment of the present invention;

[0078] Figure 3 This is a schematic diagram of a three-dimensional surface (oblique view) of the individualized driving style comfort field in an embodiment of the present invention, wherein the vertical coordinate axis represents the comfort value;

[0079] Figure 4 This is a schematic diagram of the contour lines and friction circle boundaries of the individualized driving style comfort field in an embodiment of the present invention. Detailed Implementation

[0080] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0081] Secondly, the present invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of the present invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not according to the usual scale. Furthermore, the schematic diagrams are merely examples and should not limit the scope of protection of the present invention. In addition, actual fabrication should include three-dimensional spatial dimensions of length, width, and depth.

[0082] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0083] This invention utilizes the target driver's historical driving data to construct an individualized comfort field within the vehicle dynamics constraint space. During online vehicle operation, it outputs real-time comfort values ​​consistent with the driver's style, which can be used for control optimization constraints, planning cost adjustment, stability protection, or safety intervention. It can be applied to scenarios such as intelligent vehicle trajectory planning, lateral and longitudinal control, stability management, and safety intervention.

[0084] Specifically, a method for constructing a personalized driving style comfort field includes the following steps:

[0085] S1: Historical Data Acquisition, Driver Identification and Grouping: Acquire historical driving data files corresponding to multiple driving segments, extract driver identifiers from the filenames, and group the historical driving data files according to the driver identifiers; for a target driver, concatenate multiple historical driving data files corresponding to that driver into a unified dataset to improve sample coverage and individualized modeling stability. In a preferred embodiment, the driver identifier is determined by a leading letter prefix in the filename.

[0086] S2: State and Action Feature Extraction and Coordinate Unification: Extract vehicle dynamics state vectors and driving action vectors from each driving record. In a preferred embodiment, the state vector can be represented as:

[0087]

[0088] in, Indicates longitudinal velocity. Indicates lateral velocity. Indicates longitudinal acceleration. Indicates lateral acceleration. Indicates yaw rate. Indicates the steering wheel angle. This represents the steering wheel torque. The action vector can be represented as:

[0089]

[0090] For missing fields, default values ​​can be used to fill in the missing information. Considering that the definitions of the original data and the target comfort field coordinate system in the lateral acceleration direction may differ, this invention performs a unified sign correction on the lateral acceleration to ensure that offline modeling, visualization, and online querying use a consistent definition of the dynamic direction.

[0091] S3: Construction of an individualized comfort field based on inverse reinforcement learning and two-dimensional occupancy density modulation: The historical longitudinal and lateral accelerations of the target driver are mapped onto a two-dimensional longitudinal-lateral acceleration plane. The two-dimensional plane is divided into grids, the occupancy density of driving behavior in each grid is statistically analyzed, and a smoothed behavior distribution model is constructed. At the same time, physical features such as speed, acceleration energy, friction circle utilization, over-limit quantity, and over-limit penalty are extracted from the target driver's state-action samples. Individualized reward weights are then generated through maximum entropy inverse reinforcement learning.

[0092] The two-dimensional occupancy density model can be expressed as:

[0093]

[0094] in, Indicates the first Number of samples in each grid cell and These represent the grid indices for the longitudinal and lateral acceleration directions, respectively. This represents the total number of historical samples. and They represent the first The longitudinal and lateral accelerations of the historical samples, Indicates the first Two-dimensional grid cells.

[0095] To obtain a continuous and smooth individualized comfort surface, the aforementioned two-dimensional occupancy density model is smoothed to obtain a smoothed density model. In a preferred embodiment, a one-dimensional separable convolution kernel, executed separately along the longitudinal and transverse directions, is used to smooth the two-dimensional density histogram twice. Furthermore, to achieve continuous density estimation at any query point, bilinear interpolation or equivalent local interpolation is used to obtain the sampling density for the density values ​​between grid nodes.

[0096] Simultaneously, a reward function is constructed based on state-action features:

[0097]

[0098] in, For feature vectors, This is the individualized reward weight vector obtained through inverse reinforcement learning. The input state vector, This is the driving motion vector.

[0099] S4: Comfort Mapping Based on Physical Prior Constraints and Data-Driven Modulation: After obtaining the driver's individualized reward function and behavior density distribution, the normalized reward is mapped into an individualized comfort field by combining the vehicle friction circle or adhesion limit constraints. The current dynamic utilization rate is defined as:

[0100]

[0101] in, Indicates the current query point The corresponding kinetic utilization rate Indicates longitudinal acceleration. Indicates lateral acceleration. This refers to the road surface adhesion coefficient or equivalent friction coefficient. This is the acceleration due to gravity. In a preferred embodiment, a preset constant is used; more preferably, .

[0102] For any query point, first calculate the basic reward value according to the reward function, and then normalize it:

[0103]

[0104] in, For query point The corresponding original reward value, and These represent longitudinal acceleration and lateral acceleration, respectively. and These are the lower and upper bounds estimated based on the reward range of the training samples, respectively. This means that the result within the parentheses is restricted to the range of 0 to 1. This represents the base comfort level corresponding to the normalized reward.

[0105] Specifically, the original reward value Through individualized reward weight vector With the feature vector of the query point The inner product is calculated to obtain:

[0106]

[0107] The feature vector is 11-dimensional, with the following components in order: longitudinal vehicle speed, absolute value of lateral velocity, longitudinal acceleration energy, lateral acceleration energy, yaw rate energy, steering wheel angle energy, absolute value of steering wheel torque, friction circle utilization rate, over-limit quantity, over-limit penalty, and normalized friction coefficient. All of these features are derived from the input state vector. And the physical constants are directly calculated analytically, without learnable parameters; among them, For query point The corresponding feature vector, For the eigenvector of the th One portion, Individualized reward weight vector The Each component; all learnable information is concentrated in the individualized reward weight vector. It is obtained by offline training from driver trajectory data using the Maximum Entropy Inverse Reinforcement Learning (MaxEnt IRL) algorithm.

[0108] The and The estimation method is as follows: traverse all training samples, and for each sample... Substituting into the above inner product, the reward value is calculated. Take the minimum value of all samples. and maximum value And each of them extends outwards by a 5% margin; among which, The training sample number. and The first The longitudinal and lateral accelerations of the sample. For the first The original reward value corresponding to each sample:

[0109]

[0110] The purpose of this extended margin is to prevent the normalization result from being exactly 0 or 1 when the reward value of the query point falls exactly on the boundary of the training set, thus numerically reserving a reasonable mapping margin for conditions outside the training distribution. If the reward value distribution is extremely concentrated (range smaller than...), If ), it degenerates into a fixed interval. This is to prevent numerical anomalies caused by division by zero.

[0111] After normalization, then The operation truncates the result to This achieves a basic level of comfort. This means that the result within the parentheses is restricted to the range of 0 to 1. Reflecting the driver's relative preference for this acceleration combination under the current operating condition: operating conditions that frequently appear in the training data and are highly evaluated by the reward function. Conditions that are close to 1, but deviate from the usual driving range or are penalized by the reward function correspond to... Close to 0.

[0112] Furthermore, to prevent the sparse region from collapsing directly to extremely low values ​​inside the friction circle, the sampling density is mapped to behavioral comfort through logarithmic compression:

[0113]

[0114] in, The behavioral comfort level is obtained by logarithmic compression of the sampling density. For the smoothed density model at the query point Sampling density at that location This represents the maximum density value of the smoothed density model.

[0115] To further reduce the numerical contrast between high and low density, contrast compression is applied to behavioral comfort:

[0116]

[0117] in, To compare compression indexes, The behavioral comfort level is obtained by logarithmic compression of the sampling density. To assess behavioral comfort after compression; in a preferred embodiment, More preferably, .

[0118] When the current operating condition is within the boundary of the friction circle, in order to ensure high comfort in the high-density area, non-zero support in the low-density area, and a gradual decrease in comfort as utilization increases, it is preferable to construct an internal support base that is related to utilization:

[0119]

[0120] in, The internal support base is located within the boundary of the friction circle. For kinetic utilization rate, This indicates taking the smaller value within the parentheses. When the current operating condition is outside the friction circle boundary, the excess value is defined as follows:

[0121]

[0122] in, For exceeding the limit, For kinetic utilization rate, This indicates that only the positive excess portion is taken when the utilization rate exceeds 1. Furthermore, a small-scale support term is preferably constructed outside the boundary.

[0123]

[0124] in, For minor support outside the boundary, For exceeding the limit, It is an exponential function. Based on this, a unified density modulation term is obtained:

[0125]

[0126] in, According to and The resulting unified support base Indicates the area within the boundary of the friction circle. Indicates the area outside the friction circle boundary. For density modulation terms, To assess behavioral comfort after comparative compression, an out-of-bounds smooth exponential decay term is constructed:

[0127]

[0128] in, The term represents the exponential decay term of the out-of-boundary smoothing. For exceeding the limit, It is an exponential function. The final individualized comfort value can be expressed as a continuous function determined by the normalized reward, density modulation term, and out-of-bounds attenuation term:

[0129]

[0130] in, For query point Individualized comfort values ​​at the location, To normalize the base comfort level corresponding to the reward. For density modulation terms, For the attenuation term outside the boundary, To find the maximum median value on the grid, which is used for peak normalization of the final comfort field, This indicates that the result within the parentheses is restricted to the range of 0 to 1.

[0131] Thus, the range of values ​​is obtained. The individualized driving style comfort field within the system, where a higher comfort value indicates that the current operating condition is more in line with the driver's usual driving style and is closer to or even exceeds the vehicle's traction limits; a lower comfort value indicates that the current operating condition deviates more from the driver's usual driving style, or is closer to or even exceeds the vehicle's traction limits; corresponding risk values... Represented as:

[0132] .in, For query point Risk value at the location, Same as the aforementioned individualized comfort values.

[0133] The above coefficients can be calibrated according to vehicle type, tires and control requirements, but they all belong to the technical concept of "continuous mapping of reward learning, behavior density modulation and dynamic constraint coupling" described in this invention.

[0134] S5: Online Comfort Assessment and Control Application: During online operation, the current vehicle state is acquired, the corresponding longitudinal and lateral accelerations are extracted, and these are projected onto a pre-constructed individualized comfort field. The current comfort value is obtained through grid lookup, bilinear interpolation, or equivalent local mapping methods. This comfort value can be further converted into a risk value and used as a cost term for trajectory planning, a controller weight adjustment term, a constraint boundary correction term, a stability protection trigger condition, or a human-machine interaction warning signal.

[0135] Furthermore, the individualized field is constructed using the current preferred embodiment, employing an inverse reinforcement learning reward function with physical prior constraints, and combining a two-dimensional occupancy density model and friction circle boundary constraints to construct a continuous comfort field.

[0136] A personalized driving style comfort field construction system includes:

[0137] Historical driving data acquisition module: used to acquire the historical driving data of the target driver, the historical driving data including at least one or more parameters among vehicle longitudinal speed, lateral speed, longitudinal acceleration, lateral acceleration, yaw rate, steering wheel angle and steering wheel torque;

[0138] Data preprocessing and driver grouping module: This module is used to perform field alignment, numerical completion, coordinate system unification, and driver grouping on the raw driving data to obtain the state sequence and action sequence corresponding to the target driver.

[0139] Individualized comfort field construction module: used to map the target driver's historical driving data into the vehicle dynamics constraint space, and construct an individualized comfort field based on the driver's historical behavior distribution;

[0140] Online risk assessment module: used to determine the corresponding real-time comfort value or comfort level based on the current state's position in the individualized comfort level when the vehicle is running online;

[0141] Control application interface module: used to provide the real-time comfort value or comfort level to the vehicle trajectory planning module, lateral control module, longitudinal control module, stability control module or safety intervention module.

[0142] Example:

[0143] This embodiment takes multiple historical driving data files corresponding to the same target driver as input. The historical driving data files can be in CSV format, and each record includes at least the fields time, vehicle_u, vehicle_v, vehicle_ax, vehicle_ay, vehicle_r, vehicle_ang_sw_cabin, and vehicle_TorqueSW.

[0144] First, multiple historical driving data files are read from the data directory. Driver identifiers are extracted from the leading letter prefix of the filenames and automatically grouped according to the same driver identifier. The multiple data files corresponding to the target driver are then sequentially concatenated into a unified data table to form an offline training sample set, thereby expanding the sample coverage of that driver under acceleration, braking, and steering conditions.

[0145] Secondly, extract the seven-dimensional state vector and the two-dimensional action vector from the unified data table. The state vector is... The action vector is .in, Indicates longitudinal velocity. Indicates lateral velocity. Indicates longitudinal acceleration. Indicates lateral acceleration. Indicates yaw rate. Indicates the steering wheel angle. Indicates steering wheel torque, superscript This indicates vector transpose. When individual fields are missing, zeros are added to fill the missing fields. For cases where the sign of the lateral acceleration in the original data is opposite to the coordinate system definition of the comfort field, the lateral acceleration sign is flipped only once at the preprocessing entry point, thus ensuring that the coordinate systems used for offline modeling, plotting, and online querying are completely consistent.

[0146] Then, the longitudinal and lateral accelerations of the target driver were mapped onto a two-dimensional longitudinal-lateral acceleration plane. Each coordinate axis was divided into 121 grid cells within the range of [-15, 15] m / s². The number of samples in each grid was counted, and the two-dimensional histogram was subjected to two separable smoothing processes to eliminate the blocky appearance caused by the discrete grid and retain the main distribution trends. The smoothed density model was used to reflect the frequency of the target driver's behavior under different dynamic conditions.

[0147] During the reward learning phase, an 11-dimensional physical feature vector is constructed for each sample. These features include longitudinal vehicle speed, absolute lateral velocity, longitudinal acceleration energy, lateral acceleration energy, yaw rate energy, steering wheel angle energy, absolute steering wheel torque, friction circle utilization rate, excess capacity, excess capacity penalty, and normalized friction coefficient. Offline training is performed using the maximum entropy inverse reinforcement learning algorithm, with a preferred learning rate of 0.01 and a maximum of 200 iterations. Non-positive constraints are applied to the weights corresponding to the cost-based features to obtain an individualized reward weight vector that aligns with the target driver's historical preferences.

[0148] For any query point The dynamic utilization rate is calculated primarily based on tire-road adhesion.

[0149]

[0150] in, This indicates the dynamic utilization rate corresponding to the current query point. Indicates longitudinal acceleration. Indicates lateral acceleration. The road surface adhesion coefficient is preferably set to 0.8 in this embodiment. The gravitational acceleration is 9.81 m / s². When the utilization rate is not greater than 1, it indicates that the current working condition is within the boundary of the friction circle; when the utilization rate is greater than 1, it indicates that the current working condition is close to or exceeds the adhesion limit, and an external attenuation constraint needs to be introduced in the comfort mapping.

[0151] The reward values ​​of the training samples are iterated through to find the minimum and maximum values, and a 5% margin is added to both ends of this range to obtain the reward normalization interval. For any query point, the basic reward value is first calculated and normalized to the basic comfort level. Then, the smoothed occupancy density is subjected to logarithmic compression and 0.5 power-law comparative compression to obtain the behavioral comfort level, in order to avoid oversaturation or too rapid collapse between high-density and low-density regions.

[0152] Furthermore, a support base related to utilization is introduced inside the friction circle, while a small continuous support term and an exponential decay term are introduced outside the friction circle. This ensures that the comfort field maintains a continuous transition near the boundary of the friction circle and continuously decreases after exceeding the adhesion limit. The final comfort level is calculated using the following formula:

[0153]

[0154] in, For query point Individualized comfort values ​​at the location, To normalize the base comfort level corresponding to the reward. The density modulation term is obtained from the historical behavior density. For the attenuation term outside the boundary, To query the peak normalization coefficient on the grid, This means cropping the result within the parentheses to the range of 0 to 1. After this step, you will obtain an individualized driving style comfort field with values ​​ranging from 0 to 1.

[0155] In this embodiment, the offline computation stage can generate a complete comfort field on the grid and output it. Figures 1 to 4 The results shown indicate that during the online operation phase, it is only necessary to obtain the vehicle's current longitudinal and lateral acceleration in real time. The current comfort value can be output through grid lookup or bilinear interpolation, and then provided to the trajectory planning, longitudinal and lateral control, stability control, or safety intervention modules.

[0156] Depend on Figure 1 It can be seen that the historical samples of the target driver are mainly concentrated inside the friction circle, and the dense sample area is located near small longitudinal acceleration and small lateral acceleration. Compared with the positive acceleration area, the distribution of negative longitudinal acceleration is more obvious, indicating that the driver has richer historical samples under braking conditions.

[0157] Depend on Figure 2 and Figure 3 It can be seen that the individualized comfort field forms a single-peak continuous high-value area near the origin, and its distribution along the longitudinal acceleration direction is wider than that along the lateral acceleration direction. This indicates that the target driver has a high tolerance for small longitudinal acceleration and deceleration conditions, but is more sensitive to larger lateral acceleration conditions. Figure 3 The vertical axis in the graph represents the comfort value, which ranges from 0 to 1.

[0158] Depend on Figure 4 It can be seen that the comfort contour lines are generally located inside the friction circle and gradually decrease with increasing distance from the origin, with high value areas and... Figure 1 The high-density historical sample regions in the model basically correspond to the model, indicating that the comfort field constructed by this invention can effectively inherit the historical driving style of the target driver.

[0159] Furthermore, combined Figures 2 to 4 It can also be seen that the comfort level does not change abruptly when approaching the friction circle boundary, but rather decreases smoothly; when the operating conditions exceed the friction circle boundary, the comfort level continues to decay to a low value region, indicating that the present invention takes into account both the expression of individual driver preferences and the physical constraints of vehicle dynamics.

[0160] Although the present invention has been described above with reference to embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of the invention. In particular, as long as there is no structural conflict, the features in the disclosed embodiments can be combined with each other in any manner. The lack of an exhaustive description of these combinations in this specification is merely for the sake of brevity and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A method for constructing a personalized driving style comfort field, characterized in that, Includes the following steps: S1: Historical data collection, driver identification and grouping; S2: State and action feature extraction and coordinate unification: Extract vehicle dynamics state vector and driving action vector from each driving record, and fill in missing fields with default values; S3: Construction of an individualized comfort field based on inverse reinforcement learning and a two-dimensional occupancy density model: The historical longitudinal and lateral accelerations of the target driver are mapped onto a two-dimensional longitudinal-lateral acceleration plane. The two-dimensional plane is divided into grids, the occupancy density of driving behavior in each grid is calculated, and a smoothed behavior distribution model is constructed. At the same time, physical features are extracted from the target driver's state-action samples, and the reward weights are individualized through maximum entropy inverse reinforcement learning. The two-dimensional occupancy density model is represented as: in, Indicates the first Number of samples in each grid cell and These represent the grid indices for the longitudinal and lateral acceleration directions, respectively. This represents the total number of historical samples. and They represent the first The longitudinal and lateral accelerations of the historical samples, Indicates the first Two-dimensional grid cells; the above two-dimensional occupancy density model is smoothed to obtain a smoothed density model; the two-dimensional density histogram is smoothed twice using one-dimensional separable convolution kernels executed along the vertical and horizontal directions respectively; the sampling density is obtained by bilinear interpolation or equivalent local interpolation for the density values ​​between grid nodes; Constructing a reward function based on state-action features: in, For feature vectors, This is the individualized reward weight vector obtained through inverse reinforcement learning. The input state vector, For driving action vectors; S4: Comfort mapping based on physical prior constraints and data-driven modulation: After obtaining the driver's individualized reward function and behavior density distribution, the normalized reward is mapped into an individualized comfort field by combining the vehicle friction circle or adhesion limit constraint. For any query point, first calculate the basic reward value according to the reward function, and then normalize it: in, For query point The corresponding original reward value, and These represent longitudinal acceleration and lateral acceleration, respectively. and These are the lower and upper bounds estimated based on the reward range of the training samples, respectively. This means that the result within the parentheses is restricted to the range of 0 to 1. The base comfort level corresponding to the normalized reward; Original reward value Through individualized reward weight vector With the feature vector of the query point The inner product is calculated to obtain: The feature vector is 11-dimensional, with the following components in order: longitudinal vehicle speed, absolute value of lateral velocity, longitudinal acceleration energy, lateral acceleration energy, yaw rate energy, steering wheel angle energy, absolute value of steering wheel torque, friction circle utilization rate, over-limit quantity, over-limit penalty, and normalized friction coefficient; among which, For query point The corresponding feature vector, For the eigenvector of the th One portion, Individualized reward weight vector The Each component; all of the above features are derived from the input state vector. Physical constants are directly calculated analytically, without learnable parameters; all learnable information is concentrated in the individualized reward weight vector. It is obtained by offline training of driver trajectory data using the maximum entropy inverse reinforcement learning algorithm; To avoid the sparse region collapsing directly to extremely low values ​​inside the friction circle, the sampling density is mapped to behavioral comfort through logarithmic compression: in, The behavioral comfort level is obtained by logarithmic compression of the sampling density. For the smoothed density model at the query point Sampling density at that location This represents the maximum density value of the smoothed density model. To further reduce the numerical contrast between high and low density, contrast compression is applied to behavioral comfort: in, To compare compression indexes, The behavioral comfort level is obtained by logarithmic compression of the sampling density. For the behavioral comfort after comparison and compression; When the current operating condition is within the boundary of the friction circle, construct an internal support base with utilization-related parameters: in, The internal support base is located within the boundary of the friction circle. For kinetic utilization rate, This indicates taking the smaller value within the parentheses; when the current operating condition is outside the friction circle boundary, the over-limit is defined as follows: in, For exceeding the limit, For kinetic utilization rate, This indicates that only the positive over-limit portion is taken when the utilization rate exceeds 1; and a small support term is constructed outside the boundary: in, For minor support outside the boundary, For exceeding the limit, It is an exponential function; from this, a unified density modulation term is obtained: in, According to and The resulting unified support base Indicates the area within the boundary of the friction circle. Indicates the area outside the friction circle boundary. For density modulation terms, To assess behavioral comfort after compression, an out-of-bounds smooth exponential decay term is constructed: in, The term represents an exponentially decaying term that smooths outwards from the boundary. For exceeding the limit; the final individualized comfort value is expressed as a continuous function determined by the normalized reward, density modulation term, and out-of-bounds attenuation term: in, For query point Individualized comfort values ​​at the location, To normalize the base comfort level corresponding to the reward. For density modulation terms, For the boundary attenuation term, To find the maximum median value on the grid, which is used for peak normalization of the final comfort field, This means that the result within the parentheses is restricted to the range of 0 to 1; Thus, the range of values ​​is obtained. The individualized driving style comfort field within the system, where a higher comfort value indicates that the current operating condition is more in line with the driver's usual driving style and is closer to or even exceeds the vehicle's traction limits; a lower comfort value indicates that the current operating condition deviates more from the driver's usual driving style, or is closer to or even exceeds the vehicle's traction limits; corresponding risk values... Represented as: in, For query point Risk value at the location; S5: Online comfort assessment and control application: Obtain the current state of the vehicle, extract the corresponding longitudinal and lateral accelerations, project them into the constructed individualized comfort field, and obtain the current comfort value; the comfort value can be converted into a risk value and used as a cost item for trajectory planning, a controller weight adjustment item, a constraint boundary correction item, a stability protection trigger condition, or a human-machine interaction warning signal.

2. The method for constructing a personalized driving style comfort field according to claim 1, characterized in that, The specific method of S1 is to obtain historical driving data files corresponding to multiple driving segments, extract the driver identifier from the file name, and group the historical driving data files according to the driver identifier; For each target driver, multiple historical driving data files are concatenated into a unified dataset, with the driver identifier determined by the leading letter prefix in the filename.

3. The method for constructing a personalized driving style comfort field according to claim 1, characterized in that, In S2, the vehicle dynamics state vector is represented as: in, Indicates longitudinal velocity. Indicates lateral velocity. Indicates longitudinal acceleration. Indicates lateral acceleration. Indicates yaw rate. Indicates the steering wheel angle. Indicates steering wheel torque; The driving motion vector is represented as: 。 4. The method for constructing a personalized driving style comfort field according to claim 1, characterized in that, The specific method of S4 is as follows: define the current dynamic utilization rate as follows: in, Indicates the current query point The corresponding kinetic utilization rate, Indicates longitudinal acceleration. Indicates lateral acceleration. This refers to the road surface adhesion coefficient or equivalent friction coefficient. This is the acceleration due to gravity.

5. The method for constructing a personalized driving style comfort field according to claim 4, characterized in that, The and The estimation method is as follows: traverse all training samples, and for each sample... Substituting into the above inner product, we obtain the reward value. Take the minimum value of all samples. and maximum value And each of them extends outwards by a 5% margin; among which, The training sample number. and The first The longitudinal and lateral accelerations of the sample. For the first The original reward value corresponding to each sample: If the reward value distribution is extremely concentrated, it degenerates into a fixed interval. This is to prevent numerical anomalies caused by division by zero. After normalization, then The operation truncates the result to To achieve basic comfort ;in, This means that the result within the parentheses is restricted to the range of 0 to 1. Reflecting the driver's relative preference for this acceleration combination under the current operating condition: operating conditions that frequently appear in the training data and are highly evaluated by the reward function. The corresponding conditions are close to 1, but deviate from the usual driving range or are penalized by the reward function. Close to 0.

6. The method for constructing a personalized driving style comfort field according to claim 1, characterized in that, In S5, the current comfort value is obtained through grid lookup, bilinear interpolation, or equivalent local mapping; an inverse reinforcement learning reward function with physical prior constraints is used, and a continuous comfort field is constructed by combining a two-dimensional occupancy density model and friction circle boundary constraints.

7. A personalized driving style comfort field construction system, used to implement the personalized driving style comfort field construction method according to any one of claims 1-6, characterized in that, include: Historical driving data acquisition module: used to acquire the historical driving data of the target driver, the historical driving data including at least one or more parameters among vehicle longitudinal speed, lateral speed, longitudinal acceleration, lateral acceleration, yaw rate, steering wheel angle and steering wheel torque; Data preprocessing and driver grouping module: This module is used to perform field alignment, numerical completion, coordinate system unification, and driver grouping on the raw driving data to obtain the state sequence and action sequence corresponding to the target driver. Individualized comfort field construction module: used to map the target driver's historical driving data into the vehicle dynamics constraint space, and construct an individualized comfort field based on the driver's historical behavior distribution; Online risk assessment module: used to determine the corresponding real-time comfort value or comfort level based on the current position of the individualized comfort level when the vehicle is running online; Control application interface module: used to provide the real-time comfort value or comfort level to the vehicle trajectory planning module, lateral control module, longitudinal control module, stability control module or safety intervention module.

Citation Information

Patent Citations

  • Automatic driving human-like safety self-evolution method and system based on data mechanism fusion

    CN116300850A

  • Unmanned vehicle path planning method, system, equipment and medium

    CN121590587A