Vehicle following track generation method and device, terminal equipment and storage medium

By obtaining the target driving style and adjacent vehicle data, generating state space, and evaluating the trajectory with the reward function, the problem of lack of personalized driving style adaptation in the existing technology is solved, and the safety and adaptability of the vehicle following is improved.

CN120482024APending Publication Date: 2025-08-15HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510592462.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, the vehicle track generation method lacks personalized driving style adaptation, ignores factors such as horizontally adjacent vehicles and complex traffic environments, resulting in low safety.

Method used

By obtaining the driving data of adjacent vehicles of the target driving style type, vertical and horizontal directions, generating state space, using the target reward function to evaluate candidate follow-up trajectories, selecting the optimal trajectory, fitting the vehicle position with the driving style prediction model and the polynomial model, comprehensively considering driving safety, efficiency and comfort.

Benefits of technology

It realizes personalized driving style adaptation, improves the safety and adaptability of vehicle follow-up, reduces manual intervention caused by style mismatch, and enhances driving safety and traffic efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120482024A_ABST
    Figure CN120482024A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle following track generation method and device, terminal equipment and a storage medium, and the method can fully consider the personalized demands of a driver, and can employ a corresponding target reward function based on a target driving style type, so that when a candidate following track is generated and the reward function value is evaluated, the vehicle following track can be rapidly generated. If the driving style is not matched, the track conforming to the driving style tends to be selected, so that manual intervention caused by style mismatching can be reduced, and the driving safety is improved; furthermore, the driving data of the target vehicle and the driving data of the adjacent vehicle longitudinally arranged with the target vehicle can be acquired, and the driving data of the adjacent vehicle transversely arranged with the target vehicle can be acquired, so that the dynamic interaction of the longitudinal and transverse adjacent vehicles is comprehensively considered, and the dynamic interaction between the longitudinal and transverse adjacent vehicles is comprehensively considered by integrating the multi-dimensional data. A more comprehensive state space can be constructed, the actual traffic environment where the vehicle is located can be reflected more accurately, and the safety and adaptability of vehicle following are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of vehicle following technology, and in particular to a method, device, terminal equipment and storage medium for generating a vehicle following trajectory. Background Art

[0002] Car-following trajectory generation is the process of determining the specific path and speed changes of a vehicle while following the vehicle ahead. Accurate car-following trajectory generation is crucial in the fields of autonomous driving and intelligent transportation. It ensures vehicle safety, prevents collisions, improves road efficiency, and reduces traffic congestion.

[0003] Traditional methods for generating car-following trajectories often employ unified rules or models, often based on assumptions and simplifications, using mathematical formulas to describe the vehicle's car-following behavior. Existing technologies fail to consider the behavioral characteristics of individual drivers and focus solely on the driving data between the vehicle and its longitudinal neighbors. This results in a lack of personalized adaptation to driving styles when generating car-following trajectories. They also ignore factors such as lateral neighbors and complex traffic environments, potentially leading to frequent sudden braking and acceleration, increasing the risk of collision and compromising the safety of car-following. Summary of the Invention

[0004] Embodiments of the present invention provide a method, apparatus, terminal device, and storage medium for generating a vehicle following trajectory, which can not only achieve personalized driving style adaptation, but also comprehensively consider the dynamic interaction of longitudinally and laterally adjacent vehicles, thereby generating a safer, more efficient, and more comfortable following trajectory in a complex traffic environment, thereby improving the safety of vehicle following and solving the problems existing in the prior art of lacking personalized driving style adaptation and ignoring laterally adjacent vehicles and complex traffic environment factors.

[0005] An embodiment of the present invention provides a method for generating a vehicle following trajectory, comprising:

[0006] Obtaining a target driving style type for characterizing the current driver's behavioral characteristics when operating a target vehicle; wherein the target driving style type corresponds to a target reward function for measuring the vehicle's driving safety, driving efficiency, and speed variation;

[0007] Acquire first driving data corresponding to a target vehicle in a current time period, second driving data corresponding to a first adjacent vehicle arranged longitudinally to the target vehicle in the current time period, and third driving data corresponding to a second adjacent vehicle arranged transversely to the target vehicle in the current time period;

[0008] Using the first driving data, the second driving data, and the third driving data as a state space, generating an action space containing a plurality of candidate car-following trajectories based on the first driving data in the state space; and generating a target reward function value corresponding to each candidate car-following trajectory in the action space based on the state space and a target reward function;

[0009] The candidate following trajectory with the largest target reward function value is used as the target following trajectory of the target vehicle in the next time period.

[0010] Preferably, the first driving data includes: a current speed of the target vehicle, an acceleration of the target vehicle, and a first vehicle position of the target vehicle;

[0011] The step of generating an action space including a plurality of candidate car-following trajectories based on the first driving data in the state space includes:

[0012] Taking the current speed as the center and according to a preset speed value threshold, generating a speed sampling interval containing a plurality of speeds to be selected;

[0013] For each to-be-selected speed in the speed sampling interval, different parameters of a preset polynomial model are solved based on the current speed, acceleration, first vehicle position, the to-be-selected speed, and a preset acceleration reference value; each solved parameter and the to-be-selected speed are then substituted into the polynomial model to generate a candidate vehicle position corresponding to each to-be-selected speed; wherein the polynomial model is used to generate a vehicle position corresponding to the target vehicle at the to-be-selected speed; each parameter is a different coefficient in the polynomial model;

[0014] For each candidate vehicle position, a candidate following trajectory is generated based on the candidate vehicle position and the first vehicle position; wherein the candidate following trajectory includes: a candidate speed, a candidate acceleration, and a candidate vehicle position:

[0015] Generate an action space based on each candidate car-following trajectory.

[0016] Preferably, the second driving data includes: a second vehicle position of a first adjacent vehicle; the third driving data includes: a third vehicle position of a second adjacent vehicle; the target reward function includes: a first sub-function for measuring the driving safety of the vehicle, a second sub-function for measuring the driving efficiency of the vehicle, and a third sub-function for measuring the speed change amplitude of the vehicle;

[0017] Generating a target reward function value corresponding to each candidate car-following trajectory in the action space according to the state space and the target reward function includes:

[0018] For each candidate car-following trajectory, generating a driving safety value of the target vehicle based on the candidate speed, the candidate vehicle position, the second vehicle position of the first adjacent vehicle, the third vehicle position of the second adjacent vehicle, and the first sub-function;

[0019] generating a driving efficiency value of the target vehicle according to the candidate speed and the second sub-function;

[0020] generating a speed change amplitude value of the target vehicle according to the candidate acceleration and the third sub-function;

[0021] A target reward function value corresponding to a candidate car-following trajectory is generated according to the driving safety value, the driving efficiency value, and the speed change amplitude value.

[0022] Preferably, the

[0023] The first sub-function, the second function and the third function correspond to the first weight coefficient, the second weight coefficient and the third weight coefficient respectively;

[0024] The method for generating a vehicle following trajectory further includes:

[0025] According to the maximum likelihood estimation method, a first weight coefficient, a second weight coefficient, and a third weight coefficient in the target reward function are determined.

[0026] Preferably, the process of generating the target driving style type includes:

[0027] Acquire a plurality of historical following trajectory data corresponding to a historical vehicle driven by the current driver; wherein each of the historical following trajectory data includes: an average speed of the historical vehicle, an average acceleration of the historical vehicle, position data of the historical vehicle, and position data of a preceding vehicle in front of the historical vehicle;

[0028] Each historical following trajectory data is input into a preset driving style prediction model, so that the driving style prediction model extracts driving characteristics for characterizing the driving state of the historical vehicle based on the average speed of the historical vehicle and the average acceleration of the historical vehicle in each historical following trajectory data; extracts following characteristics for characterizing the following behavior characteristics between the historical vehicle and the preceding vehicle based on the average speed of the historical vehicle, the position data of the historical vehicle and the position data of the preceding vehicle; and generates a target driving style type corresponding to the current driver based on the driving characteristics and the following characteristics.

[0029] Preferably, the following characteristics include: a spacing characteristic for characterizing the spatial distance maintained by the driver from the vehicle ahead during the following process, and a time distance characteristic for characterizing the degree of time buffering of the driver for longitudinal following;

[0030] The extracting of the following feature for characterizing the following behavior between the historical vehicle and the preceding vehicle based on the average speed of the historical vehicle, the position data of the historical vehicle, and the position data of the preceding vehicle includes:

[0031] Generate a spacing feature representing the spatial distance between the driver and the preceding vehicle during the following process based on historical vehicle position data and preceding vehicle position data;

[0032] A time-distance feature is generated based on the average speed of historical vehicles, the position data of historical vehicles, and the position data of the vehicle ahead, which is used to characterize the time buffer degree of the driver for longitudinal following.

[0033] Preferably, the training process of the driving style prediction model includes:

[0034] Each driver sample and the historical car-following trajectory data sample corresponding to each driver sample are used as training samples;

[0035] Taking each training sample and the actual driving style type corresponding to each training sample as input and the predicted driving style type of each training sample as output, the driving style prediction model to be trained is iteratively trained until the model converges to generate a preset driving style prediction model.

[0036] Based on the above method embodiments, the present invention provides corresponding device embodiments.

[0037] An embodiment of the present invention provides a device for generating a vehicle following trajectory, comprising: a driving style type acquisition module, a driving data acquisition module, a reward function value generation module, and a following trajectory determination module;

[0038] The driving style type acquisition module is configured to acquire a target driving style type that characterizes the current driver's behavioral characteristics when operating a target vehicle; wherein the target driving style type corresponds to a target reward function that measures the vehicle's driving safety, driving efficiency, and speed variation;

[0039] The driving data acquisition module is used to acquire first driving data corresponding to the target vehicle in the current time period, second driving data corresponding to the first adjacent vehicle arranged longitudinally to the target vehicle in the current time period, and third driving data corresponding to the second adjacent vehicle arranged transversely to the target vehicle in the current time period;

[0040] The reward function value generation module is configured to use the first driving data, the second driving data, and the third driving data as a state space, generate an action space containing a plurality of candidate car-following trajectories based on the first driving data in the state space, and generate a target reward function value corresponding to each candidate car-following trajectory in the action space based on the state space and the target reward function;

[0041] The car-following trajectory determination module is configured to use the candidate car-following trajectory with the largest target reward function value as the target car-following trajectory of the target vehicle in the next time period.

[0042] Based on the above method embodiments, the present invention provides corresponding terminal device embodiments.

[0043] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the method for generating a vehicle following trajectory described in the above embodiment of the invention.

[0044] Based on the above method embodiments, the present invention provides corresponding storage medium embodiments.

[0045] Another embodiment of the present invention provides a storage medium, wherein the computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the method for generating a vehicle following trajectory described in the above-mentioned embodiment of the invention.

[0046] The following beneficial effects are achieved by implementing the present invention:

[0047] Embodiments of the present invention provide a method, apparatus, terminal device, and storage medium for generating a vehicle following trajectory. The present invention can obtain a target driving style type used to characterize the current driver's behavioral characteristics when operating a target vehicle. Different driving styles can correspond to different target reward functions. Therefore, the present invention can fully consider the driver's personalized needs. Because a corresponding target reward function can be used based on the target driving style type, when generating candidate following trajectories and evaluating their reward function values, trajectories that match the driving style are favored. This can reduce manual intervention caused by style mismatches and thereby improve driving safety. Furthermore, the present invention not only obtains first driving data corresponding to the target vehicle in the current time period, second driving data corresponding to the first adjacent vehicle arranged longitudinally with the target vehicle in the current time period, but also obtains third driving data corresponding to the second adjacent vehicle arranged transversely with the target vehicle in the current time period, thereby accounting for the influence of lateral adjacent vehicles. By integrating these multi-dimensional data, a more comprehensive state space can be constructed that more accurately reflects the actual traffic environment in which the vehicle is located, thereby improving the safety and adaptability of vehicle following. Compared with the existing technology, the present invention can not only achieve personalized driving style adaptation, but also comprehensively consider the dynamic interaction of longitudinal and lateral adjacent vehicles, generating a safer, more efficient and comfortable following trajectory in complex traffic environments, thereby improving the safety of vehicle following. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a flowchart of a method for generating a vehicle following trajectory provided by one embodiment of the present invention.

[0049] Figure 2 3 is a schematic diagram of the execution structure of a method for generating a vehicle following trajectory provided by one embodiment of the present invention.

[0050] Figure 3 This is a structural block diagram of a driving style clustering module for implementing driving style classification provided by an embodiment of the present invention.

[0051] Figure 4 FIG. 4 is a structural diagram of a candidate trajectory generation module for generating candidate car-following trajectories provided in one embodiment of the present invention.

[0052] Figure 5 1 is a structural diagram of a reward function learning module for generating a reward function provided by an embodiment of the present invention.

[0053] Figure 6 It is a structural schematic diagram of a device for generating a vehicle following trajectory provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments

[0055] The examples are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present invention.

[0056] like Figure 1 As shown, in order to solve the problems in the prior art of lacking personalized driving style adaptation and ignoring factors such as laterally adjacent vehicles and complex traffic environments, an embodiment of the present invention provides a method for generating a vehicle following trajectory, including:

[0057] Step S1: Obtaining a target driving style type for characterizing the current driver's behavioral characteristics when operating a target vehicle; wherein the target driving style type corresponds to a target reward function for measuring the vehicle's driving safety, driving efficiency, and speed variation;

[0058] Step S2: Acquire first driving data corresponding to the target vehicle in the current time period, second driving data corresponding to the first adjacent vehicle arranged longitudinally to the target vehicle in the current time period, and third driving data corresponding to the second adjacent vehicle arranged transversely to the target vehicle in the current time period;

[0059] Step S3: Using the first driving data, the second driving data, and the third driving data as a state space, generating an action space containing a plurality of candidate car-following trajectories based on the first driving data in the state space; and generating a target reward function value corresponding to each candidate car-following trajectory in the action space based on the state space and the target reward function;

[0060] Step S4: The candidate car-following trajectory with the largest target reward function value is used as the target car-following trajectory of the target vehicle in the next time period.

[0061] Regarding step S1, in a preferred embodiment, when deriving the target vehicle's following trajectory in the next time period, the present invention may first obtain the driving style of the current driver of the target vehicle and select a target reward function corresponding to the current driver's driving style;

[0062] Schematically, the target reward function can measure the vehicle's driving safety, driving efficiency, and speed variation. The present invention can start from the underlying logic of human driver thinking patterns, consider driving safety, efficiency, and comfort to design and learn the final target reward function, and use this as the basis for following trajectory decision-making, so that the target following trajectory finally selected in the next time period can not only meet the current driver's driving style, but also meet driving safety, efficiency, and comfort. When generating the trajectory, the target reward function will give priority to meeting the core requirements of the current driver's corresponding style (such as the aggressive type allows a slightly closer following distance to improve efficiency, and the conservative type forces an increase in the safe distance to reduce the risk of collision), so that the following trajectory is more in line with the driver's habits, reduces human intervention and operational conflicts, and improves driving safety and user acceptance from the source.

[0063] In a preferred embodiment, Figure 2 The system execution architecture for vehicle following trajectory generation is demonstrated, which is divided into two main parts: offline training and online execution. The generation of vehicle following trajectory is achieved through the collaboration of different modules.

[0064] First, in the offline training phase, the filter in the data preprocessing module filters the raw vehicle trajectory data, removing noise and making the data smoother and more accurate, providing a reliable data foundation for subsequent analysis. For example, this removes abnormal fluctuations in data caused by vehicle sensor vibration. The following trajectory extractor in the data preprocessing module extracts valid following trajectory data from the filtered data based on specific following trajectory criteria (such as following the same preceding vehicle for 30 consecutive seconds while maintaining a certain distance, or driving in a specific lane).

[0065] Furthermore, the driving style clustering module can use the extracted car-following trajectory data to extract feature values from the perspectives of the distance between following vehicles, headway, average value and standard deviation of vehicle speed and acceleration. Figure 3 As shown, the driving style clustering module can extract driving style feature value vectors according to the feature extractor in the driving style clustering module. Based on these feature vectors, the unsupervised learning K-means algorithm is used to cluster the driver's driving style into categories such as conservative, normal, and aggressive, and output the classification results of the driving style.

[0066] For the reward function learning module, it can generate a target reward function for measuring vehicle driving safety, driving efficiency and speed change. For example, through methods such as maximum entropy inverse reinforcement learning, the parameters of the reward function are learned and optimized using expert trajectory data to obtain a suitable reward function.

[0067] For the data communication module (offline), it can transmit the driving style type results obtained by the driving style clustering module and the reward function obtained by the reward function learning module to the data communication module of the online execution part, providing basic parameters for online execution.

[0068] Next, for the online execution part, the autonomous perception module can perceive the driving status of the vehicle itself and surrounding vehicles (such as the first adjacent vehicle arranged longitudinally and the second adjacent vehicle arranged horizontally) in real time, including speed, position, acceleration and other information, and output real-time observation status.

[0069] For the candidate trajectory generation module, the real-time observation state output by the autonomous perception module can be used as the state space. According to the vehicle's own driving data (first driving data) in the state space, an action space containing several candidate following trajectories is generated. Then, the target reward function value corresponding to each candidate following trajectory in the action space is calculated by combining the state space and the reward function obtained by offline training.

[0070] Finally, the candidate following trajectory with the largest target reward function value can be selected from all candidate following trajectories as the target following trajectory of the target vehicle in the next time period.

[0071] It is understandable that the data communication module (online) can receive the driving style type results and reward function transmitted by the offline training part, and at the same time feed back the relevant data during the online execution process to the offline training part for subsequent optimization and adjustment.

[0072] Therefore, the present invention can cluster driving styles, enabling vehicles to select appropriate car-following strategies based on different driver styles (conservative, average, aggressive), providing a personalized driving experience. Furthermore, a reward function can be used to comprehensively consider driving safety, driving efficiency, and speed variation, ensuring both safety and efficiency during car-following, thereby improving overall traffic operation efficiency. Furthermore, the online execution component can perceive the vehicle's state and surrounding environment in real time, rapidly generating and selecting appropriate car-following trajectories, enabling the vehicle to adapt to ever-changing traffic scenarios and enhancing driving stability and safety.

[0073] In a preferred embodiment, the process of generating the target driving style type of the current driver includes:

[0074] Acquire a plurality of historical following trajectory data corresponding to a historical vehicle driven by the current driver; wherein each of the historical following trajectory data includes: an average speed of the historical vehicle, an average acceleration of the historical vehicle, position data of the historical vehicle, and position data of a preceding vehicle in front of the historical vehicle;

[0075] Each historical following trajectory data is input into a preset driving style prediction model, so that the driving style prediction model extracts driving characteristics for characterizing the driving state of the historical vehicle based on the average speed of the historical vehicle and the average acceleration of the historical vehicle in each historical following trajectory data; extracts following characteristics for characterizing the following behavior characteristics between the historical vehicle and the preceding vehicle based on the average speed of the historical vehicle, the position data of the historical vehicle and the position data of the preceding vehicle; and generates a target driving style type corresponding to the current driver based on the driving characteristics and the following characteristics.

[0076] Furthermore, the following characteristics may include: a spacing characteristic for characterizing the spatial distance maintained between the driver and the preceding vehicle during the following process, and a time distance characteristic for characterizing the degree of time buffering of the driver for longitudinal following;

[0077] Then, when extracting the following feature for characterizing the following behavior between the historical vehicle and the preceding vehicle based on the average speed of the historical vehicle, the position data of the historical vehicle, and the position data of the preceding vehicle, the following is specifically included:

[0078] Generate a spacing feature representing the spatial distance between the driver and the preceding vehicle during the following process based on historical vehicle position data and preceding vehicle position data;

[0079] A time-distance feature is generated based on the average speed of historical vehicles, the position data of historical vehicles, and the position data of the vehicle ahead, which is used to characterize the time buffer degree of the driver for longitudinal following.

[0080] Illustratively, the driving style prediction model of this embodiment of the present invention extracts driving characteristics and car-following characteristics based on historical car-following trajectory data. Driving characteristics (based on average speed and average acceleration) can reflect the historical vehicle's driving state, such as whether the vehicle was accelerating, decelerating, or traveling at a constant speed. Furthermore, car-following characteristics reflect the car-following behavior relationship between the historical vehicle and the preceding vehicle.

[0081] Finally, based on the extracted driving and following characteristics, the model can generate a target driving style type corresponding to the current driver. Different driving style types correspond to different driving behavior patterns, such as conservative, normal, and aggressive.

[0082] Specifically, generating spacing features from historical vehicle and leading vehicle position data can intuitively reflect the spatial distance a driver maintains between themselves and the vehicle ahead during car-following. Drivers with different driving styles have different spacing preferences when following a vehicle. For example, conservative drivers typically maintain a larger gap, while aggressive drivers may follow closer. By extracting spacing features from historical car-following trajectory data, embodiments of the present invention can better understand drivers' safety awareness and driving habits.

[0083] Specifically, time-distance features are generated based on historical average vehicle speed, historical vehicle position data, and the position data of the preceding vehicle. These time-distance features take speed into account and more comprehensively reflect the driver's dynamic behavior during vehicle-following. For example, even if the spatial distance between two vehicles is the same, the driver's actual reaction time will vary due to different speeds. Therefore, time-distance features can help the model determine the driver's confidence in following the vehicle at different speeds and provide a reference for generating following trajectories that better suit the driver's style and safety requirements.

[0084] Therefore, the present invention can use a driving style prediction model to analyze and extract features from several historical car-following trajectory data of the current driver, thereby classifying the driving style corresponding to the current driver's driving habits, so as to more accurately reflect the driver's real behavior and preferences, avoiding the limitations of subjective judgment and fixed rules in traditional methods, and making the generation of subsequent car-following trajectories more scientific and reasonable.

[0085] Illustratively, the training process of the driving style prediction model adopted by the present invention includes:

[0086] Each driver sample and the historical car-following trajectory data sample corresponding to each driver sample are used as training samples;

[0087] Taking each training sample and the actual driving style type corresponding to each training sample as input and the predicted driving style type of each training sample as output, the driving style prediction model to be trained is iteratively trained until the model converges to generate a preset driving style prediction model.

[0088] It is understood that the present invention can utilize a large number of known driver samples and their corresponding historical car-following trajectory data samples, combined with the actual driving style types of these samples, to enable the trained driving style prediction model to learn the mapping relationship between historical car-following trajectory data and driving style types. Through continuous iterative training, the model can accurately predict the corresponding driving style type based on the input historical car-following trajectory data, thereby generating a driving style prediction model that can classify driving style types.

[0089] In a preferred embodiment, the present invention can be based on the collected natural driving data set D of the driver, and use the rolling median filter to smooth the trajectory within the 5-second window, and then use the Savitzky-Golay filter to further smooth and fill the missing values. The pre-processed trajectory data sample of each driver is composed of a series of coordinate position points, and then:

[0090] ζ i ={x j ,y j}; i=1,…,N; j=1,…,T;

[0091] Among them, i is the driver number and j is the time step.

[0092] Indicatively, the pre-processed trajectory data samples of each driver, that is, the historical car-following trajectory data samples corresponding to each driver sample, should meet the following conditions:

[0093] (1) For 30 consecutive seconds, the autonomous vehicle follows the same preceding vehicle, maintaining a following distance of 50 meters without any lane-changing behavior;

[0094] (2) The autonomous vehicle drives in the middle two lanes to ensure that there are adjacent vehicles on both sides;

[0095] (3) Vehicles within a distance of 10 m are considered as adjacent vehicles that may affect the vehicle's following behavior.

[0096] Understandably, the requirement to follow the same vehicle for 30 consecutive seconds, maintain a following distance of 50 meters, and avoid lane changes ensures that all collected samples fall within relatively standard following scenarios. This allows the following behaviors of different drivers to be compared and analyzed under similar conditions, avoiding interference from complex factors such as frequent changes in the following object, fluctuating distances, and lane changes. This allows model training to focus on pure differences in following behavior styles rather than interference from different scenarios.

[0097] Limiting the autonomous vehicle to the two middle lanes and ensuring adjacent vehicles on both sides ensures consistency in the driving style characteristics learned by the model. This is because different lane positions may face different traffic conditions and driving strategies (for example, the outer lane may be more commonly used for entering and exiting the main road, while the inner lane may be faster). Staying in the middle lane can reduce the interference caused by different lane positions, allowing the model to more accurately capture the characteristics of the driving style itself.

[0098] When collecting samples, identifying vehicles within 10 meters of the vehicle ahead and behind as adjacent vehicles that have an impact on car-following behavior helps accurately extract features related to this behavior. For example, when calculating key driving style indicators such as headway and headway, the model can accurately identify reference objects, avoiding the inclusion of vehicles that are too far away and have minimal actual impact on car-following behavior. This allows the model to be trained based on more accurate and valid information, improving the accuracy of driving style predictions.

[0099] In a preferred embodiment, the present invention clusters driving styles into three categories: conservative, normal, and aggressive. The algorithm principle is to divide N observations into K sets S = {S1, S2, ..., S K}, thereby minimizing the intra-cluster sum of squares. The formula is as follows:

[0100]

[0101] Among them, μ i is the mean of the points in the cluster.

[0102] Ultimately, using the K-means algorithm, each driver can be clustered into three categories: conservative, average, and aggressive. Each driving style has its own unique characteristics and personality. Statistics show that the mean speed and acceleration values for each driving style vary, corresponding to their type.

[0103] Through the aforementioned training process, the driving style prediction model gradually adjusts its parameters and optimizes its prediction capabilities, ultimately forming a pre-set model that can accurately predict driving style types based on historical car-following trajectory data. This allows the driving style prediction model to be used in subsequent applications to quickly and accurately predict a driver's driving style type based on newly acquired historical car-following trajectory data, providing a basis for generating vehicle-following trajectories and enabling personalized, safe, and efficient car-following decisions.

[0104] Regarding steps S2 and S3, in a preferred embodiment, the first driving data includes: the current speed of the target vehicle, the acceleration of the target vehicle, and the first vehicle position of the target vehicle; when fitting a plurality of candidate car-following trajectories, the present invention specifically includes:

[0105] Taking the current speed as the center and according to a preset speed value threshold, generating a speed sampling interval containing a plurality of speeds to be selected;

[0106] For each to-be-selected speed in the speed sampling interval, different parameters of a preset polynomial model are solved based on the current speed, acceleration, first vehicle position, the to-be-selected speed, and a preset acceleration reference value; each solved parameter and the to-be-selected speed are then substituted into the polynomial model to generate a candidate vehicle position corresponding to each to-be-selected speed; wherein the polynomial model is used to generate a vehicle position corresponding to the target vehicle at the to-be-selected speed; each parameter is a different coefficient in the polynomial model;

[0107] For each candidate vehicle position, a candidate following trajectory is generated based on the candidate vehicle position and the first vehicle position; wherein the candidate following trajectory includes: a candidate speed, a candidate acceleration, and a candidate vehicle position:

[0108] Generate an action space based on each candidate car-following trajectory.

[0109] Schematically, the above process generates multiple possible candidate car-following trajectories for the target vehicle, allowing subsequent evaluation and selection of the optimal candidate trajectory based on the target reward function. Specifically, by processing the target vehicle's current speed, acceleration, and position information, using speed sampling intervals and a polynomial model, a series of candidate vehicle positions are generated. These candidate vehicle positions are then used to generate candidate car-following trajectories containing speed, acceleration, and position information. Ultimately, all candidate car-following trajectories are combined into an action space.

[0110] Specifically, first the current state of the target vehicle is: the current speed is v o , the acceleration is a o , the coordinate point of the first vehicle position is (x o ,y o ), the time step is 0.1s; further, the speed sampling to be selected can be performed based on the current state of the target vehicle. Since the steering wheel rotation of the following behavior can be ignored, it is only necessary to obtain the target point speed and acceleration value to fit the target trajectory. Therefore, based on the set prediction time window W (that is, the time window corresponding to the next time period), the speed sampling interval containing several speeds to be selected is v t ∈[v o -5,v o +5], and the preset acceleration reference value can be taken as a t =0.

[0111] Based on the speed v to be selected t And the preset acceleration reference value a t , we can solve the different parameters in the preset polynomial model, and the preset polynomial model is a quartic polynomial, then:

[0112] x(τ)=a0+a1τ+a2τ 2 +a3τ 3 +a4τ 4 ;

[0113] Where τ is the time step; x(τ) represents the position coordinates of the target vehicle on the trajectory as time τ changes; {a0, a1, …, a4} are the unknown parameters of the polynomial equation, which can be calculated based on a given initial state (for example, the current speed is v o , the acceleration is a o , the coordinate point of the first vehicle position is (x o ,y o ), the time step is 0.1s) and the target state (such as the speed v to be selected t And the preset acceleration reference value a t ), which can solve different parameters in the preset polynomial model, as shown below:

[0114]

[0115] After obtaining the values of each parameter, we can calculate the value of each parameter based on the solved a0, a1, ..., a4 and the speed v to be selected. t Based on the polynomial model, the candidate vehicle position x(τ) corresponding to each speed to be selected is generated, so that the candidate vehicle position x(τ) and the first vehicle position (i.e. (x o ,y o )), and obtain the candidate car-following trajectory.

[0116] Therefore, the above process maps the target vehicle's current operating position to a candidate car-following trajectory. The candidate car-following trajectory generation process in this embodiment of the present invention utilizes a polynomial solution, which aligns with the laws of vehicle motion. It can be understood that a vehicle's driving process can be viewed as a process in which position, velocity, and acceleration change over time. A quartic polynomial is a good fit for the vehicle's motion trajectory during car-following, describing the vehicle's position changes over time. Furthermore, by differentiating the polynomial, expressions for velocity and acceleration can be derived, accurately reflecting the vehicle's dynamic characteristics during car-following. For example, various operating conditions, such as acceleration, deceleration, and constant speed, can be simulated using appropriate polynomial parameters.

[0117] like Figure 4 As shown in the figure, the generation process of candidate car-following trajectories is as follows:

[0118] Real-time status observation: Real-time monitoring and acquisition of the vehicle's current status (such as speed, acceleration, position, etc.) and surrounding environment.

[0119] State feature extraction: The real-time observed state is input into the "state feature extractor" to extract key features for generating trajectories (such as the vehicle's speed, the distance to the preceding vehicle, and other core information).

[0120] Candidate target sampling: Based on the extracted state features, certain rules are used to generate candidate target sampling results containing several parameters such as speed to be selected, providing different target options for subsequent trajectory generation.

[0121] Polynomial fitting: Using the polynomial fitting method, combined with the candidate target sampling results, the trajectory information such as the vehicle position change corresponding to each target is generated.

[0122] Forming a set of candidate car-following trajectories: All trajectories generated by polynomial fitting are aggregated to form a set of candidate car-following trajectories containing multiple possibilities, providing a basis for subsequent selection of the optimal car-following trajectory.

[0123] In a preferred embodiment, the vehicle following behavior of the present invention can be modeled as a Markov decision process (MDP), which can be defined as {S, A, T, r, γ}, where s t ∈S is the surrounding state observed by the vehicle at time t, including the distance to the preceding vehicle, the preceding vehicle speed, and the vehicle speed. t ∈A is a series of following actions taken by the ego vehicle according to the observed state, T is the state transition probability matrix, r is the reward that the ego vehicle can obtain from the environment, and γ is the discount rate of future rewards.

[0124] Therefore, in an embodiment of the present invention, the vehicle following behavior can be modeled as the vehicle following behavior of the present invention. After being modeled as a Markov decision process, several candidate following trajectories can be obtained through the steps of state observation, action space generation, state transition probability calculation, reward function design, candidate trajectory generation, and trajectory evaluation and selection.

[0125] In a preferred embodiment, the present invention can design a reward function based on safety, efficiency (corresponding to driving efficiency), and comfort (corresponding to the speed variation). Therefore, in order to better measure the safety, efficiency, and comfort of each candidate car-following trajectory, it is necessary to take into account the position data of vehicles adjacent to the target vehicle (such as longitudinal and lateral vehicles). Thus, the following is true:

[0126] The second driving data includes: a second vehicle position of a first adjacent vehicle; the third driving data includes: a third vehicle position of a second adjacent vehicle; the target reward function includes: a first sub-function for measuring the driving safety of the vehicle, a second sub-function for measuring the driving efficiency of the vehicle, and a third sub-function for measuring the speed change amplitude of the vehicle;

[0127] It can be understood that, when constructing the first sub-function for measuring the driving safety of a vehicle, the embodiment of the present invention can perform calculations based on the second vehicle position of the first adjacent vehicle in the longitudinal direction, the third vehicle position of the second adjacent vehicle in the transverse direction, and the current driving position of the vehicle to measure whether the vehicle is within a safe distance range from the adjacent vehicles when traveling to a certain position point using a certain trajectory.

[0128] Furthermore, when constructing the second sub-function for measuring the driving efficiency of a vehicle, the embodiment of the present invention, since driving efficiency can usually be considered from aspects such as the time it takes for the vehicle to reach the destination and whether it frequently stops during driving, by considering the candidate speeds of the candidate following trajectories, the driving efficiency value of the target vehicle under a certain candidate following trajectory can be derived.

[0129] By using the first sub-function and the second sub-function, when generating car-following trajectories, trajectories that enable the target vehicle to reach the destination quickly and are less obstructed by adjacent vehicles can be selected.

[0130] Furthermore, in an embodiment of the present invention, a third sub-function is constructed to measure the speed variation amplitude of a vehicle. Since the speed variation amplitude mainly reflects the smoothness of the vehicle's driving, excessive speed variation will affect the comfort of passengers and may also increase energy consumption and vehicle wear. The third sub-function can be constructed by calculating the candidate acceleration of the target vehicle under a certain candidate following trajectory.

[0131] Through the above-mentioned third sub-function, when generating car-following trajectories, the model will give priority to those trajectories with relatively stable speed changes. When combined with the first, second and third sub-functions, trajectories that can enable the target vehicle to reach the destination quickly, with less obstruction from adjacent vehicles, and with relatively stable speed changes can be selected, thereby ensuring the safety, efficiency and comfort of the car-following trajectory and generating a car-following trajectory that better meets actual needs.

[0132] Specifically, in a preferred embodiment, the present invention can design a reward function that mimics the underlying logic of human thinking, starting from the three important factors that human drivers consider when driving, to design the reward function:

[0133] a. Safety: An eigenvalue expression related to the headway between vehicles is constructed to measure driving safety. In addition to the most commonly used longitudinal safety, this invention also considers the vehicle's lateral safety, namely the lateral headway between vehicles in adjacent lanes, to characterize the probability of adjacent vehicles cutting in.

[0134] b. Efficiency: Speed is used to characterize the efficiency of vehicle driving. A faster speed indicates a more efficient driving behavior.

[0135] c. Comfort: The comfort of vehicle driving is closely related to the amplitude of acceleration changes and is also an important indicator of driving smoothness. Therefore, the first-order derivative of acceleration is used as a characteristic indicator to measure comfort.

[0136] After completing the design of each characteristic index, the reward function can be designed as the weighted sum of each characteristic index (i.e., each sub-function).

[0137] Specifically, regarding the safety factor of the first sub-function, including the longitudinal safety factor and the lateral safety factor, we have:

[0138] The longitudinal safety factor is exponentially related to the headway at each time step:

[0139]

[0140] Among them, x lv (t) represents the longitudinal position coordinate of the preceding vehicle at time t, x ego (t) represents the longitudinal position coordinate of the vehicle at time t, v ego (t) represents the speed of the vehicle at time t, and ε = 0.001 is a parameter to protect the denominator from being zero.

[0141] The lateral safety factor is exponentially related to the tendency of adjacent vehicles to cut in at each time step:

[0142]

[0143] Among them, y sv (t) represents the lateral position coordinate of the adjacent vehicle at time t;

[0144] y ego (t) represents the lateral position coordinate of the vehicle at time t.

[0145] Furthermore, regarding the efficiency-related coefficient of the second function, the efficiency-related characteristic reflects the efficiency of the trajectory, so it can be represented by the speed at time t:

[0146] f efficiency (t)=|v ego (t)|;

[0147] Furthermore, regarding the comfort correlation coefficient of the third function, the first-order derivative of acceleration can be used to measure comfort, which is:

[0148]

[0149] Among them, a ego (t) represents the acceleration value of the vehicle at time t.

[0150] Therefore, the design of the characteristic indicators (i.e., each sub-function) of the reward function has been completed. It is necessary to further define how to construct the reward function. The reward function can be constructed as a linear combination function of the characteristic vector and the parameter vector, then:

[0151] r i =∑ t∈T ω i f(s t );

[0152] Among them, ω i is the parameter vector that measures the weight of each feature, f(s t ) is the characteristic index at time t, that is, the different function values corresponding to the first sub-function, the second function and the third function at time t, which can be the driving safety value, the driving efficiency value and the speed change amplitude value.

[0153] Furthermore, the weight value of each sub-function (i.e. parameter vector) can be estimated by the maximum likelihood estimation method, then:

[0154] The first sub-function, the second function and the third function correspond to the first weight coefficient, the second weight coefficient and the third weight coefficient respectively; then:

[0155] According to the maximum likelihood estimation method, a first weight coefficient, a second weight coefficient, and a third weight coefficient in the target reward function are determined.

[0156] Schematically, the embodiment of the present invention can imitate the expert trajectory to calibrate the reward function parameters, and then based on the obtained reward function, the selection probability of each trajectory can be calculated. Specifically, the maximum entropy inverse reinforcement learning (Max-EntIRL) method can be used to calibrate the unknown parameters in the reward function. Before this, the eigenvalues need to be standardized. According to the maximum entropy inverse reinforcement learning method, the probability of the trajectory can be calculated by the following formula:

[0157]

[0158] Where Z(ω)=∑ζ′ exp(∑ t r ω (t)) is the partition function.

[0159] Then, the parameter vector is calculated using maximum likelihood estimation:

[0160]

[0161] Where ζ represents the driver's following trajectory.

[0162] Therefore, the present invention can use the car-following trajectories of different driving styles as simulation objects, learn the reward functions suitable for different driving styles, and thus derive the target reward functions corresponding to different drivers, so as to score each candidate car-following trajectory and derive the target reward function value corresponding to the candidate car-following trajectory.

[0163] In real-world driving, there are many uncertainties, such as sudden lane changes by other vehicles and sudden changes in road conditions. The target reward function takes into account multiple factors, such as driving safety, driving efficiency, and speed fluctuations. By comparing the target reward function values of different candidate car-following trajectories, we can intuitively determine which trajectory is superior and make accurate decisions.

[0164] In a preferred embodiment, Figure 5 As shown in Figure 2, the figure shows the processing process of the reward function learning module, which is as follows:

[0165] Input the car-following trajectory data into the reward function learning module;

[0166] The input car-following trajectory data is processed by the human-like reward function feature extractor to extract the human-like reward function features.

[0167] The expert trajectory characteristics and the initial value of the reward function are obtained, and the human-like reward function characteristics, the expert trajectory characteristics and the initial value of the reward function are input into the maximum entropy inverse reinforcement learning module for processing. Through maximum entropy inverse reinforcement learning, the reward function learning result is finally generated.

[0168] Therefore, the above process extracts the features of the car-following trajectory data, combines the expert trajectory features and the initial value of the reward function, and uses the maximum entropy inverse reinforcement learning method to complete the learning and optimization of the reward function, providing a reasonable reward function basis for the subsequent vehicle following trajectory generation.

[0169] In step S4, a trajectory that best matches a human's driving style can be selected from a multitude of possible trajectories. In actual application, when the target vehicle is driving on the road, a series of future trajectory target point sampling results are calculated based on the current state. Multiple candidate car-following trajectories are then fitted using polynomials. Then, based on reward functions for different driving styles, reward values are calculated for these candidate trajectories. The trajectory with the highest reward value is selected as the one most likely to be chosen by a human driver under the current driving style, and this trajectory is executed, achieving personalized human-like car-following.

[0170] like Figure 6 As shown, based on the above-mentioned various embodiments of the method for generating the vehicle following trajectory, the present invention provides a corresponding embodiment of the device;

[0171] An embodiment of the present invention provides a device for generating a vehicle following trajectory, comprising: a driving style type acquisition module, a driving data acquisition module, a reward function value generation module, and a following trajectory determination module;

[0172] The driving style type acquisition module is configured to acquire a target driving style type that characterizes the current driver's behavioral characteristics when operating a target vehicle; wherein the target driving style type corresponds to a target reward function that measures the vehicle's driving safety, driving efficiency, and speed variation;

[0173] The driving data acquisition module is used to acquire first driving data corresponding to the target vehicle in the current time period, second driving data corresponding to the first adjacent vehicle arranged longitudinally to the target vehicle in the current time period, and third driving data corresponding to the second adjacent vehicle arranged transversely to the target vehicle in the current time period;

[0174] The reward function value generation module is configured to use the first driving data, the second driving data, and the third driving data as a state space, generate an action space containing a plurality of candidate car-following trajectories based on the first driving data in the state space, and generate a target reward function value corresponding to each candidate car-following trajectory in the action space based on the state space and the target reward function;

[0175] The car-following trajectory determination module is configured to use the candidate car-following trajectory with the largest target reward function value as the target car-following trajectory of the target vehicle in the next time period.

[0176] It should be noted that the device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without paying any creative effort.

[0177] Those skilled in the art can clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0178] Based on the above-mentioned various embodiments of the method for generating a vehicle following trajectory, the present invention provides corresponding embodiments of terminal equipment.

[0179] An embodiment of the present invention provides a terminal device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements a method for generating a vehicle following trajectory as described in any method embodiment of the present invention.

[0180] The terminal device may be a computing terminal device such as a desktop computer, a notebook computer, a palmtop computer, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.

[0181] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, connecting various parts of the entire terminal device using various interfaces and lines.

[0182] The memory can be used to store the computer program, and the processor implements various functions of the terminal device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device or other volatile solid-state storage device.

[0183] Based on the above-mentioned various embodiments of the method for generating a vehicle-following trajectory, the present invention provides corresponding embodiments of a storage medium.

[0184] An embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a method for generating a vehicle following trajectory as described in any method embodiment of the present invention.

[0185] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0186] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for generating a vehicle following trajectory, characterized in that: include: Obtaining a target driving style type for characterizing the current driver's behavioral characteristics when operating a target vehicle; wherein the target driving style type corresponds to a target reward function for measuring the vehicle's driving safety, driving efficiency, and speed variation; Acquire first driving data corresponding to a target vehicle in a current time period, second driving data corresponding to a first adjacent vehicle arranged longitudinally to the target vehicle in the current time period, and third driving data corresponding to a second adjacent vehicle arranged transversely to the target vehicle in the current time period; Using the first driving data, the second driving data, and the third driving data as a state space, generating an action space containing a plurality of candidate car-following trajectories based on the first driving data in the state space; and generating a target reward function value corresponding to each candidate car-following trajectory in the action space based on the state space and a target reward function; The candidate following trajectory with the largest target reward function value is used as the target following trajectory of the target vehicle in the next time period.

2. The method for generating a vehicle following trajectory according to claim 1, wherein: The first driving data includes: a current speed of the target vehicle, an acceleration of the target vehicle, and a first vehicle position of the target vehicle; The step of generating an action space including a plurality of candidate car-following trajectories based on the first driving data in the state space includes: Taking the current speed as the center and according to a preset speed value threshold, generating a speed sampling interval containing a plurality of speeds to be selected; For each to-be-selected speed in the speed sampling interval, different parameters of a preset polynomial model are solved based on the current speed, acceleration, first vehicle position, the to-be-selected speed, and a preset acceleration reference value; each solved parameter and the to-be-selected speed are then substituted into the polynomial model to generate a candidate vehicle position corresponding to each to-be-selected speed; wherein the polynomial model is used to generate a vehicle position corresponding to the target vehicle at the to-be-selected speed; each parameter is a different coefficient in the polynomial model; For each candidate vehicle position, a candidate following trajectory is generated based on the candidate vehicle position and the first vehicle position; wherein the candidate following trajectory includes: a candidate speed, a candidate acceleration, and a candidate vehicle position: Generate an action space based on each candidate car-following trajectory.

3. The method for generating a vehicle following trajectory according to claim 2, wherein: The second driving data includes: a second vehicle position of a first adjacent vehicle; the third driving data includes: a third vehicle position of a second adjacent vehicle; the target reward function includes: a first sub-function for measuring the driving safety of the vehicle, a second sub-function for measuring the driving efficiency of the vehicle, and a third sub-function for measuring the speed change amplitude of the vehicle; Generating a target reward function value corresponding to each candidate car-following trajectory in the action space according to the state space and the target reward function includes: For each candidate car-following trajectory, generating a driving safety value of the target vehicle based on the candidate speed, the candidate vehicle position, the second vehicle position of the first adjacent vehicle, the third vehicle position of the second adjacent vehicle, and the first sub-function; generating a driving efficiency value of the target vehicle according to the candidate speed and the second sub-function; generating a speed change amplitude value of the target vehicle according to the candidate acceleration and the third sub-function; A target reward function value corresponding to a candidate car-following trajectory is generated according to the driving safety value, the driving efficiency value, and the speed change amplitude value.

4. The method for generating a vehicle following trajectory according to claim 3, wherein: The first sub-function, the second function and the third function correspond to the first weight coefficient, the second weight coefficient and the third weight coefficient respectively; The method for generating a vehicle following trajectory further includes: According to the maximum likelihood estimation method, a first weight coefficient, a second weight coefficient, and a third weight coefficient in the target reward function are determined.

5. The method for generating a vehicle following trajectory according to claim 4, wherein: The process of generating the target driving style type includes: Acquire a plurality of historical following trajectory data corresponding to a historical vehicle driven by the current driver; wherein each of the historical following trajectory data includes: an average speed of the historical vehicle, an average acceleration of the historical vehicle, position data of the historical vehicle, and position data of a preceding vehicle in front of the historical vehicle; Each historical following trajectory data is input into a preset driving style prediction model, so that the driving style prediction model extracts driving characteristics for characterizing the driving state of the historical vehicle based on the average speed of the historical vehicle and the average acceleration of the historical vehicle in each historical following trajectory data; extracts following characteristics for characterizing the following behavior characteristics between the historical vehicle and the preceding vehicle based on the average speed of the historical vehicle, the position data of the historical vehicle and the position data of the preceding vehicle; and generates a target driving style type corresponding to the current driver based on the driving characteristics and the following characteristics.

6. The method for generating a vehicle following trajectory according to claim 5, wherein: The following characteristics include: a spacing characteristic used to characterize the spatial distance maintained by the driver from the vehicle ahead during the following process, and a time distance characteristic used to characterize the degree of time buffer for the driver to longitudinally follow the vehicle ahead; The extracting of the following feature for characterizing the following behavior between the historical vehicle and the preceding vehicle based on the average speed of the historical vehicle, the position data of the historical vehicle, and the position data of the preceding vehicle includes: Generate a spacing feature representing the spatial distance between the driver and the preceding vehicle during the following process based on historical vehicle position data and preceding vehicle position data; A time-distance feature is generated based on the average speed of historical vehicles, the position data of historical vehicles, and the position data of the vehicle ahead, which is used to characterize the time buffer degree of the driver for longitudinal following.

7. The method for generating a vehicle following trajectory according to claim 6, wherein: The training process of the driving style prediction model includes: Each driver sample and the historical car-following trajectory data sample corresponding to each driver sample are used as training samples; Taking each training sample and the actual driving style type corresponding to each training sample as input and the predicted driving style type of each training sample as output, the driving style prediction model to be trained is iteratively trained until the model converges to generate a preset driving style prediction model.

8. A device for generating a vehicle following trajectory, characterized in that: include: Driving style type acquisition module, driving data acquisition module, reward function value generation module and car-following trajectory determination module; The driving style type acquisition module is configured to acquire a target driving style type that characterizes the current driver's behavioral characteristics when operating a target vehicle; wherein the target driving style type corresponds to a target reward function that measures the vehicle's driving safety, driving efficiency, and speed variation; The driving data acquisition module is used to acquire first driving data corresponding to the target vehicle in the current time period, second driving data corresponding to the first adjacent vehicle arranged longitudinally to the target vehicle in the current time period, and third driving data corresponding to the second adjacent vehicle arranged transversely to the target vehicle in the current time period; The reward function value generation module is configured to use the first driving data, the second driving data, and the third driving data as a state space, generate an action space containing a plurality of candidate car-following trajectories based on the first driving data in the state space, and generate a target reward function value corresponding to each candidate car-following trajectory in the action space based on the state space and the target reward function; The car-following trajectory determination module is configured to use the candidate car-following trajectory with the largest target reward function value as the target car-following trajectory of the target vehicle in the next time period.

9. A terminal device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for generating a vehicle following trajectory according to any one of claims 1 to 7 is implemented.

10. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute the method for generating a vehicle following trajectory according to any one of claims 1 to 7.

Citation Information

Cited By

  • A driving behavior modeling method and system considering driving style

    CN122693423A