Vehicle Control Method, Device, Vehicle, Storage Medium and Program Product
Through the target action model trained based on user historical driving data, control parameters that meet user habits are generated, which solves the problem that the autonomous driving system cannot meet personalized driving habits and improves the autonomous driving experience.
Patent Information
- Application Number
- CN202411367755.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-09-27
AI Technical Summary
The autonomous driving system cannot meet the personalized driving habits of all users, resulting in a poor autonomous driving experience.
Through the target action model trained based on user historical driving data, target control parameters that conform to user driving habits are generated to control the movement of the vehicle.
It improves the autonomous driving experience, reduces user manual intervention, and is more in line with users' driving habits.
Smart Images

Figure CN118928373B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of vehicle control, and particularly to a vehicle control method, apparatus, vehicle, storage medium, and program product. Background Art
[0002] As an important direction for the future development of the automotive industry, autonomous driving technology is accelerating its development globally. With the development of autonomous driving technology, users have higher expectations for the experience of autonomous vehicles.
[0003] In related technologies, the autonomous driving system uses general driving parameters and cannot meet the personalized driving habits of all users. For example, for the control of adaptive cruise, the vehicle is controlled by a fixed number of following distances. Summary of the Invention
[0004] To overcome the problems existing in related technologies, the present disclosure provides a vehicle control method, apparatus, vehicle, storage medium, and program product to determine target control parameters based on the vehicle control habits of users, so as to better simulate the driving habits of users and improve the autonomous driving experience of users.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a vehicle control method, including:
[0006] Determine a target action corresponding to a target driving scenario;
[0007] Process target data through a target action model corresponding to the target action to obtain target control parameters, where the target data is data related to the execution of the target action, and the target action model is obtained by training a basic model based on a plurality of sample data, and the sample data is historical data related to the target action during the historical driving process of the target user;
[0008] Control the vehicle according to the target control parameters.
[0009] Optionally, the target action model is obtained through the following steps:
[0010] Obtain a plurality of sample data related to the target action during the historical driving process of the target user, where the sample data includes sample correlation parameters and actual control parameters;
[0011] Perform multiple rounds of iterative training on the basic model through the sample correlation parameters in the plurality of sample data;
[0012] After each round of training, obtain the predicted control parameters corresponding to this round of training;
[0013] Obtain the prediction loss corresponding to this round of training through the prediction control parameters corresponding to this round of training and the actual control parameters in the sample data corresponding to this round of training;
[0014] Optimize the basic model through the prediction loss corresponding to this round of training;
[0015] When the basic model meets the preset conditions, stop training to obtain the target action model.
[0016] Optionally, the target action includes a turn signal operation, the target action model includes a turn signal model, and the target data includes a target road condition and a target driving scenario;
[0017] The process of using the target action model corresponding to the target action to process the target data to obtain target control parameters includes:
[0018] Process the target road condition and the target driving scenario through the turn signal model to obtain the turn signal control parameters corresponding to the turn signal operation. The turn signal model is obtained by training a basic turn signal model based on multiple sample turn signal data. The sample turn signal data is the sample road condition, sample driving scenario, and actual turn signal parameters related to the turn signal operation of the target user during historical driving.
[0019] Optionally, the target action includes a pedal operation, the target action model includes a pedal model, and the target data includes a target speed and a target traffic condition;
[0020] The process of using the target action model corresponding to the target action to process the target data to obtain target control parameters includes:
[0021] Process the target speed and the target traffic condition through the pedal model to obtain the pedal control parameters corresponding to the pedal operation. The pedal model is obtained by training a basic pedal model based on multiple sample pedal data. The sample pedal data is the sample speed, sample traffic condition, and actual pedal parameters related to the pedal operation of the target user during historical driving.
[0022] Optionally, the target action includes a speed adjustment operation, the target action model includes a speed adjustment model, and the target data includes a target road section and a target traffic condition;
[0023] The process of using the target action model corresponding to the target action to process the target data to obtain target control parameters includes:
[0024] Processing the target road section and the target traffic condition through the speed adjustment model to obtain speed control parameters corresponding to the speed adjustment operation, where the speed adjustment model is obtained by training a basic speed model based on multiple sample speed data, and the sample speed data is sample road sections, sample traffic conditions, and actual speed parameters related to the speed adjustment operation during the historical driving of the target user.
[0025] Optionally, the target action includes a following distance adjustment operation, the target action model includes a following distance adjustment model, and the target data includes a target speed and a target traffic condition;
[0026] Processing the target data through the target action model corresponding to the target action to obtain target control parameters, including:
[0027] Processing the target speed and the target traffic condition through the following distance adjustment model to obtain following distance control parameters corresponding to the following distance adjustment operation, where the following distance adjustment model is obtained by training a basic distance model based on multiple sample distance data, and the sample distance data is sample speeds, sample traffic conditions, and actual distance parameters related to the following distance adjustment operation during the historical driving of the target user.
[0028] Optionally, the target action includes an obstacle avoidance operation, the target action model includes an obstacle avoidance model, and the target data includes a target obstacle and a target vehicle state;
[0029] Processing the target data through the target action model corresponding to the target action to obtain target control parameters, including:
[0030] Processing the target obstacle and the target vehicle state through the obstacle avoidance model to obtain an obstacle avoidance trajectory corresponding to the obstacle avoidance operation, where the obstacle avoidance model is obtained by training a basic obstacle avoidance model based on multiple sample obstacle avoidance data, and the sample obstacle avoidance data is sample obstacles, sample vehicle states, and actual obstacle avoidance trajectories related to the obstacle avoidance operation during the historical driving of the target user.
[0031] Optionally, determining the target action corresponding to the target driving scenario includes:
[0032] Determining a target control strategy according to vehicle driving parameters and environmental parameters;
[0033] Determining the target driving scenario according to the target control strategy;
[0034] Based on the association relationship between the driving scenario and the execution action, determining the target action corresponding to the target driving scenario, and one target driving scenario corresponds to at least one target action.
[0035] Optionally, determining a target control strategy according to vehicle driving parameters and environmental parameters includes:
[0036] Processing the vehicle driving parameters and the environmental parameters through a target policy model to obtain the target control strategy, where the target policy model is obtained by training a basic policy model based on a plurality of sample policy data, and the sample policy data is historical data related to intervention actions of the vehicle with the target user in the autonomous driving mode.
[0037] According to a second aspect of the embodiments of the present disclosure, a vehicle control device is provided, including:
[0038] A determination module configured to determine a target action corresponding to a target driving scenario;
[0039] A first acquisition module configured to process target data through a target action model corresponding to the target action to obtain target control parameters, where the target data is data related to executing the target action, and the target action model is obtained by training a basic model based on a plurality of sample data, and the sample data is historical data related to the target action of the target user in the historical driving process;
[0040] A control module configured to control the vehicle according to the target control parameters.
[0041] According to a third aspect of the embodiments of the present disclosure, a vehicle is provided, including:
[0042] A processor;
[0043] A memory for storing instructions executable by the processor;
[0044] Wherein, the processor is configured to execute the steps of the vehicle control method provided in the first aspect of the present disclosure when executed.
[0045] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the steps of the vehicle control method provided in the first aspect of the present disclosure are implemented.
[0046] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps of the vehicle control method provided in the first aspect of the present disclosure are implemented.
[0047] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0048] First, determine the target action corresponding to the target driving scenario, and then process the target data through the target action model corresponding to the target action to obtain the target control parameter. The target data is data related to the execution of the target action, and the target action model is obtained by training a basic model based on multiple sample data. The sample data is historical data related to the target action during the target user's historical driving process. Then, control the vehicle according to the target control parameter, so as to obtain a vehicle control method that better conforms to the user's driving habits.
[0049] Among them, by using the historical data related to the target action during the target user's historical driving process as sample data to train the basic model, a target action model corresponding to the target action and conforming to the user's driving habits can be obtained. Then, when the vehicle needs to perform the corresponding target action, the target action model processes the data related to the execution of the target action, and the target control parameter conforming to the user's driving habits can be obtained. By controlling the vehicle with the target control parameter, the user's driving habits can be better simulated, the user's manual intervention can be minimized, and the user's autonomous driving experience can be improved.
[0050] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0051] The accompanying drawings here are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0052] Figure 1 It is a schematic diagram of a scenario of a vehicle control method shown according to an exemplary embodiment.
[0053] Figure 2 It is a flowchart of a vehicle control method shown according to an exemplary embodiment.
[0054] Figure 3 It is a block diagram of a vehicle control device shown according to an exemplary embodiment.
[0055] Figure 4 It is a block diagram of a vehicle shown according to an exemplary embodiment. Detailed Description of the Embodiments
[0056] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are only examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0057] It should be noted that all actions of obtaining signals, information or data in this disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining authorization from the owner of the corresponding device.
[0058] In the related art, the autonomous driving system uses general driving parameters and cannot meet the personalized driving habits of all users. For example, for the control of adaptive cruise, the vehicle is controlled by a fixed number of following distances.
[0059] In view of the above technical problems, the embodiments of the present disclosure provide a vehicle control method, device, vehicle, storage medium and program product. By using the historical data related to the target action during the historical driving process of the target user as sample data to train the basic model, a target action model corresponding to the target action and conforming to the driving habits of the user can be obtained. Furthermore, when the vehicle needs to perform the corresponding target action, the data related to the execution of the target action is processed through the target action model, and then the target control parameters conforming to the driving habits of the user can be obtained, and the vehicle is controlled by the target control parameters, so as to better simulate the driving habits of the user, minimize the manual intervention of the user, and improve the user's autonomous driving experience.
[0060] Figure 1 is a schematic diagram of a scenario of a vehicle control method shown according to an exemplary embodiment, as Figure 1 shown, the vehicle driven by the target user, that is, this vehicle, is traveling on the road and may encounter driving scenarios such as lane change, overtaking lane change, lane keeping, turning, obstacle avoidance, etc.
[0061] Figure 2 is a flowchart of a vehicle control method shown according to an exemplary embodiment, as Figure 2 shown, which may include the following steps.
[0062] In step S201, determine the target action corresponding to the target driving scenario.
[0063] During the driving process of the vehicle, various different driving scenarios will be encountered, such as lane change, overtaking lane change, lane keeping, turning, obstacle avoidance, etc. Each driving scenario requires at least one action to be completed. The association relationship between different driving scenarios and execution actions can be established in advance, and then, when the target driving scenario is determined, the target action corresponding to the target driving scenario is determined through this association relationship.
[0064] In step S202, the target data is processed by the target action model corresponding to the target action to obtain target control parameters. The target data is data related to the execution of the target action, and the target action model is obtained by training a basic model based on multiple sample data. The sample data is historical data related to the target action during the historical driving process of the target user.
[0065] In this embodiment, by using the historical data related to the target action during the historical driving process of the target user as sample data to train the basic model, a target action model corresponding to the target action and conforming to the user's driving habits can be obtained. Furthermore, when the vehicle needs to perform the corresponding target action, by processing the data related to the execution of the target action through the target action model, target control parameters conforming to the user's driving habits can be obtained, so as to better simulate the user's driving habits. Among them, the basic model can be MILE (Model - Based Imitation Learning for Urban Driving, an imitation learning framework).
[0066] In step S203, the vehicle is controlled according to the target control parameters.
[0067] In this embodiment, the vehicle is controlled by the target control parameters, so as to better simulate the user's driving habits, minimize the user's manual intervention, and improve the user's autonomous driving experience.
[0068] In a possible implementation manner, the target action model is obtained through the following steps:
[0069] Obtain multiple sample data related to the target action during the historical driving process of the target user. The sample data includes sample correlation parameters and actual control parameters; perform multiple rounds of iterative training on the basic model through the sample correlation parameters in the multiple sample data; after each round of training, obtain the predicted control parameters corresponding to this round of training; obtain the predicted loss corresponding to this round of training through the predicted control parameters corresponding to this round of training and the actual control parameters in the sample data of this round of training; optimize the basic model through the predicted loss corresponding to this round of training; when the basic model meets the preset conditions, stop training to obtain the target action model.
[0070] In this embodiment, the basic model can be iteratively trained multiple times through multiple sample data to continuously optimize the basic model, so that the basic model can meet the preset conditions and the output results can conform to the user's driving habits, thereby obtaining the target action model.
[0071] Specifically, the basic model can be trained with sample correlation parameters in multiple sample data. During each round of training, input the multiple sample correlation parameters corresponding to this round of training into the basic model, and the predicted control parameters corresponding to this round of training can be output through the basic model. Then, based on the predicted control parameters corresponding to this round of training and the actual control parameters in the sample data corresponding to this round of training, the predicted loss corresponding to this round of training can be obtained. The distance between the predicted control parameters corresponding to this round of training and the actual control parameters in the sample data corresponding to this round of training can be calculated, and this distance can be used as the corresponding predicted loss. And based on the gradient descent algorithm, the parameters in the basic model can be optimized through this predicted loss. Each round of training can be optimized once until the basic model meets the preset conditions. The preset conditions can be that the number of training rounds reaches the preset number of rounds, or the basic model converges. When the basic model meets the preset conditions, the training can be stopped, and the basic model after the last optimization can be determined as the target action model. Target recognition can also be performed on the user. Considering that the driving habits of different users may be different, for the same target action, the corresponding historical data may also be different. For the vehicle driven by the target user, the historical data related to the target action during the target user's historical driving process needs to be used as sample data to train the corresponding target action model.
[0072] By training the basic model with the relevant data during the user's manual driving of the vehicle when performing the target action, a target action model corresponding to the target action that conforms to the user's driving habits can be trained, so that the target control parameters that conform to the user's driving habits can be obtained based on this target action model to control the vehicle.
[0073] In a possible implementation manner, the target action includes a turn signal operation, the target action model includes a turn signal model, and the target data includes the target road conditions and the target driving scenario;
[0074] Processing the target data through the target action model corresponding to the target action to obtain the target control parameters, including:
[0075] Processing the target road conditions and the target driving scenario through the turn signal model to obtain the turn signal control parameters corresponding to the turn signal operation. The turn signal model is obtained by training the basic turn signal model with multiple sample turn data. The sample turn data is the sample road conditions, sample driving scenarios, and actual turn signal parameters related to the turn signal operation during the target user's historical driving process.
[0076] In this embodiment, the target action may include the operation of the turn signal, which can be used in scenarios such as turning, lane changing, obstacle avoidance, and overtaking lane changing. The basic turn signal model can be trained based on the historical data related to the turn signal operation of the target user during the historical driving process. The specific training method can refer to the training method for obtaining the target action model mentioned above, which will not be elaborated here. Among them, the sample road conditions and sample driving scenarios are the sample correlation parameters, and the actual turn signal parameters are the actual control parameters. The target road conditions can be one of multiple road grades.
[0077] Among them, the actual turn signal parameters can be the parameters related to the operation of the turn signal by the target user during driving, which may include the operation L of using the turn signal, such as operating the left turn or right turn, the operation timing is T, the frequency is F, and the turn signal duration is D. Then the turn signal operation of the user can be represented by the function g(L, T, F, D).
[0078] The turn signal operation simulated by the autonomous driving system for the user's operation can be expressed by the following function:
[0079] L_auto = g(L_user)×K
[0080] Among them, K is an adjustment coefficient, and different target data correspond to different values of K. L_user represents the user's operation habits, that is, (L, T, F, D), and L_auto is the operation simulated by the system, that is, the turn signal control parameter.
[0081] By processing the target road conditions and target driving scenarios through the turn signal model, the turn signal control parameters corresponding to the turn signal operation can be obtained. The turn signal control parameters can include the operation of using the turn signal, the operation timing, the frequency, and the turn signal duration.
[0082] In a possible embodiment, the target action includes the pedal operation, the target action model includes the pedal model, and the target data includes the target speed and target traffic conditions;
[0083] By processing the target data through the target action model corresponding to the target action, the target control parameters are obtained, including:
[0084] By processing the target speed and target traffic conditions through the pedal model, the pedal control parameters corresponding to the pedal operation are obtained. The pedal model is obtained by training the basic pedal model based on multiple sample pedal data. The sample pedal data is the sample speed, sample traffic conditions, and actual pedal parameters related to the pedal operation of the target user during the historical driving process.
[0085] In this embodiment, the target action may include pedal operation, which can be used in scenarios such as turning, lane changing, obstacle avoidance, lane keeping, and overtaking lane changing. The basic pedal model can be trained based on the historical data related to pedal operation during the target user's historical driving process. The specific training method can refer to the above-mentioned training method for obtaining the target action model, which will not be elaborated here. Among them, the sample speed and sample traffic conditions are the sample correlation parameters, and the actual pedal parameters are the actual control parameters.
[0086] The pedal includes an accelerator pedal and a brake pedal. The actual pedal parameters can be the operation details of the target user on the accelerator and brake pedals when accelerating or decelerating, such as pedal depth, smoothness of operation, and reaction time. The target speed is the speed when performing the pedal operation, and the target traffic condition can be one of a red light or the deceleration of the vehicle ahead.
[0087] The depth of the accelerator pedal is P_throttle, the depth of the brake pedal is P_brake, the smoothness of operation is S, and the reaction time is R. The user's pedal operation can be described by the function h(P_throttle, P_brake, S, R).
[0088] The simulation of the user's pedal operation by the vehicle's autonomous driving system can be expressed by the following function:
[0089] P_auto = h(P_user)×β
[0090] Where β is an adaptation coefficient, and different target data correspond to different values of β. P_user represents the user's pedal operation habit, that is, (P_throttle, P_brake, S, R), and P_auto is the operation intensity simulated by the system, that is, the pedal control parameter.
[0091] By processing the target speed and target traffic conditions through the pedal model, the pedal control parameters corresponding to the pedal operation can be obtained. The pedal control parameters can include pedal depth, smoothness of operation, and reaction time.
[0092] In a possible implementation manner, the target action includes a speed adjustment operation, the target action model includes a speed adjustment model, and the target data includes a target road section and a target traffic condition;
[0093] By processing the target data through the target action model corresponding to the target action, target control parameters are obtained, including:
[0094] The target road section and target traffic conditions are processed by a speed adjustment model to obtain speed control parameters corresponding to speed adjustment operations. The speed adjustment model is obtained by training a basic speed model based on multiple sample speed data. The sample speed data are the sample road sections, sample traffic conditions, and actual speed parameters related to speed adjustment operations during the target user's historical driving process.
[0095] In this embodiment, the target action may include a speed adjustment operation, which can be used in scenarios such as turning, lane changing, obstacle avoidance, lane keeping, and overtaking lane changing. And the speed adjustment can be completed in combination with pedal operations. The basic speed model can be trained based on the historical data related to speed adjustment operations during the target user's historical driving process. The specific training method can refer to the training method for obtaining the target action model above and will not be elaborated here. Among them, the sample road section and sample traffic conditions are the sample correlation parameters, and the actual speed parameter is the actual control parameter.
[0096] The actual speed parameter can be the speed after speed adjustment when the target user is driving manually or in autonomous driving. The target road section is the road section where the user is when making a speed adjustment, such as an urban road or a highway. The target traffic condition can be the traffic condition where the user is when making a speed adjustment, such as congestion or smooth traffic.
[0097] Through a machine learning algorithm, the basic speed model is trained to learn the speed preferences of the target user under different road sections and different traffic states, and a speed adjustment model is obtained to obtain the speed that the target user may expect under specific road conditions, that is, the speed control parameter, such as the cruise speed during adaptive cruise. This enables the vehicle to automatically adjust to the user's preferred speed without explicit instructions, improving the user's driving experience.
[0098] The speed preference of the target user is V_pref, and the speed selection function under different road sections L and different traffic conditions C is v(V_pref, L, C).
[0099] The speed adjustment operation of the vehicle's autonomous driving system simulating the user can be expressed by the following function:
[0100] V_cruise = v(V_pref, L, C) + δ
[0101] Among them, δ is a deviation value adjusted according to historical data, and different target data correspond to different values of δ, which is used for fine-tuning to be closer to the speed preference of the target user. V_cruise is the speed control parameter.
[0102] Process the target road section and target traffic conditions through a speed adjustment model to obtain speed control parameters corresponding to speed adjustment operations. The speed control parameters may include the desired speed.
[0103] In a possible implementation, the target action includes a following distance adjustment operation, the target action model includes a following distance adjustment model, and the target data includes the target speed and target traffic conditions;
[0104] Process the target data through the target action model corresponding to the target action to obtain target control parameters, including:
[0105] Process the target speed and target traffic conditions through the following distance adjustment model to obtain following distance control parameters corresponding to the following distance adjustment operation. The following distance adjustment model is obtained by training a basic speed model based on multiple sample distance data. The sample distance data is the sample speed, sample traffic conditions, and actual distance parameters related to the following distance adjustment operation during the target user's historical driving process.
[0106] In this implementation, the target action may include a following distance adjustment operation, and the following distance adjustment operation can be used in the scenario of lane keeping. The basic speed model can be trained based on the historical data related to the following distance adjustment operation during the target user's historical driving process. The specific training method can refer to the training method for obtaining the target action model above and will not be elaborated here. Among them, the sample speed and sample traffic conditions are the sample correlation parameters, and the actual distance parameter is the actual control parameter.
[0107] The actual distance parameter can be the speed after the following distance adjustment by the target user in the case of manual driving or autonomous driving. The target speed is the speed when the user makes a following distance adjustment, and the target traffic condition can be the traffic condition when the user makes a following distance adjustment, such as congestion or smooth traffic.
[0108] Through a machine learning algorithm, train the basic distance model to learn the following distance preference of the target user at different speeds and different traffic states, and obtain the following distance adjustment model, so as to obtain the following distance that the target user may expect under specific road conditions, that is, the following distance control parameter, such as the following distance during adaptive cruise. This enables the vehicle to automatically adjust to the user's preferred following distance without explicit instructions, improving the user's driving experience.
[0109] Process the target speed and target traffic conditions through the following distance adjustment model to obtain the following distance control parameters corresponding to the following distance adjustment operation. The following distance control parameter can be the desired following distance.
[0110] In a possible implementation, the target action includes an obstacle avoidance operation, the target action model includes an obstacle avoidance model, and the target data includes target obstacles and target vehicle states;
[0111] Processing the target data through the target action model corresponding to the target action to obtain target control parameters, including:
[0112] Processing the target obstacles and target vehicle states through the obstacle avoidance model to obtain an obstacle avoidance trajectory corresponding to the obstacle avoidance operation. The obstacle avoidance model is obtained by training a basic obstacle avoidance model based on multiple sample obstacle avoidance data. The sample obstacle avoidance data is the sample obstacles, sample vehicle states, and actual obstacle avoidance trajectories related to the obstacle avoidance operation during the target user's historical driving process.
[0113] In this implementation, the target action may include an obstacle avoidance operation, and the obstacle avoidance operation can be used in the scenario of obstacle avoidance. The basic obstacle avoidance model can be trained based on the historical data related to the obstacle avoidance operation during the target user's historical driving process. The specific training method can refer to the training method for obtaining the target action model above and will not be elaborated here. Among them, the sample obstacles and sample vehicle states are the sample correlation parameters, and the actual obstacle avoidance trajectory is the actual control parameter.
[0114] The actual obstacle avoidance trajectory can be the trajectory when the target user is driving manually and performing obstacle avoidance. The target obstacle is the obstacle that the user avoids during obstacle avoidance, such as road damage or obstacles on the road. The target vehicle state can be the state of the vehicle when the user performs obstacle avoidance, such as vehicle speed (v_vehicle), steering wheel angle (θ_steering), brake pressure (p_brake), and accelerator pedal position (p_accel), etc.
[0115] Through machine learning algorithms, train the basic obstacle avoidance model to learn the obstacle avoidance preferences of the target user under different obstacles and different vehicle states, and obtain the obstacle avoidance model, so as to obtain the obstacle avoidance trajectory that the target user may expect under specific road conditions, such as the obstacle avoidance trajectory when encountering a stationary stone or a stationary vehicle on the road. This enables the vehicle to automatically adjust to the user-preferred obstacle avoidance trajectory without explicit instructions, improving the user's driving experience. During the training process, an ideal obstacle avoidance trajectory (T_ideal) and the time interval (t_reaction) corresponding to the time from when the user detects an obstacle to when they start taking action can also be added to ensure the safety of autonomous driving, so as to conform to the user's obstacle avoidance habits as much as possible while ensuring driving safety.
[0116] In a possible implementation, determining the target action corresponding to the target driving scenario includes:
[0117] Determine a target control strategy based on vehicle driving parameters and environmental parameters; determine a target driving scenario according to the target control strategy; based on the association relationship between the driving scenario and the execution action, determine the target action corresponding to the target driving scenario, and one target driving scenario corresponds to at least one target action.
[0118] In this embodiment, the scenarios during vehicle driving can be divided into multiple scenarios, such as lane change, overtaking lane change, lane keeping, turning, obstacle avoidance and other driving scenarios. Each driving scenario requires at least one action to complete. The association relationship between different driving scenarios and execution actions can be established in advance. The vehicle driving parameters and environmental parameters can be obtained in advance to determine the target control strategy. The vehicle driving parameters may include driving speed, acceleration, steering wheel angle, brake pressure, accelerator pedal position, etc. The environmental parameters may include the lane where the vehicle is located, road conditions, the distances from the vehicle to the vehicle in front and behind, the road section where the vehicle is located, the position coordinates of the vehicle, and the position coordinates of the vehicles around the vehicle, etc. One target driving scenario corresponds to at least one target action. For example, when the target driving scenario is lane change, overtaking lane change or turning, the corresponding target actions may include turn signal operation, pedal operation and speed adjustment operation; when the target driving scenario is lane keeping, the corresponding target actions may include pedal operation, speed adjustment operation or following distance adjustment operation; when the target driving scenario is obstacle avoidance, the corresponding target action may include obstacle avoidance operation.
[0119] After determining the target control strategy, the target driving scenario can be determined according to the state of the current vehicle and the state of the vehicle corresponding to the target control strategy, and then based on the association relationship between the driving scenario and the execution action, the target action corresponding to the target driving scenario can be determined.
[0120] In a possible implementation manner, determining the target control strategy according to the vehicle driving parameters and environmental parameters includes:
[0121] Process the vehicle driving parameters and environmental parameters through a target policy model to obtain the target control strategy. The target policy model is obtained by training a basic policy model based on multiple sample policy data, and the sample policy data is historical data related to the intervention actions of the target user when the vehicle is in the autonomous driving mode.
[0122] In this embodiment, a plurality of sample policy data related to intervention actions with the target user in the autonomous driving mode can be obtained. The sample data includes sample policy parameters and actual policy parameters. The basic policy model is trained iteratively for multiple rounds through the sample policy parameters in the plurality of sample policy data. After each round of training, the predicted policy parameters corresponding to this round of training are obtained. Through the predicted policy parameters corresponding to this round of training and the actual policy parameters in the sample policy data corresponding to this round of training, the predicted policy loss corresponding to this round of training is obtained. The basic policy model is optimized through the predicted policy loss corresponding to this round of training. When the basic policy model meets the preset conditions, the training is stopped to obtain the target policy model. Among them, the sample policy parameters may include sample driving parameters and sample environmental parameters. The basic policy model can also be MILE.
[0123] The trained target policy model can take the driving parameters and environmental parameters as inputs and output control strategies, so as to simulate the vehicle control strategies of the target user in different situations. By processing the vehicle driving parameters and environmental parameters through the target policy model, the target control strategy can be obtained.
[0124] Based on the manual intervention of the target user in the autonomous driving mode, for example, suddenly controlling the steering wheel or changing the vehicle speed, etc., the driving parameters and environmental parameters of these events can be recorded, such as context information, including but not limited to speed, surrounding vehicle position coordinates, road conditions, and vehicle states before and after the intervention.
[0125] Using machine learning algorithms, the vehicle can analyze the patterns of these intervention behaviors and identify possible driving preferences or dissatisfaction with the performance of the automatic system behind them. In this way, the system can adjust the autonomous driving strategy to better simulate the user's driving habits and reduce the need for future interventions.
[0126] Let the intervention action be An, the vehicle states before and after the intervention be S_pre and S_post, and the relevant context information, such as speed V, surrounding vehicle position coordinates P, and road conditions C be Context. The vehicle adjustment parameter can be expressed as a function f(An, S_pre, S_post, V, P, Context).
[0127] The target control strategy P_auto^n can be transformed through a function in the following form:
[0128]
[0129] Among them, α is the learning rate, represents the gradient of the relevant parameters with respect to the intervention behavior function, and P_auto is the current control strategy.
[0130] In a possible implementation, when the target actions are multiple actions, the target control parameters corresponding to each target action can be determined respectively, and the target control parameters corresponding to all target actions are determined as the total control parameters. Then, the vehicle is controlled by the total control parameters to achieve multiple target actions in the target driving scenario.
[0131] Figure 3 It is a block diagram of a vehicle control device shown according to an exemplary embodiment. Refer to Figure 3 As shown, the vehicle control device 300 includes a determination module 301, a first acquisition module 302, and a control module 303.
[0132] The determination module 301 is configured to determine the target actions corresponding to the target driving scenario;
[0133] The first acquisition module 302 is configured to process the target data through the target action model corresponding to the target action to obtain the target control parameters. The target data is the data related to the execution of the target action, and the target action model is obtained by training a basic model based on multiple sample data. The sample data is the historical data related to the target action during the historical driving of the target user;
[0134] The control module 303 is configured to control the vehicle according to the target control parameters.
[0135] Optionally, the vehicle control device 300 further includes:
[0136] An acquisition module, configured to acquire multiple sample data related to the target action during the historical driving of the target user. The sample data includes sample association parameters and actual control parameters;
[0137] A training module, configured to perform multiple rounds of iterative training on the basic model through the sample association parameters in the multiple sample data;
[0138] A second acquisition module, configured to obtain the predicted control parameters corresponding to this round of training after each round of training;
[0139] A third acquisition module, configured to obtain the predicted loss corresponding to this round of training through the predicted control parameters corresponding to this round of training and the actual control parameters in the sample data corresponding to this round of training;
[0140] An optimization module, configured to optimize the basic model through the predicted loss corresponding to this round of training;
[0141] A fourth acquisition module, configured to stop training and obtain the target action model when the basic model meets the preset conditions.
[0142] Optionally, the target action includes a turn signal operation, the target action model includes a turn signal model, and the target data includes a target road condition and a target driving scenario;
[0143] The first obtaining module 302 includes:
[0144] A first obtaining sub-module, configured to process the target road condition and the target driving scenario through the turn signal model to obtain a turn signal control parameter corresponding to the turn signal operation, where the turn signal model is obtained by training a basic turn signal model based on a plurality of sample turn data, and the sample turn data is the sample road condition, the sample driving scenario, and the actual turn signal parameter related to the turn signal operation during the historical driving process of the target user.
[0145] Optionally, the target action includes a pedal operation, the target action model includes a pedal model, and the target data includes a target speed and a target traffic condition;
[0146] The first obtaining module 302 includes:
[0147] A second obtaining sub-module, configured to process the target speed and the target traffic condition through the pedal model to obtain a pedal control parameter corresponding to the pedal operation, where the pedal model is obtained by training a basic pedal model based on a plurality of sample pedal data, and the sample pedal data is the sample speed, the sample traffic condition, and the actual pedal parameter related to the pedal operation during the historical driving process of the target user.
[0148] Optionally, the target action includes a speed adjustment operation, the target action model includes a speed adjustment model, and the target data includes a target road section and a target traffic condition;
[0149] The first obtaining module 302 includes:
[0150] A third obtaining sub-module, configured to process the target road section and the target traffic condition through the speed adjustment model to obtain a speed control parameter corresponding to the speed adjustment operation, where the speed adjustment model is obtained by training a basic speed model based on a plurality of sample speed data, and the sample speed data is the sample road section, the sample traffic condition, and the actual speed parameter related to the speed adjustment operation during the historical driving process of the target user.
[0151] Optionally, the target action includes a following distance adjustment operation, the target action model includes a following distance adjustment model, and the target data includes a target speed and a target traffic condition;
[0152] The first obtaining module 302 includes:
[0153] A fourth obtaining sub-module, configured to process the target speed and the target traffic condition through the following-distance adjustment model to obtain a following-distance control parameter corresponding to the following-distance adjustment operation, where the following-distance adjustment model is obtained by training a basic distance model based on a plurality of sample distance data, and the sample distance data is sample speed, sample traffic condition, and actual distance parameter related to the following-distance adjustment operation of a target user during a historical driving process.
[0154] Optionally, the target action includes an obstacle avoidance operation, the target action model includes an obstacle avoidance model, and the target data includes a target obstacle and a target vehicle state;
[0155] The first obtaining module 302 includes:
[0156] A fifth obtaining sub-module, configured to process the target obstacle and the target vehicle state through the obstacle avoidance model to obtain an obstacle avoidance trajectory corresponding to the obstacle avoidance operation, where the obstacle avoidance model is obtained by training a basic obstacle avoidance model based on a plurality of sample obstacle avoidance data, and the sample obstacle avoidance data is sample obstacle, sample vehicle state, and actual obstacle avoidance trajectory related to the obstacle avoidance operation of a target user during a historical driving process.
[0157] Optionally, the determining module 301 includes:
[0158] A first determining sub-module, configured to determine a target control strategy according to vehicle driving parameters and environmental parameters;
[0159] A second determining sub-module, configured to determine the target driving scenario according to the target control strategy, where the target driving scenario includes one of overtaking and lane changing, lane keeping, and obstacle avoidance;
[0160] A third determining sub-module, configured to determine a target action corresponding to the target driving scenario based on an association relationship between a driving scenario and an execution action, and one target driving scenario corresponds to at least one target action.
[0161] Optionally, the first determining sub-module includes:
[0162] An obtaining unit, configured to process the vehicle driving parameters and the environmental parameters through a target policy model to obtain the target control strategy, where the target policy model is obtained by training a basic policy model based on a plurality of sample policy data, and the sample policy data is historical data related to an intervention action of the target user when the vehicle is in an automatic driving mode.
[0163] Regarding the vehicle control device 300 in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0164] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the vehicle control method provided by the present disclosure are implemented.
[0165] Figure 4 is a block diagram of a vehicle shown according to an exemplary embodiment. For example, the vehicle 400 may be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 400 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.
[0166] Referring to Figure 4 , the vehicle 400 may include various subsystems. For example, the infotainment system 410, the perception system 420, the decision control system 430, the drive system 440, and the computing platform 450. Among them, the vehicle 400 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 400 may be interconnected by wired or wireless means.
[0167] In some embodiments, the infotainment system 410 may include a communication system, an entertainment system, a navigation system, and the like.
[0168] The perception system 420 may include several sensors for sensing information about the environment around the vehicle 400. For example, the perception system 420 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter wave radar, ultrasonic radar, and a camera device.
[0169] The decision control system 430 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.
[0170] The drive system 440 may include components that provide power motion for the vehicle 400. In one embodiment, the drive system 440 may include an engine, an energy source, a powertrain, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine can convert the energy provided by the energy source into mechanical energy.
[0171] Some or all functions of the vehicle 400 are controlled by the computing platform 450. The computing platform 450 may include at least one processor 451 and a memory 452, and the processor 451 may execute instructions 453 stored in the memory 452.
[0172] The processor 451 can be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.
[0173] The memory 452 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0174] In addition to the instructions 453, the memory 452 can also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 452 can be used by the computing platform 450.
[0175] In the embodiments of the present disclosure, the processor 451 can execute the instructions 453 to complete all or part of the steps of the above-described vehicle control method.
[0176] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program that can be executed by a programmable device. The computer program has a code portion for executing the above-described vehicle control method when executed by the programmable device.
[0177] Those skilled in the art can also understand that the various illustrative logical blocks and steps listed in the embodiments of the present application can be implemented by electronic hardware, computer software, or a combination of both. Whether such a function is implemented by hardware or software depends on the specific application and the design requirements of the entire system. For each specific application, those skilled in the art can use various methods to implement the described function, but such implementation should not be construed as exceeding the scope of protection of the embodiments of the present application.
[0178] It should be understood that unless otherwise specifically stated, the features of the various embodiments of the present disclosure described herein can be combined with each other.
[0179] Although terms such as "first", "second", and "third" may be used herein to describe various components, parts, regions, layers, or sections, these components, parts, regions, layers, or sections are not limited to these terms. On the contrary, these terms are only used to distinguish one component, part, region, layer, or section from another. Thus, the first component, part, region, layer, or section mentioned in the examples described herein may also be referred to as the second component, part, region, layer, or section without departing from the teachings of the various examples. Additionally, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description herein, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0180] In addition, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. On the contrary, the use of the word exemplary is intended to present concepts in a concrete fashion. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified, or clear from the context, "X applies A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies A; X applies B; or X applies both A and B, then "X applies A or B" is satisfied under any one of the foregoing instances. Additionally, unless otherwise specified or clear from the context that it is referring to the singular form, the articles "a" and "an" as used in this application and the appended claims are generally understood to mean "one or more".
[0181] Similarly, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding the specification and the drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. Specifically with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. Additionally, although a particular feature of the present disclosure may have been disclosed with respect to only one of several implementations, such a feature may be combined with one or more other features of the other implementations as may be desired and advantageous for any given or particular application. Further, with respect to the terms "comprising", "having", "including", "containing", or variations thereof as used in the detailed description or the claims, such terms are intended to be inclusive in a manner similar to the term "including".
[0182] Other embodiments of the present disclosure will readily occur to those of ordinary skill in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known or customary techniques in the art that are not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
[0183] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes may be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A vehicle control method, characterized in that, Including: Determine a target control strategy based on vehicle driving parameters and environmental parameters; Determine a target driving scenario according to the target control strategy; Based on the association relationship between the driving scenario and the execution action, determine at least one target action corresponding to the target driving scenario, where the driving scenario includes lane change, lane keeping, turning, and obstacle avoidance, and the at least one target action includes at least one of a turn signal operation, a pedal operation, a speed adjustment operation, a following distance adjustment operation, and an obstacle avoidance operation; Process target data through a target action model corresponding to the target action to obtain target control parameters, where the target data is data related to executing the target action, and the target action model is obtained by training a basic model based on multiple sample data, and the sample data is historical data related to the target action during the historical driving process of the target user; wherein, when the target action includes a turn signal operation, the target control parameter includes a turn signal control parameter corresponding to the turn signal operation, and the turn signal control parameter includes the operation of using the turn signal, the operation timing, the frequency, and the turn signal duration. When the target action includes a pedal operation, the target control parameter includes a pedal control parameter corresponding to the pedal operation, and the pedal control parameter includes the pedal depth, the smoothness of the operation, and the reaction time; Control the vehicle according to the target control parameters; Wherein, the determining the target control strategy based on the vehicle driving parameters and the environmental parameters includes: Process the vehicle driving parameters and the environmental parameters through a target strategy model to obtain the target control strategy, where the target strategy model is obtained by training a basic strategy model based on multiple sample strategy data, and the sample strategy data is historical data related to the intervention actions of the target user in the automatic driving mode.
2. The vehicle control method according to claim 1, wherein The target action model is obtained through the following steps: Obtain multiple sample data related to the target action during the historical driving process of the target user, where the sample data includes sample association parameters and actual control parameters; Perform multiple rounds of iterative training on the basic model through the sample association parameters in the multiple sample data; After each round of training, obtain the predicted control parameter corresponding to this round of training; Obtain the predicted loss corresponding to this round of training through the predicted control parameter corresponding to this round of training and the actual control parameter in the sample data corresponding to this round of training; Optimize the basic model through the predicted loss corresponding to this round of training; Stop training when the basic model meets the preset conditions to obtain the target action model.
3. The vehicle control method according to claim 1, characterized in that The target action includes a turn signal operation, the target action model includes a turn signal model, and the target data includes a target road condition and a target driving scenario; The processing the target data through the target action model corresponding to the target action to obtain the target control parameter includes: Processing the target road conditions and the target driving scenarios through the turn signal model to obtain the turn signal control parameters corresponding to the turn signal operation. The turn signal model is obtained by training a basic turn signal model based on multiple sample turn data. The sample turn data is the sample road conditions, sample driving scenarios, and actual turn signal parameters related to the turn signal operation of the target user during historical driving.
4. The vehicle control method according to claim 1, wherein, The target action includes a pedal operation. The target action model includes a pedal model. The target data includes a target speed and target traffic conditions. Processing the target data through the target action model corresponding to the target action to obtain target control parameters, including: Processing the target speed and the target traffic conditions through the pedal model to obtain the pedal control parameters corresponding to the pedal operation. The pedal model is obtained by training a basic pedal model based on multiple sample pedal data. The sample pedal data is the sample speed, sample traffic conditions, and actual pedal parameters related to the pedal operation of the target user during historical driving.
5. The vehicle control method according to claim 1, characterized in that The target action includes a speed adjustment operation. The target action model includes a speed adjustment model. The target data includes a target road section and target traffic conditions. Processing the target data through the target action model corresponding to the target action to obtain target control parameters, including: Processing the target road section and the target traffic conditions through the speed adjustment model to obtain the speed control parameters corresponding to the speed adjustment operation. The speed adjustment model is obtained by training a basic speed model based on multiple sample speed data. The sample speed data is the sample road section, sample traffic conditions, and actual speed parameters related to the speed adjustment operation of the target user during historical driving.
6. The vehicle control method according to claim 1, wherein The target action includes a following distance adjustment operation. The target action model includes a following distance adjustment model. The target data includes a target speed and target traffic conditions. Processing the target data through the target action model corresponding to the target action to obtain target control parameters, including: Processing the target speed and the target traffic conditions through the following distance adjustment model to obtain the following distance control parameters corresponding to the following distance adjustment operation. The following distance adjustment model is obtained by training a basic distance model based on multiple sample distance data. The sample distance data is the sample speed, sample traffic conditions, and actual distance parameters related to the following distance adjustment operation of the target user during historical driving.
7. The vehicle control method according to claim 1, wherein The target action includes an obstacle avoidance operation. The target action model includes an obstacle avoidance model. The target data includes a target obstacle and target vehicle state. Processing the target data through the target action model corresponding to the target action to obtain target control parameters, including: Process the target obstacle and the target vehicle state through the obstacle avoidance model to obtain the obstacle avoidance trajectory corresponding to the obstacle avoidance operation. The obstacle avoidance model is obtained by training a basic obstacle avoidance model based on multiple sample obstacle avoidance data. The sample obstacle avoidance data is the sample obstacles, sample vehicle states, and actual obstacle avoidance trajectories related to the obstacle avoidance operation during the target user's historical driving process.
8. A vehicle control device, characterized in that, Including: A first determination sub-module configured to determine a target control strategy according to vehicle driving parameters and environmental parameters; A second determination sub-module configured to determine a target driving scenario according to the target control strategy; A third determination sub-module configured to determine at least one target action corresponding to the target driving scenario based on the association relationship between the driving scenario and the execution action. The driving scenario includes lane change, lane keeping, turning, and obstacle avoidance. The at least one target action includes at least one of a turn signal operation, a pedal operation, a speed adjustment operation, a following distance adjustment operation, and an obstacle avoidance operation; A first acquisition module configured to process target data through a target action model corresponding to the target action to obtain target control parameters. The target data is data related to the execution of the target action. The target action model is obtained by training a basic model based on multiple sample data. The sample data is historical data related to the target action during the target user's historical driving process. Wherein, when the target action includes a turn signal operation, the target control parameter includes a turn signal control parameter corresponding to the turn signal operation. The turn signal control parameter includes the operation of using the turn signal, the operation timing, the frequency, and the turn signal duration. When the target action includes a pedal operation, the target control parameter includes a pedal control parameter corresponding to the pedal operation. The pedal control parameter includes the pedal depth, the smoothness of the operation, and the reaction time; A control module configured to control the vehicle according to the target control parameters; Wherein, the first determination sub-module is specifically configured to: Process the vehicle driving parameters and the environmental parameters through a target policy model to obtain the target control strategy. The target policy model is obtained by training a basic policy model based on multiple sample policy data. The sample policy data is historical data related to the intervention actions of the target user in the automatic driving mode of the vehicle.
9. A vehicle, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the steps of the vehicle control method according to any one of claims 1 to 7 when executed.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions implement the steps of the vehicle control method according to any one of claims 1 to 7 when executed by the processor.
11. A computer program product, characterized in that, Including a computer program that implements the steps of the vehicle control method according to any one of claims 1 to 7 when executed by the processor.
Citation Information
Patent Citations
Automatic driving method and device
CN111824157A