Automatic driving behavior decision self-adaptive correction method and system based on time sequence prediction
By constructing a hybrid traffic scenario and timing prediction model for autonomous driving, using IDM and MOBIL models to simulate human driving behavior and design adaptive action correction equations, the problem that traditional correction methods cannot adaptively adjust in complex environments is solved, and real-time decision-making correction and robustness improvement of autonomous driving vehicles are achieved.
Patent Information
- Application Number
- CN202510758427.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional self-driving behavior correction methods cannot adaptively to make decisions, and cannot achieve real-time, intelligent and rapid and effective adjustments under complex traffic environments and sensor disturbances, resulting in wrong decisions and traffic accident risks.
A hybrid traffic scenario for autonomous driving is constructed, and a hybrid traffic scenario is used to simulate human driving behavior using IDM and MOBIL models, a state space, action space and reward functions are designed, state perturbation is implemented through gradient descent, adaptive action correction equations are designed, and wrong decisions are corrected using the timing prediction model, decompose into multiple timing prediction tasks and train the LSTM network model.
Real-time decision-making and correction of autonomous vehicles in complex environments is realized, the dangers caused by state perception errors are reduced, the system's robustness and adaptability are improved, and new disturbances can be adapted to.
Smart Images

Figure CN120363930A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and specifically to an adaptive correction method and system for autonomous driving behavior decision-making based on temporal prediction. Background Art
[0002] As a core component of future intelligent transportation systems, autonomous vehicles herald a revolutionary change in road safety, traffic efficiency, and reduced congestion. With the continuous development of artificial intelligence technology, reinforcement learning, as a technology capable of self-learning and adjustment in dynamic environments, has become an important decision-making optimization tool in autonomous driving systems. In autonomous driving, reinforcement learning algorithms often rely on high-quality sensor data to make decisions, and the accuracy of sensors directly affects the performance and decision-making ability of reinforcement learning agents. Unfortunately, sensors are also one of the most vulnerable components in autonomous vehicles. In complex traffic environments, sensor data may be affected by factors such as weather, light, and perturbations from other vehicles, resulting in data errors or losses. These errors may cause reinforcement learning agents to make incorrect decisions, leading to traffic accidents.
[0003] Traditional action correction methods usually rely on preset rules or static models. These methods correct abnormal behaviors of autonomous driving agents by setting fixed compensation parameters or adjustment rules. Common practices include rule judgments based on the perception system or directly adjusting control strategies according to physical models. For example, when sensor perturbations occur, the system will make a rigid adjustment according to preset rules, attempting to restore the vehicle to a predetermined safe path. However, these methods have certain limitations. Traditional correction methods are usually only effective under fixed conditions and do not have the ability to adapt. They cannot flexibly adjust decisions according to environmental changes or the degree of sensor perturbations. Especially when facing dynamic changes in the environment and unforeseen sensor perturbations in complex traffic scenarios, it is difficult to achieve real-time, intelligent, fast, and effective adaptive adjustments.
[0004] Therefore, the present invention proposes an adaptive correction method and system for autonomous driving behavior decision-making based on temporal prediction. Summary of the Invention
[0005] The purpose of the present invention is to provide an adaptive correction method and system for autonomous driving behavior decision-making based on temporal prediction to solve the problems raised in the above background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: An adaptive correction method for autonomous driving behavior decision-making based on temporal prediction, including:
[0007] Construct an autonomous driving mixed traffic scenario, which includes autonomous vehicles and human-driven vehicles;
[0008] Construct a human-driven vehicle model for human-driven vehicles in a mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making and the MOBIL model is used to control its lateral decision-making, for simulating human driving behavior;
[0009] Based on autonomous vehicles in an autonomous driving mixed traffic scenario, design a state space, an action space, and a reward function, and train a pre-constructed reinforcement learning decision-making model for autonomous vehicles according to the state space, the action space, and the reward function;
[0010] Design a state perturbation agent for the reinforcement learning decision-making model, and implement state perturbation in a gradient descent manner, so as to induce the reinforcement learning decision-making model to output incorrect decisions;
[0011] For the incorrect decisions output by the reinforcement learning decision-making model, design an adaptive action correction equation for performing correction tasks. According to the adaptive action correction equation, decompose the correction tasks into multiple time series prediction tasks, construct a time series prediction model and train it for performing multiple time series prediction tasks, so as to correct the output incorrect decisions.
[0012] Furthermore, the autonomous driving mixed traffic scenario is built using the highway-env toolkit, and the autonomous driving mixed traffic scenario is specifically the highway ramp merging scenario.
[0013] Furthermore, construct a human-driven vehicle model for human-driven vehicles in a mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making and the MOBIL model is used to control its lateral decision-making, for simulating human driving behavior, specifically as follows:
[0014] (31) Control the longitudinal decision-making through the IDM model, and the control formula is as follows:
[0015]
[0016] In the formula is the calculated vehicle acceleration, a is the maximum allowable acceleration, v is the current speed, v0 is the desired speed, s * is the desired minimum following distance, s * is the following distance between the current vehicle and the vehicle in front, b is a comfortable deceleration of the vehicle, Δv is the speed difference between the current vehicle and the vehicle in front, and T is the safe time headway of the vehicle;
[0017] (32) Control the lateral decision-making through the MOBIL model, and the control formula is as follows:
[0018]
[0019] Where The acceleration of the vehicle in front of the current vehicle in the new lane after lane change, a n The acceleration of this vehicle before lane change The predicted acceleration of the current vehicle after lane change, a c The acceleration of the current vehicle before lane change The acceleration of the vehicle in front of the current vehicle in the current lane after lane change, a o The acceleration of this vehicle before lane change, Δa th The lane change threshold, b safe The safe deceleration threshold
[0020] Furthermore, based on the autonomous vehicle in the autonomous driving mixed traffic scenario, the state space and action space are designed as follows:
[0021] (41) The state space is {s1, s2, s3, …, s N} where represents the observation information of the i-th vehicle, whether present can be observed and represent the x and y coordinates of the position when observed, v x 、v y represent the x and y coordinates of the speed when observed;
[0022] (42) The action space is A = {a1, a2, a3, a4, a5}, including 5 discrete actions, where a1, a2, a3, a4, a5 represent the 5 discrete actions of turning left, turning right, remaining unchanged, accelerating, and decelerating respectively
[0023] Furthermore, the reward function is R t = R a + R v + R s - C t where
[0024]
[0025] In the formula, v ego is the current speed of the autonomous vehicle is the average speed of surrounding vehicles, R a is the reward for reaching the end coordinates, R v is the reward for maintaining a target speed, R s is the reward for successful lane change, C t is the cost of the autonomous vehicle at time t
[0026] Furthermore, a state perturbation agent for the reinforcement learning decision model is designed, and state perturbation is implemented by means of gradient descent. The perturbation objective function is as follows:
[0027]
[0028] In the formula, δ is the state perturbation matrix, π is the policy function of the autonomous vehicle, s is the current observation state of the autonomous vehicle, and a is the action decision output by its decision-making model.
[0029] Furthermore, for the incorrect decisions output by the reinforcement learning decision-making model, an adaptive action correction equation is designed as follows:
[0030] bias = ||Δ actual - Δ pred ||
[0031]
[0032] π final = α * π rl + (1 - α) * π alt
[0033] In the formula, Δ actual is the actual difference between the current state and the previous state, Δ pred is the predicted difference between the current state and the previous state, bias is their difference, α is the adjustment factor, k, τ are variables that control the change range of the adjustment factor, which can be determined by Bayesian optimization during training, π rl is the policy output by the reinforcement learning autonomous driving agent, π alt is the conservative policy under the predicted current state. The larger the bias, the greater the degree of disturbance of the autonomous vehicle state, the closer α tends to 0, and the final decision is more inclined to use the predicted conservative policy.
[0034] Furthermore, according to the adaptive action correction equation, the correction task is decomposed into two sequential prediction tasks:
[0035] The first sequential task is to predict the difference between the current state and the previous state, providing an important basis for state changes for subsequent behavior decisions;
[0036] The second sequential task is to predict a safe conservative policy under the current state, which is used to provide a safe behavior option for the autonomous driving system in complex road conditions.
[0037] Furthermore, the sequential prediction model is constructed and trained to perform multiple sequential prediction tasks. The sequential prediction model is two LSTM network models with an attention mechanism. The specific training process is as follows:
[0038] (91) Preprocess the data set and normalize the features;
[0039] (92) Initialize the network training parameters;
[0040] (93) Divide the dataset into a training set and a validation set according to 8:2;
[0041] (94) Use the backpropagation algorithm to train two LSTM network models on the training set, and optimize the model parameters by minimizing the loss function;
[0042] The loss function is:
[0043]
[0044] Where L Diff is the loss function for predicting the difference between the current state and the previous state, L Alt is the loss function for predicting a safe and conservative policy under the current state, Δ pred and are the output values of two time series prediction models, Δ actual and are the labels in the dataset.
[0045] S125. Conduct multiple rounds of training until the error of the model converges and shows good generalization ability on the validation set.
[0046] According to the second aspect of the present invention, the present invention provides a time series prediction-based adaptive correction system for autonomous driving behavior decision-making, which is used to implement the above-mentioned time series prediction-based adaptive correction method for autonomous driving behavior decision-making, including:
[0047] A traffic scene construction module that constructs an autonomous driving mixed traffic scene, and the mixed traffic scene includes autonomous driving vehicles and human-driven vehicles;
[0048] A human-driven vehicle model construction module, which is used to construct a human-driven vehicle model for the human-driven vehicles in the mixed traffic scene, where the IDM model is used to control its longitudinal decision-making, and the MOBIL model is used to control its lateral decision-making, so as to simulate human driving behavior;
[0049] A design module, which is used to design a state space, an action space and a reward function based on the autonomous driving vehicles in the autonomous driving mixed traffic scene, and train a pre-constructed autonomous driving vehicle reinforcement learning decision model according to the state space, the action space and the reward function;
[0050] An induced output module, which is used to design a state perturbation agent for the reinforcement learning decision model, and implement state perturbation by means of gradient descent, so as to induce the reinforcement learning decision model to output wrong decisions;
[0051] A training and correction module is used to design an adaptive action correction equation for the wrong decisions output by the reinforcement learning decision model to perform the correction task. According to the adaptive action correction equation, the correction task is decomposed into multiple time-series prediction tasks, and a time-series prediction model is constructed and trained to perform multiple time-series prediction tasks, thereby correcting the wrong decisions output.
[0052] The present invention at least has the following beneficial effects:
[0053] By constructing a time-series prediction model and adaptively adjusting decisions according to the degree of state perturbation, the present invention can predict and correct the decisions of autonomous driving vehicles in real time, reduce the risks caused by state perception errors, improve the system's ability to cope with complex environments, and through dynamic adjustment, the present invention can adapt to new and unknown perturbation situations, thereby making the autonomous driving system have stronger robustness and adaptability.
[0054] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a schematic flow chart of the correction method of the present invention;
[0056] Figure 2 is a schematic structural principle diagram of the correction method of the present invention;
[0057] Figure 3 is the training curve of the state perturbation degree prediction model of the present invention;
[0058] Figure 4 is the training curve of the conservative strategy prediction model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0060] Please refer to Figure 1-2 , the present invention provides a technical solution: an adaptive correction method for autonomous driving behavior decision-making based on time-series prediction, including the following steps:
[0061] S1. Construct an autonomous driving mixed traffic scenario, which includes autonomous driving vehicles and human-driven vehicles;
[0062] For the technical solution of this embodiment, this embodiment is built using the highway-env toolkit, and a mixed traffic scenario is established, specifically the on-ramp merging scenario on a highway;
[0063] S2. Construct a human-driven vehicle model for human-driven vehicles in the mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making, and the MOBIL model is used to control its lateral decision-making to simulate human driving behavior;
[0064] (S21) Control the longitudinal decision-making through the IDM model, and the control formula is as follows:
[0065]
[0066] In the formula is the calculated vehicle acceleration, a is the maximum allowable acceleration, v is the current speed, v0 is the desired speed, s * is the desired minimum following distance, s * is the following distance between the current vehicle and the vehicle in front, b is a comfortable deceleration of the vehicle, Δv is the speed difference between the current vehicle and the vehicle in front, and T is the safe time headway of the vehicle;
[0067] (S22) Control the lateral decision-making through the MOBIL model, and the control formula is as follows:
[0068]
[0069] Where is the acceleration of the vehicle in front of the current vehicle in the new lane after lane change, a n is the acceleration of this vehicle before lane change, is the predicted acceleration of the current vehicle after lane change, a c is the acceleration of the current vehicle before lane change, is the acceleration of the vehicle in front of the current vehicle in the current lane after lane change, a o is the acceleration of this vehicle before lane change, Δa th is the lane change threshold, b safe is the safe deceleration threshold;
[0070] S3. Based on the autonomous driving vehicles in the autonomous driving mixed traffic scenario, design the state space, action space, and reward function, and train the pre-constructed autonomous driving vehicle reinforcement learning decision-making model according to the state space, action space, and reward function;
[0071] (S31) The state space is {s1, s2, s3, …, s N} where represents the observation information of the i-th vehicle, whether present can be observed, and represent the x and y coordinates of the position when observed, v x 、v y represent the x and y coordinates of the velocity when observed;
[0072] (S32) The action space is A = {a1, a2, a3, a4, a5}, including 5 discrete actions, where a1, a2, a3, a4, a5 represent 5 discrete actions of turning left, turning right, staying unchanged, accelerating, and decelerating respectively.
[0073] (S33) The reward function is R t = R a + R v + R s - C t , where
[0074]
[0075] In the formula, v ego is the current speed of the autonomous vehicle, is the average speed of surrounding vehicles, R a is the reward for reaching the end coordinates, R v is the reward for maintaining a target speed, R s is the reward for successful lane change, C t is the cost of the autonomous vehicle at time t;
[0076] S4. Design a state perturbation agent for the reinforcement learning decision model, and implement state perturbation in the way of gradient descent, so as to induce the reinforcement learning decision model to output wrong decisions;
[0077] The perturbation objective function is specifically as follows:
[0078]
[0079] In the formula, δ is the state perturbation matrix, π is the policy function of the autonomous vehicle, s is the current observation state of the autonomous vehicle, a is the action decision output by its decision model. By maximizing the difference in policy output before and after agent perturbation, the effect of simulating dangerous perception errors in reality is achieved;
[0080] S5. For the wrong decisions output by the reinforcement learning decision model, design an adaptive action correction equation for performing correction tasks. According to the adaptive action correction equation, decompose the correction task into multiple time series prediction tasks, construct a time series prediction model and train it for performing multiple time series prediction tasks, so as to correct the wrong decisions output;
[0081] (S51) Design an adaptive action correction equation:
[0082] bias = ||Δ actual -Δ pred ||
[0083]
[0084] π final = α * π rl + (1 - α) * π alt
[0085] In the formula, Δ actual is the actual difference between the current state and the previous state, Δ pred is the predicted difference between the current state and the previous state, bias is the difference between them, α is the adjustment factor, k, τ are variables that control the change range of the adjustment factor, which can be determined by Bayesian optimization during training, π rl is the policy output by the reinforcement learning autonomous driving agent, π alt is the conservative policy in the predicted current state. The larger the bias, the greater the degree of disturbance of the autonomous driving vehicle state, the closer α tends to 0, and the final decision is more inclined to use the predicted conservative policy;
[0086] (S52) According to the adaptive action correction equation, decompose the correction task into two sequential prediction tasks:
[0087] The first sequential task is to predict the difference between the current state and the previous state, providing an important basis for state changes for subsequent behavior decisions; the second sequential task is to predict a safe conservative policy in the current state. The prediction of this policy aims to provide a safe and robust behavior choice for the autonomous driving system in complex road conditions;
[0088] (S53) Design two LSTM sequential prediction models with attention mechanisms, and the specific design is as follows:
[0089] For the first sequential task, the designed LSTM sequential prediction model incorporates an attention mechanism, which can pay more attention to the historical state information crucial for predicting the current state difference. By weighting the key information in the input sequence, the accuracy and effectiveness of state difference prediction are improved; for the model of the second sequential task, the introduction of the attention mechanism enables the model to focus on the key factors affecting the safety and conservativeness of the policy when processing sequential data related to the safe conservative policy, thus generating a safe conservative policy prediction that better meets the actual needs;
[0090] (S54) Construct a training data set, and the composition of the data set is as follows:
[0091] The dataset for the first temporal task consists of trajectories generated by the reinforcement learning model. These trajectories contain the behavioral data of the autonomous driving system in different scenarios, providing rich training materials for predicting the difference between the current state and the previous state. The dataset for the second temporal task consists of two parts. One part is the low-cost trajectories generated by the reinforcement learning model, and the other part is the conservative trajectories manually controlled. The low-cost trajectories generated by the reinforcement learning model reflect the relatively safe behavior of the system while pursuing a certain efficiency, and the conservative trajectories manually controlled incorporate the safe and conservative behavior patterns in human driving experience. The combination of the two provides comprehensive and diverse sample data for the model training of the second temporal task, helping to improve the model's prediction ability for safe and conservative strategies.
[0092] (S55) Construct a temporal prediction model and train it for performing multiple temporal prediction tasks. The temporal prediction model is two LSTM network models with attention mechanisms. The specific training process is as follows:
[0093] S55.1. Preprocess the dataset and normalize the features.
[0094] S55.2. Initialize the network training parameters.
[0095] S55.3. Divide 80% of the dataset into the training set and the remaining 20% into the validation set.
[0096] S55.4. Use the backpropagation algorithm to train the two LSTM network models on the training set and optimize the model parameters by minimizing the loss function.
[0097] The loss function is:
[0098]
[0099] where L Diff is the loss function for predicting the difference between the current state and the previous state, L Alt is the loss function for predicting a safe and conservative strategy in the current state, Δ pred and are the output values of the two temporal prediction models, Δ actual and are the labels in the dataset.
[0100] S55.5. Conduct multiple rounds of training until the error of the model converges and shows good generalization ability on the validation set.
[0101] Figure 3It is the training curve of the state disturbance degree prediction model in this embodiment. A total of about 45,000 rounds were trained, and the loss decreased from 0.43 at the beginning to 0.06. This indicates that through training, the first time series prediction model can accurately predict the difference between the current state and the previous state;
[0102] Figure 4 It is the training curve of the conservative strategy prediction model in this embodiment. A total of about 130,000 rounds were trained, and the loss decreased from 0.61 at the beginning to 0.08. This indicates that through training, the second time series prediction model can accurately predict a safe conservative strategy.
[0103] In summary, the present invention can adaptively adjust decisions according to the state disturbance degree by constructing a time series prediction model, can predict and correct the decisions of autonomous driving vehicles in real time, reduce the risks caused by state perception errors, improve the ability of the system to cope with complex environments, and through dynamic adjustment, the present invention can adapt to new and unknown disturbance situations, so that the autonomous driving system has stronger robustness and adaptability.
[0104] Embodiment 2:
[0105] This embodiment provides an autonomous driving behavior decision adaptive correction system based on time series prediction, which is used to implement the above-mentioned autonomous driving behavior decision adaptive correction method based on time series prediction, and includes:
[0106] A traffic scene construction module, which constructs an autonomous driving mixed traffic scene, and the mixed traffic scene includes autonomous driving vehicles and human-driven vehicles;
[0107] A human-driven vehicle model construction module, which is used to construct a human-driven vehicle model for the human-driven vehicles in the mixed traffic scene, where the IDM model is used to control its longitudinal decision-making, and the MOBIL model is used to control its lateral decision-making, so as to simulate human driving behavior;
[0108] A design module, which is used to design a state space, an action space and a reward function based on the autonomous driving vehicles in the autonomous driving mixed traffic scene, and train the pre-constructed autonomous driving vehicle reinforcement learning decision model according to the state space, the action space and the reward function;
[0109] An induced output module, which is used to design a state disturbance agent for the reinforcement learning decision model, and implement state disturbance in a gradient descent manner, so as to induce the reinforcement learning decision model to output wrong decisions;
[0110] A training and correction module is used to design an adaptive action correction equation for the wrong decisions output by the reinforcement learning decision-making model to perform correction tasks. According to the adaptive action correction equation, the correction task is decomposed into multiple time-series prediction tasks, and a time-series prediction model is constructed and trained to perform multiple time-series prediction tasks, thereby correcting the wrong decisions output.
[0111] Specifically, the above traffic scene construction module, human-driven vehicle model construction module, design module, induced output module, and training and correction module can be embedded in a computer processing system. The computer, based on the above-provided method for self-adaptive correction of autonomous driving behavior decisions based on time-series prediction, calls the above modules to complete the task of correcting wrong behavior decisions in autonomous driving; the above traffic scene construction module, human-driven vehicle model construction module, design module, induced output module, and training and correction module can perform operations according to the specific steps given by the above method for self-adaptive correction of autonomous driving behavior decisions based on time-series prediction.
[0112] It should be noted that it should be understood that the division of each module of the above system is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; they can also be partially implemented in the form of software called by processing elements and partially implemented in the form of hardware. For example, the traffic scene construction module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code and called and executed by a certain processing element of the above device to perform the functions of the above signal processing module. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together or can be independently implemented. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through the hardware integrated logic circuit or software-form instructions in the processor element.
[0113] For example, the above-mentioned modules may be one or more integrated circuits configured to implement the above methods, such as: one or more Application Specific Integrated Circuits (ASICs), or one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain above-mentioned module is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0114] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0115] For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. When an element is referred to as being "assembled on", "mounted on", "fixed to" or "disposed on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used herein are for illustrative purposes only and do not represent the only implementation.
[0116] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
[0117] In the description of this specification, the description referring to terms such as "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in connection with that embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
Claims
1. An adaptive correction method for autonomous driving behavior decision-making based on time series prediction, characterized in that, Including the following steps: Construct an autonomous driving mixed traffic scenario, which includes autonomous vehicles and human-driven vehicles; Construct a human-driven vehicle model for the human-driven vehicles in the mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making and the MOBIL model is used to control its lateral decision-making, for simulating human driving behavior; Based on the autonomous vehicles in the autonomous driving mixed traffic scenario, design the state space, action space and reward function, and train the pre-constructed autonomous vehicle reinforcement learning decision-making model according to the state space, action space and reward function; Design a state perturbation agent for the reinforcement learning decision-making model, and implement state perturbation by means of gradient descent, so as to induce the reinforcement learning decision-making model to output wrong decisions; For the wrong decisions output by the reinforcement learning decision-making model, design an adaptive action correction equation for performing the correction task. According to the adaptive action correction equation, decompose the correction task into multiple time series prediction tasks, construct a time series prediction model and train it for performing multiple time series prediction tasks, so as to correct the output wrong decisions.
2. The adaptive correction method for autonomous driving behavior decision-making based on timing prediction according to claim 1, wherein: The autonomous driving mixed traffic scenario is built using the highway-env toolkit, and the autonomous driving mixed traffic scenario is specifically the highway ramp merging scenario.
3. The adaptive correction method for autonomous driving behavior decision-making based on time series prediction according to claim 2, characterized in that: Construct a human-driven vehicle model for the human-driven vehicles in the mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making and the MOBIL model is used to control its lateral decision-making, for simulating human driving behavior, specifically as follows: (31) Control the longitudinal decision-making through the IDM model, and the control formula is as follows: where is the calculated vehicle acceleration, a is the maximum allowable acceleration, v is the current speed, v0 is the desired speed, s * is the desired minimum following distance, s * is the following distance between the current vehicle and the vehicle ahead, b is a comfortable deceleration of the vehicle, Δv is the speed difference between the current vehicle and the vehicle ahead, and T is the safe time headway of the vehicle; (32) Control the lateral decision-making through the MOBIL model, and the control formula is as follows: Among them is the acceleration of the vehicle in front of the current vehicle in the new lane after lane change, a n is the acceleration of this vehicle before lane change is the predicted acceleration of the current vehicle after lane change, a c is the acceleration of the current vehicle before lane change is the acceleration of the vehicle in front of the current vehicle in the current lane after lane change, a o is the acceleration of this vehicle before lane change, Δa th is the lane change threshold, b safe is the safe deceleration threshold 4. The adaptive correction method for autonomous driving behavior decision-making based on time series prediction according to claim 2, wherein: Based on the autonomous vehicles in the autonomous driving mixed traffic scenario, design the state space and action space, and train the pre-constructed autonomous vehicle reinforcement learning decision-making model according to the state space, action space and reward function, specifically as follows: (41) The state space is {s1, s2, s3, …, s N}}, where represents the observation information of the i-th vehicle, whether present can be observed, and represents the x and y coordinates of the position when observed, v x and v y represent the x and y coordinates of the velocity when observed; (42) The action space is A = {a1, a2, a3, a4, a5}, including 5 discrete actions, where a1, a2, a3, a4, a5 represent 5 discrete actions of turning left, turning right, remaining unchanged, accelerating, and decelerating respectively.
5. The method for adaptively correcting the decision-making of an autonomous driving behavior based on time series prediction according to claim 4, wherein: The reward function is R t = R a + R v + R s - C t , where where v ego is the current speed of the autonomous vehicle, is the average speed of surrounding vehicles, R a is the reward for reaching the end coordinates, R v is the reward for maintaining a target speed, R s is the reward for successful lane change, C t is the cost of the autonomous vehicle at time t.
6. The adaptive correction method for autonomous driving behavior decision-making based on time series prediction according to claim 1, wherein Design a state perturbation agent for the reinforcement learning decision-making model, and implement state perturbation by means of gradient descent. The perturbation objective function is specifically as follows: In the formula, δ is the state perturbation matrix, π is the policy function of the autonomous vehicle, s is the current observation state of the autonomous vehicle, and a is the action decision output by its decision-making model.
7. The method for adaptively correcting the decision-making of an autonomous driving behavior based on time series prediction according to claim 1, wherein For the wrong decisions output by the reinforcement learning decision-making model, design an adaptive action correction equation, specifically as follows: bias = ||Δ actual -Δ pred || π final =α*π rl +(1-α) * π alt where, Δ actual is the actual difference between the current state and the previous state, Δ pred is the predicted difference between the current state and the previous state, bias is the difference between them, α is the adjustment factor, k, τ are variables that control the change range of the adjustment factor, which can be determined by Bayesian optimization during training, π rl is the policy output by the reinforcement learning autonomous driving agent, π alt is the conservative policy in the predicted current state. The larger the bias, the greater the degree of disturbance of the autonomous driving vehicle state, the closer α tends to 0, and the final decision is more inclined to use the predicted conservative policy.
8. The adaptive correction method for autonomous driving behavior decision-making based on timing prediction according to claim 7, characterized in that, According to the adaptive action correction equation, decompose the correction task into two time series prediction tasks: The first time series task is to predict the difference between the current state and the previous state, providing an important basis for state changes for subsequent behavior decisions; The second time series task is to predict a safe and conservative policy under the current state, for providing a safe behavior choice for the autonomous driving system in complex road conditions.
9. The self - adaptive correction method for autonomous driving behavior decision - making based on time - series prediction according to claim 8, wherein: The above-mentioned time series prediction model is constructed and trained to perform multiple time series prediction tasks. The time series prediction model is two LSTM network models with attention mechanisms. The specific training process is as follows: (91) Preprocess the data set and normalize the features; (92) Initialize the network training parameters; (93) Divide the data set into a training set and a validation set at a ratio of 8:2; (94) Use the backpropagation algorithm to train the two LSTM network models on the training set, and optimize the model parameters by minimizing the loss function; The loss function is: where L Diff is the loss function for predicting the difference between the current state and the previous state, L Alt is the loss function for predicting a safe and conservative policy in the current state, Δ pred and are the output values of two time series prediction models, Δ actual and are the labels in the dataset. S125. Conduct multiple rounds of training until the error of the model converges and shows good generalization ability on the validation set.
10. An adaptive correction system for autonomous driving behavior decision-making based on time series prediction, which is used to implement the adaptive correction method for autonomous driving behavior decision-making based on time series prediction according to any one of claims 1 to 9, characterized in that, It includes: A traffic scenario construction module that constructs an autonomous driving mixed traffic scenario, which includes autonomous driving vehicles and human-driven vehicles; A human-driven vehicle model construction module used to construct a human-driven vehicle model for the human-driven vehicles in the mixed traffic scenario, where the IDM model is used to control its longitudinal decision-making and the MOBIL model is used to control its lateral decision-making to simulate human driving behavior; A design module that designs the state space, action space, and reward function based on the autonomous driving vehicles in the autonomous driving mixed traffic scenario, and trains the pre-constructed autonomous driving vehicle reinforcement learning decision model according to the state space, action space, and reward function; An induced output module that designs a state perturbation agent for the reinforcement learning decision model and implements state perturbation in a gradient descent manner to induce the reinforcement learning decision model to output incorrect decisions; A training correction module that designs an adaptive action correction equation for the incorrect decisions output by the reinforcement learning decision model to perform the correction task. According to the adaptive action correction equation, the correction task is decomposed into multiple time series prediction tasks, and a time series prediction model is constructed and trained to perform multiple time series prediction tasks, thereby correcting the incorrect decisions output.