A Lane-Changing Decision-Making Method Considering Reinforcement Learning of Driving Style Perceptual Meta-learning

CN122275891BActive Publication Date: 2026-09-01JILIN UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610746854.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-09-01
Estimated Expiration
2046-05-28

AI Technical Summary

Technical Problem

[0007]针对现有技术中的上述不足,本发明提供的一种考虑驾驶风格感知元强化学习的换道决策方法解决了传统强化学习换道模型难以适应交通密度动态变化、无法感知并适配周边车辆驾驶风格、跨场景泛化能力与决策鲁棒性不足的问题

Benefits of technology

(1)本发明提供了一种考虑驾驶风格感知元强化学习的换道决策方法,构建了一个驾驶风格感知与交通密度自适应的元强化学习换道决策框架。首先利用执行驾驶任务所需的全维度状态信息构建智能体观测对象的特征和驾驶风格类别,为策略学习提供紧凑且具有物理含义的风格表示,使智能体能够在混合驾驶人交通流中更准确地理解和预测周车行为,从而生成更稳健、自适应性更强的换道策略。随后采用MAML构建单智能体博弈换道模型,将不同交通密度与不同风格组合视为多个相关任务,使模型在跨任务训练中学习通用策略结构,并在新场景中通过少量梯度更新即可快速适应。该机制有效提升了换道决策在密度变化与行为差异显著环境下的泛化能力、稳定性与环境适应性,解决了传统强化学习换道模型难以适应交通密度动态变化、无法感知并适配周边车辆驾驶风格、跨场景泛化能力与决策鲁棒性不足的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122275891B_ABST
    Figure CN122275891B_ABST
Patent Text Reader

Abstract

This invention discloses a lane-changing decision-making method considering driving style perception meta-reinforcement learning, belonging to the field of autonomous driving lane-changing decision-making technology. The method includes: collecting full-dimensional state information required to perform driving tasks, preprocessing it to generate features of the observed objects, and establishing a DQN model based on Markov decision processes; loading meta-parameters obtained from meta-reinforcement learning as the initial model parameters of the DQN model, training the DQN model through meta-reinforcement learning, and establishing a single-agent game-theoretic lane-changing model; and using the single-agent game-theoretic lane-changing model to select lane-changing decision schemes based on the input full-dimensional state information. This invention can produce safe, smooth, and efficient lane-changing behavior close to human driving habits in low, medium, and high density scenarios, and outperforms baseline models such as DDQN and DuelingDQN in terms of cumulative reward, average vehicle speed, safety, and comfort.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autonomous driving lane-changing decision technology, specifically relating to a lane-changing decision method that considers reinforcement learning of driving style perception elements. Background Technology

[0002] With the rapid development of autonomous driving technology, lane-changing decisions have become one of the major challenges in achieving safe and efficient travel. As highway autonomous driving technology advances, lane-changing decisions are gradually evolving from rule-based methods to data-driven intelligent decision-making models.

[0003] Currently, most intelligent decision-making models applied to lane-changing decisions are based on deep learning. However, deep learning methods rely on offline training with static datasets and lack online interaction and trial-and-error updates with the environment. In highly dynamic and interactive high-speed traffic environments, they are unable to handle continuous state spaces and real-time changing vehicle behavior, resulting in significant limitations in adaptability.

[0004] In recent years, reinforcement learning (RL) has been widely applied in the field of autonomous driving decision-making. By continuously interacting with the environment to learn optimal policies, reinforcement learning can handle complex decision-making problems such as continuous control and multi-objective balancing, overcoming the limitations of traditional deep learning's insufficient adaptability. However, traditional reinforcement learning models are typically trained under fixed or single traffic densities, and the learned policies are highly dependent on the state distribution of the training environment. When traffic density changes, factors that need to be considered when changing lanes, such as safety gaps, traffic flow fluctuations, and surrounding vehicle reactions, vary in distribution under different traffic density characteristics, making it difficult for the model to maintain stable policy performance.

[0005] Meta-learning was proposed to address the demands of complex and dynamic tasks. Its core idea is "learning how to learn quickly," meaning that by training on multiple related tasks, the model can extract common structures across tasks, allowing it to adapt quickly to new tasks with only minor updates. Based on this idea, meta-reinforcement learning introduces meta-learning mechanisms into the reinforcement learning framework. Methods like MAML (Model-Agnostic Meta-Learning) do not directly learn the optimal policy for a specific task, but rather learn an initial policy with good transferability, enabling it to achieve good performance on different tasks with minimal gradient updates. This "rapid adaptation" is highly suitable for autonomous driving environments with constantly changing traffic densities. Meta-reinforcement learning can learn shared patterns across multi-density tasks, thus obtaining meta-policies with cross-scenario adaptability. When traffic density changes, the agent only needs minimal updates to adjust to a suitable policy, significantly improving generalization, stability, and learning efficiency. Therefore, meta-reinforcement learning provides an effective path to solve the core problem of "single-density training and cross-density instability" in lane-changing decisions, enabling models to better cope with dynamic density changes in real-world traffic environments.

[0006] While meta-reinforcement learning can address the adaptability issues arising from changes in traffic density, driving style remains a core factor influencing vehicle behavior patterns in autonomous driving lane-changing decisions. Different drivers exhibit significant differences in speed selection, acceleration / deceleration habits, following distance preferences, and lane-changing tendencies. These differences directly determine the interaction methods and risk levels of vehicles within traffic flow. Numerous studies have shown that accurately characterizing the driving styles of surrounding vehicles not only improves the accuracy of predicting future vehicle behavior but also helps autonomous driving systems more rationally assess the safety and feasibility of lane-changing actions, thereby improving decision-making quality and traffic efficiency. Therefore, introducing driving style modeling into the decision-making model is a crucial prerequisite for achieving safe, stable, and personalized autonomous driving. However, a significant limitation in previous driving style research is the assumption that the driving styles of surrounding vehicles are stable and unchanging over time. In most existing methods, surrounding vehicles are typically treated as "homogeneous drivers" with fixed behavioral patterns, and their speed preferences, acceleration / deceleration habits, following distances, and lane-changing tendencies are simplified to static characteristics. However, in real traffic flow, each vehicle has an independent and dynamically changing driving style, which adjusts in real time according to traffic density, driving context, and surrounding pressure. Without explicitly modeling these style differences and their dynamics, the agent will struggle to accurately predict the future trajectory of surrounding vehicles and cope with the environmental uncertainties brought about by heterogeneous drivers, thereby reducing the safety and robustness of decision-making. Summary of the Invention

[0007] To address the aforementioned shortcomings in existing technologies, this invention provides a lane-changing decision-making method that incorporates reinforcement learning based on driving style perception. This method solves the problems of traditional reinforcement learning lane-changing models, which struggle to adapt to dynamic changes in traffic density, cannot perceive and adapt to the driving styles of surrounding vehicles, and lack cross-scenario generalization ability and decision robustness.

[0008] To achieve the aforementioned objectives, the technical solution adopted by this invention is as follows: a lane-changing decision-making method considering reinforcement learning of driving style perception elements, comprising the following steps: S1. Collect all-dimensional state information required to perform driving tasks, preprocess it to generate the characteristics and driving style of the observed objects, and establish a DQN model based on Markov decision process; S2. Load the meta-parameters obtained from meta-reinforcement learning as the initial model parameters of the DQN model, perform meta-reinforcement learning training on the DQN model, and establish a single-agent game-theoretic lane-changing model. S3. The lane-changing decision scheme is selected based on the input full-dimensional state information through a single-agent game-theoretic lane-changing model.

[0009] Furthermore: In S1, the method for generating the features of the observed object is specifically as follows: Filtering, fusing, and feature extraction of full-dimensional state information generate features of the observed object. In the formula The lateral position relative to the vehicle. The longitudinal position relative to the vehicle. The lateral speed is relative to the vehicle's own speed. The longitudinal speed relative to the vehicle itself; The specific method for generating driving styles is as follows: By statistically analyzing the historical trajectory data of surrounding vehicles, the historical trajectory data is clustered and classified to extract driving styles, which include aggressive, normal, and cautious driving styles. In S1, the specific method for establishing the DQN model based on the Markov decision process is as follows: S11. Define the state space: Construct the state space for highway lane-changing decisions using the characteristics of the observed objects as the states; S12. Define the action space: The action space of each agent consists of a set of high-level driving actions, including acceleration, deceleration, constant speed, left lane change and right lane change.

[0010] Furthermore: In S2, the specific method for obtaining the meta-parameters from meta-reinforcement learning is as follows: S21. Divide different traffic flow density scenarios into several tasks; S22. Based on the task partitioning, output the task adaptation parameters for each task through an inner loop; S23. Aggregate gradient information across all tasks, update meta-parameters, and construct the meta-parameters obtained from meta-reinforcement learning.

[0011] Furthermore: In S21, the expression for dividing different traffic flow density scenarios into several tasks is as follows: In the formula, For the first One task, For tasks with different traffic flow densities, The driving task distribution is such that each task corresponds to a traffic flow scenario with different density, which is used to construct the corresponding lane change decision scenario in the simulation environment.

[0012] Furthermore, in S22, the method for constructing the meta-parameters obtained from meta-reinforcement learning includes the following sub-steps: S221. Set the maximum number of iterations and initialize the number of iterations; S222. At the current iteration number, the agent interacts with the task environment using the current policy to collect a batch of experience data as an experience replay dataset. The experience replay dataset includes state, action, immediate reward, and subsequent state transition information; S223. Based on the empirical replay dataset, calculate the TD target using the temporal difference concept in reinforcement learning, and measure the current... The difference between the estimated value and the TD objective is used to construct the loss function for the current task. S224. Along the negative gradient direction of the loss function with respect to the current task adaptation parameters, perform a gradient update with a preset in-task learning rate to obtain the updated task adaptation parameters. S225. Determine if the current iteration count is equal to the maximum iteration count. If not, increment the current iteration count by 1 and return to S222. If yes, output the final task adaptation parameters.

[0013] Furthermore: In S223, the TD target is calculated. The specific expression is: In the formula, In the state Execute action Instant rewards In order to take new action The new state that is obtained later In order to be in The maximum expected return that can be obtained by taking different actions. t For time, For the first The next iteration, the... Task adaptation parameters for each task k This represents the current iteration number. The discount factor for the reward; In S223, calculate the... Task loss function The specific expression is: In the formula, For time state, In the state Next action At that time, by the first The next iteration, the... Task adaptation parameters for each task The corresponding action value output by the Q network, For from the first The second iteration Experience replay dataset for each task Mathematical expectation operation for sampling and transferring samples.

[0014] Furthermore: In S224, we obtain the first... The next iteration, the... Adaptation parameters for each task The specific expression is: In the formula, The learning rate within the task. For the loss function on the th The next iteration, the... Adaptation parameters for each task The gradient operator is used to calculate the direction and step size of parameter updates.

[0015] Furthermore: S23 includes the following sub-steps: S231. Set the driving task distribution, meta-parameters, and meta-learning rate; S232. Sample a batch of tasks from the driving task distribution, and for each task, use the experience playback dataset. The loss function for the current task is calculated, and the loss of all tasks in the current batch is aggregated. The gradient of the loss with respect to the meta-parameters is calculated, and then the meta-parameters are updated using the gradient. S233. Repeat S232 until the meta-parameters converge, generating the meta-parameters obtained from meta-reinforcement learning.

[0016] Furthermore: In S232, the updated meta-parameters are obtained. The specific expression is: In the formula, For meta-parameters, The gradient of the meta-parameters, The meta-learning rate, For the first Task In the The meta-loss function at the next iteration; In the formula, This is the validation set.

[0017] Furthermore: In S2, the DQN model is trained using meta-reinforcement learning, with a total reward... The specific expression is: In the formula, For collision penalties, For speed bonus items, For right-lane bonus items, For lane change bonus items, For driving style rewards; In the formula, This is the collision reward weighting coefficient. This is a collision status identifier variable; when a vehicle collision occurs, ,otherwise ; In the formula, For speed reward weighting coefficient, For the amplitude limiting function, For the longitudinal speed of the vehicle, For the minimum speed, This represents the maximum speed. In the formula, The lane position reward weighting coefficient, This is the index value of the lane the vehicle is currently in. This refers to the total number of lanes on the road. In the formula, The reward weighting coefficient for lane-changing behavior. Used as a variable to identify lane-changing behavior; In the formula, The penalty weighting coefficient is the average driving style of surrounding vehicles. The penalty weighting coefficient for cautious driving style of surrounding vehicles, The penalty weighting coefficient for aggressive driving styles of surrounding vehicles. This is a style category identifier for general drivers. This is a style category identifier for cautious drivers. This is a style category identifier for aggressive drivers, and its value is determined based on the degree of conflict between the driving styles of surrounding vehicles and the current lane-changing decision. This is the transpose symbol.

[0018] The beneficial effects of this invention are as follows: (1) This invention provides a lane-changing decision-making method that considers driving style perception meta-reinforcement learning, and constructs a driving style perception and traffic density adaptation meta-reinforcement learning lane-changing decision-making framework. First, the features and driving style categories of the agent's observed objects are constructed using the full-dimensional state information required to perform the driving task, providing a compact and physically meaningful style representation for policy learning. This enables the agent to more accurately understand and predict the behavior of surrounding vehicles in mixed driver traffic flow, thereby generating a more robust and adaptive lane-changing strategy. Subsequently, a single-agent game-theoretic lane-changing model is constructed using MAML, treating different traffic densities and different style combinations as multiple related tasks. This allows the model to learn a general policy structure during cross-task training and quickly adapt to new scenarios with a small number of gradient updates. This mechanism effectively improves the generalization ability, stability, and environmental adaptability of lane-changing decisions in environments with significant density changes and behavioral differences, solving the problems of traditional reinforcement learning lane-changing models being unable to adapt to dynamic changes in traffic density, unable to perceive and adapt to the driving styles of surrounding vehicles, and lacking cross-scenario generalization ability and decision robustness.

[0019] (2) The model extracts driving style labels by performing cluster analysis on real driving data and introduces driving behavior differences into the decision model to improve its accuracy in predicting weekly driving behavior.

[0020] (3) Experimental results show that the proposed method can produce safe, smooth and efficient lane-changing behavior that is close to human driving habits in low, medium and high density scenarios. It is also superior to baseline models such as DDQN and DuelingDQN in terms of cumulative reward, average speed, safety and comfort. Attached Figure Description

[0021] Figure 1 This is a flowchart of a lane-changing decision-making method that considers reinforcement learning of driving style perception elements according to the present invention.

[0022] Figure 2 This is the structure of the DQN algorithm.

[0023] Figure 3 This is a schematic diagram of a meta-reinforcement learning framework.

[0024] Figure 4 This is a schematic diagram of a simulation experiment.

[0025] Figure 5 A comparison chart of average velocity convergence curves in low-density environments.

[0026] Figure 6 Comparison of cumulative reward convergence curves for low-density environments.

[0027] Figure 7 A comparison chart of convergence curves for mission duration in low-density environments.

[0028] Figure 8 This is a comparison chart of the average velocity convergence curves in a medium-density environment.

[0029] Figure 9 Comparison of cumulative reward convergence curves for medium-density environments.

[0030] Figure 10 This is a comparison chart of the convergence curves for mission duration in medium-density environments.

[0031] Figure 11 A comparison chart of average velocity convergence curves in high-density environments.

[0032] Figure 12 Comparison of cumulative reward convergence curves for high-density environments.

[0033] Figure 13 A comparison chart of convergence curves for high-density environmental mission duration. Detailed Implementation

[0034] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0035] Example 1: like Figure 1 As shown, in one embodiment of the present invention, a lane-changing decision-making method considering driving style perception reinforcement learning includes the following steps: S1. Collect all-dimensional state information required to perform driving tasks, preprocess it to generate features of the observed objects, and establish a DQN model based on Markov decision process; S2. Load the meta-parameters obtained from meta-reinforcement learning as the initial model parameters of the DQN model, perform meta-reinforcement learning training on the DQN model, and establish a single-agent game-theoretic lane-changing model. S3. The lane-changing decision scheme is selected based on the input full-dimensional state information through a single-agent game-theoretic lane-changing model.

[0036] The principle of the method of this invention is as follows: In S1, the onboard perception system collects full-dimensional state information required to perform driving tasks. This embodiment assumes that the vehicle is equipped with complete multi-source sensors and core electronic control units. The perception system integrates visual sensors, millimeter-wave radar, and other devices, enabling accurate collection of the state information of vehicles in front and behind within a multi-lane range. Based on the sensor performance threshold, the information perception range is limited to 150 meters. Subsequently, the onboard data processing unit filters, fuses, and extracts features from the original sensor signals: on the one hand, it transforms motion parameters into structured feature parameters required for decision-making; on the other hand, it extracts quantitative indicators of driving styles by statistically analyzing the behavior sequences of surrounding vehicles, classifying them into aggressive, robust, and cautious types, and completes style classification through clustering algorithms.

[0037] In S2, during the meta-reinforcement learning training phase of the DQN model, the reinforcement learning decision model does not start from scratch. Instead, it loads meta-parameters obtained through meta-reinforcement learning training, which possess cross-task generalization capabilities, as the initial weights of its network. This step ensures that the model starts from a high-quality "general knowledge" base point when facing new environments, rather than being randomly initialized. After constructing the decision-oriented state information database, the system further integrates the key state features and driving style classification results of the vehicle and surrounding vehicles, and after normalization preprocessing, inputs them into the initialized reinforcement learning decision model.

[0038] In S3, the single-agent game-theoretic lane-changing model selects lane-changing decision schemes based on the preprocessed features of the input. This lane-changing decision scheme serves as the core output of the upper-level decision module and is then passed to the control unit built into the lower-level simulation environment to execute specific vehicle motion control.

[0039] like Figure 2 As shown, in this embodiment, the focus of the present invention is on lane-changing decision-making; therefore, the algorithm of the present invention uses the DQN model, while the traditional When learning to handle the high-dimensional continuous state space of highway lane-changing decisions, the state space of highway lane-changing decisions contains multi-dimensional continuous features such as the vehicle's position / speed and the positions / speeds of surrounding vehicles. Traditional tables... Learning methods are completely unsuitable for such high-dimensional scenarios. They suffer from excessively high state-action space dimensions and tabular structures. The problem of functions being unable to be stored and updated. To solve this problem, Deepin... The network introduces deep neural networks into reinforcement learning, for The function is approximated and optimized.

[0040] In S1, the specific method for generating the features of the observed object is as follows: Filtering, fusing, and feature extraction of full-dimensional state information generate features of the observed object. In the formula The lateral position relative to the vehicle. The longitudinal position relative to the vehicle. The lateral speed is relative to the vehicle's own speed. The longitudinal speed relative to the vehicle itself; The specific method for generating driving styles is as follows: By statistically analyzing the historical trajectory data of surrounding vehicles, clustering and classifying the historical trajectory data, extracting driving styles and using them as explicit environmental feature inputs, the intelligent agent can more accurately understand and predict the behavior of surrounding vehicles in mixed driver traffic flow, thereby generating more robust and adaptive lane-changing strategies. Among them, driving styles include aggressive, general, and cautious. In S1, the specific method for establishing the DQN model based on the Markov decision process is as follows: S11. Define the state space: Construct the state space for highway lane-changing decisions using the characteristics of the observed objects as states. (Intelligent agent) state space Defined as a dimension The matrix, where This indicates the number of detectable neighboring vehicles. The number of features represents the current motion state of surrounding vehicles. This state definition can effectively characterize the dynamic changes of the local traffic environment and provide input information for agent decision-making.

[0041] S12. Define the action space: the action space for each agent. It consists of a set of high-level driving actions, including acceleration, deceleration, constant speed mode, left lane change, and right lane change.

[0042] DQN successfully combines deep learning and reinforcement learning, achieving effective policy learning in high-dimensional and complex environments and significantly improving performance in autonomous driving tasks. However, traditional DQN still belongs to the reinforcement learning method under the assumption of a single task and a single environment, and its learning process is highly dependent on the data distribution in a specific training scenario. When environmental features such as traffic density change, the trained policy often needs to be retrained or have a large number of parameters updated, showing insufficient adaptability.

[0043] To enhance the generalization and rapid adaptation capabilities of the bill of lading agent game-theoretic lane-changing model under different traffic environments, this invention introduces a meta-reinforcement learning framework. The meta-reinforcement learning process is as follows: Figure 3 As shown, it includes meta-task division, inner loop stage, outer loop stage, and meta-test stage.

[0044] In S2, the specific method for obtaining the meta-parameters from meta-reinforcement learning is as follows: S21. Divide different traffic flow density scenarios into several tasks; S22. Based on the task partitioning, output the task adaptation parameters for each task through an inner loop; S23. Aggregate gradient information across all tasks, update meta-parameters, and construct the meta-parameters obtained from meta-reinforcement learning.

[0045] In S21, the expression for dividing different traffic flow density scenarios into several tasks is as follows: In the formula, For the first One task, For tasks with different traffic flow densities, For driving task distribution, each task corresponds to a traffic flow scenario with different densities, used to construct the corresponding lane-changing decision scenarios in the simulation environment. In this embodiment, the algorithm begins by using the target task... Exclusive parameters Initialize to meta-parameters, i.e. This ensures that the learning process of the autonomous driving lane-changing decision model begins with a policy base point that has good generality, rather than random initialization.

[0046] In S22, the method for constructing the meta-parameters obtained from meta-reinforcement learning includes the following steps: S221. Set the maximum number of iterations and initialize the number of iterations; S222. At the current iteration number, the agent interacts with the task environment using the current policy to collect a batch of experience data as an experience replay dataset. The experience replay dataset includes state, action, immediate reward and subsequent state transition information, which is the data foundation for evaluating and updating execution strategies; S223. Based on the empirical replay dataset, calculate the TD objective (temporal difference objective) using the temporal difference concept in reinforcement learning, by measuring the current... The difference between the estimated value and the TD objective is used to construct the loss function for the current task. S224. Along the negative gradient direction of the loss function with respect to the current task adaptation parameters, perform a gradient update with a preset in-task learning rate to obtain the updated task adaptation parameters. S225. Determine if the current iteration count is equal to the maximum iteration count. If not, increment the current iteration count by 1 and return to S222. If yes, output the final task adaptation parameters.

[0047] In S223, the TD target is calculated. The specific expression is: In the formula, In the state Execute action Instant rewards In order to take new action The new state that is obtained later In order to be in The maximum expected return that can be obtained by taking different actions. t For time, For the first The next iteration, the... The task adaptation parameters for each task are used to define the action value function for the current iteration phase. , k This represents the current iteration number. As a discount factor for rewards, This is used to weigh the importance of future rewards. The larger the agent, the more it values ​​long-term returns; TD goals This represents the current state-action pair. The target value estimation under the existing strategy incorporates immediate rewards and discounted expectations of future returns.

[0048] In S223, calculate the... Task loss function The specific expression is: In the formula, For time state, In the state Next action At that time, by the first The next iteration, the... Task adaptation parameters for each task The corresponding action value output by the Q network, For from the first The second iteration Experience replay dataset for each task Mathematical expectation operation for mid-sample transfer. Loss function. It is usually expressed in the form of mean squared error, and its mathematical expectation is based on the collected empirical playback dataset. The calculations are performed to minimize the error between the predicted and target values.

[0049] In S224, the updated task adaptation parameters are obtained. The specific expression is: In the formula, The learning rate within the task. For the loss function on the th The next iteration, the... Adaptation parameters for each task The gradient operator is used to calculate the direction and step size of parameter updates and is the core symbol for gradient descent updates within meta-reinforcement learning tasks.

[0050] In S225, after completing all iterations, the algorithm outputs the final task adaptation parameters. At this point, the parameters have changed from the general meta-parameters. Evolved into task-oriented The optimized version.

[0051] The goal of this invention is to learn a global meta-parameter. This enables it to quickly adapt to task distribution. The scenarios covered have varying densities. Ensure the learned meta-knowledge has broad generalization ability. The outer loop iteratively executes the following steps: S231. Set the driving task distribution, meta-parameters, and meta-learning rate; S232. Sample a batch of tasks from the driving task distribution, and for each task, use the experience playback dataset. The loss function for the current task is calculated, and the loss of all tasks in the current batch is aggregated. The gradient of the loss with respect to the meta-parameters is calculated, and then the meta-parameters are updated using the gradient. S233. Repeat S232 until the meta-parameters converge, generating the meta-parameters obtained from meta-reinforcement learning.

[0052] In S232, the updated meta-parameters are obtained. The specific expression is: In the formula, For meta-parameters, The gradient of the meta-parameters, The meta-learning rate, For the first Task In the The meta-loss function at the next iteration depends on the in-task update process: firstly, it utilizes the in-task loss... Gradient descent is applied to the initial parameters to obtain the updated task adaptation parameters. Then, based on this parameter, the prediction error is calculated on the task's validation data, which is the meta-loss. The meta-loss measures the generalization performance after task adaptation and is the target of outer-layer optimization of the meta-parameters, used to update the global meta-parameters. . No. Task In the Meta-loss function at the next iteration The specific expression is: In the formula, For the validation set. Meta-loss function. It is calculated on the validation set of the same task, for example, a task in a low-density scene. Use the updated parameters The test was conducted on low-density scene data that had never been seen before. If the lane-changing decision remains safe and efficient, the meta-loss is low; if it only performs well on the training set but fails on the validation set, the meta-loss is high.

[0053] During the meta-testing phase, the intelligent agent vehicle uses meta-parameters. As initial parameters, in the new environment In this system, strategies can be implemented with minimal interaction with the environment. This process simulates the intelligence of "learning how to learn": the model utilizes its existing meta-knowledge to quickly understand the basic rules of the new environment and fine-tune its strategies to adapt to the specific requirements of the new task. Ultimately, the agent can implement strategies based on these rapid adaptations. It makes safe, efficient and scenario-appropriate lane-changing decisions within the environment.

[0054] In this embodiment, to achieve a balance between safety, driving efficiency, and reasonable lane changing, the present invention designs a multi-objective total reward. In S2, the DQN model is trained using meta-reinforcement learning, and the total reward... The specific expression is: In the formula, For collision penalties, For speed bonus items, For right-lane bonus items, For lane change bonus items, For driving style rewards; In the formula, This is the collision reward weighting coefficient. Therefore, when a vehicle collision occurs, ,otherwise This item is used to constrain the intelligent agent to maintain safe driving and avoid collision risks. A penalty is imposed when a collision occurs; if no collision occurs, this item is zero.

[0055] In the formula, For speed reward weighting coefficient, This is a limiting function used to restrict the speed bonus value within a set range. For the longitudinal speed of the vehicle, For the minimum speed, This is the maximum speed; this item is based on the vehicle's longitudinal speed. Within the set range The linear normalized value within the range is used as a reward to incentivize vehicles to maintain a higher and more stable speed within a safe range, thereby improving traffic efficiency.

[0056] In the formula, The lane position reward weighting coefficient, This is the index value of the lane the vehicle is currently in. This represents the total number of lanes on the road. The further to the right a vehicle is positioned, the higher the reward, to encourage vehicles to follow the right-hand driving principle. If the right-lane incentive is not enabled, its weight can be set to zero.

[0057] In the formula, The reward weighting coefficient for lane-changing behavior. This is a variable that identifies lane-changing behavior; an agent receives an extra reward when it successfully completes a lane change; otherwise, this value is zero. This variable is used to encourage reasonable lane-changing operations and improve traffic efficiency while ensuring safety.

[0058] In the formula, The penalty weighting coefficient is the average driving style of surrounding vehicles. The penalty weighting coefficient for cautious driving style of surrounding vehicles, The penalty weighting coefficient for aggressive driving styles of surrounding vehicles. This is a style category identifier for general drivers. This is a style category identifier for cautious drivers. This is a style category identifier for aggressive drivers, and its value is determined based on the degree of conflict between the driving styles of surrounding vehicles and the current lane-changing decision. This is the transpose symbol. This term guides the model to make safe lane-changing decisions that adapt to the driving styles of surrounding vehicles by penalizing the driving styles of vehicles that conflict with the vehicle's lane-changing behavior.

[0059] Example 2: To verify the effectiveness of the method of the present invention, this embodiment provides a k-means driver style clustering experiment based on NANI: This embodiment utilizes the HighD dataset in the clustering stage. The HighD dataset was created and released by the Automotive Engineering Institute team at RWTH Aachen University in Germany, aiming to provide real-world data support for the research and validation of highly automated driving systems. This dataset was collected using drones at six different locations along German highways, accumulating driving trajectory information for over 110,000 vehicles, covering a total distance of approximately 45,000 kilometers and a total duration of 16.5 hours. It possesses significant advantages such as large scale, realism, and high accuracy.

[0060] To determine the optimal number of clusters, this paper uses three commonly used clustering evaluation metrics, SC, CH, and DBI, to systematically evaluate the clustering effect. The quantitative evaluation results are shown in Table 1.

[0061] Table 1 Quantitative Indicator Results From the quantitative indicators, the evaluation results for clusters with a number of clusters in the range of 2 to 7 are as follows: When k=3, the silhouette coefficient (SC) reaches its highest value (0.388), indicating that the intra-cluster compactness and inter-cluster separation are optimal, which helps to clearly distinguish different driving styles. The Calinski-Harabasz index (CH) is relatively high (3483.02), indicating that the ratio of inter-cluster variance to intra-cluster variance is large, and the cluster structure is relatively obvious. At the same time, the Davies-Bouldin index (DBI) is 1.203 at k=3. Although it is not the lowest value, it is significantly lower than that at k=2, indicating that the overall performance of intra-cluster compactness and inter-cluster dispersion is good. This paper ultimately selects k=3 as the optimal number of clusters.

[0062] This paper selects key features such as vehicle longitudinal acceleration and following time interval (THW). These features have significant physical meaning. To accurately reflect the differences in driving styles, firstly, trajectory data of all trucks in the dataset were removed to avoid interference from vehicle type differences. Secondly, to more accurately capture the driving environment and interaction during lane-changing decisions, only valid data with at least one vehicle present in the three directions of front, left front, and left rear were retained. Furthermore, the original data had high dimensionality. To reduce data redundancy and improve the efficiency and accuracy of clustering, this paper uses Principal Component Analysis (PCA) to reduce the dimensionality of the data. Through PCA, seven principal components were finally extracted, with a cumulative variance contribution rate of 0.96, indicating that the dimensionality-reduced principal components fully preserved the main feature information of the original data. The clustering results are shown in Table 2.

[0063] Table 2. Clustering Feature Analysis of Driving Styles Based on the above clustering results and combined with actual driving behavior characteristics, the three clusters are defined interpretively: Cluster 0: The average longitudinal acceleration a_mean=0.20, the driver has a moderate acceleration level, and the average headway thw_mean=1.34, the following time interval is small. This type of driver tends to balance safety and efficiency and can be defined as a general driver. Cluster 1: The average longitudinal acceleration a_mean=0.12, the driver's acceleration level is low, and the average headway thw_mean=3.16, the following time distance is the largest, showing a clear safety tendency, and can be defined as a cautious driver; Cluster 2: The average longitudinal acceleration a_mean=0.41, the driver has the highest acceleration, and the average headway thw_mean=2.52, the following time interval is also small, showing obvious aggressive driving characteristics, so it is defined as an aggressive driver.

[0064] After completing the unsupervised driving style clustering based on NANI-K-means, this paper further constructs a supervised classification model of Support Vector Machine (SVM) based on the clustering results to achieve rapid identification of the driving style of new vehicles, enabling online recognition of driving styles. Specifically, the clustering labels obtained above are regarded as "pseudo-labels" of three driving styles, and the feature vectors after standardization and PCA dimensionality reduction are used as input to train a multi-class SVM classifier to learn the nonlinear mapping relationship between the feature space and the three driving styles of "cautious, general, and aggressive".

[0065] In the offline phase, a sample set was first constructed using the PCA features of all samples and their corresponding K-means clustering labels. The dataset was then divided into training and validation sets in an 8:2 ratio, and a multi-class SVM model was constructed using the radial basis function (RBF). After model training, the overall classification accuracy, precision, recall, and F1 score for each class were calculated on the validation set. The results are shown in Table 3, and a confusion matrix was plotted to analyze the confusion between different driving styles. Evaluation results show that the SVM model can effectively distinguish between cautious, general, and aggressive drivers, and the classification performance is highly consistent with the clustering results, indicating that the style labels obtained by K-means clustering have a certain degree of stability and separability.

[0066] Table 3. Validation results computed on the validation set. In the online recognition phase, this paper embeds the trained SVM driving style recognition model into the highway-env simulation environment, enabling the agent to perceive the driving style of surrounding vehicles in real time within each decision cycle. When a new vehicle enters the observation range, key behavioral features such as longitudinal velocity, longitudinal acceleration, lateral velocity, and THW are first extracted based on its short-time series trajectory information. These features are then standardized and transformed using the same process as in the offline phase to ensure consistency in the feature space. Subsequently, the dimensionality-reduced features are input into the pre-trained SVM classifier to obtain the vehicle's driving style category label (e.g., cautious, general, or aggressive). Due to the extremely low inference overhead of SVM, this recognition process can run stably within the control cycle of the environment simulation, achieving real-time judgment of surrounding vehicle behavior patterns. This provides accurate and timely prior information on driving style for meta-reinforcement learning lane-changing decisions, significantly improving the adaptability and safety of the strategy in heterogeneous traffic environments.

[0067] Example 3: To verify the effectiveness of the method of the present invention, this embodiment provides a reinforcement learning simulation experiment: I. Experimental Platform and Scenario Configuration: like Figure 4 As shown, the simulation experiment uses the highway-env platform, an open-source high-fidelity autonomous driving simulation platform that supports multi-lane highway scenarios, personalized traffic participant parameter configuration, and fine-grained dynamic trajectory control. The simulation scenario in this paper is set as a three-lane highway, with a single lane width of 3.75m and a total length of 1000m. The initial number of simulated vehicles is 3. During the simulation, blue represents the autonomous vehicle, green represents vehicles with a cautious driving style, yellow represents vehicles with a normal driving style, and red represents vehicles with an aggressive driving style. Medium density, high density, and low density are set to vehicle density values ​​of 0.2, 1, and 2, respectively. II. Comparison Model and Experimental Procedure: This invention compares the proposed single-agent game-theoretic lane-changing model MDQN with the following baselines and models that have been verified to perform well: (1) DDQN (Double DQN): Traditional DQN in calculating the target When calculating values, use the same network to simultaneously "select actions" and "calculate". "Value" can easily lead to systematic overestimation. The core idea of ​​Double DQN is to decouple action selection from action evaluation: (2) DuDQN (Dueling Double DQN): DuDQN introduces a Dueling structure on top of DDQN. On one hand, it decomposes the Q-function into state value and action advantage, enhancing the ability to learn state value. On the other hand, the target value is still constructed using the "selection-evaluation separation" method of Double DQN: the online network selects actions, and the target network evaluates the Q-value. This alleviates the overestimation problem and leverages the advantages of the Dueling structure in state value estimation, making lane-changing decisions more stable and precise in complex traffic scenarios.

[0068] III. Reinforcement Learning Parameter Settings To ensure the stable operation of the single-agent reinforcement learning environment and the repeatability of the training process, the present invention sets the environmental parameters as shown in Table 4. The main parameters are explained below: Table 4 Main Parameters IV. Comparison of Convergence Curves: (1) Low-Density Environment: Average speed: such as Figure 5 As shown, in low-density environments, MDQN can maintain a high and stable speed, demonstrating good acceleration and deceleration control capabilities. Compared to DDQN and DuDQN, MDQN exhibits less speed fluctuation, indicating that it can better cope with traffic flow in low-density environments and maintain smooth lane-changing operations.

[0069] Cumulative Rewards: such as Figure 6 As shown, MDQN accumulates higher rewards than DDQN and DuDQN in low-density environments, indicating that it can make lane-changing decisions more efficiently, and its decision-making process is safer and more effective. DuDQN and DDQN, on the other hand, experience slower reward growth in this low-density scenario.

[0070] Task Duration: e.g. Figure 7 As shown, MDQN has the shortest task duration in low-density environments, enabling it to quickly complete lane-changing tasks. In contrast, DDQN and DuDQN have longer task durations, reflecting their lower decision-making efficiency in low-density environments.

[0071] (2) Medium-Density Environment: Average speed: such as Figure 8 As shown, MDQN maintains high stability and speed in medium-density environments. Although the speed of all algorithms decreases to varying degrees with increasing density, MDQN is better able to maintain higher speed and adapt to the challenges posed by medium-density traffic flow.

[0072] Cumulative Rewards: such as Figure 9 As shown, MDQN continues to outperform DDQN and DuDQN in terms of accumulated rewards under medium-density conditions. MDQN is able to find more efficient lane-changing decisions in more complex environments, enabling it to consistently achieve high rewards in medium-density traffic flows. Conversely, DDQN and DuDQN accumulate rewards more slowly when facing medium-density traffic, reflecting their inadequacy in handling such density changes.

[0073] Task Duration: e.g. Figure 10 As shown, MDQN exhibits shorter task duration in medium-density environments, enabling it to complete lane-changing decisions more quickly. In contrast, DDQN and DuDQN have longer task durations, meaning their decision-making efficiency is lower than MDQN when dealing with medium-density environments.

[0074] (3) High-Density Environment: Average speed: such as Figure 11 As shown, in high-density environments, MDQN maintains a relatively stable speed and avoids excessive deceleration or stagnation. High-density traffic flow causes a significant decrease in speed for all algorithms, but MDQN manages to maintain vehicle efficiency well while controlling speed.

[0075] Cumulative Rewards: such as Figure 12 As shown, under high-density conditions, MDQN's reward accumulation performance still outperforms DDQN and DuDQN, indicating that it can make accurate lane-changing decisions quickly in complex high-density traffic environments, thereby obtaining more rewards. DDQN and DuDQN, on the other hand, face greater challenges, with relatively slower reward growth.

[0076] Task Duration: e.g. Figure 13 As shown, MDQN maintains the shortest task duration even in high-density environments. While task durations generally increase with density, MDQN maintains high decision-making efficiency and can quickly complete lane-switching tasks. In contrast, DDQN and DuDQN have longer decision-making times in high-density environments, reducing overall task efficiency.

[0077] V. Quantitative Indicator Analysis: (1) Reasonable number of movements and reasonable proportion of movements: Table 5. Reasonable Number of Movements and Reasonable Proportion of Movements Reasonable action count: Reasonable action is defined as action that does not reduce cumulative reward. As shown in Table 5, the number of reasonable actions of MDQN is significantly higher than that of DDQN and DuelingDQN under all density conditions. In particular, in low-density and medium-density environments, the number of reasonable actions of MDQN has a more stable growth, indicating that it can make better lane-changing decisions in these environments.

[0078] For example, in low-density environments, the reasonable action number of MDQN is 79392, which is much higher than that of DDQN (76661) and DuelingDQN (76300).

[0079] Reasonable action ratio: MDQN has the highest reasonable action ratio across all density environments, demonstrating its higher accuracy and efficiency in lane-changing decisions. MDQN's reasonable action ratios far exceed those of other algorithms in low density (98.26%), medium density (99.92%), and high density (99.95%) environments.

[0080] In contrast, DDQN's reasonable action ratio drops to 91.51% in high-density environments, indicating its lower decision-making efficiency in high-density scenarios.

[0081] (2) Total reward value: Table 6 Total Reward Value MDQN's reward performance: As shown in Table 6, MDQN's total reward value is significantly higher than DDQN and DuelingDQN at all densities. Especially in low-density environments, MDQN's reward value is as high as 95628.67, which is significantly higher than DDQN's 74687.36 and DuelingDQN's 76300.

[0082] In medium-density and high-density environments, the reward value of MDQN also remained at a high level, at 78688.74 and 78429.73 respectively, demonstrating its adaptability and decision optimization capabilities under different densities.

[0083] (3) Average speed: Table 7 Average Speed MDQN speed performance: As shown in Table 7, MDQN has a relatively high and stable average speed in all density environments, especially in medium and high density scenarios, where MDQN performs better.

[0084] In low-density environments, MDQN has an average velocity of 28.58, slightly higher than DuDQN (28.25) and DDQN (27.50).

[0085] In medium-density and high-density environments, MDQN achieved speeds of 29.63 and 29.35, respectively, significantly outperforming the other two algorithms. This indicates that MDQN better addresses the impact of density variations on speed control.

[0086] Performance of other algorithms: In comparison, DDQN has a slightly lower average speed, especially in low-density and high-density environments, showing its weaker ability to adjust speed in dynamic environments.

[0087] In the description of this invention, the above are merely preferred embodiments and are not intended to limit the scope of protection of this invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A lane-changing decision-making method considering reinforcement learning of driving style perception, characterized in that, Includes the following steps: S1. Collect all-dimensional state information required to perform driving tasks, preprocess it to generate the characteristics and driving style of the observed objects, and establish a DQN model based on Markov decision process; S2. Load the meta-parameters obtained from meta-reinforcement learning as the initial model parameters of the DQN model, perform meta-reinforcement learning training on the DQN model, and establish a single-agent game-theoretic lane-changing model. S3. The lane-changing decision scheme is selected based on the full-dimensional state information input by the single-agent game-theoretic lane-changing model. In S2, the specific method for obtaining the meta-parameters from meta-reinforcement learning is as follows: S21. Divide different traffic flow density scenarios into several tasks; S22. Based on the task partitioning, output the task adaptation parameters for each task through an inner loop; S23. Aggregate gradient information across all tasks, update meta-parameters, and construct meta-parameters obtained from meta-reinforcement learning. S23 includes the following steps: S231. Set the driving task distribution, meta-parameters, and meta-learning rate; S232. Sample a batch of tasks from the driving task distribution, and for each task, use the experience playback dataset. The loss function for the current task is calculated, and the loss of all tasks in the current batch is aggregated. The gradient of the loss with respect to the meta-parameters is calculated, and then the meta-parameters are updated using the gradient. S233, Repeat S232 until the meta-parameters converge, and generate the meta-parameters obtained from meta-reinforcement learning; In S232, the updated meta-parameters are obtained. The specific expression is: In the formula, For meta-parameters, The gradient of the meta-parameters, The meta-learning rate, For the first Task In the The meta-loss function at the next iteration; In the formula, For the validation set, In the state Next action At that time, by the first The next iteration, the... Task adaptation parameters for each task The corresponding action value output by the Q network, For TD objectives; In S2, the DQN model is trained using meta-reinforcement learning, with a total reward of The specific expression is: In the formula, For collision penalties, For speed bonus items, For right-lane bonus items, For lane change bonus items, For driving style rewards; In the formula, This is the collision reward weighting coefficient. This is a collision status identifier variable; when a vehicle collision occurs, ,otherwise ; In the formula, For speed reward weighting coefficient, For the amplitude limiting function, For the longitudinal speed of the vehicle, For the minimum speed, This represents the maximum speed. In the formula, The lane position reward weighting coefficient, This is the index value of the lane the vehicle is currently in. This refers to the total number of lanes on the road. In the formula, The reward weighting coefficient for lane-changing behavior. Used as a variable to identify lane-changing behavior; In the formula, The penalty weighting coefficient is the average driving style of surrounding vehicles. The penalty weighting coefficient for cautious driving style of surrounding vehicles, The penalty weighting coefficient for aggressive driving styles of surrounding vehicles. This is a style category identifier for general drivers. This is a style category identifier for cautious drivers. This is a style category identifier for aggressive drivers, and its value is determined based on the degree of conflict between the driving styles of surrounding vehicles and the current lane-changing decision. This is the transpose symbol.

2. The lane-changing decision-making method considering driving style perceptual reinforcement learning according to claim 1, characterized in that, In S1, the specific method for generating the features of the observed object is as follows: Filtering, fusing, and feature extraction of full-dimensional state information generate features of the observed object. In the formula The lateral position relative to the vehicle. The longitudinal position relative to the vehicle. The lateral speed is relative to the vehicle's own speed. The longitudinal speed relative to the vehicle itself; The specific method for generating driving styles is as follows: By statistically analyzing the historical trajectory data of surrounding vehicles, the historical trajectory data is clustered and classified to extract driving styles, which include aggressive, normal, and cautious driving styles. In S1, the specific method for establishing the DQN model based on the Markov decision process is as follows: S11. Define the state space: Construct the state space for highway lane-changing decisions using the characteristics of the observed objects as the states; S12. Define the action space: The action space of each agent consists of a set of high-level driving actions, including acceleration, deceleration, constant speed, left lane change and right lane change.

3. The lane-changing decision-making method considering driving style perceptual reinforcement learning according to claim 1, characterized in that, In S21, the expression for dividing different traffic flow density scenarios into several tasks is as follows: In the formula, For the first One task, For tasks with different traffic flow densities, The driving task distribution is such that each task corresponds to a traffic flow scenario with different density, which is used to construct the corresponding lane change decision scenario in the simulation environment.

4. The lane-changing decision-making method considering driving style perceptual reinforcement learning according to claim 3, characterized in that, In S22, the method for constructing the meta-parameters obtained from meta-reinforcement learning includes the following steps: S221. Set the maximum number of iterations and initialize the number of iterations; S222. At the current iteration number, the agent interacts with the task environment using the current policy to collect a batch of experience data as an experience replay dataset. The experience replay dataset includes state, action, immediate reward and subsequent state transition information. S223. Based on the empirical replay dataset, calculate the TD target using the temporal difference concept in reinforcement learning, and measure the current... The difference between the estimated value and the TD objective is used to construct the loss function for the current task. S224. Along the negative gradient direction of the loss function with respect to the current task adaptation parameters, perform a gradient update with a preset in-task learning rate to obtain the updated task adaptation parameters. S225. Determine if the current iteration count is equal to the maximum iteration count. If not, increment the current iteration count by 1 and return to S222. If yes, output the final task adaptation parameters.

5. The lane-changing decision-making method considering driving style perceptual reinforcement learning according to claim 3, characterized in that, In S223, the TD target is calculated. The specific expression is: In the formula, In the state Execute action Instant rewards In order to take new action The new state that is obtained later In order to be in The maximum expected return that can be obtained by taking different actions. t For time, For the first The next iteration, the... Task adaptation parameters for each task k This represents the current iteration number. The discount factor for the reward; In S223, calculate the... Task loss function The specific expression is: In the formula, For time state, In the state Next action At that time, by the first The next iteration, the... Task adaptation parameters for each task The corresponding action value output by the Q network, For from the first The second iteration Experience replay dataset for each task Mathematical expectation operation for sampling and transferring samples.

6. The lane-changing decision-making method considering driving style perceptual reinforcement learning according to claim 5, characterized in that, In S224, we obtain the first... The next iteration, the... Adaptation parameters for each task The specific expression is: In the formula, The learning rate within the task. For the loss function on the th The next iteration, the... Adaptation parameters for each task The gradient operator is used to calculate the direction and step size of parameter updates.

Citation Information

Patent Citations

  • Vehicle lane changing decision model training method and vehicle lane changing decision method

    CN118569097A

  • Expressway lane changing decision-making method and system considering driving styles of surrounding vehicles

    CN118953411A

  • Intelligent highway lane changing method for autonomous vehicle based on reinforcement learning

    CN119568155A