Multi-mode perception and DRL-MPC fused vehicle lane changing control method

The multi-modal perception and DRL-MPC fusion method addresses the robustness issues in vehicle lane changing by integrating environmental perception, driver intent recognition, and trajectory prediction to enhance safety and adaptability in complex traffic conditions.

CN120308122AActive Publication Date: 2025-07-15CHANGCHUN UNIV OF TECH

Patent Information

Application Number
CN202510740280.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-15
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The existing DRL-MPC fusion vehicle lane change control method is difficult to ensure robustness and generalization in complex traffic scenarios, especially in lateral behavior and multi-vehicle interaction environments.

Method used

The method of fusion of multimodal perception and DRL-MPC is adopted to extract BEV feature information through the environment perception module, combine driving intention recognition and multimodal trajectory prediction to generate multimodal trajectory information, and the DRL decision module generates control quantities such as expected vehicle speed and angle, and finally the MPC module solves the optimal control action.

Benefits of technology

It improves the safety, forward-looking and generalization capabilities of lane change control of autonomous driving vehicles, and enhances the robustness of vehicle lane change control and generalization of strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120308122A_ABST
    Figure CN120308122A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode perception and DRL-MPC fused vehicle lane changing control method, which is used for guaranteeing the safety of vehicle lane changing driving in a complex traffic environment. The invention relates to the field of intelligent driving. The system comprises an upper layer and a lower layer, the upper layer comprises an environment sensing module, a driving intention recognition module and a multi-modal trajectory prediction module, and the lower layer comprises a DRL decision module and an MPC module. The upper layer uses an environment perception module to extract traffic environment information from a driving environment to generate BEV feature information, a driving intention recognition module generates a driving intention probability according to the BEV feature information, and a multi-modal trajectory prediction module generates multi-modal trajectory information according to the BEV feature information and the driving intention probability; and the lower-layer DRL decision module generates an expected vehicle speed, an expected front wheel steering angle, a target weight and a control quantity weight according to the environment state information and the multi-modal trajectory information, and finally, an optimal control action is solved and determined by the MPC module, so that lane changing control of the automatic driving vehicle is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent driving, and particularly relates to a vehicle lane-changing control method integrating multi-modal perception and DRL-MPC. Background Art

[0002] With the rapid development of automotive intelligence and networking, intelligent driving technology has gradually become an important way to improve vehicle safety and energy efficiency. Under this background, the stability of vehicle lane-changing control, as one of the core challenges of intelligent driving, directly affects driving safety. Model Predictive Control (MPC) is widely used in autonomous driving lane-changing control by predicting the future state of the vehicle and optimizing the control sequence. However, MPC relies on physical models, and the system computational complexity increases with the increase of the prediction step length and constraint complexity, making it difficult to respond to complex road environments in real time. Deep Reinforcement Learning (DRL) can generate control strategies through environmental interaction learning, but due to insufficient policy generalization, there is a risk of out-of-control in new scenarios. When an autonomous vehicle performs motion control, the advantages of MPC considering safety constraints and actuator physical constraints can be fully utilized, taking MPC as a safety shield to ensure that the control actions output by DRL can guarantee the driving safety of the vehicle. In addition, the ability of DRL to continuously interact with the environment for global optimization can be fully utilized to provide prior control guidance for MPC, significantly reducing the computational burden of online solution of MPC. Therefore, the vehicle lane-changing control method integrating DRL-MPC can effectively respond to complex dynamic traffic scenarios and achieve stable and efficient lane-changing control.

[0003] Currently, in the aspect of vehicle lane-changing control integrating DRL and MPC, Patent CN117922567A conducts research on longitudinal trajectory error based on MPC. The DRL module real-time collects the states such as the spacing, speed, and acceleration error of each vehicle in the formation, online learns and adjusts the control weights in MPC, realizes the multi-objective optimization of the desired acceleration of the intelligent connected vehicles in the formation and converts it into control commands at the lower layer, and realizes the safe and stable cooperative control of the queue. However, this method only focuses on longitudinal dynamics, lacks in-depth modeling of lateral behavior and road semantics, and it is difficult to ensure robustness in complex traffic scenarios; Patent CN118605160A designs a quadratic performance index including trajectory error and stability error based on the MPC framework, and the weight matrix is updated in real time by the Deep Deterministic Policy Gradient (DDPG) algorithm to coordinate lateral tracking accuracy and lateral stability. However, this method does not consider multi-source environmental information from the BEV perspective and the driving intention of the driver, and it is difficult to adapt to complex traffic scenarios; Patent CN117360544A solves the initial optimal steering angle sequence based on MPC, and the prediction result of MPC is adjusted and optimized in real time by combining DRL with the tracking error, significantly improving the lateral control accuracy and anti-interference ability. However, this method uses a simplified two-degree-of-freedom bicycle model and fails to fully utilize the complex interaction relationships between vehicles, and it is easy to have potential safety hazards in the complex traffic environment of multi-vehicle interaction. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a vehicle lane-changing control method integrating multi-modal perception and DRL-MPC. This method extracts traffic environment information from the driving environment by the hierarchical Transformer in the environment perception module to generate BEV feature information, the driving intention recognition module outputs the driving intention probability based on the BEV feature information, and the multi-modal trajectory prediction module generates multi-modal trajectory information according to the driving intention probability and BEV feature information; the DRL decision module generates the desired vehicle speed, desired front wheel steering angle, target weight, and control quantity weight end-to-end, and finally the MPC optimizes and solves online to determine the optimal control action, realizing the lane-changing control of autonomous vehicles. This method effectively improves the forward-looking ability, risk assessment ability of autonomous vehicles to select lane-changing schemes, and generalization ability to complex scenarios while ensuring safety.

[0005] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0006] The present invention relates to a vehicle lane-changing control method integrating multi-modal perception and DRL-MPC. This method includes two layers. The upper layer includes a multi-modal environment perception module, a driving intention recognition module, and a multi-modal trajectory prediction module. The lower layer includes a DRL decision-making module and an MPC module. The upper layer uses the environment perception module to extract traffic environment information from the driving environment to generate BEV feature information. The driving intention recognition module generates a driving intention probability based on the BEV feature information. The multi-modal trajectory prediction module generates multi-modal trajectory information based on the BEV feature information and the driving intention probability and transmits it to the lower layer. The lower layer DRL decision-making module generates an expected vehicle speed, an expected front wheel steering angle, a target weight, and a control quantity weight based on the environment state information and the multi-modal trajectory information. Finally, the MPC module solves to determine the optimal control action to achieve the lane-changing control of the autonomous vehicle.

[0007] The method includes the following steps:

[0008] Step 1, Traffic environment information perception:

[0009] The environment perception module extracts BEV feature information from the driving environment through a hierarchical Transformer network and uses a dynamic sparse attention mechanism to focus on key areas.

[0010] Step 2, Driving intention recognition:

[0011] The driving intention recognition module generates the surrounding vehicle driving intention probability based on the BEV feature information. Specifically, the driving intention recognition module receives the BEV feature information generated by the environment perception module and inputs it into the graph neural network in the driving intention recognition module to construct the interaction relationship between the agent and the road information. It uses Transformer to output the intention time series features and score the intention time series features. At the same time, a mask suppression term is introduced, and finally, the driving intention probability P is obtained by normalization. The scoring of the driving intention recognition module is defined as in Equations (1) to (3).

[0012] Z i =W (i) h t +b (i) (1),

[0013]

[0014]

[0015] where Z i is the intention score, h t is the input feature, W and b are learnable parameters, β is the mask penalty intensity parameter, ρ is the non-linear suppression coefficient, is the intention score after introducing the mask suppression term, m iis the mask identifier, and ε is the probability smoothing factor to prevent division by zero.

[0016] Step 3, Multimodal Trajectory Prediction:

[0017] The multimodal trajectory prediction module generates multimodal trajectory information based on the driving intention probability and BEV feature information; specifically, the Transformer in the multimodal trajectory prediction module is used to process the BEV feature information, the future trajectory point sequence is generated through the LSTM decoder, the semantic score and geometric score of each trajectory are calculated, and finally the trajectory confidence Conf K is obtained. The multimodal trajectory confidence scoring is defined as in Equations (4) to (7),

[0018]

[0019]

[0020]

[0021]

[0022] where, is the semantic consistency score, w i is a learnable parameter, h k is the state feature of the k-th trajectory, is the geometric matching degree score, is the feature of the corresponding area in the BEV space, IOU is the intersection over union of the trajectory and the lane line, λ is the road deviation penalty factor, H is the indicator function, τ is the threshold, and H is 1 when IOU ≤ τ, otherwise H is 0.

[0023] Step 4, Design the DRL Decision Module:

[0024] Step 4.1, State Space Design:

[0025] The state space of the lower-level DRL decision module is defined as s t =[X ego , X sur , r road , r topo , where X ego is the ego-vehicle state information, X sur is the surrounding-vehicle state information, r road is the road curvature, r topo is the road topology; the action space of the DRL decision module is defined as v f is the desired vehicle speed, δ f is the desired front wheel angle, is the target weight, and ψ is the control quantity weight.

[0026] Step 4.2, Reward Function Design:

[0027] The present invention designs a reward function based on safety and comfort, as shown in Equations (8) to (10).

[0028] r t = r safe + r comfort (8),

[0029]

[0030]

[0031] In the formula, r safe is the safety reward function, r comfort is the comfort reward function, p1, p2, p3 are constants, ΔT d is the change in driving torque, and F b is the braking force.

[0032] Step 5, Design the MPC Module:

[0033] Step 5.1, Establish the Prediction Model:

[0034] The prediction model is established using the vehicle kinematic model, vehicle dynamics model, and non-linear tire model, as shown in Equation (11).

[0035]

[0036] In the formula, x and y are the longitudinal and lateral position coordinates of the host vehicle, v x and v y are the longitudinal and lateral vehicle speeds respectively, θ is the heading angle of the host vehicle, F y,f is the lateral force of the front wheel, F y,r is the lateral force of the rear wheel, m is the vehicle mass, γ is the vehicle yaw angular velocity, a and b are the distances from the vehicle center of mass to the front and rear axles, and I z is the vehicle yaw moment of inertia.

[0037] Step 5.2, Model Discretization:

[0038] The above model is discretized using the fourth-order Runge-Kutta method, and the sampling time is T s , and the input-output relationship of the incremental discrete system model can be expressed, as shown in Equation (12).

[0039]

[0040] In the formula, u(s) = [δ f ,

[0041] Among them, ξ(s) is the state variable of the vehicle at the current s moment, u(s) and Δu(s) are the control input and control input increment of the system at the current s moment respectively, and y c (s) is the predicted output of the system at the current s moment, and C is the coefficient matrix used to determine the number of the system predicted output.

[0042] Step 5.3: Optimization problem construction:

[0043] The optimization problem construction takes into account the ride comfort and control action stability, imposes hard constraints on the control actions, and defines the optimization objectives of the MPC module, such as Equations (13) to (15),

[0044]

[0045]

[0046] L(x(k), u(k)) = ω r (v(k) - v f ) 2 + ω r (δ(k) - δ f ) 2 + ω c P(k)(15),

[0047] Among them, J is the objective function representing the weighted combination of the speed and front wheel angle tracking deviations, x = [v h , δ h is the system state variable, v h (N + 1) is the vehicle speed at the (N + 1)-th moment, δ h (N + 1) is the front wheel angle at the (N + 1)-th moment, ω r is the dynamic performance parameter, ω c is the comfort weight parameter, and P(k) in the objective function represents the comfort performance function, such as Equation (16),

[0048] P(k) = (T f (k)I g (k) - T f (k - 1)I g (k - 1)) 2 + F b 2 (k)(16),

[0049] Among them, T f (k) is the drive motor torque at the k-th moment, I g (k) is the transmission gear ratio at the k-th moment, F b(k) is the braking force at the k-th moment.

[0050] Step 5.4, Dynamic optimization and solution:

[0051] Dynamic optimization is to solve the optimization problem to obtain the optimal control action u0 at the current moment, including the front wheel angle δ o , and the longitudinal vehicle speed v0; finally, input the optimal control action u o into the autonomous vehicle to achieve a lane change operation.

[0052] The beneficial effects of the present invention are as follows: The present invention relates to a vehicle lane change control method integrating multi-modal perception and DRL-MPC. This method includes two layers. The upper layer constructs the traffic environment information in the driving environment into BEV feature information from a bird's-eye view, and generates multi-modal trajectory information in combination with the driving intention recognition module and the multi-modal trajectory prediction module. On this basis, the DRL decision-making module interacts with the environment to output the desired vehicle speed, desired front wheel angle, target weight, and control quantity weight. Finally, the MPC solution determines the optimal control action to achieve the lane change control of the autonomous vehicle. This method can improve the generalization and robustness of the vehicle lane change control strategy while ensuring safety, and has good application prospects in the field of intelligent vehicle lane change control. Brief Description of the Drawings

[0053] Figure 1 is the hierarchical control schematic diagram of DRL-MPC based on multi-modal perception of the present invention. Detailed Embodiment

[0054] The present invention will be described in detail below with reference to the drawings.

[0055] The present invention proposes a vehicle lane change control method integrating multi-modal perception and DRL-MPC. Referring to Figure 1 for illustration, it specifically includes the following steps:

[0056] Step 1, Environment perception module:

[0057] The environment perception module is responsible for extracting traffic environment information from the driving environment and generating BEV feature information. The BEV feature information includes obstacle type, vehicle position, speed, acceleration, and road environment characteristics. The environment perception module consists of a hierarchical Transformer structure. The lower layer extracts geometric features, the middle layer extracts semantic features, and the upper layer uses a dynamic sparse attention mechanism to focus on key areas, finally forming an efficient and concise feature representation system. By collecting traffic environment information in the driving environment in real time, the system can provide accurate data support for the control algorithm.

[0058] Step 2, Driving intention recognition module:

[0059] The driving intention recognition module is used to obtain the surrounding vehicle driving intention probability based on the BEV feature information. Specifically, the driving intention recognition module receives the BEV feature information generated by the environmental perception module and inputs it into the graph neural network in the driving intention recognition module to construct the interaction relationship between the agent and the road information. It uses Transformer to output the intention time series features and score the intention time series features. At the same time, a mask suppression term is introduced, and finally the driving intention probability P is obtained through normalization. The scoring of the driving intention recognition module is defined by Equations (17) to (19).

[0060] Z i =W (i) h t +b (i) (17),

[0061]

[0062]

[0063] Among them, Z i is the intention score, h t is the intention time series feature, W and b are learnable parameters, β is the mask penalty strength parameter, ρ is the non-linear suppression coefficient, is the intention score after introducing the mask suppression term, m i is the mask identifier, and ε is the probability smoothing factor to prevent division by zero.

[0064] Step 3, Multimodal Trajectory Prediction Module:

[0065] The multimodal trajectory prediction module generates multimodal trajectory information based on the driving intention probability and BEV feature information. Specifically, it uses the Transformer in the multimodal trajectory prediction module to process the BEV feature information to form spatio-temporal dependencies, generates the future trajectory point sequence through the LSTM decoder, then calculates the semantic score and geometric score of each trajectory, and finally obtains the trajectory confidence Conf K , and the scoring of the multimodal trajectory confidence is defined by Equations (20) to (23).

[0066]

[0067]

[0068]

[0069]

[0070] Among them, is the semantic consistency score, w i is the learnable parameter, h kis the state feature of the k-th trajectory, is the geometric matching degree score, is the feature of the corresponding area in the BEV space, IOU is the intersection over union of the trajectory and the lane line, λ is the road deviation penalty factor, H is the indicator function, τ is the threshold, H is 1 when IOU ≤ τ, otherwise H is 0.

[0071] Step 4: Design the DRL decision-making module:

[0072] Step 4.1: State space design:

[0073] The state space of the DRL decision-making module is defined as s t = [X ego , X sur , r road , r topo , where X ego is the ego-vehicle state information, X sur is the surrounding-vehicle state information, r road is the road curvature, r topo is the road topology; the action space of the DRL decision-making module is defined as v f is the desired vehicle speed, δ f is the desired front wheel angle, is the target weight, ψ is the control quantity weight.

[0074] Step 4.2: Action space design:

[0075] The action space of the DRL decision-making module is defined as This action outputs the desired front wheel angle, desired vehicle speed, target weight, and control quantity weight of the vehicle through the deep reinforcement learning algorithm to ensure that the vehicle can achieve good lane-changing effects in different driving scenarios. The outputs of the front wheel angle and vehicle speed are dynamically adjusted according to the ego-vehicle state information X ego , surrounding-vehicle state information X sur , road curvature r road and road topology r topo , where the ego-vehicle state information X ego includes vehicle speed, vehicle position coordinates, and front wheel angle, and the surrounding-vehicle state information X sur includes vehicle speed, vehicle position coordinates, and driving intention probability. The target weight and control quantity weight ψ are dynamically adjusted according to the speed error, front wheel angle error, road curvature r road and surrounding-vehicle state information X sur .

[0076] Step 4.3: Reward function design:

[0077] The present invention designs a reward function based on safety and comfort, such as Equations (24) to (26).

[0078] r t = r safe + r comfort (24),

[0079]

[0080]

[0081] In the formula, r safe is the safety reward function, r comfort is the comfort reward function, p1 is 100, p2 is 200, p3 is 100, ΔT d is the change in drive torque, F b is the braking force.

[0082] Step 4.4, Experience replay:

[0083] In the deep reinforcement learning method, the experience replay mechanism is used to improve the sample utilization rate and accelerate policy convergence. This mechanism breaks the temporal correlation between data by storing and re-sampling historical interaction data, thereby enhancing the stability and generalization ability of training. In this solution, an experience pool is defined, and the training data at each moment is stored in this pool. When the number of experiences in the experience pool reaches a preset value, a certain number of experience samples are randomly drawn from the experience pool to form a batch for training. Specifically, at the k-th moment of the t-th round of training, N experiences are randomly selected from the experience pool to form an experience batch for training, where N is set to 64.

[0084] Step 4.5, Select a deep reinforcement learning algorithm:

[0085] Different reinforcement learning methods have their own advantages and disadvantages. In Step 4 of the present invention, the SAC algorithm is selected as the deep reinforcement learning algorithm. The SAC algorithm can select the optimal action in the continuous action space and, by introducing the maximum entropy mechanism, makes the policy learned by the intelligent agent more random, thereby helping to improve the generalization ability of the control policy.

[0086] Step 4.6, DRL decision module training:

[0087] Set the regularization parameter α = 0.01, the soft update parameter τ = 0.003, the discount factor γ = 0.95, and initialize the neural network parameters in the SAC algorithm, including the Actor network parameter θ actor 、the Critic1 network parameter θ critic1 、the Critic2 network parameter θ critic2 、the target Critic1 network parameter θtargetcritic1 and the target Critic2 network parameters θ targetcritic2 . In the t-th training episode, obtain the environmental state information from the driving environment, i.e., the ego-vehicle state, the surrounding vehicle state, the road topology, and the road curvature, and form the current state s in step 4.1 t . This state is input into the Actor network, and after global optimization, the action a in step 4.2 is output t , and this action is within the predetermined action space range. Execute the action a t . After that, the information collection module returns the reward value r according to the current state t and the new state s t+1 . Then, the state s t , the action a t , the reward value r t and the new state s t+1 are stored in the experience pool in step 4.4. When the number of experiences in the experience pool reaches 512, randomly extract 128 experiences from the pool. Input these experiences into the Critic1 network, the Critic2 network, the target Critic1 network, and the target Critic2 network, and calculate Q critic1 , Q critic2 , the target value Q targetcritic1 , and Q targetcritic2 , and calculate the temporal difference (TD) target value using Equation (27).

[0088] η t = r t + γ min{Q targetcritic1 , Q targetcritic2}- α log π(a t+1 |s t+1 )(27),

[0089]

[0090]

[0091] By minimizing the loss functions in Equations (28) and (29), the Critic1 network and the Critic2 network are updated.

[0092] Sample actions using the reparameterization method

[0093]

[0094] Update the Actor network using the loss function in Equation (30).

[0095] The target Critic1 network and the target Critic2 network are updated using Equations (31) and (32),

[0096] τQ critic1 +(1 - τ)Q targetcritic1 →Q targetcritic1 (31),

[0097] τQ critic2 +(1 - τ)Q targetcritic1 →Q targetcritic1 (32),

[0098] The above training process will be continuously iterated until the algorithm converges. After the algorithm converges, the optimal network parameters are selected and loaded into the Actor network, thus completing the training process of the reinforcement learning module in the present invention.

[0099] Step 5: Design the MPC module:

[0100] Step 5.1: Establish a prediction model:

[0101] The prediction model is established using the vehicle kinematic model, vehicle dynamics model, and nonlinear tire model, as shown in Equation (33),

[0102]

[0103] where F y,f is the lateral force of the front wheel, F y,r is the lateral force of the rear wheel, m is the vehicle mass, γ is the vehicle yaw angular velocity, a and b are the distances from the vehicle center of mass to the front and rear axles, and I z is the vehicle yaw moment of inertia.

[0104] The calculation of the tire lateral force is as shown in Equation (34),

[0105] F y = μD y sin(C y arctan(B y α - E y (B y α - arctan(B y α)))) + S vy (34),

[0106] where F y is the tire lateral force, μ is the road adhesion coefficient, α is the tire slip angle, B y is the tire characteristic curve stiffness factor, C y is the tire characteristic curve shape factor, D y is the tire characteristic curve peak factor, E y is the tire characteristic curve curvature factor, S vy is the tire characteristic curve vertical offset, except for C yExcept for this, other tire characteristic curve parameters are related to the vertical load of the tire.

[0107] The equivalent sideslip angles of the tires on the front and rear axles of the vehicle, as shown in Equations (35) and (36),

[0108]

[0109]

[0110] where α f , α r respectively represent the equivalent sideslip angles of the tires on the front and rear axles.

[0111] The vertical loads of the tires on the front and rear axles of the vehicle, as shown in Equations (37) and (38),

[0112]

[0113]

[0114] Using the fourth-order Runge-Kutta method to discretize the above model with a sampling time of T s , the input-output relationship of the incremental discrete system model can be expressed, as shown in Equation (39),

[0115]

[0116] In the formula, u(s) = [δ f ,

[0117] where ξ(s) is the state variable of the vehicle at the current s moment, u(s) and Δu(s) are the control input and control input increment of the system at the current s moment respectively, and y c (s) is the predicted output of the system at the current s moment, and C is the coefficient matrix used to determine the number of system predicted outputs.

[0118] The construction of the optimization problem takes into account the ride comfort and the control action stability, imposes hard constraints on the control actions, and defines the optimization objectives of the MPC module, as shown in Equations (40) to (42),

[0119]

[0120]

[0121] L(x(k), u(k)) = ω r (v(k) - v f ) 2 + ω r (δ(k) - δ f )2 +ω c P(k)(42),

[0122] where J is the objective function representing the weighted combination of speed and front wheel angle tracking deviation, x = [v h , δ h is the system state variable, v h (N + 1) is the vehicle speed at the (N + 1)-th moment, δ h (N + 1) is the front wheel angle at the (N + 1)-th moment, ω r is the dynamic performance parameter, ω c is the comfort weight parameter, and P(k) in the objective function represents the comfort performance function, as shown in Equation (43),

[0123] P(k) = (T f (k)I g (k) - T f (k - 1)I g (k - 1)) 2 + F b 2 (k)(43),

[0124] where T f (k) is the driving motor torque at the k-th moment, I g (k) is the transmission gear ratio at the k-th moment, F b (k) is the braking force at the k-th moment.

[0125] Dynamic optimization is to solve the optimization problem to obtain the optimal control action u0 at the current moment, including the front wheel angle δ o , and the longitudinal vehicle speed v0; finally, input the optimal control action u o into the autonomous driving vehicle to implement the lane change operation.

[0126] In summary, the present invention relates to a vehicle lane change control method integrating multi-modal perception and DRL-MPC. This method constructs the traffic environment information in the driving environment into BEV feature information from a bird's-eye view, and combines the driving intention recognition module and the multi-modal trajectory prediction module to generate multi-modal trajectory information. On this basis, the DRL decision-making module interacts with the environment to output the desired vehicle speed, desired front wheel angle, target weight, and control quantity weight, and at the same time determines the optimal control action based on the MPC solution to achieve the lane change control of the autonomous driving vehicle. This method can improve the generalization and robustness of the vehicle lane change control strategy on the premise of ensuring safety, and has good application prospects in the field of intelligent vehicle lane change control.

Claims

1. A vehicle lane-changing control method integrating multi-modal perception and DRL-MPC, characterized in that: The method includes two layers, the upper layer includes an environmental perception module, a driving intention recognition module, and a multi-modal trajectory prediction module, and the lower layer includes a DRL decision-making module and an MPC module; the environmental perception module in the upper layer extracts traffic environment information from the driving environment through a hierarchical Transformer network with dynamic sparse attention to generate BEV feature information, the driving intention recognition module processes the BEV feature information through a graph neural network and a Transformer and introduces a masking mechanism to generate driving intention probabilities, and the multi-modal trajectory prediction module processes the BEV feature information and driving intention probabilities according to a Transformer and an LSTM and introduces a scoring mechanism to generate multi-modal trajectory information; The DRL decision-making module in the lower layer generates an expected vehicle speed, an expected front wheel angle, a target weight, and a control quantity weight according to the environmental state information and the multi-modal trajectory information, and finally the MPC module solves to determine the optimal control action to achieve the lane-changing control of the autonomous vehicle.

2. The vehicle lane-changing control method integrating multimodal perception and DRL-MPC according to claim 1, characterized in that: The environmental perception module in the upper layer extracts traffic environment information from the driving environment through a hierarchical Transformer network to generate BEV feature information, and adopts a dynamic sparse attention mechanism to focus on key areas; The driving intention recognition module in the upper layer receives the BEV feature information generated by the environmental perception module and inputs it into the graph neural network in the driving intention recognition module to construct the interaction relationship between the agent and the road information, uses the Transformer to output the intention time-series features and score the intention time-series features, and at the same time introduces a masking suppression term, and finally normalizes to obtain the driving intention probability P. The scoring of the driving intention recognition module is defined as in Equations (1) to (3). Z i = W (i) h t + b (i) (1) Among them, Z i is the intention score, h t is the input feature, W and b are learnable parameters, β is the mask penalty intensity parameter, ρ is the non-linear suppression coefficient, is the intention score after introducing the mask suppression term, m i is the mask flag, ε is the probability smoothing factor to prevent division by zero; The upper multi-modal trajectory prediction module processes BEV feature information and driving intention probability according to Transformer, uses the LSTM decoder to generate a sequence of future trajectory points, calculates the semantic score and geometric score of each trajectory, and finally obtains the trajectory confidence Conf K , and the definition of the multi-modal trajectory confidence score is as shown in Equations (4) to (7). Among them, is the semantic consistency score, w i is a learnable parameter, h k is the state feature of the k-th trajectory, is the geometric matching score, is the feature of the corresponding region in the BEV space, IOU is the intersection over union of the trajectory and the lane line, λ is the road deviation penalty factor, H is the indicator function, τ is the threshold, H is 1 when IOU ≤ τ, otherwise H is 0.

3. A vehicle lane-changing control method integrating multimodal perception and DRL-MPC according to claim 1, characterized in that: The lower-level DRL decision-making module generates the desired vehicle speed, desired front-wheel steering angle, target weight, and control quantity weight based on the environmental state information and multi-modal trajectory information; the state space of the DRL decision-making module is defined as s t =[X ego ,X sur ,r road ,r topo , where X ego is the ego-vehicle state information, X sur is the surrounding-vehicle state information, r road is the road curvature, r topo is the road topology; the action space of the DRL decision-making module is defined as v f is the desired vehicle speed, δ f is the desired front-wheel steering angle, is the target weight, ψ is the control quantity weight; the present invention designs a reward function based on safety and comfort, as shown in Equations (8) to (10), r t =r safe +r comfort (8), where r safe is the safety reward function, r comfort is the comfort reward function, p1, p2, p3 are constants, and ΔT d is the change in driving torque, and F b is the braking force; The MPC module in the lower layer outputs the optimal control action using the expected vehicle speed, the expected front wheel angle, the target weight, and the control quantity weight output by the DRL decision-making module; the MPC module includes three parts: a prediction model, an optimization problem construction, and a dynamic optimization. The prediction model is established using a vehicle kinematic model, a vehicle dynamics model, and a non-linear tire model, as shown in Equation (11). where x and y are the longitudinal and lateral position coordinates of the vehicle itself, v x and v y are the longitudinal and lateral vehicle speeds respectively, θ is the heading angle of the vehicle itself, F y,f is the lateral force of the front wheels, F y,r is the lateral force of the rear wheels, m is the vehicle mass, γ is the yaw angular velocity of the vehicle, a and b are the distances from the vehicle's center of mass to the front and rear axles, I z is the yaw moment of inertia of the vehicle; The above model is discretized using the fourth-order Runge-Kutta method, and the sampling time is T s , and the input-output relationship of the incremental discrete system model can be expressed as in Equation (12). wherein, u(s) = [δ f , Among them, ξ(s) is the state variable of the vehicle at the current time s, u(s) and Δu(s) are the control input and the control input increment of the system at the current time s respectively, and y c (s) is the predicted output of the system at the current time s, and C is the coefficient matrix used to determine the number of the predicted outputs of the system; The optimization problem construction takes into account the ride comfort and the stability of the control action, imposes hard constraints on the control action, and defines the optimization objective of the MPC module, as shown in Equations (13) to (15). L(x(k), u(k)) = ω r (v(k) - v f ) 2 + ω r (δ(k) - δ f ) 2 + ω c P(k)(15), where J is the objective function representing the weighted combination of the speed and the front wheel steering angle tracking deviation, and x = [v h , δ h is the system state variable, v h (N + 1) is the vehicle speed at the (N + 1)-th moment, and δ h (N + 1) is the front wheel steering angle at the (N + 1)-th moment, ω r is the dynamic performance parameter, ω c is the comfort weight parameter, and P(k) in the objective function represents the comfort performance function, as shown in Equation (16). P(k) = (T f (k)I g (k) - T f (k - 1)I g (k - 1)) 2 + F b 2 (k)(16), Among them, T f (k) is the driving motor torque at the k-th moment, I g (k) is the transmission gear ratio at the k-th moment, F b (k) is the braking force at the k-th moment; Dynamic optimization is to solve the optimization problem to obtain the optimal control action u0 at the current moment, including the front wheel steering angle δ o , and the longitudinal vehicle speed v0; finally, input the optimal control action u o into the autonomous vehicle to implement the lane change operation.

Citation Information

Patent Citations

  • Rotor-helicopter-borne navigation device based on strapdown inertial navigation system

    CN110260862A

  • Intelligent vehicle lane changing decision-making method and system for LSTM trajectory prediction

    CN117325865A

  • Vehicle driving state judgment method based on dynamic threshold equation

    CN118247306A

  • Intelligent driving optimization control method integrating global optimization and safety protection

    CN118618404A

  • Path tracking control method fusing reinforcement learning adaptive preview

    CN119472689A

Cited By

  • Method and system for generating driving track of autonomous vehicle

    CN120646020A