A reinforcement learning method for automatic driving ramp merging based on MPC guidance
By using an MPC-guided reinforcement learning method, combining environmental information and vehicle state to generate reward and feasible set projection intervals, and optimizing the policy network, the problem of balancing safety and efficiency in autonomous driving ramp merging is solved, improving model prediction accuracy and reducing solution latency sensitivity, and mitigating reward signal confusion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- EAST CHINA JIAOTONG UNIVERSITY
- Filing Date
- 2026-02-24
- Publication Date
- 2026-04-17
AI Technical Summary
Existing reinforcement learning methods struggle to balance safety and efficiency in autonomous driving ramp merging processes, and suffer from issues such as model predictive control being sensitive to errors, solution delays, distribution offset effects, and confusion between single-channel reward-mixed tasks and safety signals.
We employ a reinforcement learning method based on MPC guidance. By collecting environmental information and vehicle operation status information of ramp merging, we generate MPC guidance rewards and feasible set projection intervals. We then fuse the MPC guidance rewards with environmental task rewards to optimize the initial multi-critic policy network, generate the final executable safety actions, improve the model's prediction accuracy, and reduce the sensitivity to solution latency.
It improves model prediction accuracy, reduces sensitivity to solution latency, reduces the impact of distribution offset, solves the gradient interference and credit allocation confusion problems of single-channel reward mixed tasks and security signals, and achieves an explicit balance between security and efficiency.
Smart Images

Figure CN121697638B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, and in particular to a reinforcement learning-based autonomous driving ramp merging method guided by MPC. Background Technology
[0002] Ramp merging requires autonomous vehicles to maintain traffic efficiency while meeting safety constraints. Existing reinforcement learning methods typically require a large number of interactions, making it difficult to achieve a reliable balance between safety and efficiency. Existing methods mainly include:
[0003] Rule-based, heuristic, and classical control methods (such as IDM / MOBIL) lack flexibility under dense traffic or atypical driver behavior; Model predictive control (MPC) and game theory methods, while able to handle dynamic constraints, are sensitive to model errors and solution delays; Imitation learning and offline reinforcement learning are susceptible to distribution shifts and lack explicit efficiency drivers; In reinforcement learning and safety reinforcement learning methods, single-channel rewards mixed with task and safety signals lead to gradient interference and confusion in credit assignment.
[0004] Therefore, there is an urgent need for a control method that can explicitly balance safety and efficiency, and has high sample efficiency and robustness. Summary of the Invention
[0005] Therefore, the purpose of this invention is to provide a reinforcement learning-guided autonomous driving ramp merging method based on MPC to address the shortcomings of existing technologies.
[0006] To achieve the above objectives, this invention provides a reinforcement learning-guided autonomous driving ramp merging method based on MPC guidance, the method comprising:
[0007] Collect environmental information and vehicle operation status information of ramp merging, and generate MPC guidance reward and feasible set projection interval based on the environmental information and vehicle operation status information;
[0008] Based on the current driving stage, the MPC guidance reward and the environmental task reward are integrated to obtain the single-step reward. The reward signal corresponding to the single-step reward is optimized and decomposed into the target task reward and the target safety cost.
[0009] The initial multi-commentator policy network is optimized and updated based on the target task reward and the target safety cost to obtain an optimized action policy network. Based on the current state of the master vehicle, the original action is generated through the optimized action policy network, and the original action is mapped to the feasible set projection interval to determine the final executable safety action. The initial multi-commentator policy network is constructed by explicitly introducing task commentators and safety commentators into the policy network structure of the near-end policy optimization. The initial multi-commentator policy network includes policy head parameters, task commentator parameters, and safety commentator parameters.
[0010] The beneficial effects of this invention are as follows: By collecting environmental information and vehicle operating status information of ramp merging, and generating MPC guidance rewards and feasible set projection intervals based on the environmental information and vehicle operating status information, and then fusing the MPC guidance rewards and environmental task rewards based on the current driving stage to obtain a single-step reward, the reward signal corresponding to the single-step reward is optimized and decomposed into target task reward and target safety cost. Then, the initial multi-critic policy network is optimized and updated based on the target task reward and target safety cost to obtain an optimized action policy network. Based on the current state of the master vehicle, the original action is generated through the optimized action policy network, and the original action is mapped to the feasible set projection interval to determine the final executable safety action. This invention differs from existing technologies in that it improves the model prediction accuracy and reduces the sensitivity to solution delay, improves the problems of being susceptible to distribution shift and lacking explicit efficiency drive, and also improves the problem of gradient interference and credit allocation confusion caused by mixing task and safety signals in a single channel reward.
[0011] Furthermore, the step of generating the MPC guidance reward and feasible set projection interval based on the environmental information and the vehicle operating status information includes:
[0012] Based on the environmental information and the vehicle operating status information, a reference control sequence for the master vehicle is generated within a preset prediction time domain, and the MPC guidance reward is calculated based on the reference control sequence.
[0013] The control inputs and state trajectories corresponding to the reference control sequence are projected onto a preset feasible set to obtain the feasible set projection interval.
[0014] Furthermore, the environmental information includes the lateral deviation and flight path deviation of the master vehicle, and the step of generating the reference control sequence of the master vehicle within the preset prediction time domain includes:
[0015] Based on the lateral deviation and the flight path deviation, a corresponding discrete linearized model is constructed. With vehicle dynamics constraints and road geometry constraints, and with the tracking error and control quantity minimization as the optimization function of the discrete linearized model, the optimization function is solved to obtain the reference control sequence.
[0016] The expression for the discrete linearized model is shown below:
[0017]
[0018] in, This represents the two-dimensional state of the main vehicle at time t+1, which is the difference between the lateral error and the heading error. Represents the state transition matrix. This represents the two-dimensional state of the main vehicle at time t, which is the difference between the lateral error and the heading error. Represents the control input matrix. This represents the control quantity of the master vehicle at time t;
[0019] in:
[0020]
[0021] ,
[0022] The expression for the optimization function is as follows:
[0023]
[0024] in, Represents the optimization function. This represents the sum of the predicted state variables from time t to time t+n. This represents the sum of the predicted control quantities from time t to time t+n. and Both are weight matrices. Indicates a minute amount of time. The weighted L2 norm of the state variables. The weighted L2 norm of the control quantity.
[0025] Furthermore, the step of calculating the MPC bootstrapping reward based on the reference control sequence includes:
[0026] Based on the current longitudinal speed and wheelbase of the main vehicle, the first control variable in the reference control sequence is mapped to the MPC expert steering angle at the current moment, in combination with the vehicle's physical characteristics.
[0027] The RL agent strategy outputs the policy steering angle under the same state, and constructs the MPC guided reward based on the MPC expert steering angle and the policy steering angle. The same state is the state of the master vehicle under the control of the first control quantity in the reference control sequence.
[0028] Furthermore, the expression for the MPC-guided reward is as follows:
[0029]
[0030] in, This indicates the MPC introductory reward. Indicates the same state. Represents a set of control actions. Indicates the strategy turning angle, Indicates the MPC expert steering angle;
[0031] The expression for the single-step reward is as follows:
[0032]
[0033] in, Indicates a single-step reward. Indicates the weight of the MPC bootstrapping reward. This indicates the task reward. As a safety reward, For security reward weighting, To guide the weights.
[0034] Furthermore, the step of merging the MPC guidance reward with the environment task reward to obtain the single-step reward includes:
[0035] Based on the vehicle operation status information, different parameter indicators and rewards are calculated for each stage of the merging process. The rewards for each parameter indicator are then mixed and processed by stage weighting and risk gating to obtain the task reward.
[0036] Security rewards and security costs are generated based on security indicators. The single-step reward is obtained by reconstructing the MPC guidance reward, the task reward, the security reward, and the security cost.
[0037] Furthermore, the step of optimizing and updating the initial multi-commenter policy network based on the target task reward and the target security cost to obtain the optimized action policy network includes:
[0038] The target security cost and the target task reward are collected synchronously over a fixed time span, and the target task reward and the target security cost are written into the task buffer and the security buffer, respectively.
[0039] The target task reward and the target security cost are estimated by time-series difference estimation using generalized advantage estimation to obtain different advantage data. Each advantage data is then normalized to obtain corresponding normalized data.
[0040] Based on the near-end policy optimization training framework, the normalized data is weighted and fused using Lagrange multipliers to obtain a hybrid advantage. Based on the hybrid advantage, the policy head parameters, the task critic parameters, and the safety critic parameters are jointly updated through backpropagation and gradient pruning to obtain an optimized action policy network.
[0041] Furthermore, the policy network structure includes a policy head, which is a Gaussian random policy in a continuous action space constructed on top of shared features.
[0042] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0043] Figure 1 This is a flowchart of the MPC-guided reinforcement learning-based autonomous driving ramp merging method according to Embodiment 1 of the present invention;
[0044] Figure 2 This is a scene diagram illustrating a complex highway road simulation as exemplified in Embodiment 2 of the present invention;
[0045] Figure 3 This is a comparison chart of the training success rates of the SAC method, PPO method, and PPO-MC-MPC method in the same training environment according to Embodiment 2 of the present invention.
[0046] Figure 4 This is a comparison chart of the average training speed of the SAC method, PPO method, and PPO-MC-MPC method in the same training environment according to Embodiment 2 of the present invention.
[0047] Figure 5 A comparison of the average rotation angles trained under the same training environment for the SAC method, PPO method, and PPO-MC-MPC method of Embodiment 2 of the present invention;
[0048] Figure 6 This is a structural block diagram of the MPC-guided reinforcement learning-based autonomous driving ramp merging system according to Embodiment 3 of the present invention.
[0049] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.
[0051] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.
[0052] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0053] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following objects are in an "or" relationship. The terms "first," "second," and "third" used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.
[0054] Example 1
[0055] Please see Figure 1 The flowchart below shows the reinforcement learning-based autonomous driving ramp merging method based on MPC guidance in the first embodiment of the present invention. The method includes the following steps:
[0056] Step S101: Collect environmental information and vehicle operation status information of ramp merging, and generate MPC guidance reward and feasible set projection interval based on the environmental information and vehicle operation status information;
[0057] The environmental information includes the relative distances between the vehicle and surrounding vehicles. Lateral deviation of the main vehicle relative to the center line of the target lane and the deviation of the vehicle's heading from the lane , These represent the distances of the main vehicle to the nearest vehicles in front and behind it relative to the target lane; the vehicle operating status information includes the main vehicle's x-coordinate. with the longitudinal coordinate of the main vehicle Longitudinal velocity Horizontal position Heading angle acceleration accelerometer and front wheel cornering .
[0058] It should be noted that the feasible set projection interval will serve as a safety boundary for subsequent action decisions.
[0059] Step S102: Based on the current driving stage, the MPC guidance reward and the environmental task reward are merged to obtain a single-step reward, and the reward signal corresponding to the single-step reward is optimized and decomposed into target task reward and target safety cost;
[0060] In this embodiment, the vehicle ramp merging process is divided into four driving stages, namely the ramp approach stage. 1. Initial ramp merging phase During the ramp merging phase and the driving phase after merging The merging process is marked as .
[0061] Because the same strategy has different objective priorities at different driving stages, without stage-adaptive reward organization, the trained strategy will either be hesitant to merge, overly aggressive in merging, or unstable. Therefore, to ensure that the decision-making for ramp merging is more appropriate when the vehicle is in different driving stages, the same strategy is used to make different responses. It should be noted that by obtaining the current position of the driver vehicle, i.e., the environmental information of ramp merging on the mobile phone, the road segment is defined, and then the current position of the driver vehicle is used to correspond to the position of the current road segment to determine which driving stage the driver vehicle is in.
[0062] Step S103: Optimize and update the initial multi-commenter policy network based on the target task reward and the target safety cost to obtain an optimized action policy network. Based on the current state of the master vehicle, generate original actions through the optimized action policy network and map the original actions to the feasible set projection interval to determine the final executable safety action.
[0063] Specifically, in the policy network structure for near-end policy optimization, task commentators and security commentators are introduced to form an initial multi-commentator policy network. The optimized action policy network generates original actions, which are then projected onto the feasible set projection interval to determine the final executable security action. Specifically, the policy head... Output a primitive action based on the current state. (Such as the desired acceleration and steering angle), the original action is sent to the safety action layer, which accesses the projection range of the feasible set. and the original action Projecting the data into the feasible set projection area yields the final executable safety action. .
[0064] Through the above steps, environmental information and vehicle operating status information of ramp merging are collected. Based on the environmental information and vehicle operating status information, MPC guidance reward and feasible set projection interval are generated. Then, based on the current driving stage, the MPC guidance reward and environmental task reward are fused to obtain a single-step reward. The reward signal corresponding to the single-step reward is optimized and decomposed into target task reward and target safety cost. Then, based on the target task reward and target safety cost, the initial multi-critic policy network is optimized and updated to obtain an optimized action policy network. Based on the current state of the master vehicle, the original action is generated through the optimized action policy network and mapped to the feasible set projection interval to determine the final executable safety action. Unlike existing technologies, this method improves the model prediction accuracy and reduces the sensitivity to solution delay. It also improves the problems of being susceptible to distribution offset and lacking explicit efficiency drive. Furthermore, it improves the problem of gradient interference and credit allocation confusion caused by mixing task and safety signals in a single channel reward.
[0065] Furthermore, the step of generating the MPC guidance reward and feasible set projection interval based on the environmental information and the vehicle operating status information includes:
[0066] Based on the environmental information and the vehicle operating status information, a reference control sequence for the master vehicle is generated within a preset prediction time domain, and the MPC guidance reward is calculated based on the reference control sequence.
[0067] The step of generating the reference control sequence of the master vehicle within the preset prediction time domain includes:
[0068] Based on the lateral deviation and the flight path deviation, a corresponding discrete linearized model is constructed. With vehicle dynamics constraints and road geometry constraints, and with the tracking error and control quantity minimization as the optimization function of the discrete linearized model, the optimization function is solved to obtain the reference control sequence.
[0069] It should be noted that a vehicle kinematic model, i.e., a discrete linearized model, is first established. The state space and dynamics employ a two-dimensional state model of lateral error and heading error. , The lateral deviation of the main vehicle relative to the centerline of the target lane. The deviation between the main vehicle's heading and the target lane's heading; This represents the two-dimensional state of the vehicle at time t, which is the lateral error minus the heading error; the control variable is the equivalent front wheel steering angle. , This represents the control quantity of the vehicle at time t.
[0070] The expression for the discrete linearized model is shown below:
[0071]
[0072] in, This represents the two-dimensional state of the main vehicle at time t+1, which is the difference between the lateral error and the heading error. Represents the state transition matrix. This represents the two-dimensional state of the main vehicle at time t, which is the difference between the lateral error and the heading error. Represents the control input matrix. This represents the control quantity of the master vehicle at time t.
[0073] Furthermore, the vehicle dynamics constraint is: steering angle limit. Steering speed limit These are the upper and lower limits of the steering angle, respectively; the road geometric constraints are: to avoid excessive lateral deviation of the vehicle, the lateral error and heading error are limited to meet the safety lane boundary. , . in, These are the maximum values of the lateral error and the heading error, respectively. Based on lane width and safety margin settings, it adaptively shrinks to [a certain size] when there is a risk of collision. , .
[0074] Solving the above constrained optimization problem yields an optimal control sequence, which serves as the reference control sequence. , It is a collection of the first The vector of control quantities from time t to time t+n.
[0075] The control inputs and state trajectories corresponding to the reference control sequence are projected onto a preset feasible set to obtain the feasible set projection interval.
[0076] Among them, hard constraints jointly define the feasible set projection interval of actions in the current state. , Included in the reference control sequence The state of the vehicle under the action, the state trajectory can be the state space based on Markov decision modeling, or it can be defined by vehicle dynamics modeling. In addition, the preset feasible set is based on common sense or verified consensus, as well as environmental constraints.
[0077] Furthermore, the expression for the optimization function is as follows:
[0078]
[0079] in, Represents the optimization function. This represents the sum of the predicted state variables from time t to time t+n. This represents the sum of the predicted control quantities from time t to time t+n. and Both are weight matrices. The weighted L2 norm of the state variables. The weighted L2 norm of the control quantity.
[0080] Furthermore, the step of calculating the MPC bootstrapping reward based on the reference control sequence includes:
[0081] Based on the current longitudinal speed and wheelbase of the main vehicle, the first control variable in the reference control sequence is mapped to the MPC expert steering angle at the current moment, in combination with the vehicle's physical characteristics.
[0082] The RL agent strategy outputs the policy steering angle under the same state, and constructs the MPC guided reward based on the MPC expert steering angle and the policy steering angle. The same state is the state of the master vehicle under the control of the first control quantity in the reference control sequence.
[0083] Among them, the introduced state and control hard constraints theoretically give the current state. Projection range of the feasible set of the next action In implementation, these constraints correspond to the clipping of lateral error, heading error, and rate of change of steering, and together with the safety action layer in the environment, define the safe action space of the strategy. Therefore, after obtaining the reference control sequence... Subsequently, this invention does not directly use the entire MPC optimal control sequence to control the vehicle, but instead uses it as an expert reference for the reinforcement learning strategy, i.e., the RL agent strategy. Specifically, firstly, based on the vehicle's current longitudinal speed and wheelbase, the first control variable of the optimal sequence is... The MPC expert steering angle is mapped to the vehicle's physical characteristics at the current moment. Meanwhile, the RL agent's policy remains the same in the same state. Downward output strategy steering angle MPC-guided rewards are constructed by comparing the two.
[0084] It should be noted that when the turning angle output by the reinforcement learning policy is exactly the same as the turning angle of the MPC expert, the MPC guidance reward is close to 1; as the deviation increases, the MPC guidance reward monotonically decays to 0, thereby encouraging the reinforcement learning policy to mimic the safe turning behavior given by the MPC in key confluence scenarios.
[0085] Furthermore, the expression for the MPC-guided reward is as follows:
[0086]
[0087] in, This indicates the MPC introductory reward. Indicates the same state. Represents a set of control actions. Indicates the strategy turning angle, Indicates the MPC expert steering angle;
[0088] The expression for the total reward in reinforcement learning is as follows:
[0089]
[0090] in, Indicates the total reward. Indicates the weight of the MPC bootstrapping reward. Indicates environmental task rewards;
[0091]
[0092] in, To retain the final weight, The preset attenuation period length, This represents the current number of training steps. It gradually decreases over time during training but is increased during ramp merging and high-risk situations to enhance the safety guidance of MPC in the early stages of training and in dangerous scenarios, while preserving the autonomy and flexibility of the strategy in the later stages of training.
[0093] Furthermore, the step of merging the MPC guidance reward with the environment task reward to obtain the single-step reward includes:
[0094] Based on the vehicle's operating status, different parameter indicators and rewards are calculated for each stage of the merging process. The rewards for each parameter indicator are then mixed and processed using stage weights and risk gating to obtain the task reward.
[0095] Security rewards and security costs are generated based on security indicators. The single-step reward is obtained by reconstructing the MPC guidance reward, the task reward, the security reward, and the security cost.
[0096] Specifically, generating an environment task reward based on the task reward and the security reward, and merging the environment task reward and the MPC guidance reward, involves inputting the environment task reward and the MPC guidance reward into the formula for the total reward of reinforcement learning, thereby obtaining the single-step reward.
[0097] Then, the single-step reward is explicitly decomposed into three parts: environmental task reward branch, safety reward and safety cost branch, and MPC guided reward branch. Risk gating and stage-aware weights are used to coordinate the branches, and the next step is carried out based on the coordinated data.
[0098] at discrete time step The single-step reward output by the reward signal optimization module is defined as follows:
[0099] ;
[0100] Simultaneously, the safety cost of a single step is output as the cost signal of the safety constraint algorithm. This single-step safety cost is defined as:
[0101] ;
[0102] in: To reward tasks, the vehicle's performance in terms of progress, efficiency, and comfort is described. For safety rewards, depict safety conditions such as TTC, distance between vehicles, and obstacles; This serves as a reward for MPC (Multi-Level Marketing) and is used to measure the consistency between policy actions and MPC actions. To guide weights, It varies with the merging phase and time; For security reward weighting, ; For safety and cost considerations, it is constructed from safety indicators such as TTC, vehicle distance, and collision, and is used for constrained reinforcement learning.
[0103] Task Rewards This is used to characterize "how well it is being driven," including multiple indicators such as progress, efficiency, comfort, and merging readiness. A risk gating coefficient is used to suppress excessive task-driven behavior under high-risk conditions. The expression for the task's reward branch is shown below:
[0104]
[0105] in These are the upper and lower bounds for task rewards, respectively. To limit the reward value Functions between This is the round step size penalty coefficient. Basic task rewards, As a phased reward, A reward is given for the main vehicle stopping. The merging process is marked as... The stage weight vector is:
[0106]
[0107] in , , , These are the weights of each reward at different merging stages;
[0108] Then at time Basic task rewards The expression is as follows:
[0109]
[0110] The expression for progress rewards is as follows:
[0111]
[0112] in, The slope , This is the amplification factor during the merging phase. tanh is the hyperbolic tangent function. This is a dimensionless deviation. The expression is as follows:
[0113]
[0114] in, It is a very small positive number. For single-step forward increment, , and These represent the forward increments at the current time step and the previous time step, respectively. For the desired incremental progress:
[0115]
[0116] in, This is the proportionality coefficient. , For simulating step size, The desired vehicle speed set for the mission;
[0117] Rewards for efficiency The expression is as follows:
[0118]
[0119] in, The slope coefficient, , This is the inflection point. , For speed ratio, , For longitudinal velocity, For the desired vehicle speed;
[0120] For comfort rewards, The expression is as follows:
[0121]
[0122] in, Lane keeping reward To control smoothness bonuses, These are the weights for lane keeping bonus and control smoothness bonus, respectively.
[0123]
[0124] in, For horizontal scoring, Scoring is given based on the course. To use different weight combinations at different stages;
[0125]
[0126] in, Half the width of the lane. The lateral offset distance between the main vehicle and the lane centerline;
[0127]
[0128] in, The maximum allowable heading deviation, The difference between the main vehicle's heading and the lane tangent angle;
[0129]
[0130] in, To control rewards, In order to control costs, The expression is as follows:
[0131]
[0132] Among them, weight Adjustments are made based on whether the current location is in a sensitive area such as the end of a ramp. This represents the single-step change in the steering angle between adjacent time steps. This represents the single-step change in longitudinal acceleration between adjacent time steps. This refers to the impact amount.
[0133] Let the longitudinal acceleration and steering angle of the current time step and the previous time step be respectively... , ,but:
[0134]
[0135] .
[0136] As a reward for merging preparation, The expression is as follows:
[0137]
[0138] in, , , These are lateral alignment, speed matching, and sufficient gap. The weights for lateral alignment, speed matching, and sufficient gap are respectively. .
[0139] Horizontal alignment The expression is as follows:
[0140]
[0141] in, The upper limit of the allowed lateral offset, This is the absolute value of the lateral offset distance.
[0142] Speed matching The expression is as follows:
[0143]
[0144] in, The target traffic flow speed is obtained by weighting the speeds of vehicles in front of and behind the target lane. To allow for speed deviation, The current speed of the main vehicle.
[0145] Sufficient gap The expression is as follows:
[0146]
[0147] The evaluation process assesses whether lane merging is safe. If lane merging is safe, the clearance adequacy value is assigned as 1; otherwise, it is assigned as 0.
[0148] The risk gating function is calculated by comprehensively considering TTC, distance between vehicles, safe distance, and remaining ramp distance. It then constructs three sub-factors using Sigmoid normalization. ,get:
[0149]
[0150] Where TTC is the time to collision calculated from the relative distance and relative velocity.
[0151] The safety rewards branch explicitly quantifies accident risks and outputs both rewards and costs.
[0152] Security rewards consist of both terminal penalties and continuous shaping criteria.
[0153]
[0154] in: As a penalty and reward system for terminal safety, if the main vehicle is involved in a collision, then... Equal to the reward for collision, <0; if the main vehicle leaves the road boundary, then It's equivalent to the reward for leaving the road. <0; other cases All are 0. The minimum acceptable safety reward. The maximum acceptable reward for safety.
[0155] The continuous safety shaping reward, in the absence of a terminal event, is obtained by weighting the following continuous components:
[0156]
[0157] in:
[0158]
[0159] To construct a smooth penalty reward using TTC values and collision risk metrics, Weights for smooth penalty rewards;
[0160] The obstacle reward is designed in segments based on the distance between the main vehicle and the obstacle at the end of the ramp. The weight of the obstacle reward;
[0161] To construct a dynamic distance bonus based on the ratio of the nearest vehicle distance to the dynamic safe distance, The weight for dynamic vehicle distance reward;
[0162] All are calculated in real time from the state observed from the environment.
[0163] Security Costs Used as a constraint in policy optimization, it takes the form of "linear superposition of multiple risks + truncation":
[0164]
[0165] in,
[0166] The cost is based on TTC; the smaller the TTC, the higher the cost.
[0167] Costs are based on vehicle distance; the smaller the vehicle distance and the closer it is to the safe distance, the higher the cost.
[0168] The value is 1 if a collision occurs, and 0 otherwise.
[0169] Costs associated with prolonged low-speed or short-distance towing;
[0170] The cost of merging failure risk at the end of a ramp is related to the remaining distance of the ramp and the current speed.
[0171] All these quantities can be obtained from the current TTC, vehicle distance, speed, stage marker, and freeze detector output.
[0172] As a weight for the cost of TTC, As a weight of the cost of vehicle distance, The maximum acceptable safety cost. For time steps.
[0173] MPC bootstrap rewards are used to provide a soft constraint on policies to "security experts".
[0174] The current policy turning angle given by the RL agent policy is: The MPC expert steering angle calculated under the same conditions is: ,Right now
[0175] MPC Introductory Rewards for:
[0176]
[0177] in, For the maximum permissible steering angle, The slope constant; guiding weight Determined by both the stage base value and time decay;
[0178]
[0179] in, The basic weights for different merging stages; This is the time decay factor, belonging to (0, 1), ensuring that the dependence on MPC decreases as training progresses; They are respectively The maximum and minimum values that can be taken.
[0180] In summary, at each time step First, based on information such as the vehicle's longitudinal position, speed, lateral offset, control changes, and merging phase, indicators such as progress, efficiency, comfort, and merging readiness are calculated. After a hybrid processing involving phase weighting and risk gating, the task reward is obtained. Safety rewards are then generated using safety metrics such as TTC, distance to other vehicles, distance to obstacles, and collision markers. With security costs Finally, the MPC-guided reward is constructed using the MPC expert steering angle and the policy steering angle given by the RL agent policy. And set guidance weights according to stages and time. The final single-step reward is output through the overall formula:
[0181] .
[0182] Furthermore, the step of optimizing and updating the initial multi-commenter policy network based on the target task reward and the target security cost to obtain the optimized action policy network includes:
[0183] The target security cost and the target task reward are collected synchronously over a fixed time span, and the target task reward and the target security cost are written into the task buffer and the security buffer, respectively.
[0184] Among them, in the first Step, environment output current observation Features are obtained through the shared representation layer. The following was obtained through forward computation using a multi-commenter strategy network:
[0185]
[0186] in, This represents the logarithmic probability of the current action under the old policy (the previous version of the Actor policy network's parameters), used for subsequent PPO-CLIP probability ratio calculations. Current state The conditional probability below should be noted; it should be pointed out that the old strategy generates actions. The neural network model (parameters). The old strategy part refers to the Actor network version used to collect this data before the parameters were updated.
[0187] Then the action Input the ramp merging environment; the environment updates its state based on the dynamic model and the behavior of other vehicles; return to the next time-instance observation. And the following content designed in the preceding steps:
[0188] Task Rewards Calculations based on TTC, minimum distance between vehicles, obstacle distance, collision events, etc.
[0189] Safety Rewards Calculations based on TTC, minimum distance between vehicles, obstacle distance, collision events, etc.
[0190] Security Costs The result is obtained by normalizing factors such as TTC over-threshold, insufficient distance, collision, freeze, and approaching the end of the ramp, and is used for safety commentator training and constraint optimization.
[0191] Risk-related auxiliary information (aux): including risk gate, TTC, front and rear vehicle spacing, merging window quality, speed matching bonus, etc.
[0192] Regarding the task optimization section, the data in this step will be divided according to... The data is written to the task buffer in the form of , where The termination marker (collision, out of bounds, successful completion, etc.). The risk and event characteristics returned by the environment are used for subsequent scenario-based weighted or statistical analysis.
[0193] For the security optimization portion, the target security cost and security commentator estimates are written into a separate security buffer:
[0194]
[0195] in, For the safety costs of this step, This is an estimate by security commentators. The values are estimated by the task commentator. To facilitate subsequent standardized processing, the safety buffer has a length of [value missing]. Stored as a flattened array of (time steps × number of environments).
[0196] When the task buffer length reaches the preset time span Afterwards, it is considered that one round of multi-commenter sampling has been completed, and the process moves on to the next step.
[0197] The target task reward and the target security cost are estimated by time-series difference estimation using generalized advantage estimation to obtain different advantage data. Each advantage data is then normalized to obtain corresponding normalized data.
[0198] After completing one sampling round, this step uses the Generalized Advantage Estimation (GAE) method to perform temporal difference estimation on the target task reward and the target security cost, respectively, to obtain two paths of advantage and reward, namely task advantage and task report, and security cost advantage and security cost report, which provide a numerical basis for subsequent Lagrange optimization.
[0199] GAE calculation for task commentators: For the task portion, a discount factor is applied. With GAE parameters For the sequence Perform standard GAE calculations to obtain task advantages. With task rewards :
[0200]
[0201]
[0202] The above calculations were performed jointly by the task buffer and the task commentator network. For the task commentator's estimate at step t+1, For the task commentator's estimate at step t, Value function for task commentators The one-step TD error (Temporal Difference Error).
[0203] Security commentator's Safe-GAE calculation: for the security cost sequence Introducing a security-specific discount factor With safety GAE parameters Construct a secure time-series difference:
[0204]
[0205] in, Value function of safety cost One step of TD error, For the safety critic's estimate at step t+1, For the security critic's estimate at step t, and back-substitute to obtain the security cost advantage:
[0206]
[0207] Safety cost-return:
[0208]
[0209] in, Slightly smaller than the task discount factor to reflect the rapid feedback characteristic of security costs. Used to weigh bias against variance.
[0210] Advantage normalization: To improve numerical stability and training efficiency, the task advantage and safety cost advantage are normalized to zero mean and unit variance, respectively. An appropriate temperature coefficient and pruning threshold are added to the safety channel to suppress the dominance of extreme samples on the gradient.
[0211]
[0212]
[0213] in, The safety advantage temperature coefficient is less than 1. The cropping threshold, This represents the average of the task advantages. The standard deviation of the task advantage. As the average of safety cost advantages, The standard deviation of the safety cost advantage.
[0214] Based on the near-end policy optimization training framework, the normalized data is weighted and fused using Lagrange multipliers to obtain a hybrid advantage. Based on the hybrid advantage, the policy head parameters, the task critic parameters, and the safety critic parameters are jointly updated through backpropagation and gradient pruning to obtain an optimized action policy network.
[0215] In the single PPO-CLIP (Proximal Policy Optimization) training framework, the task advantage and the security cost advantage are normalized separately, and then the two are weighted and fused based on Lagrange multipliers to obtain a hybrid advantage. Then, the hybrid advantage is used as an update signal to synchronously update the policy head parameters of the policy network and the network parameters of the task commentator and the security commentator, thereby obtaining the optimized action policy.
[0216] Probability ratio and mixed advantage construction: In each training mini-batch, the current log probability is recalculated using the policy network. Log probabilities of the old strategy during sampling Construct probability ratios:
[0217]
[0218] Advantages of the mission Safety and cost advantages via Lagrange multipliers The combination yields mixed advantages:
[0219]
[0220] in, For probability ratios, For the sake of mixed advantages, To amplify the impact based on recent collision rates, when the collision frequency exceeds a threshold, the weight of safety items is increased in the short term to strengthen safety constraints.
[0221] PPO-CLIP policy loss construction: By introducing the mixed advantage into the PPO-CLIP objective function, the policy head loss is obtained:
[0222]
[0223] in, For the strategy head loss, This is the shear range hyperparameter. The mathematical expectation operator is used to perform empirical averaging on the sample data at each time step in the current sampling batch.
[0224] The aforementioned loss is mathematically equivalent to maximizing the expectation of "task advantage minus weighted safety cost advantage" in the policy gradient, thereby naturally achieving a soft constraint on safety costs.
[0225] Entropy regularization term construction: To maintain policy exploration capability, a policy entropy regularization term is introduced. :
[0226]
[0227] in, The entropy coefficient, Let be the action distribution entropy. This prevents the strategy from prematurely converging to a local optimum, and is especially helpful in exploring more safe and reasonable traffic behaviors in complex traffic scenarios.
[0228] Joint loss and adaptive weights for task and security value functions: Constructing temporal difference residuals for task reward and security cost reward respectively:
[0229]
[0230]
[0231] The temporal difference residual of the task reward. The time-series differential residuals are used as a safety cost incentive, and the exponential moving variance of the two residuals is estimated. Adaptive allocation of value loss weights:
[0232]
[0233]
[0234] After normalization, guarantee , Weighting the value loss of task rewards. The value loss is weighted according to the security cost reward, and a minimum weight lower bound is set for each branch to prevent any value head from being ignored. The final result is the joint value loss. :
[0235]
[0236] This design ensures a dynamic balance in the convergence speed of the two value functions—task and safety—during the training process.
[0237] Total Loss and Parameter Update: The multi-commenter optimization module uses a weighted sum of policy loss, entropy regularization, and joint value loss to form the total objective. The total loss function is shown below:
[0238]
[0239] The total loss function is used to train the multi-commentator policy network, and the policy head parameters are adjusted through backpropagation and gradient pruning. Task commentator parameters Security commentator parameters Synchronous updates enable optimization of the multi-commentator strategy network.
[0240] The Lagrange multiplier adaptive update and safety constraint tracking step aims to ensure that the agent automatically meets a preset safety cost ceiling during long-term training. This step adjusts the Lagrange multipliers based on cost-return statistics output by safety critics. Adaptive updates are performed to enable online adjustment of the average security cost.
[0241] Batch average security cost assessment: Statistical analysis of security cost return in each training batch. The average value is used to obtain the current estimated average security cost. The cost estimate is then smoothed using an exponential moving average to obtain an EMA (exponential moving average) cost estimate. This reduces the impact of fluctuations in individual batches.
[0242] Freezing perception target cost correction: Utilizing the "freezing rate" (the proportion of agents whose speed is too low and progress stagnates due to excessive conservatism) fed back by the environment at the end of the round, if the observed freezing rate is higher than the expected threshold, then the original target safety cost is adjusted accordingly. By adding a certain amount of slack to the base, the effective target cost can be obtained. :
[0243]
[0244] in, This is the freeze correction coefficient. The vehicle freeze rate for the entire round at the current moment. The target is the overall vehicle freeze rate for the entire round. This design automatically relaxes safety constraints when collisions are reduced but the vehicle is overly conservative, preventing the strategy from being stuck in a low-speed "self-preservation" state for an extended period.
[0245] Lagrange multiplier update rule: The Lagrange multipliers are updated using an approximate gradient ascent method.
[0246] in, To update the obtained Lagrange multipliers, The maximum allowable value for the Lagrange multiplier. The values of the Lagrange multipliers before the update. For learning rate, To cut off the multiplier at Projection operator on.
[0247] When the average security cost is higher than the target value Increase this, thereby strengthening the penalty for safety cost advantages;
[0248] When the average security cost is low and the freeze rate is not high Gradually reduce the amount of data to restore the efficiency of the strategy while ensuring safety.
[0249] In summary, by unifying the strategy head parameters, task commentator parameters, and safety commentator parameters into the same optimization framework, and through multiple rounds of "interactive sampling - advantage estimation - Lagrange weighting - parameter update", the system gradually converges to a vehicle control strategy that balances efficiency and safety, thus serving as the optimized action strategy network.
[0250] It should be noted that, through iterative execution, the multi-commentator optimization module utilizes information from both task commentators and safety commentators during the strategy evaluation and strategy improvement phases. Within a single PPO-CLIP framework, it achieves joint optimization of task rewards and safety costs, ultimately resulting in an automatic merging control strategy that can balance traffic efficiency and safety constraints in ramp scenarios.
[0251] The method further includes:
[0252] In the policy network structure for near-end policy optimization, two value branches, task critic and security critic, are introduced to form a multi-critic policy network. The task critic network is a task value function built on shared features. The policy network structure includes a policy head, which is a Gaussian random policy in a continuous action space built on shared features.
[0253] Among them, the construction of the multi-commentator strategy network:
[0254] In a unified policy network structure, two value branches, task commentator and security commentator, are explicitly introduced to form a multi-commentator policy network. Specifically, this includes:
[0255]
[0256] in, It is a multilayer perceptron. This is the raw observation vector acquired at the current moment, containing state information of the main vehicle and surrounding neighboring vehicles. Its parameters are specified. This presentation layer is shared between the policy header and the two commentator headers, ensuring that all three are optimized within the same state semantic space.
[0257] Strategy head (Actor) construction, in shared features Construct a Gaussian random policy in a continuous action space:
[0258]
[0259] in, Represents a multidimensional Gaussian distribution. This is the mean vector of the action (such as steering angle, longitudinal acceleration). It is a diagonal covariance matrix. These are the policy parameters. The sampled actions. As the control output of the agent in the current state. In the current state Below, the parameters are The strategy generates actions The conditional probability distribution.
[0260] Task commentator network built on shared features Construct the task value function above This is used to estimate the cumulative task reward with discounts related to driving efficiency, merging completion, and comfort. The network output is a scalar used to calculate the task advantage and guide the policy to update in a direction that improves task performance.
[0261]
[0262] in, For task commentator network parameters.
[0263] Security commentator network construction, introducing a security commentator network parallel to the task commentator network. This is used to estimate the cumulative safety cost due to factors such as TTC, insufficient following distance, collision, and freezing. The network output is a scalar, but it represents the "expected future safety cost," providing a foundation for subsequent calculations of Lagrangian constraints and safety advantages.
[0264]
[0265] in, Parameters for safety commentators.
[0266] Through the construction of the aforementioned task commentator network and security commentator network, a multi-commentator policy network consisting of "shared representation + policy head + task value head + security value head" is formed under a unified structure, laying the foundation for subsequent joint optimization.
[0267] Furthermore, the method also includes:
[0268] The environmental information and vehicle operating status information are classified and integrated according to the Markov decision process to form a decision process modeling result that includes state space, action space, state transition relationship and immediate reward.
[0269] The data required for a Markov decision process include the state space, action space, and reward function.
[0270] Specifically, a state space is constructed based on the relative motion characteristics of the main vehicle's operating state and surrounding vehicles, and its form is as follows:
[0271]
[0272] in, For state space, The main characteristics of the vehicle's motion state This refers to the movement characteristics of surrounding vehicles. Considering the surrounding environment, the action space is constructed based on the vehicle's main acceleration and the front wheel steering angle, as follows:
[0273]
[0274] in, For the action space, For acceleration, This refers to the steering angle of the front wheels.
[0275] The instant reward function is based on security rewards. It consists of forward reward, efficiency reward, comfort reward, survival reward, freeze reward and MPC guidance reward.
[0276] It should be noted that the Markov decision process runs continuously throughout the merging process to provide data support for each step.
[0277] Example 2
[0278] like Figure 2 The diagram illustrates an exemplary scenario of a complex highway road simulation. The invention describes a reinforcement learning-based autonomous driving ramp merging method guided by MPC, which is defined as the PPO-MC-MPC method. The controlled autonomous vehicle merges from the ramp onto the main road. There are 8 surrounding vehicles on the main road, each of which is randomly generated with a speed range between 10 m / s and 15 m / s.
[0279] in, Figures 3 to 5 This paper describes a comparison of experimental results for the SAC reinforcement learning method, the PPO reinforcement learning method, and the PPO-MC-MPC method. The experimental results show that the autonomous driving ramp merging method based on physical information-driven multi-commentator proximal policy optimization reinforcement learning outperforms the SAC and PPO reinforcement learning methods in terms of success rate, average speed, and average turning angle. Figure 3 This chart compares the training success rates of the SAC, PPO, and PPO-MC-MPC methods under the same training environment. It shows that the PPO-MC-MPC method has already begun to converge to its maximum success rate by the 1000th iteration, while the SAC and PPO methods have not yet begun to converge. Furthermore, the final success rate of the PPO-MC-MPC method is higher than that of the SAC and PPO methods. This demonstrates that the PPO-MC-MPC method has higher safety and training efficiency than the SAC and PPO methods. Figure 4 The chart compares the average training speeds of the SAC, PPO, and PPO-MC-MPC methods under the same training environment. It can be seen that the average speed of the PPO-MC-MPC method is consistently slightly higher than that of the SAC and PPO methods. This proves that the merging efficiency of the PPO-MC-MPC method is higher than that of the SAC and PPO methods, and the average speed of the PPO-MC-MPC method fluctuates much less and is more stable. Figure 5 The graph compares the average rotation angles of the SAC, PPO, and PPO-MC-MPC methods under the same training environment. It can be seen that the average rotation angle of the PPO-MC-MPC method is lower than that of the SAC and PPO methods overall, proving that the control stability of the PPO-MC-MPC method is better than that of the SAC and PPO methods.
[0280] In summary, the present invention has the following beneficial effects:
[0281] By employing a dual-channel feedback and multi-commenter architecture, scale conflicts and gradient interference between task and safety signals are eliminated, thereby improving learning stability and interpretability.
[0282] MPC guidance and feasible set projection suppress high-risk exploration in early training, improving sample efficiency and physical security;
[0283] To achieve "explicitly adjustable" safety boundaries, the Lagrange dual variable provides an adjustable safety budget "knob," enabling an explicit balance between safety and efficiency.
[0284] In the extended Highway-Env benchmark test, the method of this invention outperformed the baseline method in terms of success rate, merging efficiency, and control stability.
[0285] Example 3
[0286] Please see Figure 6 The diagram below shows the structural block diagram of the MPC-guided reinforcement learning-based autonomous driving ramp merging system according to the second embodiment of the present invention. The system includes:
[0287] The generation module 10 is used to collect environmental information and vehicle operation status information of ramp merging, and generate MPC guidance reward and feasible set projection interval based on the environmental information and vehicle operation status information.
[0288] The fusion and decomposition module 20 is used to fuse the MPC guidance reward and the environmental task reward based on the current driving stage to obtain a single-step reward, and to optimize and decompose the reward signal corresponding to the single-step reward into a target task reward and a target safety cost.
[0289] The decision module 30 is used to optimize and update the initial multi-commenter policy network based on the target task reward and the target safety cost to obtain an optimized action policy network. Based on the current state of the master vehicle, the module generates original actions through the optimized action policy network and maps the original actions to the feasible set projection interval to determine the final executable safety action.
[0290] In practical implementation, environmental information and vehicle operating status information of ramp merging are collected, and MPC guidance rewards and feasible set projection intervals are generated based on the environmental information and vehicle operating status information. Then, based on the current driving stage, the MPC guidance rewards and environmental task rewards are fused to obtain a single-step reward. The reward signal corresponding to the single-step reward is optimized and decomposed into target task reward and target safety cost. Then, the initial multi-critic policy network is optimized and updated based on the target task reward and target safety cost to obtain an optimized action policy network. Based on the current state of the master vehicle, the original action is generated through the optimized action policy network and mapped to the feasible set projection interval to determine the final executable safety action. Unlike existing technologies, this method improves the model prediction accuracy and reduces the sensitivity to solution delay. It also improves the problems of being susceptible to distribution shift and lacking explicit efficiency drivers. Furthermore, it improves the problem of gradient interference and credit allocation confusion caused by mixing task and safety signals in a single channel reward.
[0291] Furthermore, the generation module 10 includes:
[0292] The calculation unit is used to generate a reference control sequence for the master vehicle within a preset prediction time domain based on the environmental information and the vehicle operating status information, and to calculate the MPC guidance reward based on the reference control sequence.
[0293] The projection unit is used to project the control input and state trajectory corresponding to the reference control sequence onto a preset feasible set to obtain the feasible set projection interval.
[0294] Furthermore, the environmental information includes the lateral deviation and flight path deviation of the main vehicle, and the calculation unit includes:
[0295] A sub-unit is constructed and solved to build a corresponding discrete linearized model based on the lateral deviation and the flight path deviation, with vehicle dynamics constraints and road geometry constraints, and to construct the optimization function of the discrete linearized model by minimizing the tracking error and control quantity, and to solve the optimization function to obtain the reference control sequence;
[0296] The expression for the discrete linearized model is shown below:
[0297]
[0298] in, This represents the two-dimensional state of the main vehicle at time t+1, which is the difference between the lateral error and the heading error. Represents the state transition matrix. This represents the two-dimensional state of the main vehicle at time t, which is the difference between the lateral error and the heading error. Represents the control input matrix. This represents the control quantity of the master vehicle at time t;
[0299] in:
[0300]
[0301] ,
[0302] The expression for the optimization function is as follows:
[0303]
[0304] in, Represents the optimization function. This represents the sum of the predicted state variables from time t to time t+n. This represents the sum of the predicted control quantities from time t to time t+n. and Both are weight matrices. Indicates a minute amount of time. The weighted L2 norm of the state variables. The weighted L2 norm of the control quantity.
[0305] Furthermore, the computing unit includes:
[0306] The mapping subunit is used to map the first control quantity in the reference control sequence into the MPC expert steering angle at the current moment, based on the current longitudinal speed and wheelbase of the master vehicle and the vehicle's physical characteristics.
[0307] A subunit is constructed, which outputs a policy steering angle under the same state through the RL agent policy. An MPC-guided reward is constructed based on the MPC expert steering angle and the policy steering angle, wherein the same state is the state of the master vehicle under the control of the first control quantity in the reference control sequence.
[0308] Furthermore, the expression for the MPC-guided reward is as follows:
[0309]
[0310] in, This indicates the MPC introductory reward. Indicates the same state. Represents a set of control actions. Indicates the strategy turning angle, Indicates the MPC expert steering angle;
[0311] The expression for the single-step reward is as follows:
[0312]
[0313] in, Indicates a single-step reward. Indicates the weight of the MPC bootstrapping reward. This indicates the task reward. As a safety reward, For security reward weighting, To guide weights, This is a reward for MPC (Multi-Level Marketing) guidance.
[0314] Furthermore, the step of merging the MPC guidance reward with the environment task reward to obtain the single-step reward includes:
[0315] Based on the vehicle operation status information, different parameter indicators and rewards are calculated for each stage of the merging process. The rewards for each parameter indicator are then mixed and processed by stage weighting and risk gating to obtain the task reward.
[0316] Security rewards and security costs are generated based on security indicators. The single-step reward is obtained by reconstructing the MPC guidance reward, the task reward, the security reward, and the security cost.
[0317] Furthermore, the step of optimizing and updating the initial multi-commenter policy network based on the target task reward and the target security cost to obtain the optimized action policy network includes:
[0318] The target security cost and the target task reward are collected synchronously over a fixed time span, and the target task reward and the target security cost are written into the task buffer and the security buffer, respectively.
[0319] The target task reward and the target security cost are estimated by time-series difference estimation using generalized advantage estimation. The advantages and rewards of the two paths are normalized to obtain the corresponding normalized data.
[0320] Based on the near-end policy optimization training framework, the normalized data is weighted and fused using Lagrange multipliers to obtain a hybrid advantage. Based on the hybrid advantage, the policy head parameters, task commentator parameters, and safety commentator parameters are jointly updated through backpropagation and gradient pruning to obtain the optimized action policy network.
[0321] Furthermore, the method also includes:
[0322] In the policy network structure for near-end policy optimization, task critics and security critics are introduced to form an initial multi-critic policy network. The policy network structure includes a policy head, which is a Gaussian random policy in a continuous action space built on shared features.
[0323] Example 4
[0324] In the third embodiment of the present invention, based on the same inventive concept, the present invention proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the reinforcement learning-based MPC-guided autonomous driving ramp merging method of the above embodiments.
[0325] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means containing storage, communication, propagation, or transmission programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0326] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0327] The memory may include a large-capacity storage device for data or instructions. For example, and not limitingly, the memory may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory may include removable or non-removable (or fixed) media. Where appropriate, the memory may be internal or external to the data processing device. In a particular embodiment, the memory is non-volatile memory. In a particular embodiment, the memory includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), an electrically alterable read-only memory (EAROM), or flash memory, or a combination of two or more of these. Where appropriate, the RAM can be Static Random-Access Memory (SRAM) or Dynamic Random-Access Memory (DRAM). DRAM can be Fast Page Mode Dynamic Random-Access Memory (FPMDRAM), Extended Data Out Dynamic Random-Access Memory (EDODRAM), Synchronous Dynamic Random-Access Memory (SDRAM), etc.
[0328] Example 5
[0329] In the fourth embodiment of the present invention, based on the same inventive concept, the present invention proposes a terminal, the terminal comprising: a processor and a memory; the processor and the memory communicate with each other; the memory is used to store instructions; the processor is used to execute the instructions in the memory to execute the reinforcement learning-based autonomous driving ramp merging method based on MPC guidance of the above embodiment.
[0330] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0331] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0332] Without causing conflict, those skilled in the art can freely combine and use the above-mentioned additional technical features.
[0333] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A reinforcement learning-guided autonomous driving ramp merging method based on MPC guidance, characterized in that, The method includes: Collect environmental information and vehicle operation status information of ramp merging, and generate MPC guidance reward and feasible set projection interval based on the environmental information and vehicle operation status information; Based on the current driving stage, the MPC guidance reward and the environmental task reward are integrated to obtain the single-step reward. The reward signal corresponding to the single-step reward is optimized and decomposed into the target task reward and the target safety cost. The initial multi-commentator policy network is optimized and updated based on the target task reward and the target safety cost to obtain an optimized action policy network. Based on the current state of the master vehicle, the original action is generated through the optimized action policy network, and the original action is mapped to the feasible set projection interval to determine the final executable safety action. The initial multi-commentator policy network is constructed by explicitly introducing task commentators and safety commentators into the policy network structure of the near-end policy optimization. The initial multi-commentator policy network includes policy head parameters, task commentator parameters, and safety commentator parameters.
2. The MPC-guided reinforcement learning-based autonomous driving ramp merging method according to claim 1, characterized in that, The step of generating the MPC guidance reward and feasible set projection interval based on the environmental information and the vehicle operating status information includes: Based on the environmental information and the vehicle operating status information, a reference control sequence for the master vehicle is generated within a preset prediction time domain, and the MPC guidance reward is calculated based on the reference control sequence. The control inputs and state trajectories corresponding to the reference control sequence are projected onto a preset feasible set to obtain the feasible set projection interval.
3. The MPC-guided reinforcement learning-based autonomous driving ramp merging method according to claim 2, characterized in that, The environmental information includes the lateral deviation and flight path deviation of the master vehicle, and the step of generating the reference control sequence of the master vehicle within the preset prediction time domain includes: Based on the lateral deviation and the flight path deviation, a corresponding discrete linearized model is constructed. With vehicle dynamics constraints and road geometry constraints, and with the minimization of tracking error and control quantity, the optimization function of the discrete linearized model is constructed. The optimization function is solved to obtain the reference control sequence. The expression for the discrete linearized model is shown below: in, This represents the two-dimensional state of the main vehicle at time t+1, which is the difference between the lateral error and the heading error. Represents the state transition matrix. This represents the two-dimensional state of the main vehicle at time t, which is the difference between the lateral error and the heading error. Represents the control input matrix. This represents the control quantity of the master vehicle at time t; in: , The expression for the optimization function is as follows: in, Represents the optimization function. This represents the sum of the predicted state variables from time t to time t+n. This represents the sum of the predicted control quantities from time t to time t+n. and Both are weight matrices. Indicates a minute amount of time. The weighted L2 norm of the state variables. The weighted L2 norm of the control quantity.
4. The MPC-guided reinforcement learning-based autonomous driving ramp merging method according to claim 2, characterized in that, The step of calculating the MPC bootstrap reward based on the reference control sequence includes: Based on the current longitudinal speed and wheelbase of the main vehicle, the first control variable in the reference control sequence is mapped to the MPC expert steering angle at the current moment, in combination with the vehicle's physical characteristics. The RL agent strategy outputs the policy steering angle under the same state, and constructs the MPC guided reward based on the MPC expert steering angle and the policy steering angle. The same state is the state of the master vehicle under the control of the first control quantity in the reference control sequence.
5. The MPC-guided reinforcement learning-based autonomous driving ramp merging method according to claim 1, characterized in that, The expression for the MPC-guided reward is as follows: in, This indicates the MPC introductory reward. Indicates the same state. Represents a set of control actions. Indicates the strategy turning angle, Indicates the MPC expert steering angle; The expression for the single-step reward is as follows: in, Indicates a single-step reward. Indicates the weight of the MPC bootstrapping reward. This indicates the task reward. As a safety reward, For security reward weighting, To guide the weights.
6. The MPC-guided reinforcement learning-based autonomous driving ramp merging method according to claim 1, characterized in that, The step of merging the MPC guidance reward with the environment task reward to obtain the single-step reward includes: Based on the vehicle operation status information, different parameter indicators and rewards are calculated for each stage of the merging process. The rewards for each parameter indicator are then mixed and processed by stage weighting and risk gating to obtain the task reward. Security rewards and security costs are generated based on security indicators. The single-step reward is obtained by reconstructing the MPC guidance reward, the task reward, the security reward, and the security cost.
7. The MPC-guided reinforcement learning-based autonomous driving ramp merging method according to claim 1, characterized in that, The step of optimizing and updating the initial multi-commenter policy network based on the target task reward and the target security cost to obtain the optimized action policy network includes: The target security cost and the target task reward are collected synchronously over a fixed time span, and the target task reward and the target security cost are written into the task buffer and the security buffer, respectively. By performing time-series difference estimation on the target task reward and the target security cost respectively through generalized advantage estimation, different advantage data are obtained. Each advantage data is then normalized to obtain corresponding normalized data. Based on the near-end policy optimization training framework, the normalized data is weighted and fused using Lagrange multipliers to obtain a hybrid advantage. Based on the hybrid advantage, the policy head parameters, the task critic parameters, and the safety critic parameters are jointly updated through backpropagation and gradient pruning to obtain an optimized action policy network.
8. The MPC-guided reinforcement learning-based autonomous driving ramp merging method according to claim 1, characterized in that, The policy network structure includes a policy head, which is a Gaussian random policy in a continuous action space built on shared features.
Citation Information
Patent Citations
Vehicle ramp entrance confluence control method based on deep reinforcement learning
CN116215532A
Ramp driving decision control method, vehicle and storage medium
CN117261895A