Cooperative control method for mixed vehicles in spatial domain based on deep reinforcement learning for ring mountain road scene
By constructing a spatial domain deep reinforcement learning collaborative control method in the scenario of mountain roads, the problems of unstable longitudinal control and insufficient lateral tracking accuracy of convoys in mountain roads are solved, realizing high-precision path tracking and energy consumption optimization of convoys, and improving safety and comfort.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-02-12
- Publication Date
- 2026-06-02
AI Technical Summary
In mountainous road scenarios, existing vehicle queuing control methods are unable to effectively address the problems of longitudinal control instability, insufficient lateral tracking accuracy, and insufficient adhesion margin caused by slope changes. Especially in slope-curve composite conditions, speed disturbances are easily amplified, leading to a decrease in safety and comfort.
A spatial domain deep reinforcement learning collaborative control method is adopted. By constructing a vehicle lateral and longitudinal coupled dynamic model, introducing arc length variables and time inverse functions, a spatial domain lateral and longitudinal coupled dynamic model is constructed. Combined with PPO controller and multi-constraint fusion compensation, the lateral and longitudinal joint control action is output to achieve unified control of longitudinal stability and lateral accuracy.
It significantly improves the longitudinal stability and ride comfort of the fleet, achieves high-precision lateral and longitudinal coupling tracking, eliminates the cornering phenomenon in complex curves, and optimizes energy consumption, reducing fuel consumption and emissions.
Smart Images

Figure CN122135578A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent connected transportation technology, and relates to a deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountainous road scenario. Background Technology
[0002] The development of vehicle-to-everything (V2X) and autonomous driving has driven the application of vehicle platooning cooperative control technology. Through vehicle-to-vehicle / vehicle-to-infrastructure communication and onboard sensors, vehicles can acquire status information such as speed, acceleration, and heading angle of vehicles ahead, thereby maintaining a safe following distance and suppressing speed disturbance propagation longitudinally, and achieving trajectory following and cooperative maneuvering laterally. Currently, road traffic is generally in a transitional period of "mixed traffic of human-driven vehicles (HDVs) and intelligent connected vehicles (CAVs)," and the randomness and uncertainty of HDV driving behavior has become one of the main constraints on the effectiveness and safety of platooning control.
[0003] On mountain roads, vehicle longitudinal dynamics are significantly affected by the gravity component of the slope. Insufficient uphill power and frequent downhill braking cause stronger speed fluctuations. At the same time, slope-curve coupling causes lateral and longitudinal dynamics to interact, making the adhesion margin and stability constraints more stringent and increasing the complexity and uncertainty of platoon control. Therefore, integrated lateral and longitudinal collaborative control of mixed-traffic platoons for mountain road scenarios still faces significant technical challenges.
[0004] Existing control technologies can be broadly categorized into three types: First, model-based longitudinal following / cooperative control methods, including ACC / CACC, robust control, adaptive control, and model predictive control. These methods typically rely on accurate vehicle dynamics and disturbance modeling, and make simplified assumptions about slope, adhesion conditions, and communication delays during the design phase. When facing long longitudinal slopes and slope-curve composite road sections, the constraints of slope changes and the vehicle's available longitudinal traction / braking capacity can easily lead to a decline in control performance. Second, trajectory-following-based lateral control methods, including pure tracking, Stanley, and LQR / MPC. Most of these methods are decoupled from longitudinal control, making it difficult to simultaneously guarantee lateral tracking accuracy, comfort, and adhesion margin under slope-curve composite conditions. Third, deep reinforcement learning control methods, which can interactively learn control strategies in complex environments. However, if modeling and training are performed directly in the time domain, it is often necessary to handle non-stationarity caused by communication / execution delays, and there is insufficient expression of constraints on curve geometric consistency and queue alignment.
[0005] The aforementioned existing technologies suffer from the following technical problems: In steep slope environments, significant gradient changes occur. If the gradient is unavailable or has a large error, longitudinal control is prone to issues such as speed loss on uphill sections, speeding on downhill sections, and frequent braking. Speed disturbances are amplified within the platoon, leading to brake thermal load risks. Optimizing longitudinal following or lateral tracking separately ignores the effects of lateral-longitudinal coupling, making it difficult to simultaneously address corner suppression, adhesion margin, and longitudinal stability when curves and slopes overlap. HDV random driving behavior couples with gradient disturbances, causing speed fluctuations to propagate and amplify backward, resulting in decreased comfort and increased safety risks for rear vehicles. In complex mountainous conditions, constraints such as predictive safety (e.g., TTC), adhesion margin, and brake thermal load need to be considered simultaneously; existing methods often struggle to form a unified and feasible control framework. Summary of the Invention
[0006] In view of this, the purpose of this invention is to provide a deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A spatial domain deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario includes the following steps: S1: Establish the scenario and coordinate system of mixed traffic vehicles on the ring road; construct the temporal and longitudinal coupled dynamics of vehicles and introduce equivalent control inputs, introduce arc length variables and define time inverse functions, and construct a spatial domain temporal and longitudinal coupled dynamics model; S2: Calculate the slope and its confidence level; predict the HDV trajectory of the navigator vehicle; construct the spatial domain delay cooperative error; S3: Construct a spatial domain PPO controller, define a comprehensive reward function, and output the combined horizontal and vertical control actions; S4: Perform multi-constraint fusion compensation and motion saturation projection execution.
[0008] Furthermore, step S1, which involves establishing a mixed-traffic vehicle scenario and coordinate system for the mountain road, includes: Constructing a mixed-traffic vehicle fleet: (1) in, This refers to the number of connected autonomous vehicles; vehicle number 0 indicates the navigator is driving, and the vehicle number... Indicates the first Connected autonomous vehicles; Establish arc length coordinates along the road centerline. And define spatially discrete nodes: (2) in, The initial arc length, For spatial step size, For spatial indexing, Set the total number of steps; define the road curvature function. , indicating the position of the arc length The geometric curvature of the road at that location. Extracted from high-precision maps, vehicle-road cooperative infrastructure, or vehicle-mounted perception; Set the time domain sampling time and perform time sampling; (3) in, For the sensing / control sampling period, For time indexing.
[0009] Furthermore, step S1, which involves constructing the vehicle's temporal longitudinal and transverse coupled dynamics and introducing equivalent control inputs, includes: For any vehicle Define the time-domain state in the vehicle coordinate system, including longitudinal velocity. lateral velocity yaw rate and heading angle ; Introduce equivalent control inputs, including longitudinal equivalent inputs. Traction and braking forces can be synthesized and normalized; lateral / steering equivalent inputs Due to the steering angle Obtained by mapping the lateral force channel of the tire; The temporal and longitudinal coupled dynamics of the vehicle can be written as: (4) (5) (6) in, For vehicle quality, Let be the yaw moment of inertia of the vehicle about its center of mass. This is the distance from the center of gravity to the front axle. The resultant force of longitudinal resistance includes rolling resistance, air resistance, and transmission loss; This is the resultant force of the lateral forces on the tires; This refers to additional yaw terms or unmodeled terms caused by lateral forces; unmodeled portions are incorporated into noise or disturbance terms.
[0010] Furthermore, the introduction of the arc length variable and the definition of the inverse time function in step S1 includes: Define vehicle Arc length progress And define the velocity modulus: (7) in For vehicle speed modulus; definition Indicates vehicle Reaching the arc length position The time, and satisfying: (8).
[0011] Furthermore, step S1, which involves constructing a spatial domain horizontal and vertical coupled dynamic model, includes the following steps: any differentiable variable The derivative in the spatial domain satisfies: (9) Define the spatial domain state vector: (10) in For arrival time, For heading angle, , For lateral and longitudinal velocities, This refers to the yaw rate; The spatial domain dynamics model is obtained from equations (4) to (9): (11) (12) (13) (14) (15) in , The vehicle is located at the arc length position. The equivalent longitudinal and lateral control inputs.
[0012] Furthermore, the calculation of the slope and its confidence level in step S2 specifically includes: Define the longitudinal slope angle of the road Its spatial domain form is denoted as The goal is to sample at each time step. Obtain slope estimate And covariance, and map to spatial nodes get ; Define the longitudinal resistance model: (16) in For vehicle speed, , , For calibration parameters, Units are ; At any moment Acquisition of IMU longitudinal acceleration measurement Vehicle speed measurement Traction estimation With braking force estimation Vehicle weight is recorded as Gravitational acceleration ; Define the filter state: (17) in The slope angle, For IMU longitudinal acceleration bias; The system equations are constructed using a random walk model: (18) in , , The standard deviation of process noise; Establish dynamic consistency measurement: (19) in To measure noise, To measure the standard deviation of noise; Definition of measurement Measurement function: (20) Jacobi's construction is as follows: (twenty one) Measurement noise covariance ; Slope prediction using EKF recursion: (twenty two) The updated formula is as follows: (twenty three) (twenty four) (25) The output is: (26) Define cumulative arc length: (27) For spatial nodes Take the satisfied index Obtained using linear interpolation: (28) Confidence level is output using slope variance: (29) in The confidence level is used to adjust the slope feedforward and safety constraint weights, with the variance constant serving as the reference.
[0013] Furthermore, step S2, which involves predicting the trajectory of the HDV (High-Density Vehicle), includes: In time The set of vehicles selected for observation ahead is: (30) in The number of vehicles used for prediction, where 0 represents HDV; obtain the observation window. Discrete data within: (31) in For vehicles The projection position along the road, For longitudinal velocity, The length of the observation window; Output HDV Future Predicted trajectory: (32) in To predict the number of steps; Spacing between adjacent vehicles Define: (33) set up To prevent division by zero, define density estimation: (34) Select Calculate the average speed of the vehicle: (35) Get traffic: (36) Select time interval Construct two sets of states: (37) set up The shock wave velocity is estimated to be: (38) For upstream vehicles Its distance from HDV is: (39) set up To prevent zero transmission, the propagation delay is: (40) Candidate predictions are obtained by shifting the upstream velocity sequence according to the time delay: (41) Integrating to obtain position candidates: (42) Define fusion weights satisfy , The fusion speed is predicted to be: (43) Weights are based on propagation latency and data quality. The fusion weights are constructed as follows: (44) in The attenuation coefficient is used to predict the final location: (45) Furthermore, the construction of spatial domain delay cooperative error in step S2 includes: For each vehicle Establish an arc-length index cache: (46) in For vehicles At the arc length node The state tuple; At adjacent times , Known , For a given node Find satisfaction: (47) Define interpolation coefficients: (48) The arrival time is: (49) Arbitrary state quantity The interpolation value is: (50) Will Write to the cache and share via V2X broadcast or edge side; The car behind At the node Read the preceding vehicle of , used to construct cooperative error and reinforcement learning state; Set the desired time interval constant ; At the same arc length node Defined as: (51) in To account for the time of arrival error, the heading error is defined as: (52) in To express the geometric consistency of a curve, the derivative of the lateral error is defined as: (53) in Used to express the derivative of lateral error with respect to arc length; To construct reinforcement learning states, an analytical form is used: (54) (55)
[0014] Furthermore, step S3 specifically includes the following steps: S31: Construct spatial domain PPO controller input states: Define curvature window length The curvature input is a sequence: (56) For each CAV At the node Construct the state vector: (57) in To and Corresponding time index; S32: The spatial domain PPO controller outputs combined lateral and longitudinal actions and performs slope feedforward correction. Policy Network by For continuous input and output actions: (58) in For vertical equivalent input, This is the lateral / steering equivalent input; Set upper and lower bounds: (59) in Determined by the vehicle's power / steering capabilities; To compensate for the slope gravity term, the effective longitudinal input is defined as follows: (60) in For vehicle quality, It is the acceleration due to gravity; S23: Reward Function Construction, Sample Collection, and PPO Training Updates Define an instant reward for each CAV: (61) in: As weight, It is a Euclidean norm; TTC penalty is to set a minimum safety distance. Threshold , Prevention of zero constant ; based on the predicted trajectory Calculate the minimum TTC: (62) in To predict relative distance, For predicting speed; the penalty term is: (63) The attachment penalty is: (64) Braking thermal penalty defines the braking thermal state for long downhill driving conditions. ,set up , , , : (65) (66) The PPO advantage is estimated as follows: Let the discount factor be... GAE parameters Value Network Output state value; The TD error is: (67) The advantages are: (68) PPO shearing update: Set shearing coefficient The ratio of the new strategy to the old strategy is: (69) The strategy update objective is: (70) Value networks minimize Update, in which For return estimates; Parameter-sharing multi-agent (PPO) approach: All CAVs share policy parameters. With value parameters Each vehicle, in its own way This generates samples and aggregates and updates them to improve the scalability and generalization capabilities of the fleet size.
[0015] Furthermore, step S4 specifically includes: Spatial Domain Recursion: Using first-order Euler's method to recursively apply spatial domain dynamics from... Recursively to : (71) in Determined by equations (11)-(15); Executor mapping: Mapped to traction / braking force commands or longitudinal acceleration commands: (72) Will Mapped to corner command: (73) in , For implementable inverse model / lookup table / online identification mapping functions, The unit is rad; Input saturated projection: Define the saturation function and execute: (74) Adhesion margin constraint: Define the adhesion coefficient And estimate the lateral acceleration: (75) longitudinal acceleration Depend on Or dynamic estimation; set friction circle constraints: (76) If this constraint is violated, using motion scaling projection will... Scaling back to the feasible region.
[0016] The beneficial effects of this invention are as follows: Significantly improves longitudinal stability and ride comfort: Through deep reinforcement learning strategies, it effectively and actively suppresses speed disturbances in human-driven vehicles, preventing disturbance amplification and maintaining high platoon stability.
[0017] Achieving high-precision lateral and longitudinal coupling tracking and corner elimination: Extending the spatial domain delay time-distance strategy to lateral and longitudinal coupling control fundamentally solves the lateral safety problem of corner cutting phenomenon in complex curve scenarios for mixed-traffic vehicles.
[0018] Achieve energy efficiency optimization: Incorporate energy consumption targets into the control framework, and effectively reduce fuel consumption and emissions of the fleet during operation by suppressing speed fluctuations and optimizing driving strategies.
[0019] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 Flowchart of a deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario; Figure 2 A flowchart for deep reinforcement learning collaborative control of mixed-traffic vehicles in a mountain road scenario; Figure 3 This is a scene depicting a vehicle queue on a curved road section. Figure 4 The curve showing the change in the vehicle platoon trajectory; Figure 5 vehicle heading angle Curve showing the variation with arc length; Figure 6 (a) represents the longitudinal velocity. The curves showing the change in velocity with arc length, (b) represents the longitudinal velocity. Curve showing change over time; Figure 7 In the middle (a), angular velocity is... The curves showing the change in arc length, (b) representing the angular velocity. Curve showing change over time; Figure 8 (a) represents the longitudinal error. The curves showing the variation with arc length, (b) represents the lateral error. Curve showing the change with arc length. Detailed Implementation
[0021] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0022] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0023] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.
[0024] Example 1: This invention provides a spatial domain deep reinforcement learning cooperative control method for mixed-vehicle traffic in mountainous road scenarios. The convoy consists of a lead HDV and several follower CAVs. This invention transforms the vehicle's lateral and longitudinal coupled dynamics into a spatial domain model; it constructs a delayed cooperative error in the spatial domain to ensure alignment of the convoy at the same arc length and avoid cornering on curves; simultaneously, it sets up an online slope estimation module to output the longitudinal slope angle as longitudinal feedforward and safety assessment input, and a trajectory prediction module to output the HDV's future trajectory as look-ahead disturbance information and a basis for predictive safety assessment; finally, all CAVs adopt a spatial domain PPO deep reinforcement learning controller to directly output joint lateral and longitudinal control actions, achieving convergence of longitudinal arrival time error and lateral geometric consistency, and jointly optimizing energy consumption, comfort, and safety constraints.
[0025] The overall architecture of this method includes the following parts: complex road orientation environment modeling and spatial domain coordinate transformation, online slope estimation and dynamic confidence assessment, navigator trajectory prediction and disturbance look-ahead perception, spatial domain delay collaborative error construction and same arc length state alignment, horizontal and vertical integrated joint control based on deep reinforcement learning (PPO), multi-constraint fusion compensation and action saturation projection execution.
[0026] like Figures 1-2 As shown, this method includes the following steps: Step 1: Establish the scene and coordinate system for mixed-traffic vehicles on the mountain road 1) (Fleet definition) Construct a mixed-traffic fleet set (1) in, This refers to the number of connected autonomous vehicles; vehicle number 0 indicates the navigator is driving, and the vehicle number... Indicates the first A connected autonomous vehicle.
[0027] 2) (Road coordinates) Establish arc length coordinates on the road centerline. And set spatial discrete nodes (2) in, The initial arc length, For spatial step size, For spatial indexing, This represents the total number of steps. Simultaneously, the road curvature function is defined. , indicating the position of the arc length The geometric curvature of the road at that location. It can be extracted from high-precision maps, vehicle-road cooperative infrastructure, or vehicle-mounted perception.
[0028] 3) (Time Sampling) Set the time domain sampling time. (3) in, For the sensing / control sampling period, For time indexing.
[0029] Step 2: Construct the vehicle's temporal and longitudinal coupled dynamics and introduce equivalent control inputs. For any vehicle Define the time-domain state in the vehicle coordinate system: longitudinal velocity. lateral velocity yaw rate and heading angle To uniformly describe the longitudinal traction / braking and lateral steering channels, an equivalent control input is introduced: The longitudinal equivalent input can be achieved by synthesizing and normalizing the traction force and braking force; Lateral / steering equivalent input, which can be obtained from the steering angle. It is obtained by mapping the lateral force channel of the tire.
[0030] The temporal and longitudinal coupled dynamics of the vehicle can be written as: (4) (5) (6) in, For vehicle quality, Let be the yaw moment of inertia of the vehicle about its center of mass. This is the distance from the center of gravity to the front axle. The resultant force of longitudinal resistance includes rolling resistance, air resistance, and transmission losses. This is the resultant force of the lateral forces on the tires; This refers to additional yaw terms or unmodeled terms caused by lateral forces. The aforementioned unmodeled portions can be incorporated into the noise or disturbance terms for processing.
[0031] Step 3: Introduce the arc length variable and define the inverse time function. 1) (Arc Length Progress) Define the vehicle Arc length progress And define the velocity modulus (7) in The vehicle speed modulus.
[0032] 2) Definition of (inverse time function) Indicates vehicle Reaching the arc length position The time, and satisfy (8) Step 4: Construct a spatial domain horizontal and vertical coupled dynamic model 1) (Spatial domain derivative) Any differentiable variable The derivative in the spatial domain satisfies (9) 2) (Definition of Spatial Domain State) Define the spatial domain state vector. (10) in For arrival time, For heading angle, , For lateral and longitudinal velocities, ω represents the yaw rate.
[0033] 3) (Spatial Domain Dynamics) From equations (4)-(9), we obtain: (11) (12) (13) (14) (15) in , The vehicle is located at the arc length position. The equivalent longitudinal and lateral control inputs.
[0034] Step 5: Online Slope Estimation Module 1) (Definition of slope) Define the longitudinal slope angle of a road. Its spatial domain form is denoted as The goal of this step is to sample at each time point. Obtain slope estimate And covariance, and map to spatial nodes get .
[0035] 2) (Resistance Model) Define the longitudinal resistance model: (16) in For vehicle speed, , , For calibration parameters, Units are .
[0036] 3) (Measurement Input) at time Acquisition: IMU longitudinal acceleration measurement Vehicle speed measurement Traction estimation With braking force estimation Vehicle weight is recorded as Gravitational acceleration .
[0037] 4) (EKF State Definition) Define the filter state: (17) in The slope angle, For IMU longitudinal acceleration bias.
[0038] 5) (System Equations) Employ a random walk model: (18) in , , This represents the standard deviation of process noise.
[0039] 6) (Measurement Equation: Kinetic Consistency) Establish kinetic consistency measurement: (19) in To measure noise, This is to measure the standard deviation of noise.
[0040] Definition of measurement Measurement function: (20) Jacobi: (twenty one) Measurement noise covariance .
[0041] 7) (EKF recursive) prediction: (twenty two) renew: (twenty three) (twenty four) (25) Output: (26) 8) (Time-to-spatial domain mapping) Define arc length accumulation: (27) For spatial nodes Take the satisfied index Obtained using linear interpolation: (28) 9) (Confidence Output) Output confidence scores as slope variance: (29) in This serves as the reference variance constant. The confidence level is used to adjust the slope feedforward and safety constraint weights.
[0042] Step 6: Trajectory Prediction Module 1) (Observation set) in time Select the set of vehicles to observe ahead: (30) in The number of vehicles used for prediction, where 0 represents HDV. Obtain the observation window. Discrete data within: (31) in For vehicles The projection position along the road, For longitudinal velocity, This is the length of the observation window.
[0043] 2) (Predictive Output Definition) Output HDV future Predicted trajectory: (32) in To predict the number of steps.
[0044] 3) (Local density and flow estimation) for the distance between adjacent vehicles definition: (33) To avoid division by zero, set Defining density estimation: (34) Select Calculate the average speed of the vehicle: (35) Get traffic: (36) 4) (Shock wave velocity estimation) Selecting time intervals Construct two sets of states: (37) set up The shock wave velocity is estimated to be: (38) 5) (Propagation delay and sequence shift) For upstream vehicles Its distance from HDV: (39) set up Propagation delay: (40) Candidate predictions are obtained by shifting the upstream velocity sequence according to the time delay: (41) And integrate to obtain the position candidates: (42) 6) (Multi-vehicle fusion) Define fusion weights satisfy , Fusion speed prediction: (43) Weights can be determined by propagation delay and data quality. ,structure: (44) in This represents the attenuation coefficient. Final position prediction: (45) Step 7: Align the arc-length index cache with the same arc length: 1) (Cache definition) For each vehicle Establish an arc-length index cache: (46) in For vehicles At the arc length node The state tuple.
[0045] 2) (Arc length interpolation writing) at adjacent time points , Known , For a given node Find satisfaction (47) Define interpolation coefficients: (48) The arrival time is: (49) Arbitrary state quantity The interpolation value is: (50) Will Write to the cache and share via V2X broadcast or edge side.
[0046] 3) (Reading the same arc length) The following vehicle At the node Read the preceding vehicle of It is used to construct cooperative error and reinforcement learning state.
[0047] Step 8: Construct spatial domain delay cooperative error 1) (Desired Time Interval) Set the desired time interval constant. .
[0048] 2) (Error definition) At the same arc length node definition: (51) in To account for the time of arrival error, the heading error is defined as: (52) in To express the geometric consistency of a curve, the derivative of the lateral error is defined as: (53) in Used to express the derivative of lateral error with respect to arc length.
[0049] 3) The derivative term can be used in analytical form to construct the reinforcement learning state: (54) (55)
[0050] Step 9: Construct the spatial domain PPO controller input state 1) (Curvature Window) Defines the length of the curvature window. The curvature input is a sequence: (56) 2) (PPO state vector) for each CAV At the node Construct the state vector: (57) in To and The corresponding time index.
[0051] Step 10: The spatial domain PPO controller outputs the combined lateral and longitudinal actions and performs slope feedforward correction. 1) (Action Definition) Policy Network by For continuous input and output actions: (58) in For vertical equivalent input, This is the lateral / steering equivalent input.
[0052] 2) (Motion Constraints) Set upper and lower bounds: (59) in Determined by the vehicle's power / steering capabilities.
[0053] 3) (Slope feedforward correction) is used to compensate for the slope gravity term, defining the effective longitudinal input: (60) in For vehicle quality, This is the acceleration due to gravity.
[0054] Step 11: Reward function construction, sample collection, and PPO training update 1) (Instant Reward) Define an instant reward for each CAV: (61) in: As weight, It is a Euclidean norm.
[0055] 2) (TTC Penalty) Set minimum safe distance Threshold , Prevention of zero constant From the predicted trajectory Calculate the minimum TTC: (62) in To predict relative distance, For speed prediction. Penalty: (63) 3) (Attached penalty): (64) 4) (Brake thermal penalty): Defines the brake thermal state for long downhill driving conditions. ,set up , , , : (65) (66) 5) (PPO Advantage Estimation) Set a discount factor GAE parameters Value Network Output the state value.
[0056] TD error: (67) Advantages: (68) 6) (PPO Shearing Update) Set the shearing coefficient Comparison of old and new strategies: (69) Strategy update objectives: (70) Value networks minimize Update, in which For return estimates.
[0057] 7) (Parameter Sharing Training) Preferred method is Parameter Sharing Multi-Agent (PPO): All CAVs share policy parameters. With value parameters Each vehicle, in its own way This generates samples and aggregates and updates them to improve the scalability and generalization capabilities of the fleet size.
[0058] Step 12: Spatial Domain Recursion, Actuator Mapping, and Constraint Projection 1) (Spatial Domain Recursion) Using first-order Euler equations to recursively apply spatial domain dynamics from... Recursively to : (71) in Determined by equations (11)-(15).
[0059] 2) (Actuator mapping) will Mapped to traction / braking force commands or longitudinal acceleration commands: (72) Will Mapped to corner command: (73) in , For implementable inverse model / lookup table / online identification mapping functions, The unit is rad.
[0060] 3) (Input saturation projection) Define the saturation function and execute: (74) 4) (Adhesion Margin Constraint) Define the adhesion coefficient And estimate the lateral acceleration: (75) longitudinal acceleration can be Or dynamic estimation. Set friction circle constraints: (76) If this constraint is violated, motion scaling projection can be used. Scaling back to the feasible region.
[0061] Example 2: This invention aims to implement vehicle platoon control simulation using Matlab to verify the effectiveness of the proposed integrated longitudinal and lateral cooperative control method based on a spatial domain time delay strategy in eliminating cornering phenomena under complex conditions. The control flow is as follows: Figure 2 As shown. This simulation scenario is set as a continuous curve condition, consisting of straight sections, large-curvature circular arc sections, and gradient sections, to comprehensively evaluate the dynamic performance of the control strategy under lateral and longitudinal coupling. For example... Figure 3 As shown, the simulation queue consists of one lead vehicle (man-driven) and four follower vehicles (intelligent connected vehicles), forming a mixed vehicle queue. The lead vehicle introduces random speed fluctuations as a disturbance source. The first CAV following behind directly acquires the HDV state using onboard sensors and performs look-ahead disturbance perception. The remaining CAVs acquire the state of the preceding vehicle via V2X communication and achieve asynchronous data synchronization by combining the spatial domain time inverse function. This simulation case focuses on the following CAVs using the feedback control law based on spatial domain PPO reinforcement learning proposed in this invention, demonstrating its ability to eliminate cornering phenomena, path tracking accuracy, and longitudinal compensation performance in mountainous road environments, while satisfying safety constraints such as adhesion margin and brake thermal load.
[0062] The initial conditions for the vehicle queue are set as follows: Initial velocity:
[0063] Initial longitudinal position:
[0064] Initial horizontal position:
[0065] Initial arc length position:
[0066] Initial delay interval:
[0067] To evaluate the effectiveness of the proposed control strategy in solving the tangent problem, this embodiment simulates a composite trajectory consisting of straight line segments and circular arc segments. Therefore, the longitudinal velocity of the leading vehicle... It can be constructed as follows: (77) The transverse angular velocity can be constructed as follows: (78) Figure 4 The image shows a vehicle queue employing the proposed spatial domain control strategy. The trajectory in the coordinate system. The trajectory includes straight line segments and circular arc segments, forming a complex "S"-shaped curve. As shown in the figure, all four vehicles can accurately follow the trajectory, and the trajectories of the following vehicles highly overlap with those of the lead vehicle. In particular, when passing through the area with the greatest curvature change, the vehicles did not deviate from the predetermined path or cut corners, verifying the effectiveness of lateral and longitudinal coupling control in geometric path tracking.
[0068] Figure 5 The figure shows the vehicle's heading angle. With arc length The heading angles of all vehicles exhibit a consistent trend, rapidly increasing and decreasing as they enter and exit curves. Due to the time delay, the heading angle response of each following vehicle lags slightly in the spatial domain, but its peak value and rate of change remain highly similar to the lead vehicle, ensuring smooth platooning.
[0069] Figure 6 Figures (a) and (b) show the changes in the vehicle's longitudinal velocity in the arc length coordinate system and the time coordinate system, respectively. Figure 6 (a) (arc-length coordinate system): When the lead vehicle enters a curve, its speed is reduced according to a preset curve to ensure lateral stability, and it accelerates when exiting the curve. The peak speed response of all following vehicles is consistent with it, reflecting the steady-state tracking performance of the controller. Figure 6 (b) (Time coordinate system): This figure clearly illustrates the characteristics of the time-distance delay strategy. The speed changes of the following vehicles are staggered sequentially, meaning that the acceleration and deceleration processes of the following vehicles are delayed relative to the vehicles in front, thus ensuring a constant time interval within the queue.
[0070] Figure 7 Images (a) and (b) show the vehicle's angular velocity, respectively. The changes in the arc length coordinate system and the time coordinate system. Figure 7 In (a) (arc-length coordinate system): the angular velocity of the vehicle changes in a pulse-like manner when entering and leaving the curve, and the non-zero value of the angular velocity corresponds to the curvature of the curve segment. The angular velocity responses of all vehicles are basically consistent, indicating that the lateral control has good synchronization. Figure 7 (b) (Time coordinate system): Similarly, this figure intuitively shows the time delay of angular velocity change, which is a manifestation of the effective coupling of lateral motion and longitudinal time delay strategy.
[0071] Figure 8 The figure shows the longitudinal error. and lateral error With arc length The curve showing the change. Figure 8(a) (Longitudinal error): The longitudinal error exhibits brief transient fluctuations during vehicle acceleration and deceleration, but the absolute value of the error remains at an extremely low level and quickly converges to zero, which proves the ability of the proposed control strategy to accurately maintain the desired time delay. Figure 8 (b) (lateral error): The lateral error remains at an extremely low absolute value throughout the entire trajectory segment, including the curvature-changing curves. This result strongly demonstrates that the proposed lateral-longitudinal coupling control successfully eliminates the cornering phenomenon in mixed vehicle platoons under complex curve conditions, ensuring high-precision path tracking performance.
[0072] Simulation results clearly demonstrate that the horizontal and vertical coupling control method based on the time delay strategy proposed in this invention can effectively and stably achieve high-precision tracking and anti-angle control of vehicle platoons on time-varying curvature trajectories.
[0073] Example 3: An electronic device, comprising a memory and a processor; The memory is used to store computer programs; The processor is configured to implement the method described in Embodiment 1 when executing the computer program.
[0074] Example 4: A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in Embodiment 1.
[0075] Example 5: A computer program product includes a computer program that, when executed by a processor, implements the method described in Example 1.
[0076] In the above embodiments, the reference to "this embodiment" in the specification indicates that a specific feature, structure, or characteristic described in connection with the embodiment is included in at least some embodiments, but not necessarily all embodiments. Multiple appearances of "this embodiment" do not necessarily refer to the same embodiment.
[0077] In the above embodiments, although the invention has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory structures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed. The embodiments of the invention are intended to cover all such substitutions, modifications, and variations falling within the broad scope of the appended claims.
[0078] As will be understood by those skilled in the art, the computer-readable storage medium described in this embodiment allows for the implementation of all or part of the steps in the above method embodiments by computer program-related hardware. The aforementioned computer program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0079] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication between them. The memory is used to store computer programs, the communication interface is used to perform communication, and the processor and the transceiver are used to run the computer programs, so that the electronic terminal performs the steps of the above method.
[0080] In this embodiment, the memory may include random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.
[0081] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0082] This invention can be used in a wide range of general-purpose or special-purpose computing system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.
[0083] This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario, characterized by: Includes the following steps: S1: Establish the scenario and coordinate system of mixed traffic vehicles on the ring road; construct the temporal and longitudinal coupled dynamics of vehicles and introduce equivalent control inputs, introduce arc length variables and define time inverse functions, and construct a spatial domain temporal and longitudinal coupled dynamics model; S2: Calculate the slope and its confidence level; predict the trajectory of the HDV (High-Density Vehicle). Construct spatial domain delay cooperative error; S3: Construct a spatial domain PPO controller, define a comprehensive reward function, and output the combined horizontal and vertical control actions; S4: Perform multi-constraint fusion compensation and motion saturation projection execution.
2. The spatial domain deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario as described in claim 1, characterized in that: Step S1, which establishes the mixed-traffic vehicle scenario and coordinate system for the mountain road, includes: Constructing a mixed-traffic vehicle fleet: (1) in, This refers to the number of connected autonomous vehicles; vehicle number 0 indicates the navigator is driving, and the vehicle number... Indicates the first Connected autonomous vehicles; Establish arc length coordinates along the road centerline. And define spatially discrete nodes: (2) in, The initial arc length, For spatial step size, For spatial indexing, Set the total number of steps; define the road curvature function. , indicating the position of the arc length The geometric curvature of the road at that location. Extracted from high-precision maps, vehicle-road cooperative infrastructure, or vehicle-mounted perception; Set the time domain sampling time and perform time sampling; (3) in, For the sensing / control sampling period, For time indexing.
3. The spatial domain deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario as described in claim 2, characterized in that: Step S1, which involves constructing the vehicle's temporal longitudinal and transverse coupled dynamics and introducing equivalent control inputs, includes: For any vehicle Define the time-domain state in the vehicle coordinate system, including longitudinal velocity. lateral velocity yaw rate and heading angle ; Introduce equivalent control inputs, including longitudinal equivalent inputs. Traction and braking forces can be synthesized and normalized; lateral / steering equivalent inputs Due to the steering angle Obtained by mapping the lateral force channel of the tire; The temporal and longitudinal coupled dynamics of the vehicle can be written as: (4) (5) (6) in, For vehicle quality, Let be the yaw moment of inertia of the vehicle about its center of mass. This is the distance from the center of gravity to the front axle. The resultant force of longitudinal resistance includes rolling resistance, air resistance, and transmission loss; This is the resultant force of the lateral forces on the tires; This refers to additional yaw terms or unmodeled terms caused by lateral forces; unmodeled portions are incorporated into noise or disturbance terms.
4. The spatial domain deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario as described in claim 3, characterized in that: The introduction of the arc length variable and the definition of the inverse time function in step S1 includes: Define vehicle Arc length progress And define the velocity modulus: (7) in For vehicle speed modulus; definition Indicates vehicle Reaching the arc length position The time, and satisfying: (8)。 5. The spatial domain deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario as described in claim 4, characterized in that: Step S1, which involves constructing a spatial domain coupled dynamic model, includes the following steps: any differentiable variable The derivative in the spatial domain satisfies: (9) Define the spatial domain state vector: (10) in For arrival time, For heading angle, , For lateral and longitudinal velocities, This refers to the yaw rate; The spatial domain dynamics model is obtained from equations (4) to (9): (11) (12) (13) (14) (15) in , The vehicle is located at the arc length position. The equivalent longitudinal and lateral control inputs.
6. The spatial domain deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario as described in claim 1, characterized in that: The calculation of the slope and its confidence level in step S2 specifically includes: Define the longitudinal slope angle of the road Its spatial domain form is denoted as The goal is to sample at each time step. Obtain slope estimate And covariance, and map to spatial nodes get ; Define the longitudinal resistance model: (16) in For vehicle speed, , , For calibration parameters, Units are ; At any moment Acquisition of IMU longitudinal acceleration measurement Vehicle speed measurement Traction estimation With braking force estimation Vehicle weight is recorded as Gravitational acceleration ; Define the filter state: (17) in The slope angle, For IMU longitudinal acceleration bias; The system equations are constructed using a random walk model: (18) in , , The standard deviation of process noise; Establish dynamic consistency measurement: (19) in To measure noise, To measure the standard deviation of noise; Definition of measurement Measurement function: (20) Jacobi's construction is as follows: (21) Measurement noise covariance ; Slope prediction using EKF recursion: (22) The updated formula is as follows: (23) (24) (25) The output is: (26) Define cumulative arc length: (27) For spatial nodes Take the satisfied index Obtained using linear interpolation: (28) Confidence level is output using slope variance: (29) in The confidence level is used to adjust the slope feedforward and safety constraint weights, with the variance constant serving as the reference.
7. The spatial domain deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario as described in claim 1, characterized in that: Step S2, which involves predicting the trajectory of the HDV (High-Depth Vehicle), includes: In time The set of vehicles selected for observation ahead is: (30) in The number of vehicles used for prediction, where 0 represents HDV; obtain the observation window. Discrete data within: (31) in For vehicles The projection position along the road, For longitudinal velocity, The length of the observation window; Output HDV Future Predicted trajectory: (32) in To predict the number of steps; Spacing between adjacent vehicles Define: (33) set up To prevent division by zero, define density estimation: (34) Select Calculate the average speed of the vehicle: (35) Get traffic: (36) Select time interval Construct two sets of states: (37) set up The shock wave velocity is estimated to be: (38) For upstream vehicles Its distance from HDV is: (39) set up To prevent zero transmission, the propagation delay is: (40) Candidate predictions are obtained by shifting the upstream velocity sequence according to the time delay: (41) Integrating to obtain position candidates: (42) Define fusion weights satisfy , The fusion speed is predicted to be: (43) Weights are based on propagation latency and data quality. The fusion weights are constructed as follows: (44) in The attenuation coefficient is used to predict the final location: (45)。 8. The spatial domain deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario as described in claim 1, characterized in that: Step S2, which involves constructing the spatial domain delay cooperative error, includes: For each vehicle Establish an arc-length index cache: (46) in For vehicles At the arc length node The state tuple; At adjacent times , Known , For a given node Find satisfaction: (47) Define interpolation coefficients: (48) The arrival time is: (49) Arbitrary state quantity The interpolation value is: (50) Will Write to the cache and share via V2X broadcast or edge side; The car behind At the node Read the preceding vehicle of , used to construct cooperative error and reinforcement learning state; Set the desired time interval constant ; At the same arc length node Defined as: (51) in To account for the time of arrival error, the heading error is defined as: (52) in To express the geometric consistency of a curve, the derivative of the lateral error is defined as: (53) in Used to express the derivative of lateral error with respect to arc length; To construct reinforcement learning states, an analytical form is used: (54) (55)。 9. The spatial domain deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario according to claim 1, characterized in that: Step S3 specifically includes the following steps: S31: Construct spatial domain PPO controller input states: Define curvature window length The curvature input is a sequence: (56) For each CAV At the node Construct the state vector: (57) in To and Corresponding time index; S32: The spatial domain PPO controller outputs combined lateral and longitudinal actions and performs slope feedforward correction. Policy Network by For continuous input and output actions: (58) in For vertical equivalent input, This is the lateral / steering equivalent input; Set upper and lower bounds: (59) in Determined by the vehicle's power / steering capabilities; To compensate for the slope gravity term, the effective longitudinal input is defined as follows: (60) in For vehicle quality, It is the acceleration due to gravity; S23: Reward Function Construction, Sample Collection, and PPO Training Updates Define an instant reward for each CAV: (61) in: As weight, It is a Euclidean norm; TTC penalty is to set a minimum safety distance. Threshold , Prevention of zero constant ; based on the predicted trajectory Calculate the minimum TTC: (62) in To predict relative distance, For predicting speed; the penalty term is: (63) The attachment penalty is: (64) Braking thermal penalty defines the braking thermal state for long downhill driving conditions. ,set up , , , : (65) (66) The PPO advantage is estimated as follows: Let the discount factor be... GAE parameters Value network Output state value; The TD error is: (67) The advantages are: (68) PPO shearing update: Set shearing coefficient The ratio of the new strategy to the old strategy is: (69) The strategy update objective is: (70) Value networks minimize Update, in which For return estimates; Parameter-sharing multi-agent (PPO) approach: All CAVs share policy parameters. With value parameters Each vehicle, in its own way This generates samples and aggregates and updates them to improve the scalability and generalization capabilities of the fleet size.
10. The spatial domain deep reinforcement learning collaborative control method for mixed-traffic vehicles in a mountain road scenario according to claim 1, characterized in that: Step S4 specifically includes: Spatial Domain Recursion: Using first-order Euler's method to recursively apply spatial domain dynamics from... Recursively to : (71) in Determined by equations (11)-(15); Executor mapping: Mapped to traction / braking force commands or longitudinal acceleration commands: (72) Will Mapped to cornering instructions: (73) in , For implementable inverse model / lookup table / online identification mapping functions, The unit is rad; Input saturated projection: Define the saturation function and execute: (74) Adhesion margin constraint: Define the adhesion coefficient And estimate the lateral acceleration: (75) longitudinal acceleration Depend on Or dynamic estimation; set friction circle constraints: (76) If this constraint is violated, using motion scaling projection will... Scaling back to the feasible region.