A method and system for longitudinal vehicle platoon coordination control considering brake fade
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-08-11
AI Technical Summary
然而,现有技术尚缺乏一种面向车辆队列、能够将受热衰退影响车辆及其后续车辆作为相互耦合决策单元统一处理的协调控制机制,难以在满足防追尾安全要求的同时兼顾队列效率与运行平稳性
(1)本发明将制动热状态演化及热衰退风险预测引入车辆队列纵向防追尾控制,构建面向未来短时工况的前瞻性热风险评估机制;基于制动器温度信息,对未来短时制动器温度进行保守上限预测,并据此估计车辆未来可用最大制动减速度;同时结合车间距和相对速度,评估未来避免追尾所需的最小制动需求,从而实现制动能力与制动需求在统一预测时域内的前瞻性比较,提高热衰退工况下风险识别的及时性。
Smart Images

Figure CN122354513B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of road vehicle control systems, and specifically relates to a longitudinal vehicle platoon coordination control method and system that takes into account brake heat fade. Background Technology
[0002] With the development of intelligent connected vehicles, vehicle-road cooperative systems, and vehicle platooning technology, longitudinal following control has become one of the key technologies for improving road traffic efficiency and driving safety. Existing vehicle platooning control, automatic following control, and cooperative adaptive cruise control technologies typically use vehicle speed, relative speed, vehicle spacing, or time distance as the main control variables to achieve longitudinal stability of the vehicle platoon by maintaining a preset target spacing or time distance.
[0003] In existing technologies, most research on longitudinal control of vehicle platoons is based on the premise that the vehicle's braking capacity is stable and available, that is, assuming that the vehicle has braking capacity matching its current motion state at any given time. On this basis, the controller generates longitudinal control commands according to control laws based on fixed safety distances, fixed safety time intervals, or relative motion relationships. This type of technical solution can meet basic car-following and platooning requirements under normal operating conditions. However, under conditions such as long downhill slopes, heavy loads, frequent braking, or continuous braking, the vehicle's braking system may experience thermal fade due to temperature rise, leading to a decrease in actual available braking capacity. In such cases, if traditional control strategies based on fixed time intervals or fixed safety distances are still used, it is possible that the vehicle's geometric spacing superficially meets the requirements, but the actual braking capacity is insufficient to support rear-end collision prevention, thus significantly increasing the risk of rear-end collisions within the platoon.
[0004] Existing research and applications on brake fade mainly focus on brake structure design optimization, brake thermal management, auxiliary brake distribution, cooling method improvement, and single-vehicle braking performance compensation. These technologies primarily focus on reducing the degree of fade or improving braking performance at the single-vehicle level, rather than uniformly modeling brake fade as a key safety constraint in the longitudinal movement of vehicle platoons. Especially in vehicle platoon scenarios, current technologies generally lack an assessment mechanism that can uniformly characterize the vehicle's thermal state, future available braking capacity, and braking requirements to avoid rear-end collisions, making it difficult to dynamically adjust the safe distance between vehicles based on the risk of fade.
[0005] Furthermore, when vehicles are in a platoon, if a vehicle experiences increased risk due to brake fade, relying solely on its own local adjustments is often insufficient to mitigate the risk in a timely manner. Subsequent vehicles typically can only react passively after the actual distance between vehicles has significantly decreased or the speed difference has increased significantly, easily causing the risk to propagate or even amplify along the platoon. Even though some existing vehicle platoon coordination control methods consider information exchange and coordinated adjustment among multiple vehicles, their control objectives are mostly focused on platoon stability, following accuracy, and traffic efficiency, with insufficient consideration given to the decrease in available braking capacity caused by brake fade and its impact on the platoon's rear-end collision prevention safety boundary.
[0006] Furthermore, in vehicle platooning scenarios, the safety spacing adjustment of an individual vehicle often affects the control requirements and operational status of its following vehicles, resulting in significant front-to-back coupling between vehicles. Therefore, controlling the risk of brake fade requires not only identifying the current risk status of an individual vehicle but also coordinating the target spacing and longitudinal control behavior of related vehicles at the platoon scale. However, current technologies lack a coordinated control mechanism for vehicle platoons that can treat vehicles affected by brake fade and their following vehicles as mutually coupled decision-making units, making it difficult to simultaneously meet rear-end collision prevention safety requirements while also ensuring platoon efficiency and operational stability.
[0007] Therefore, how to fully consider the impact of brake fade on the future available braking capacity of vehicles in the scenario of vehicle platooning, construct a risk characterization index that can uniformly reflect the relationship between "current or future available braking capacity" and "braking requirements to avoid rear-end collisions", and realize platoon-level dynamic safety distance coordinated adjustment and longitudinal control based on the coupling relationship between vehicles has become a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0008] In view of the shortcomings and deficiencies of the existing technology, the purpose of this invention is to provide a longitudinal vehicle platoon coordination control method and system that takes into account brake heat fade. The method evaluates the rear-end collision prevention capability of vehicles in the current heat fade state in real time under conditions such as long downhill, heavy load or continuous braking, and calculates the optimal longitudinal control command for each vehicle through a multi-agent reinforcement learning method. The method coordinates global risk information in the platoon in a unified manner, thereby reducing the probability of rear-end collision and suppressing the spread of risk in the platoon.
[0009] To achieve the above objectives, the present invention adopts the following technical solution: A longitudinal vehicle platoon coordination control method considering brake heat fade, the method comprising the following steps: Step S1. Obtain the vehicle's running status information and its following relationship information with the vehicle in front in the vehicle queue; Step S2. Based on the acquired data, determine whether the vehicle is in a thermal fade activation state. If so, predict the conservative upper limit of the future brake temperature based on the current state, and calculate the maximum available braking deceleration and the minimum safe distance for thermal adaptation. Step S3. Based on the maximum available braking deceleration and the thermal adaptive minimum safe distance, combined with the current car-following state, evaluate the braking capacity margin and distance margin, take the minimum value of the two as the comprehensive risk margin, and map it to the risk level; Step S4. Construct a multi-agent reinforcement learning framework with centralized training and distributed execution. The state space includes vehicle operating state, braking thermal state, and safety risk information; the action space is the longitudinal control action; and the reward function is: ; in, This is a safety cost item used to quantify the safety costs incurred by a vehicle platoon during following another vehicle due to rear-end collision risk, insufficient thermal safety distance, and brake heat accumulation. For efficiency cost; For comfort-related costs, This marks the end of the entire optimization evaluation timeframe. These are the weighting coefficients; Step 5. Each vehicle outputs longitudinal control actions based on the current state space using the trained policy network; then, the longitudinal control actions of each vehicle are restricted to the control range jointly determined by the vehicle's current available braking and driving capabilities to obtain the final control command, thereby realizing longitudinal vehicle platoon coordinated control.
[0010] As a preferred embodiment of the present invention, the operating status information of the vehicle in step S1 includes the vehicle speed, vehicle acceleration, brake temperature, brake system pressure signal and continuous braking time; the following relationship information of the preceding vehicle includes the actual distance between the two vehicles and the speed of the preceding vehicle.
[0011] In a preferred embodiment of the present invention, step S2 first filters the brake temperature signal, and then determines whether the vehicle has entered a warning state based on the filtered brake temperature, the current continuous braking duration, and the vehicle's acceleration. If the vehicle has entered a warning state, it determines whether the vehicle has entered a thermal fade activation state based on the filtered brake temperature, the vehicle's acceleration, and the trigger accumulation counter. The trigger accumulation counter increments when the warning conditions are met and decrements when the warning conditions are not met. If the vehicle has entered a thermal fade activation state and the activation state is maintained for multiple cycles, a conservative upper limit prediction of the future brake temperature is made based on the current state, and the current maximum available braking deceleration and the thermal adaptive minimum safe distance are calculated.
[0012] As a preferred embodiment of the present invention, when making a conservative upper limit prediction of the future brake temperature in step S2, a thermal inertia time constant is introduced to replace the fixed equivalent heat capacity. The formula for predicting brake temperature is: ; in, : No. A conservative upper limit for the brake temperature at each sampling time; : Current brake temperature after filtering; Predict the number of interval steps. Equal to the predicted length; : No. The vehicle's current thermal inertia time constant is estimated using a recursive least squares identification algorithm based on real-time measured brake temperature and braking power. Average heat dissipation coefficient; Maximum braking power of the vehicle; Current vehicle cooling capacity; Use period length; Then, based on a conservative upper limit for brake temperature, a braking performance factor is calculated. This braking performance factor employs a piecewise function to quantify the degree to which brake fade weakens braking capability. Based on the braking performance factor, the vehicle's braking performance is determined in the [missing information - likely a specific timeframe or period]. The maximum available braking deceleration at each sampling time.
[0013] As a preferred embodiment of the present invention, in step S2, when calculating the thermally adaptive minimum safe distance, the thermal inertia relaxation factor is first used. The relative ability of a vehicle to withstand additional thermal load in the next sampling period at the current brake temperature is expressed as: ; in, Warning temperature threshold; The brake's thermal saturation temperature, i.e., the braking performance factor has dropped to its minimum value. The temperature at that time; Then, the minimum safe distance will be fixed based on the thermal inertia relaxation factor. Extended to dynamic thermal adaptive minimum safe distance Its expression is: ; in, Calibration coefficient, range of values ; Avoid positive numbers that are divided by zero; Relative velocity influence coefficient, range of values ; : Relative speed to the vehicle in front; Reference speed Speed of this vehicle; Speed of the vehicle in front.
[0014] As a preferred embodiment of the present invention, in step S3, the minimum deceleration required to avoid a rear-end collision is evaluated based on the relative motion state, and the braking capacity margin is evaluated based on the minimum deceleration required to avoid a rear-end collision and the maximum available braking deceleration; the distance margin is evaluated based on the thermally adaptive minimum safe distance and the actual distance between the current vehicle and the vehicle in front, and the minimum value of the two is taken when calculating the comprehensive risk margin.
[0015] As a preferred embodiment of the present invention, the vehicle operating state in the state space in step S4 includes the vehicle speed. The acceleration of this vehicle Relative speed to the vehicle in front and the actual distance between this vehicle and the vehicle in front. Braking thermal condition includes the temperature of the brake after filtering. Thermal inertia relaxation factor Safety risk information includes thermally adaptive minimum safe distance. Comprehensive risk margin Current maximum available braking deceleration ; The security cost term in the reward function Represented as: ; in, At the current sampling time, For fixed weighting coefficients, The dynamic weighting coefficient is determined by the risk level. Calculated; The dynamic weighting coefficients are determined by the thermal inertia relaxation factor. Calculated; The temperature cost term is expressed as: ; In the formula: This is the thermal decay exit temperature threshold. This represents the total number of vehicles in the fleet. This is the average brake temperature across all vehicles. These are the preset weight parameters.
[0016] As a preferred embodiment of the present invention, the efficiency cost term ;in, For the most economical following distance; Comfort cost ;in, These are the weighting coefficients. This represents the acceleration of the vehicle in front.
[0017] As a preferred embodiment of the present invention, the multi-agent reinforcement learning framework described in step S4 employs a multi-agent deep deterministic policy gradient algorithm.
[0018] A longitudinal vehicle platoon coordination control system considering brake thermal fade, the system being used to implement the aforementioned longitudinal vehicle platoon coordination control method considering brake thermal fade, includes a state perception module, a data processing module, a thermal fade assessment module, a thermal inertia time constant estimation module, a brake temperature prediction module, a brake performance factor estimation module, a maximum available braking deceleration estimation module, a thermal adaptive minimum safe distance estimation module, a comprehensive risk margin and risk level assessment module, a strategy network, and a constraint module. The status perception module is used to acquire the vehicle's operating status information and its following relationship information with the vehicle in front; The data processing module is used to filter the collected brake temperature. The thermal fade assessment module is used to determine whether the vehicle is in a thermal fade activation state. The thermal inertia time constant estimation module is used to estimate the thermal inertia time constant based on the real-time collected brake temperature and braking power using a recursive least squares identification algorithm. The brake temperature prediction module predicts the future brake temperature based on the vehicle's maximum braking power, current heat dissipation power, current brake temperature, and thermal inertia time constant. The braking performance factor estimation module is used to calculate the braking performance factor based on the brake temperature. The maximum available braking deceleration estimation module is used to determine the current or future maximum available braking deceleration of the vehicle based on the braking efficiency factor. The thermally adaptive minimum safe distance estimation module first quantifies the relative ability of a vehicle to withstand additional thermal load in the next sampling period at the current brake temperature using a thermal inertia relaxation factor, and then extends the fixed minimum safe distance to a dynamic thermally adaptive minimum safe distance based on the thermal inertia relaxation factor. The comprehensive risk margin and risk level assessment module evaluates the braking capacity margin and distance margin based on the maximum available braking deceleration and the thermal adaptive minimum safe distance, combined with the current car-following state. The minimum value of the two is taken as the comprehensive risk margin and mapped to the risk level. The policy network is used to output longitudinal control actions based on the current state space; The constraint module is used to limit the longitudinal control actions of each vehicle within a control range determined by the vehicle's current available braking and driving capabilities, thereby obtaining the final control command.
[0019] Advantages and beneficial effects of the present invention: (1) This invention introduces the evolution of braking thermal state and prediction of thermal fade risk into the longitudinal rear-end collision prevention control of vehicle platoons, and constructs a forward-looking thermal risk assessment mechanism for future short-term operating conditions; based on brake temperature information, a conservative upper limit prediction is made for the future short-term brake temperature, and the maximum braking deceleration available to the vehicle in the future is estimated accordingly; at the same time, combined with the vehicle spacing and relative speed, the minimum braking requirement required to avoid rear-end collisions in the future is evaluated, thereby realizing a forward-looking comparison of braking capacity and braking requirement in a unified prediction time domain, and improving the timeliness of risk identification under thermal fade conditions.
[0020] (2) The present invention constructs a hierarchical risk management mechanism based on the judgment of thermal fade risk trigger and the assessment of future braking capacity. The thermal fade risk is triggered and identified by the brake temperature threshold, the braking intensity threshold and the continuous braking time threshold. The future temperature prediction and braking capacity assessment process is initiated only when the thermal fade risk conditions are met, so as to avoid false triggering and frequent fluctuations in dynamic safety distance caused by short-term slight braking, thereby taking into account both the sensitivity of risk identification and the stability of control.
[0021] (3) The present invention realizes multi-timescale coupled evaluation and dynamic distance control of future available braking capacity and rear-end collision prevention safety distance requirements. It does not rely solely on fixed vehicle distance or fixed time distance for car-following control, but rather applies constraints to longitudinal control commands that are coordinated with future available braking capacity based on available braking capacity and thermal adaptive minimum safety distance. This transforms car-following control from traditional geometric spacing constraints to safety control oriented towards actual braking capacity changes.
[0022] (4) This invention proposes a multi-agent collaborative optimization safety control mechanism that considers the thermal state distribution characteristics of a vehicle platoon. Unlike the distributed and unidirectional risk compensation strategies in existing technologies, this invention models the vehicle platoon rear-end collision prevention control problem as a multi-agent reinforcement learning framework with centralized training and distributed execution. The local observation state of each agent includes information such as brake temperature, risk margin, and future available braking capacity. The joint action directly outputs longitudinal acceleration commands. In addition to the risk margin penalty, efficiency cost, and comfort cost, the reward function specifically introduces a platoon temperature variance penalty term to guide the agents to actively reduce the difference in brake temperature among vehicles, thereby achieving platoon-level thermal equilibrium. This architecture effectively overcomes the shortcomings of traditional distributed control, which relies only on local information and is difficult to achieve global coordination. Under the premise of ensuring that the risk of each vehicle is controllable, it achieves the optimal trade-off between platoon safety, thermal equilibrium, and traffic efficiency.
[0023] (5) Deep coupling between brake thermal fade and vehicle platoon rear-end collision prevention control is achieved. Existing vehicle platoon control methods usually assume constant braking capacity and do not consider the dynamic weakening of braking capacity due to brake thermal fade under conditions such as long downhill, heavy load, or continuous braking. This invention introduces brake temperature filtering, thermal fade state machine, conservative temperature prediction, and braking performance factor model to explicitly incorporate the physical process of thermal fade into the longitudinal control framework of vehicle platoon, enabling the system to actively adjust the following strategy under high temperature conditions and reduce the risk of rear-end collisions caused by brake overheating.
[0024] (6) An adaptive temperature prediction method based on thermal inertia time constant is proposed. Traditional temperature prediction relies on fixed equivalent heat capacity parameters, which cannot adapt to changes in different vehicles, loads, brake wear, and road slopes. This invention innovatively proposes a thermal inertia time constant. It also provides a recursive least squares online adaptive identification method, which uses real-time collected brake temperature and brake power to estimate the brakes online. This allows the estimate of the rate of temperature rise to be automatically adjusted according to the actual thermal inertia characteristics of the vehicle.
[0025] (7) A thermally adaptive minimum safe distance model was constructed to achieve preventative distance extension. Existing safe distance models are mostly fixed values or simple functions that only depend on speed and relative speed, which cannot reflect the cumulative impact of thermal fade on braking capability. This invention proposes a thermal inertia relaxation factor. It quantitatively describes the relative ability of a vehicle to withstand additional thermal loads within a predicted time period, and thus constructs a thermally adaptive minimum safety distance. This distance degrades to a fixed value when the engine is cold, without affecting normal car-following efficiency; however, it automatically and sharply increases when heat fade is severe, forcing vehicles to increase the distance in advance and avoiding dangerous conditions requiring large deceleration and braking from the outset. Furthermore, the model also considers the effects of the vehicle's own speed and forward relative speed, making its physical meaning clear, its calculations simple, and its deployment easy in engineering.
[0026] (8) A comprehensive risk index integrating braking capacity and distance margin is proposed. Existing risk assessment methods typically only compare available braking deceleration with demand deceleration, neglecting the important dimension of thermal safety distance. This invention defines two components: braking capacity margin and distance margin, and takes the minimum of the two as the comprehensive risk margin. This index simultaneously reflects the remaining capacity of the two safety dimensions of "being able to brake" and "being able to distance," and is continuous and smooth, making it suitable as input to the reinforcement learning reward function. The discretized risk level generated on this basis can be used for priority adjustment or hierarchical control in multi-agent collaborative decision-making, enabling the system to prioritize the allocation of limited safety resources to the highest-risk vehicles in the queue.
[0027] (9) A multi-agent reinforcement learning coordinated control framework for thermal fading scenarios was designed. This invention adopts a multi-agent reinforcement learning framework with centralized training and distributed execution, treating each vehicle as an agent. Its local observation state includes information such as vehicle motion parameters, brake thermal state, thermal adaptive minimum safe distance, comprehensive risk margin, and available braking capacity. The safety cost term in the reward function consists of three parts: comprehensive risk margin penalty (providing a smooth gradient), thermal safe distance violation penalty (the penalty intensity is inversely proportional to the thermal inertia relaxation factor, i.e., the more severe the thermal fading, the stronger the penalty), and brake temperature cost (suppressing excessively high proportions of high-temperature vehicles and uneven vehicle fleet temperature). The efficiency cost term only penalizes situations where the actual following distance is greater than the economical following distance, avoiding incorrectly penalizing safe behaviors when the distance increases due to thermal fading. The action space is simultaneously constrained by the maximum driving acceleration and the current maximum available braking deceleration, ensuring that the commands are physically executable. This framework enables the control strategy to autonomously learn how to actively adjust the following distance and longitudinal acceleration under different thermal states, achieving a dynamic balance between safety, efficiency, and comfort.
[0028] (10) A layered and complementary safety assurance system has been formed. This invention organically combines physical hard constraints with soft penalty guidance: the maximum available braking deceleration and thermal safety distance output in step S2 are used as hard constraints for action space pruning and low-level safety redundancy to ensure that control commands do not go out of bounds under any circumstances; the comprehensive risk margin output in step S3 is used as a soft penalty for the reward function to provide asymptotic gradients and guide the reinforcement learning strategy to actively maintain a safety margin. The two complement each other, which not only meets the functional safety requirements for reliability, but also solves the problems of sparse training and difficult convergence of pure reinforcement learning, so that the system can still maintain stable and reliable anti-rear-collision performance under extreme conditions such as long downhill slopes.
[0029] (11) Effectively suppresses the propagation of thermal fade risk in the vehicle platoon. Because the thermally adaptive minimum safe distance of this invention automatically increases with rising temperature, when a vehicle in the platoon experiences a temperature increase due to heavy load or brake wear, it will automatically increase the distance from the vehicle in front. Simultaneously, following vehicles will adjust their following behavior accordingly under the multi-agent collaborative decision-making mechanism, avoiding chain collisions caused by localized thermal fade. Furthermore, the hysteresis and state-holding mechanisms in the state machine effectively suppress risk state oscillations near the temperature threshold, preventing frequent fluctuations and propagation of risk information within the platoon. Ultimately, overall risk control and operational stability of the platoon are achieved. Attached Figure Description
[0030] The present invention will be described and illustrated in detail below with reference to the accompanying drawings and through a detailed description of the embodiments.
[0031] Figure 1 A flowchart of a longitudinal vehicle queuing coordination control method considering brake heat fade is provided for this invention. Figure 2 A block diagram of a longitudinal vehicle platooning coordination control system considering brake heat fade is provided for this invention. Detailed Implementation
[0032] The present invention will be further described in detail below with reference to examples and accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0033] Example 1:
[0034] like Figure 1 As shown, this embodiment provides a longitudinal vehicle platoon coordination control method considering brake fade, which includes the following steps: Step S1. Based on the state perception module, obtain the vehicle's operating status information and its following relationship information with the vehicle in front in the vehicle queue; wherein, the vehicle's operating status information includes at least the vehicle's speed, acceleration, brake temperature, brake system pressure signal and continuous braking time; the following relationship information with the vehicle in front includes at least the actual distance between the two vehicles and the speed of the vehicle in front.
[0035] Specifically, in this embodiment, the first vehicle in the queue... Taking a vehicle as an example, the system operates on a fixed cycle, sampling once per cycle; the state perception module periodically acquires the following physical quantities: (1) Obtain the vehicle's operating status information for this cycle: Specifically, the vehicle speed is obtained through onboard sensors. Brake temperature The acceleration of this vehicle The quality of this vehicle Braking system pressure signal .
[0036] (2) Obtain information on the following relationship with the vehicle in front: Specifically, the actual distance between this vehicle and the vehicle in front is obtained through onboard sensors. Speed of the vehicle in front .
[0037] Calculate relative speed based on the acquired data: ; Based on the braking system pressure signal The system determines whether the vehicle is braking and calculates the duration of braking. The update rule for this determination is as follows: ;
[0038] in, For the first The car in The continuous braking time at each sampling moment The braking determination threshold, This represents the sampling period length.
[0039] In this embodiment, the output of the state-aware module serves as the basic data source for all subsequent modules.
[0040] Step S2. Based on the data obtained in step S1, determine whether the vehicle is in a thermal fade activation state. If it is in a thermal fade activation state, predict the conservative upper limit of the future brake temperature based on the current state, and calculate two core safety boundary quantities: the current maximum available braking deceleration and the thermal adaptive minimum safe distance. Quantify the safety margin of the vehicle in the current thermal state from two independent dimensions: "braking capacity" and "following distance", as soft constraint quantities for subsequent multi-agent control.
[0041] Specifically, in this embodiment, step S2 includes the following steps: Step S2.1. Brake temperature pretreatment: This embodiment filters the brake temperature signal to suppress sensor noise and temperature fluctuations, ensuring stable and reliable subsequent braking capability assessment.
[0042] Specifically, a first-order inertial filter is used to smooth the brake temperature, implemented in a discrete form: ;
[0043] in, For the first Brake temperature after filtering at each sampling time; For the first Brake temperature measurement values at each sampling time; This is the filtering time constant.
[0044] Step S2.2. Determining the activation state of thermal decay: To avoid unnecessary computational overhead and state fluctuations, this embodiment employs a state activation mechanism. Subsequent conservative temperature predictions and safety boundary calculations are only executed when the thermal decay activation state is established. This mechanism ensures the system maintains high efficiency under normal operating conditions, activating thermal safety protection only when risk indicators appear.
[0045] Specifically, step S2.2 first determines the early warning state, and then determines the thermal decay activation state: When the vehicle is in normal condition, if any of the following conditions are met, the vehicle is determined to enter the warning state: ;
[0046] or: ;
[0047] or: ;
[0048] in, This represents the current duration of continuous braking. , and These are the warning temperature threshold, the warning braking deceleration threshold, and the warning duration threshold, respectively.
[0049] When the vehicle is in a warning state, if any of the following conditions are met, the vehicle is determined to have entered a heat fade activation state, and the subsequent steps of brake temperature conservative upper limit prediction and future available braking capacity assessment are initiated: ;
[0050] or: ;
[0051] or: ;
[0052] in, To activate the temperature threshold, To activate the braking intensity threshold, For the first Triggering the cumulative counter at each sampling time. The count threshold required to enter the active state.
[0053] Preferably, This forms a graded judgment process from the warning state to the activation state.
[0054] In this embodiment, the trigger cumulative counter The alert value is incremented when the warning conditions are met consecutively, and decremented when the warning conditions are not met. The update is performed as follows: ;
[0055] in, This is the upper limit value of the counter.
[0056] Furthermore, in this embodiment, when the vehicle enters the thermal fade activation state, to avoid frequent system exits from the activation state due to small fluctuations in brake temperature or braking deceleration near a threshold, the activation state needs to be maintained. One cycle, This is the minimum holding period.
[0057] In addition, this embodiment also sets out exit conditions and a hysteresis mechanism: when the minimum holding period is met... Subsequently, if the following exit conditions are met and continue to be met... Each sampling period, The number of consecutive holding cycles required to exit the determination is used to determine if the vehicle has exited the heat fade activation state. ;
[0058] ;
[0059] ;
[0060] in, , and These are the exit temperature threshold, exit braking intensity threshold, and exit duration threshold, respectively. After exiting the active state, the vehicle should be in a normal state, not a warning state. Therefore, the above parameters should meet the following requirements: .
[0061] This embodiment uses a hysteresis mechanism that separates the entry threshold (activation threshold) from the exit threshold to avoid the system frequently switching states near the threshold, thereby suppressing the oscillating propagation of risk information in the vehicle queue.
[0062] Step S2.3. Prediction of the conservative upper limit of brake temperature: When thermal degradation assessment is activated, a conservative upper limit prediction of the future short-term brake temperature is made based on the current state to obtain the temperature boundary under the worst operating conditions. Specifically, a thermal inertia time constant is introduced. To replace the traditional fixed equivalent heat capacity This allows the rate of temperature rise to adapt to the actual thermal inertia characteristics of different vehicles (such as brake wear, slope changes, etc.), and the relationship is as follows: ;in, The average heat dissipation coefficient is taken between 50 and 150 W / °C; the brake temperature prediction formula is: ;
[0063] in, : No. A conservative upper limit for the brake temperature at each sampling time; : Filtered current (number) Brake temperature at each sampling time; Predict the number of interval steps. Equal to the predicted length; : No. The vehicle's current thermal inertia time constant can be pre-calibrated to a fixed value using a test bench, or updated in real time using an online adaptive identification method. Average heat dissipation coefficient (W / °C), calibrated offline; Maximum braking power of the vehicle; Current vehicle cooling capacity.
[0064] Physical meaning: Thermal inertia time constant It quantitatively describes the response speed of brake temperature to changes in brake power. The smaller the value, the larger the temperature rise term, and the more conservative the temperature prediction. The larger the value, the more gradual the temperature rise. This improvement allows temperature prediction to adapt to the actual thermal inertia characteristics of different vehicles (such as heavy load, brake wear, and slope changes), without the need for repeated calibration of the equivalent heat capacity. .
[0065] Furthermore, to further improve the accuracy of brake temperature prediction, this embodiment provides an online adaptive identification method for the thermal inertia time constant. This method estimates the brake temperature and braking power based on real-time measurements using a recursive least squares algorithm. Specifically, the online adaptive identification method for thermal inertia time constant includes the following steps: (1) Temperature change prediction model: Based on energy balance, the rate of change of brake temperature can be expressed as: ;
[0066] in, This is the actual braking power. For ambient temperature, The average heat dissipation coefficient, This represents the temperature of the brake after filtering.
[0067] Will Substituting and discretizing, we get the first... The predicted values for the brake temperature at each sampling time are: ;
[0068] in, The predicted temperature of the brake at the (k+1)th sampling time. Let be the braking power at the k-th sampling time.
[0069] (2) Recursive least squares identification algorithm: Define the parameters to be identified Therefore, the temperature change prediction formula can be transformed into a linear regression form: ;
[0070] Among them, the regression quantity .
[0071] The measured temperature change is The RLS algorithm with a forgetting factor is used to update the data. : ;
[0072] ;
[0073] ;
[0074] in, The covariance matrix (initial values can be taken as follows) ), Forgetting factor, This represents the prediction error. Finally, the estimated value of the thermal inertia time constant is obtained: ;
[0075] in, It is a very small positive number to prevent division by zero.
[0076] In this embodiment, sampling is performed only once per sampling period, and the sampling time interval is equal to the sampling period length. Therefore, if the thermal inertia time constant is obtained using an online identification method, the system uses the thermal inertia time constant identified at the beginning of the previous period. calculate.
[0077] Furthermore, to avoid incorrect identification under non-braking conditions, this invention only activates the online adaptive identification of the identifier when the following conditions are met simultaneously: The vehicle is in a braking state; Brake temperature ;in, Pick ; Speed .
[0078] When the vehicle has been idle for an extended period or the temperature has dropped to near ambient temperature, the covariance matrix can be reset. To accelerate the reconvergence.
[0079] This method does not require offline calibration of equivalent heat capacity. It can automatically adapt to changes in vehicle load, brake wear, and different slope conditions; the computational load is extremely small, making it suitable for real-time on-board implementation; it improves the accuracy of temperature prediction, thereby improving the reliability of subsequent calculations of maximum braking deceleration and minimum safe distance.
[0080] Step S2.4. Calculation of braking efficiency factor: Based on the conservative upper limit of brake temperature, the braking performance factor is calculated to quantify the degree of brake thermal fade that weakens braking capability.
[0081] The braking efficiency factor is represented by a piecewise function: ;
[0082] in, Braking efficiency factor; : Thermal decay initiation temperature; Temperatures with a high risk of thermal degradation; : Performance attenuation coefficient; Minimum braking efficiency factor .
[0083] Step S2.5. Calculation of maximum available braking deceleration: Based on the braking efficiency factor, determine the vehicle's braking performance in the first... The maximum braking deceleration at each sampling time, which can be used as braking capacity: ;
[0084] in, Maximum braking deceleration of a vehicle at normal temperature.
[0085] It should be noted that, in this embodiment, the maximum available braking deceleration at the current moment can also be calculated using the braking efficiency factor. The difference is that the braking performance factor is no longer calculated using the predicted braking temperature, but instead uses the current filtered brake temperature. This value represents the maximum braking intensity that the vehicle can safely output under the current degree of heat fade. It will be directly used for subsequent risk margin assessment and as a lower limit constraint for acceleration commands in multi-agent decision-making.
[0086] Step S2.6. Calculation of minimum safe distance for thermal adaptation: This step is based on the filtered brake temperature obtained in step S2.1. The speed of the vehicle obtained in step S1 and relative velocity And the thermal inertia time constant obtained through online identification. Calculate the first Dynamic safe following distance of a vehicle under current hot conditions This distance serves as a soft constraint for the subsequent multi-agent collaborative decision-making module, guiding the controller to actively increase the following distance under thermal decay conditions.
[0087] Specifically, step S2.6 includes the following steps: Step S2.6.1. Calculation of thermal inertia relaxation factor: Define thermal inertia relaxation factor This is used to quantify the relative ability of a vehicle to withstand additional thermal loads in the next sampling period, given the current brake temperature. The calculation formula is: ;
[0088] in, Warning temperature threshold; The brake's thermal saturation temperature, i.e., the braking performance factor has dropped to its minimum value. The temperature at that time.
[0089] Physical meaning: When hour, This indicates that the vehicle is fully capable of withstanding additional heat loads; when the temperature rises to near When the exponential term approaches , Approaching a small positive number indicates that it can only withstand a very small additional heat load. Thermal inertia time constant. The smaller, The more sensitive the system is to temperature, the earlier it will increase the safety distance.
[0090] Step S2.6.2. Calculation of minimum safe distance for thermal adaptation: The traditional fixed minimum safety distance Extended to dynamic thermal adaptive minimum safe distance Its expression is: ;
[0091] in, Minimum safe distance (minimum parking spacing) can be taken as follows: ; Calibration factor, used to control the magnitude of the thermal correction term, with a range of values. ; : Very small positive numbers, avoid division by zero; Relative velocity influence coefficient, range of values ; Relative speed: A positive value indicates that the vehicle is faster than the vehicle in front. Reference speed, acceptable .
[0092] Feature analysis: When (When the engine is cold) ,but It does not affect normal following efficiency; when (When thermal decay is severe) It's very big. The speed of the vehicle increases dramatically, forcing vehicles to maintain a safe following distance, thus reducing the risk of rear-end collisions at the source; vehicle speed The higher the relative speed, or the greater the positive relative speed, the larger the correction term, which conforms to the physical law that a greater safety distance is required under conditions of high speed and rapid deceleration of the vehicle in front.
[0093] Step S3: Risk Status Calculation: This step, based on the maximum available braking deceleration and thermally adaptive minimum safe distance obtained in step S2, and combined with the current car-following state, calculates the... Overall risk margin of a vehicle and risk level .
[0094] Specifically, step S3 includes the following steps: Step S3.1. Assess the braking requirements to avoid a rear-end collision based on the relative motion state: assumed If the speeds of adjacent vehicles remain constant within a given time period, calculate... Minimum deceleration required to avoid a rear-end collision after a certain time. : ; in, For minimum safe distance, , is the adjustment coefficient; when At that time, the required deceleration is approximately 0.
[0095] Step S3.2. Calculation of comprehensive risk margin: Define two risk margin components: Braking capacity margin: ; In this embodiment, n is a fixed value of a system, usually taken as 10.
[0096] Distance margin: .
[0097] The overall risk margin is the minimum of the two: ;
[0098] Step S3.3. Risk Level Mapping: Based on the sign and magnitude of the comprehensive risk margin, three levels of risk (safe, warning, dangerous) are defined for priority adjustment in collaborative decision-making: ;
[0099] in, This represents the risk margin threshold.
[0100] The comprehensive risk margin and risk level output by this invention are directly input into the reward function in step S4, providing a smooth penalty gradient and guiding the reinforcement learning strategy to actively maintain a safety margin when approaching the dangerous boundary.
[0101] Step S4. Construct a multi-agent reinforcement learning framework with centralized training and distributed execution, and train it. In this step, each vehicle in the long downhill vehicle queue is treated as a multiple agent, and a multi-agent collaborative decision-making system is constructed. Let the vehicle set be: ;
[0102] Among them, the Vehicle corresponding to the An intelligent agent.
[0103] State space construction: In the Each sampling time The local observation state of an agent consists of the vehicle operating state, braking thermal state, and safety risk information obtained in steps S1 to S3, and can be represented as: ;
[0104] The global state of a multi-agent system is represented as follows: ;
[0105] Action space definition: Each agent at time Output longitudinal control action: ;
[0106] in, This represents the desired longitudinal acceleration of the vehicle. The joint action of a multi-agent system is represented as: ;
[0107] Reward function construction: In this implementation, maximizing the cumulative reward of the entire queue within the travel time is taken as the training objective. ;
[0108] in, This marks the end of the entire optimization evaluation timeframe. These are weighting coefficients used to adjust the relative importance of each item.
[0109] (1) Safety cost item Security cost item The expression used to quantify the safety costs incurred by a vehicle platoon during following maneuvers due to rear-end collision risk, insufficient thermal safety distance, and excessive brake heat buildup is: ;
[0110] in, At the current sampling time, All are weighting coefficients used to adjust the relative importance of each item; the first term in the formula is the comprehensive risk margin penalty term; when When this occurs, it indicates that the vehicle is already in or about to enter a high-risk state, and this result is proportional to... The penalty is to guide the controller to avoid entering the area. In this embodiment, The dynamic weighting coefficient is determined by the risk level. The calculation yielded: ;
[0111] in, Basic penalty coefficient, Represents the total number of vehicles in the fleet; The second item is the penalty for violating thermal safety distance, which directly penalizes the actual distance between vehicles. Less than thermal self-adaptive safety distance In this scenario, the penalty intensity is inversely proportional to the thermal inertia relaxation factor—the more severe the thermal decay and the smaller the relaxation factor, the higher the cost of violating the safety distance, thus forcing the controller to actively maintain a sufficient distance. In this embodiment, The dynamic weighting coefficients are determined by the thermal inertia relaxation factor. The calculation yielded: ;
[0112] in, The basic penalty coefficient (normal number); It is the thermal inertia relaxation factor. The higher the temperature, the lower the thermal inertia. The closer to 0; For extremely small positive numbers, avoid dividing by zero.
[0113] The third term is the temperature penalty term, used to penalize excessively high vehicle temperatures, an excessively high proportion of high-temperature vehicles, and uneven heat load distribution, suppressing brake overheating and differences in thermal state within the fleet (uneven heat load distribution). Its definition is: ; In the formula: This is the thermal decay exit temperature threshold. This represents the total number of vehicles in the fleet. This is the average brake temperature across all vehicles. For preset weight parameters, For the first The temperature of the vehicle's brakes after temperature filtering.
[0114] (2) Efficiency cost: A squared penalty is applied only when the actual following distance is greater than the most economical following distance. In this way, the system can safely increase the distance without being wrongly penalized during thermal decay, thus ensuring traffic efficiency under normal operating conditions.
[0115] ;
[0116] in, The most economical following distance (minimum following distance).
[0117] (3) Comfort cost: penalize the magnitude of acceleration and its rate of change to improve driving comfort.
[0118] ;
[0119] in, These are the weighting coefficients.
[0120] In this embodiment, a multi-agent reinforcement learning framework with centralized training and distributed execution is used to learn and optimize the vehicle platoon coordination control strategy, balancing the feasibility of distributed execution with the performance of global coordination control. During the training phase, a centralized value evaluation network optimizes the strategy using global state and joint action evaluation strategies. During the execution phase, each agent makes independent decisions based on local observations, automatically allocating the optimal safe distance and longitudinal acceleration according to the real-time thermal state and following environment of each vehicle, thereby effectively suppressing the propagation of rear-end collision risk in the platoon under thermal decay conditions such as long downhill slopes and heavy loads.
[0121] In a preferred embodiment, the multi-agent reinforcement learning framework employs the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm. Each agent constructs a policy network (Actor) to output continuous control actions based on local observation states; simultaneously, a centralized value evaluation network (Critic) is constructed to evaluate the long-term rewards of the joint actions of the multi-agents.
[0122] Regarding the policy network (Actor) design, in this embodiment, each agent constructs a policy network to generate longitudinal control commands for the vehicle based on local observation states; the input to the policy network is the sequence of local observation states within the current time and historical time windows. ;
[0123] in, Indicates the length of the historical time window. Indicates the first The agent in the th... The local observation status at each sampling time includes information such as vehicle operating status, brake temperature, risk margin, and available braking capacity.
[0124] To characterize the cumulative effect of braking temperature changes and thermal decay risk over time, the strategy network preferably adopts a neural network structure that includes a temporal feature extraction layer.
[0125] In a preferred embodiment, the policy network It includes: an input layer, a temporal feature extraction layer, several fully connected hidden layers, and an output layer.
[0126] Since the changes in vehicle braking temperature and the thermal fade process have obvious time-cumulative characteristics, the temporal feature extraction layer preferably uses a Long Short-Term Memory (LSTM) network to extract time-related features from the state sequence; several fully connected hidden layers are used to perform nonlinear mapping on the temporal features, and the number of layers and neurons can be set according to actual needs; the output layer is used to generate continuous control actions, and its output is the original control command (desired longitudinal acceleration) of the policy network. : ;
[0127] in, These are the policy network parameters.
[0128] Preferably, the output layer uses the hyperbolic tangent function (Tanh) to map the network output to a preset range.
[0129] With the above structure, the policy network can generate longitudinal control actions that meet the requirements of coordinated control based on the current state and historical evolution information of the vehicle.
[0130] Critic Network Structure Design: In this embodiment, a centralized Critic network is constructed to evaluate the long-term returns of the current control strategy based on the global state and joint actions of the multi-agent system. The inputs to the Critic network include: Multi-agent systems at all times Global state: ; Multi-agent systems at all times Joint actions: ; in, This indicates the number of agents in the vehicle queue.
[0131] In a preferred embodiment, the value evaluation network includes: an input layer, several fully connected hidden layers, and an output layer; wherein the input layer is used to receive global state and joint action information and perform feature concatenation; the fully connected hidden layers are used to extract the state coupling relationship and action coordination relationship between multiple agents; and the output layer is used to output the corresponding global action value. ;
[0132] in, These are the parameters for value assessment networks.
[0133] In this embodiment, the global action value function is used to characterize the expected cumulative discount return that can be obtained by performing the joint action in the current global state.
[0134] Preferably, the output layer adopts a linear output structure to accommodate the unbounded nature of the value function.
[0135] Through the above structure, the value evaluation network can perform an overall evaluation of the coordinated control behavior of multiple agents, providing a basis for the optimization of the policy network.
[0136] Model training process: In this embodiment, based on the policy network and value evaluation network, a multi-agent reinforcement learning method is used to train the vehicle platoon coordination control strategy. The process includes the following steps: (1) State interaction and sample generation: During the training process, each agent is constantly... Control actions are generated through a policy network based on the local observation state: ;
[0137] The actions of all agents constitute a joint action: ;
[0138] After applying the combined action to the vehicle queuing system, the global state at the next moment is obtained. and corresponding instant rewards And form state transition samples: ;
[0139] Experience replay and sample sampling: State transition samples are stored in the experience replay pool. During training, small batches of samples are randomly sampled from the experience replay pool for subsequent network parameter updates, in order to reduce sample correlation and improve training stability.
[0140] Target value construction: Based on the sampled data, construct the target value for the value evaluation network. : ;
[0141] in, The next-time reference action generated by the target policy network. As a discount factor, For target value assessment network, The parameters are used to evaluate the target value of the network.
[0142] Value assessment network update: Based on the current value assessment network output: ;
[0143] The parameters of the value assessment network are updated by minimizing the error between it and the target value: ;
[0144] Policy network update: Based on the updated value assessment network, the policy network is optimized using its evaluation results of the current policy, so that the output control actions can obtain higher long-term returns.
[0145] Specifically, the policy network parameters are updated by maximizing the output of the value evaluation network, with the optimization objective being: ;
[0146] Furthermore, to improve the stability of the training process, a target policy network is introduced. and target value assessment network The parameters are obtained from the current network parameters through a delayed update method.
[0147] Preferably, a soft update method is used to update the policy network parameters and the value evaluation network parameters.
[0148] Iterative training: Repeat the above steps until the policy network converges, thereby obtaining a vehicle platoon coordination control policy that meets the requirements of safety, efficiency and comfort.
[0149] It should be noted that the descriptions of the specific structure of the policy network and the value evaluation network, the number of network layers, the selection of activation functions, the specific training process of centralized training and distributed execution, and the method of calculating the target value in this embodiment are only used to clearly illustrate the implementation idea of this method and are not intended to be the only limitation of this method.
[0150] Without departing from the core idea of this method, those skilled in the art can replace, improve, and extend the network structure, training strategy, etc., based on existing reinforcement learning and multi-agent learning techniques. All such equivalent replacements or improvements should fall within the protection scope of this method.
[0151] Step 5. Each vehicle outputs longitudinal control actions based on the current state space using the trained policy network. Considering that the control commands output by the policy network may exceed the physical capabilities of the vehicle actuators, constraint mapping is performed on the original control commands to ensure their executability.
[0152] Specifically, the original control command is restricted to a control range determined by the vehicle's current available braking and driving capabilities to obtain the final control command: ;
[0153] in: This represents the maximum available braking deceleration after taking into account the effects of thermal fade. This indicates the vehicle's maximum driving acceleration.
[0154] Through the above constraint mapping, it is ensured that the control command always meets the braking capacity constraint and driving capacity constraint of the vehicle under the current heat fade state.
[0155] Example 2: like Figure 2 As shown, this embodiment provides a longitudinal vehicle platoon coordination control system that considers brake thermal fade. The longitudinal vehicle platoon coordination control method that considers brake thermal fade described above includes a state perception module, a data processing module, a thermal fade assessment module, a thermal inertia time constant estimation module, a brake temperature prediction module, a brake performance factor estimation module, a maximum available braking deceleration estimation module, a thermal adaptive minimum safe distance estimation module, a comprehensive risk margin and risk level assessment module, a strategy network, and a constraint module. The status perception module is used to acquire the vehicle's operating status information and its following relationship information with the vehicle in front; The data processing module is used to filter the collected brake temperature. The thermal fade assessment module is used to determine whether the vehicle is in a thermal fade activation state. The thermal inertia time constant estimation module is used to estimate the thermal inertia time constant based on the real-time collected brake temperature and braking power using a recursive least squares identification algorithm. The brake temperature prediction module predicts the future brake temperature based on the vehicle's maximum braking power, current heat dissipation power, current brake temperature, and thermal inertia time constant. The braking performance factor estimation module is used to calculate the braking performance factor based on the brake temperature. The maximum available braking deceleration estimation module is used to determine the current or future maximum available braking deceleration of the vehicle based on the braking efficiency factor. The thermally adaptive minimum safe distance estimation module first quantifies the relative ability of a vehicle to withstand additional thermal load in the next sampling period at the current brake temperature using a thermal inertia relaxation factor, and then extends the fixed minimum safe distance to a dynamic thermally adaptive minimum safe distance based on the thermal inertia relaxation factor. The comprehensive risk margin and risk level assessment module evaluates the braking capacity margin and distance margin based on the maximum available braking deceleration and the thermal adaptive minimum safe distance, combined with the current car-following state. The minimum value of the two is taken as the comprehensive risk margin and mapped to the risk level. The policy network (multi-agent reinforcement learning module) is used to output longitudinal control actions based on the current state space; The constraint module is used to limit the longitudinal control actions of each vehicle within a control range determined by the vehicle's current available braking and driving capabilities, thereby obtaining the final control command.
[0156] The present invention also provides an electronic device, comprising: one or more processors and a memory; wherein the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the above-described longitudinal vehicle queuing coordination control method considering brake heat fade.
[0157] The present invention also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the above-described longitudinal vehicle platoon coordination control method considering brake heat fade.
[0158] Those skilled in the art will understand that all or part of the functions of the various methods / modules in the above embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved.
[0159] In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the programs can also be stored in storage media such as servers, other computers, disks, optical discs, flash drives, or portable hard drives. They can be downloaded or copied to the memory of the local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be implemented.
[0160] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of protection described in the claims.
Claims
1. A longitudinal vehicle platoon coordination control method considering brake fade, characterized in that, The method includes the following steps: Step S1. Obtain the vehicle's running status information and its following relationship information with the vehicle in front in the vehicle queue; Step S2. Based on the acquired data, determine whether the vehicle is in a thermal fade activation state. If so, predict the conservative upper limit of the future brake temperature based on the current state, and calculate the maximum available braking deceleration and the minimum safe distance for thermal adaptation. Step S3. Based on the maximum available braking deceleration and the thermal adaptive minimum safe distance, combined with the current car-following state, evaluate the braking capacity margin and distance margin, take the minimum value of the two as the comprehensive risk margin, and map it to the risk level; Step S4. Construct a multi-agent reinforcement learning framework with centralized training and distributed execution. The state space includes vehicle operating state, braking thermal state, and safety risk information; the action space is the longitudinal control action; and the reward function is: ; in, This is a safety cost item used to quantify the safety costs incurred by a vehicle platoon during following another vehicle due to rear-end collision risk, insufficient thermal safety distance, and brake heat accumulation. For efficiency cost; For comfort-related costs, This marks the end of the entire optimization evaluation timeframe. These are the weighting coefficients; Step 5. Each vehicle outputs longitudinal control actions based on the current state space using the trained policy network; then, the longitudinal control actions of each vehicle are restricted to the control range jointly determined by the vehicle's current available braking and driving capabilities to obtain the final control command, thereby realizing longitudinal vehicle platoon coordinated control.
2. The longitudinal vehicle platoon coordination control method considering brake fade according to claim 1, characterized in that, In step S1, the vehicle's operating status information includes its speed, acceleration, brake temperature, braking system pressure signal, and continuous braking time; the following relationship information with the preceding vehicle includes the actual distance between the two vehicles and the speed of the preceding vehicle.
3. The longitudinal vehicle platoon coordination control method considering brake fade according to claim 1, characterized in that, Step S2 first filters the brake temperature signal. Then, based on the filtered brake temperature, the current continuous braking duration, and the vehicle's acceleration, it determines whether the vehicle has entered a warning state. If it has entered a warning state, it determines whether the vehicle has entered a thermal fade activation state based on the filtered brake temperature, the vehicle's acceleration, and the trigger accumulation counter. The trigger accumulation counter increments when the warning conditions are met and decrements when the warning conditions are not met. If the vehicle has entered a thermal fade activation state and the activation state is maintained for multiple cycles, a conservative upper limit prediction of the future brake temperature is made based on the current state, and the current maximum available braking deceleration and the thermal adaptive minimum safe distance are calculated.
4. The longitudinal vehicle platoon coordination control method considering brake fade according to claim 1, characterized in that, In step S2, when making a conservative upper limit prediction of the future brake temperature, a thermal inertia time constant is introduced to replace the fixed equivalent heat capacity. The formula for predicting brake temperature is: ; in, : No. A conservative upper limit for the brake temperature at each sampling time; : Current brake temperature after filtering; Predict the number of interval steps. Equal to the predicted length; : No. The vehicle's current thermal inertia time constant is estimated using a recursive least squares identification algorithm based on real-time measured brake temperature and braking power. Average heat dissipation coefficient; Maximum braking power of the vehicle; Current vehicle cooling capacity; Use period length; Then, based on a conservative upper limit for brake temperature, a braking performance factor is calculated. This braking performance factor employs a piecewise function to quantify the degree to which brake fade weakens braking capability. Based on the braking performance factor, the vehicle's braking performance is determined in the [missing information - likely a specific timeframe or period]. The maximum available braking deceleration at each sampling time.
5. The longitudinal vehicle platoon coordination control method considering brake fade according to claim 1, characterized in that, In step S2, when calculating the thermally adaptive minimum safe distance, the thermal inertia relaxation factor is first used. The relative ability of a vehicle to withstand additional thermal load in the next sampling period at the current brake temperature is expressed as: ; in, : Current brake temperature after filtering; : No. The vehicle's current thermal inertia time constant; Warning temperature threshold; The brake's thermal saturation temperature, i.e., the braking performance factor has dropped to its minimum value. The temperature at that time; Use period length; Then, the minimum safe distance will be fixed based on the thermal inertia relaxation factor. Extended to dynamic thermal adaptive minimum safe distance Its expression is: ; in, Calibration coefficient, range of values ; Avoid positive numbers that are divided by zero; Relative velocity influence coefficient, range of values ; : Relative speed to the vehicle in front; Reference speed Speed of this vehicle; Speed of the vehicle in front.
6. The longitudinal vehicle platoon coordination control method considering brake fade according to claim 1, characterized in that, In step S3, the minimum deceleration required to avoid a rear-end collision is assessed based on the relative motion state, and the braking capacity margin is assessed based on the minimum deceleration required to avoid a rear-end collision and the maximum available braking deceleration; the distance margin is assessed based on the thermally adaptive minimum safe distance and the actual distance between the current vehicle and the vehicle in front, and the minimum value of the two is taken when calculating the comprehensive risk margin.
7. A longitudinal vehicle platoon coordination control method considering brake fade according to claim 1, characterized in that, The vehicle operating state in the state space mentioned in step S4 includes the vehicle speed. The acceleration of this vehicle Relative speed to the vehicle in front and the actual distance between this vehicle and the vehicle in front. Braking thermal condition includes the temperature of the brake after filtering. Thermal inertia relaxation factor ; Safety risk information includes thermally adaptive minimum safe distance. Comprehensive risk margin Current maximum available braking deceleration ; The security cost term in the reward function Represented as: ; in, At the current sampling time, For fixed weighting coefficients, The dynamic weighting coefficient is determined by the risk level. Calculated; The dynamic weighting coefficients are determined by the thermal inertia relaxation factor. Calculated; The temperature cost term is expressed as: ; In the formula: This is the thermal decay exit temperature threshold. This represents the total number of vehicles in the fleet. This is the average brake temperature across all vehicles. These are preset weight parameters.
8. A longitudinal vehicle platoon coordination control method considering brake fade according to claim 7, characterized in that, Efficiency cost ;in, For the most economical following distance; Comfort cost ;in, These are the weighting coefficients. Indicates the first i The change in vehicle acceleration within adjacent sampling periods. This represents the acceleration of the vehicle in front.
9. A longitudinal vehicle platoon coordination control method considering brake fade according to claim 8, characterized in that, The multi-agent reinforcement learning framework described in step S4 employs a multi-agent deep deterministic policy gradient algorithm.
10. A longitudinal vehicle platooning coordination control system considering brake fade, characterized in that, This system is used to implement a longitudinal vehicle platoon coordination control method considering brake thermal fade as described in any one of claims 1 to 9, including a state perception module, a data processing module, a thermal fade assessment module, a thermal inertia time constant estimation module, a brake temperature prediction module, a brake performance factor estimation module, a maximum available braking deceleration estimation module, a thermal adaptive minimum safe distance estimation module, a comprehensive risk margin and risk level assessment module, a strategy network, and a constraint module. The status perception module is used to acquire the vehicle's operating status information and its following relationship information with the vehicle in front; The data processing module is used to filter the collected brake temperature. The thermal fade assessment module is used to determine whether the vehicle is in a thermal fade activation state. The thermal inertia time constant estimation module is used to estimate the thermal inertia time constant based on the real-time collected brake temperature and braking power using a recursive least squares identification algorithm. The brake temperature prediction module predicts the future brake temperature based on the vehicle's maximum braking power, current heat dissipation power, current brake temperature, and thermal inertia time constant. The braking performance factor estimation module is used to calculate the braking performance factor based on the brake temperature. The maximum available braking deceleration estimation module is used to determine the current or future maximum available braking deceleration of the vehicle based on the braking efficiency factor. The thermally adaptive minimum safe distance estimation module first quantifies the relative ability of a vehicle to withstand additional thermal load in the next sampling period at the current brake temperature using a thermal inertia relaxation factor, and then extends the fixed minimum safe distance to a dynamic thermally adaptive minimum safe distance based on the thermal inertia relaxation factor. The comprehensive risk margin and risk level assessment module evaluates the braking capacity margin and distance margin based on the maximum available braking deceleration and the thermal adaptive minimum safe distance, combined with the current car-following state. The minimum value of the two is taken as the comprehensive risk margin and mapped to the risk level. The policy network is used to output longitudinal control actions based on the current state space; The constraint module is used to limit the longitudinal control actions of each vehicle within a control range determined by the vehicle's current available braking and driving capabilities, thereby obtaining the final control command.
Citation Information
Patent Citations
Regenerative and conventional brake integrated controller and its control based on ABS for automobile
CN101073992A
Vehicle driving control method and device and storage medium
CN111216723A