A four-wheel distributed brake-by-wire cooperative control method based on learning enhancement

CN122481739BActive Publication Date: 2026-09-22JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610976686.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-22
Estimated Expiration
2046-07-02

AI Technical Summary

Benefits of technology

针对执行器响应时滞受多因素耦合影响难以建立精准解析模型,构建深度学习驱动的时滞动态映射机制,实现了时滞特性从静态估算向动态预测的范式转变,为线控制动控制框架提供准确的执行器响应特性信息,优化了现有控制框架下执行器响应时滞导致的系统控制能力受限的难题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122481739B_ABST
    Figure CN122481739B_ABST
Patent Text Reader

Abstract

The application discloses a learning enhancement-based four-wheel distributed brake-by-wire cooperative control method, and belongs to the technical field of vehicle intelligent chassis and active safety control. The method comprises three parts. The first part is precise characterization of the dynamic hysteresis characteristics of the brake-by-wire system executor, mainly including EMB dynamic hysteresis characteristic testing and data processing and EMB hysteresis characteristic characterization based on a linear attention recurrent neural network. The second part is construction of a brake-by-wire system control framework, mainly including explicit integration of the hysteresis characteristics of the executor and establishment of a distributed model predictive robust control method. The third part is a distributed model predictive robust control acceleration solving method. The method provides optimization initial values for control system solving, reduces searching radius and iteration frequency, and improves the real-time response capability of the control system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of vehicle intelligent chassis and active safety control technology, specifically involving a four-wheel distributed drive-by-wire braking cooperative control method based on learning reinforcement. Background Technology

[0002] With the deep integration of intelligent vehicles and electronic chassis control, the application of brake-by-wire systems has endowed the braking pressure at each wheel with decoupled control capabilities, significantly enhancing the vehicle's control flexibility in the chassis domain. Under emergency braking conditions, intelligent adjustment of braking pressure at each wheel can effectively shorten braking distance and improve vehicle safety. However, during braking, the actuators of brake-by-wire systems exhibit highly nonlinear and dynamic time-delay characteristics, which greatly restricts the real-time response of the chassis system and the steady-state accuracy of multi-wheel pressure collaborative control. Therefore, developing a multi-wheel distributed brake-by-wire collaborative control method that can accurately characterize and compensate for the dynamic characteristics of the actuators is crucial for fully unlocking the safety potential of vehicles.

[0003] To fully utilize the braking potential of vehicle brake-by-wire systems, especially to ensure vehicle safety during emergency braking, current research methods can be categorized into three main types: rule-based, learning-based, and optimization-based. Among these, logic thresholds are widely used in rule-based strategies. This type of method relies primarily on expert experience and pre-defined offline mapping tables, achieving high response speeds in embedded controllers with limited computing resources. However, it struggles to capture transient dynamic changes in actuators during emergency braking and exhibits poor adaptability to different operating conditions. Utilizing deep reinforcement learning (DRL) to construct intelligent controllers is a core research direction in learning-based strategies. This type of method can directly learn the optimal control law from complex data through end-to-end feature mapping. However, this method suffers from bottlenecks such as lack of interpretability and difficulty in verifying training convergence. Furthermore, its black-box nature makes it difficult for the controller to explicitly integrate the physical constraints and state boundaries of the chassis actuators, posing a risk of accidents during training scenarios. Model predictive control (MPC), as a model-based optimization control method, possesses the ability to model and coordinate multi-input multi-output systems and is suitable for handling high-dimensional coupling and nonlinear constraints. In online braking control, Multi-Wheel Braking Computation (MPC) can simultaneously optimize the dynamic performance of multi-wheel braking and makes it possible to embed the nonlinear dynamic coupling relationship of the vehicle into the optimization framework. Therefore, the MPC control framework is a promising multi-wheel braking cooperative control strategy. On the one hand, it is easier to handle hard constraints and physical boundary conditions to ensure sufficient safety than learning-based approaches; on the other hand, it achieves operational applicability through online rolling optimization. However, some issues still need further research in the application of vehicle emergency braking control to ensure the robustness and effectiveness of MPC and fully utilize the braking potential of the brake-by-wire system.

[0004] As a model-based optimization algorithm, the accuracy of vehicle state changes directly affects the accuracy of MPC's prediction of future system states and the reliability of control decision optimization. In particular, actuator delay in the brake-by-wire system leads to a significant time-domain bias in the control commands at the execution level. This actuator lag causes model mismatch in MPC, which in turn affects the accuracy of trajectory prediction and the stability of the closed-loop control during emergency braking. Currently, researchers mainly optimize the actuator delay problem from two aspects: time / frequency domain compensation representation and controller robustness design. Compensation representation methods often employ constant time delay bias compensation, Smith predictors, or equivalent clock synchronization strategies, aiming to predict the spatiotemporal bias of command execution in the time domain and correct it through feedforward or feedback methods, thus explicitly canceling the delay effect. However, this method relies on idealized prior models and fixed calibration parameters, making it difficult to characterize the nonlinear characteristics exhibited by the brake-by-wire system during emergency braking transients. In robust controller design research, strategies such as Robust Model Predictive Control (RMPC), Stochastic Model Predictive Control (SMPC), and Tube Model Predictive Control (Tube-MPC) are widely adopted. These strategies characterize dynamic time delays as system uncertainties with statistical properties or set constraints, and ensure state convergence and hard constraint safety of the closed-loop system under time-delay disturbances by solving the optimal control law online under the worst-case performance index. However, robust design comes at the cost of sacrificing the system's dynamic response performance, resulting in overly conservative control strategies under urgent braking scenarios. Therefore, accurately characterizing and compensating for the dynamic time delays of actuators in brake-by-wire systems in real time has become a key technology that urgently needs to be mastered in brake-by-wire control systems.

[0005] During emergency braking, the solution process for multidimensional control variables in brake-by-wire control systems under high-dimensional constraints is typically accompanied by high computational load. However, rapid solution and high-frequency control updates are crucial for ensuring the vehicle's response speed and stability in emergency braking scenarios. To improve the computational efficiency and real-time performance of system control, strategies combining offline and online approaches are widely used. Among these, Explicit Model Predictive Control (EMPC) significantly improves real-time control performance by transforming the online optimization problem into an offline pre-computation process, obtaining the optimal control input through a lookup table in real-time control. However, actuator delay characteristics are difficult to incorporate into polyhedral partitioning, and the discontinuity of control laws caused by region switching makes it difficult for brake-by-wire systems to dynamically balance control accuracy and output smoothness under emergency braking. Deep reinforcement learning transforms the solution process for complex control variables into an inference model of offline policy training and online function mapping, effectively improving online solution efficiency. However, this method carries unpredictable failure risks when dealing with extreme long-tail conditions, making it difficult to guarantee the safety of the entire vehicle in emergency braking scenarios. In contrast, Distributed Model Predictive Control (DMPC), through task decoupling and parallel computation, not only retains the physical intuitiveness of the MPC architecture in handling hard constraints and time delay compensation, but also provides an effective path to alleviate the pressure of centralized computing. However, the efficiency of traditional DMPC is highly limited by the iteration frequency required for consensus among subsystems; in high real-time scenarios, excessive iteration counts and communication latency often negate the performance advantages of distributed architecture. Therefore, how to significantly reduce the iterative burden of DMPC by optimizing subsystem partitioning strategies while ensuring braking coordination accuracy has become a key breakthrough in balancing algorithmic and real-time solution efficiency, thereby fully realizing the potential of brake-by-wire systems.

[0006] Therefore, existing technologies still have the following shortcomings: First, it is difficult to perform high-precision modeling and real-time characterization of the nonlinear, time-varying, and time-delay characteristics of actuators in brake-by-wire systems; second, it is difficult to effectively compensate for model mismatch caused by actuator dynamic time delays within a model predictive control framework; and third, it is difficult to achieve rapid solution and online high-frequency updates for complex optimization problems while ensuring the accuracy and safety constraints of multi-wheel cooperative braking. Based on this, it is necessary to propose a new multi-wheel distributed brake-by-wire cooperative control method to improve control accuracy, real-time performance, and vehicle safety in emergency braking scenarios. Summary of the Invention

[0007] To address the three shortcomings of existing technologies, this invention provides a learning-enhanced four-wheel distributed brake-by-wire cooperative control method.

[0008] The method is specifically as follows: S1. Steps for characterizing the dynamic hysteresis characteristics of the actuator in a brake-by-wire system: including EMB dynamic hysteresis characteristic testing and data processing, and EMB hysteresis characteristic characterization based on a linear attention recurrent neural network. S2. Steps for constructing the control framework of the brake-by-wire system: including explicit integration of actuator hysteresis characteristics and obtaining the optimal braking torque at each wheel end by establishing a distributed model predictive robust control method; S3. Steps to accelerate the solution of distributed model prediction robust control: Provide optimized initial values ​​for distributed model prediction by using the deep deterministic policy gradient algorithm, improve the efficiency of distributed model prediction, and thus improve the real-time response capability of the linear braking system.

[0009] Furthermore, the EMB dynamic hysteresis characteristic testing and data processing specifically involves: using the EMB dynamic characteristic testing platform to quantitatively test the actuator hysteresis behavior, extracting key characterization parameters that reflect the EMB hysteresis characteristics, including: response delay time, single input increment, and output deviation during loading and unloading processes; preprocessing, windowing, and feature organization of the raw data collected from the experiment to construct an EMB hysteresis characteristic sample set for time-series modeling.

[0010] Furthermore, the EMB hysteresis characteristic representation based on the linear attention recurrent neural network specifically involves selecting multidimensional variables representing input path information, internal electromechanical state information, and external operating condition information as the input feature vector of the linear attention recurrent neural network. The output vector is , Indicates time For a moment Input features First, the query vector, key vector, and value vector are obtained through linear mapping; then, a kernel function mapping is used. Linearize the attention calculation to obtain the result at time 1000. The attention state is updated, and thus a linear attention output is obtained; further, a cyclic update is introduced to recursively memorize historical dynamic information to obtain the hidden state at the current moment. The hidden state is processed by the output layer of a linear attention recurrent neural network. Hysteresis characteristics characterization results mapped to EMB When training a linear attention recurrent neural network, physical constraints, output rate of change constraints, and output magnitude constraints are introduced.

[0011] Furthermore, the explicit integration of actuator hysteresis characteristics specifically involves: in the EMB braking system control, selecting the state matrix and control matrix, and establishing corresponding state-space discrete expressions using a seven-DOF vehicle dynamics model; constructing a constrained optimal braking torque distribution model from these state-space discrete expressions; and characterizing the hysteresis characteristics of the EMB. Explicitly integrated into the controller, at any time The equivalent hysteresis compensation amount of EMB is output by a linear attention recurrent neural network based on the current control input, actuator state, and historical response information. ,according to Calculate the equivalent braking torque actually applied to the vehicle at each wheel end. ;Will Substituting the discrete state-space expression, we construct an optimal braking torque allocation model that considers the effect of actuator hysteresis.

[0012] Furthermore, the distributed model predictive robust control method is established as follows: based on the optimal braking torque allocation model that considers the effect of actuator hysteresis, a distributed model predictive control architecture is adopted to divide the vehicle braking control into multiple mutually coupled wheel-end subsystems. The optimal braking torque is solved for each subsystem, while the solution objective of each subsystem is limited to maintain the optimality of the overall objective, thereby obtaining the optimal braking torque of each wheel end.

[0013] Furthermore, providing initial optimization values ​​for distributed model prediction and solving through the deep deterministic policy gradient algorithm specifically involves: constructing the state space of the deep deterministic policy gradient algorithm. and action space The state space consists of vehicle dynamics deviation states, driving commands, and road environment, while the action space consists of the braking torque at each wheel end. A reward function for the deep deterministic policy gradient algorithm is constructed. , ,in This represents the predictive control cost function of the distributed model. This represents additional incentives; the deep deterministic policy gradient algorithm includes an Actor network and a Critic network, employing a target network soft update mechanism to recursively update the parameters of both the Actor and Critic networks; in each control cycle, the Actor network generates a suboptimal control sequence based on the real-time acquired vehicle state parameters. And directly map it to the initial value of the distributed model predictive control numerical solver.

[0014] The beneficial effects of the method described in this invention are as follows: To address the challenge of establishing accurate analytical models for actuator response delays due to the coupling effect of multiple factors, a deep learning-driven dynamic delay mapping mechanism was constructed. This mechanism enables a paradigm shift from static estimation to dynamic prediction of delay characteristics, providing accurate actuator response characteristic information for the linear control braking framework and optimizing the problem of limited system control capabilities caused by actuator response delays under existing control frameworks.

[0015] To address the impact of multi-actuator delay characteristics on control system stability during emergency braking scenarios, a robust DMPC control strategy with explicit integrated response delay is constructed. By explicitly embedding the predicted delay information into the discretized state space, the dynamic state transition relationship including the hysteresis effect is reconstructed, achieving cross-temporal and spatial alignment between control commands and actuator responses. This suppresses closed-loop oscillations in the control system caused by model mismatch and ensures deterministic safety of vehicle control under extreme scenarios.

[0016] To address the computational efficiency issue in online iterative optimization of DMPC, a fast DMPC solution mechanism based on deep reinforcement learning-guided optimization path is constructed. Deep reinforcement learning provides the DMPC solver with an initial search vector located in the neighborhood of the optimal solution. While ensuring the feasibility of recursive solution of the DMPC strategy and physical boundary constraints, this significantly reduces the search radius and iteration frequency, achieving a dynamic balance between real-time performance and deterministic safety in the control system. Attached Figure Description

[0017] Figure 1 This is a diagram illustrating the overall architecture of the method described in this embodiment of the invention. Figure 2 This is a configuration diagram of a distributed drive electric vehicle in an embodiment of the present invention; Figure 3 This is a hardware-in-the-loop platform architecture diagram in an embodiment of the present invention; Figure 4 This is a graph showing the four-wheel slip ratio data obtained using the method described in this invention in an embodiment of the invention; Figure 5 This is a graph showing the four-wheel slip ratio data using the benchmark method in this embodiment of the invention. Detailed Implementation

[0018] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0019] To address the problems of existing technologies mentioned in the background section, this invention proposes a learning-reinforcement-based distributed brake-by-wire cooperative control method, aiming to ensure the control robustness and real-time response performance of vehicles under emergency braking conditions, thereby guaranteeing the active safety of the entire vehicle. First, considering the time delay characteristics of actuator response, a deep learning-based time delay prediction model is constructed to achieve accurate representation and real-time mapping of actuator response time delay. Based on this, the predicted dynamic time delay information is explicitly embedded into the state-space equation of the DMPC (Distributed Controlled-by-Wire) system. An active compensation mechanism is constructed to optimize the model mismatch problem caused by actuator response lag, improving the robustness of the control system under extreme braking scenarios. Simultaneously, addressing the computational load problem caused by excessive iteration frequency in the DMPC distributed architecture under high-frequency control, a deep reinforcement learning-based optimization strategy is introduced. Offline training establishes a mapping relationship between operating condition features, vehicle state information, and optimal control quantities, providing an initial search vector within the globally optimal neighborhood for the online DMPC iteration process. This method guides the optimization path by optimizing the starting point of gradient search in the solution space, which significantly reduces the frequency of algorithm iterations and improves computational efficiency, while strictly ensuring the recursive feasibility and physical boundary constraints of the control algorithm, thus ensuring the vehicle's response capability and deterministic safety under emergency braking conditions.

[0020] Combined with appendix Figure 1 The proposed method is explained in three parts. The first part is the accurate characterization of the dynamic hysteresis characteristics of the actuator in the brake-by-wire system, mainly including EMB dynamic hysteresis characteristic testing and data processing, and EMB hysteresis characteristic characterization based on linear attention recurrent neural networks. The second part is the construction of the control framework for the brake-by-wire system, mainly including the explicit integration of actuator hysteresis characteristics and the establishment of a distributed model predictive robust control method. The third part is the accelerated solution method for distributed model predictive robust control, which reduces the search radius and iteration frequency by providing initial optimization values ​​for the control system solution, thereby improving the real-time response capability of the control system.

[0021] 1. Learning representation of dynamic hysteresis characteristics of brake-by-wire actuators 1.1 EMB Dynamic Hysteresis Characteristics Test and Data Processing The configuration diagram of the EMB (Electro-Mechanical Brake) brake-by-wire system is shown below. Figure 2 As shown, MCU represents a microcontroller unit, CAN_H represents the high-level signal line of the CAN bus, and CAN_L represents the low-level signal line of the CAN bus.

[0022] To accurately characterize the dynamic hysteresis characteristics of brake-by-wire actuators under emergency braking conditions, this invention takes the electromechanical brake (EMB) as the research object. Firstly, dynamic response data of the actuator under different input excitations and operating states are obtained through an EMB test platform, and key characterization parameters reflecting the EMB hysteresis characteristics are extracted. Based on this, a linear attention recurrent neural network (LARNN) is used to establish a nonlinear dynamic mapping relationship between the EMB input command, internal state, and output response, achieving high-precision learning and characterization of the EMB's dynamic hysteresis characteristics, providing model support for hysteresis compensation and predictive control in subsequent controllers.

[0023] Because the EMB (Electrical Motor Module) contains multiple complex factors, including motor electromagnetic dynamics, transmission backlash in the reduction mechanism, ball screw friction, nonlinearity of the brake pad-brake disc contact, and return elasticity, its output response typically exhibits significant time-varying delay characteristics. Therefore, this invention uses an EMB dynamic characteristic testing platform to quantitatively test the actuator's hysteresis behavior. During the test, step input, ramp input, sinusoidal sweep input, and pulse sequence input are applied to the EMB to cover the hysteresis response process under different excitation rates and target amplitudes.

[0024] The response delay time is defined as the time difference between the moment the controller issues the input command and the moment when the EMB output shows a valid change, as shown in Formula 1.

[0025] (1) In the formula, The moment when the EMB output first exceeds the preset threshold. To control the timing of command issuance, For response delay time.

[0026] The single input increment is used to characterize the net change in EMB output under one control cycle or one standard input pulse, as shown in Equation 2.

[0027] (2) In the formula, For the first The moment the input action begins. For the first The moment after the input action ends and stabilizes This represents the net increment of clamping force corresponding to this input action. The clamping force (braking clamping force) output by the EMB actuator.

[0028] Define the output deviation between the loading and unloading processes under the same input amplitude. As shown in Formula 3.

[0029] (3) In the formula, The input is the output value corresponding to the rising phase. The input is the output value corresponding to the descent phase. This represents the EMB control input at the current moment, which can be a control command such as the target motor current, target displacement, target clamping force, or target braking torque. The specific form is determined by the input interface definition of the EMB control system.

[0030] To establish a dataset suitable for dynamic learning representations, the raw data collected in the experiment were preprocessed, windowed, and feature-organized to construct a sample set of EMB hysteresis characteristics for time-series modeling. First, the acquired signals were time-synchronized, noise-filtered, and outlier-removed. Sensor signals with different sampling frequencies were uniformly resampled to obtain a multivariate synchronization sequence with a fixed sampling period. Then, a sliding time window was used to construct the sequence samples. Furthermore, the variables were normalized, and variables were set... The maximum and minimum values ​​are respectively and Then the normalized expression is shown in Formula 4.

[0031] (4) In the formula, This is the normalized data.

[0032] 1.2 Characterization of EMB Hysteresis Characteristics Based on Linear Attention Recurrent Neural Network Linear Attention Recurrent Neural Networks (LARNNs) are deep learning architectures that combine global attention mechanisms with recursive temporal processing capabilities, enabling them to accurately represent the non-analytical dynamics of EMBs. The core advantage of LARNNs lies in reducing the complexity of attention operations through kernel function approximation techniques, thus meeting the real-time requirements of automotive systems.

[0033] Considering the significant historical dependence, state coupling, and time-varying nonlinearity of the hysteresis characteristics of EMB actuators, multidimensional variables representing input path information, internal electromechanical state information, and external operating condition information are selected as input feature vectors. The output vector is As shown in Formula 5-6.

[0034] (5) (6) In the formula, This is the control input at the current moment; the control input can be the target motor current, the target actuated displacement, or the target clamping force. The length of the history window; These are internal EMB state variables, which can be selected as motor speed, motor angle, lead screw displacement, push rod displacement, current clamping force, brake disc speed, and output error at the previous moment. The rate of change of internal state of the EMB; Indicates actuator temperature; Indicates the supply voltage; This is the equivalent hysteresis compensation amount for EMB.

[0035] For time Input features First, the query vector, key vector, and value vector are obtained through linear mapping, as shown in Formula 7.

[0036] (7) In the formula, , and These represent the mapping matrices for queries, keys, and values, respectively.

[0037] To avoid the computational cost that increases quadratically with the sequence length in standard attention mechanisms, this invention employs kernel function mapping. Linearize the attention calculation at time 1000. The attention state update is shown in Equation 8-9.

[0038] (8) (9) In the formula, The cumulative value state matrix, This is the cumulative vector of the normalization term.

[0039] This yields the linear attention output, as shown in Equation 10.

[0040] (10) In the formula, To prevent tiny positive numbers with a denominator of zero.

[0041] Building upon this, to further enhance the model's ability to characterize EMB transient hysteresis and state abrupt changes, this invention further introduces a cyclic update module to recursively memorize historical dynamic information. Let... This represents the hidden state of the previous loop module, containing the EMB actuator hysteresis response, clamping force variation trend, and historical input path information within the previous control cycle; the linear attention output will be... Compared to the previous hidden state The shared input loop update module can obtain the current hidden state. As shown in Formula 11.

[0042] (11) In the formula, Update the weight matrix for the state. For bias terms, It is a non-linear activation function.

[0043] The hysteresis characteristics of mapping the hidden state to the EMB by the output layer are shown in Equation 12.

[0044] (12) In the formula For LARNN output, it represents the equivalent hysteresis compensation of the output EMB at the current moment; This is the output layer weight matrix. This is the output layer bias term.

[0045] To improve the fitting accuracy of LARNN to the physical hysteresis law of EMB, physical constraints are further introduced during network training, so that the model not only fits the data, but also satisfies the basic dynamic law of EMB. During the loading phase, as the input command continues to increase, the EMB output braking force should show an increasing trend. Therefore, a monotonicity constraint term is defined, as shown in Equation 13.

[0046] (13) In the formula The change in clamping force at adjacent time points is predicted using a linear attention recurrent neural network (LARNN).

[0047] To avoid drastic fluctuations in the prediction results that do not conform to the dynamic laws of the actuator, an output change rate constraint is introduced, as shown in Equation 14.

[0048] (14) To ensure that the prediction results meet the physical output boundaries of the EMB actuator, an output amplitude constraint is further introduced to keep the predicted clamping force within the range of the minimum and maximum clamping forces that the actuator can achieve, as shown in Formula 15.

[0049] (15) In the formula and These represent the maximum and minimum clamping forces achievable by the EMB, respectively.

[0050] Taking into account both the data fitting error and physical constraints, the total loss function of the model is defined as shown in Equation 16-17.

[0051] (16) (17) In the formula, , , and These are the weighting coefficients. For the prediction error term, This is the actual value of the output data.

[0052] By introducing physical constraints, LARNN establishes a nonlinear input-output mapping relationship of EMB while keeping the prediction results physically reasonable and dynamically smooth, thereby enhancing the model's generalization and accuracy.

[0053] 2. Distributed Model Predictive Robust Control Based on Hysteresis Characteristics of Explicitly Integrated Actuators Based on the learning and characterization results of the EMB dynamic hysteresis characteristics established in the first part, this invention further constructs a distributed model prediction robust control method that explicitly integrates actuator hysteresis characteristics. This method introduces the EMB dynamic hysteresis characteristics into the EMB braking force control process in the form of an equivalent hysteresis compensation, thereby correcting the model mismatch problem caused by the actual lag in actuator response and improving the accuracy, robustness, and real-time performance of multi-wheel braking cooperative control.

[0054] 2.1 Explicit Integration of Actuator Hysteresis Characteristics In the online control of the EMB braking system, the slip ratio value of the left front wheel is selected. Left rear wheel slip ratio value Right front wheel slip ratio value Right rear wheel slip ratio value Lateral velocity and yaw rate This is a state variable. The control variable is the torque value of the left front wheel. Right front wheel torque value Left rear wheel torque value Right rear wheel torque value The corresponding state-space discrete expression is established based on the seven-degree-of-freedom vehicle dynamics model, as shown in Equation 18.

[0055] (18) In the formula and Here is the state transition matrix. for State matrix at time 1 , , for Control matrix at time 1 , .

[0056] Under emergency braking conditions, due to the drastic fluctuations in braking torque and the complexity of road surface adhesion conditions, wheel slippage can cause significant nonlinear abrupt changes, even leading to wheel lock-up. Therefore, to prevent directional instability and braking performance degradation caused by complete wheel lock-up, and to improve directional stability and road surface adhesion utilization during braking, active braking control of the entire vehicle is implemented. This aims to fully utilize the longitudinal adhesion potential of the current road surface and achieve an optimal balance between braking distance and stability. Based on a model predictive control architecture, a constrained optimal braking torque allocation model is constructed. The solution for the braking torque in each rolling cycle is transformed into the following quadratic optimization problem: the solution for the optimal braking torque can be transformed into minimizing the quadratic cost function in each rolling cycle. As shown in formulas 19-20.

[0057] (19) (20) In the formula, Indicates the current moment Starting from the first The predicted state vector at each prediction step size The value, Indicates the first The control input increment at each prediction step (described directly below) The change in the braking torque of the four wheels between adjacent prediction steps (at any given moment) can be expressed as: ; This serves as a reference value for following the target. The penalty coefficient for changes in the control quantity; and Step size The weight matrix for each step in the process; and The upper and lower limits of the four-wheel slip ratio are mainly determined by the target slip ratio range, the road adhesion coefficient, and the anti-lock braking safety threshold; for lateral velocity and yaw rate, the upper and lower limits are mainly determined by the vehicle's lateral stability boundary, the desired yaw response, and the lateral stability constraint. and The parameters are determined by the physical output capability of the EMB actuator, the contact characteristics of the brake disc and friction pads, the available adhesion between the tire and the road surface, the vertical load at the wheel end, and the safety limits of the braking system.

[0058] Because there is a significant dynamic hysteresis between the EMB actuator outputting the target braking torque from the controller and the actual wheel-end clamping force and equivalent braking torque, directly using the nominal model for prediction can easily lead to a deviation between the predicted trajectory and the actual response trajectory, thus affecting control accuracy and system stability. Therefore, this invention explicitly integrates the output results of the LARNN hysteresis characterization model established in the first part into the controller.

[0059] At any moment The equivalent hysteresis compensation amount of EMB is output by LARNN based on the current control input, actuator status and historical response information, as shown in Formula 21. (twenty one) The equivalent braking torque actually applied to the vehicle by each wheel end is shown in Formula 22. (twenty two) Substituting this into the prediction model (Equation 18), we construct the state-space equation that considers the effect of actuator hysteresis, as shown in Equation 23. (twenty three) The state-space equations are then substituted into the optimal braking torque allocation model (Formula 19) to obtain the optimal braking torque allocation model that takes into account the effect of actuator hysteresis.

[0060] By characterizing the EMB braking pressure hysteresis using LARNN, ​​it is equivalently mapped to a negative disturbance in the torque dimension, thereby achieving proactive correction of model mismatch in the control-by-wire system. Simultaneously, through this mechanistic model and data-driven error compensation method, the solvability of the controller's mathematical structure is maintained while improving the model's accuracy in representing the actual behavior of the actuator.

[0061] 2.2 Distributed Model Prediction Robust Control Method During emergency braking, the EMBs at each wheel end of a vehicle have both localized needs for tire adhesion utilization and are coupled with the vehicle's yaw rate and lateral velocity. To achieve real-time control and vehicle braking stability while reducing the complexity of solving high-dimensional global optimization problems, this invention employs a distributed model predictive control architecture, dividing the vehicle braking control problem into multiple interconnected wheel-end subsystems. Within these subsystems, to fully utilize braking efficiency, greater emphasis is placed on slip ratio control, with the objective function shown in Equations 24-25.

[0062] (twenty four) (25) In the formula, For the first Subsystems in The slip ratio at any given time; For the first Subsystems in The change in braking torque at any given time; The target slip ratio; These are the weighting coefficients; and This is the weight coefficient matrix.

[0063] Indicates the first Minimize the quadratic cost function of each subsystem; this is the th subsystem under distributed model predictive control. Local optimal control problem of individual wheel terminal system ( (Corresponding to four independent wheels: left front, right front, left rear, and right rear) The core is to achieve accurate tracking of single-wheel slip rate under the premise of satisfying physical hard constraints, maximize the utilization rate of single-wheel road surface adhesion, and reduce the computational complexity of global optimization.

[0064] Meanwhile, in order to maintain the stability of the whole vehicle, the solution objectives of each subsystem must maintain the optimality of the overall objective, as shown in Equations 26-27.

[0065] (26) (27) In the formula, for Lateral velocity at time; for Reference lateral velocity at the given moment; for Yaw velocity at time t; for Reference yaw rate at the given moment; and This is the weight coefficient matrix. and These are the weighting coefficients. This represents the cost function for predictive control in a distributed model.

[0066] The above text This only represents the cost function for four rounds. However, This represents the local cost of integrating the four wheel terminal systems. To achieve synergistic optimization of braking performance and vehicle safety, at the cost of overall vehicle lateral stability.

[0067] It is determined by the vehicle's sideslip stability boundary, tire-road adhesion conditions, longitudinal vehicle speed, safe range of sideslip angle, and vehicle lateral stability requirements. It is determined by the vehicle's expected yaw response, tire-road adhesion limit, longitudinal vehicle speed, steering input, and lateral stability boundary.

[0068] By solving formulas 24 and 26, the optimal braking torque at each wheel end is obtained, which can maximize the utilization rate of road adhesion under emergency braking conditions, ensure that the vehicle quickly converges to a safe state under extreme conditions, and achieve a dynamic balance between braking performance and vehicle stability.

[0069] 3. Accelerated Solution Method for Distributed Model Predictive Robust Control The explicit integration of EMB hysteresis, as described in Part 2, can effectively improve the control accuracy and robustness of brake-by-wire systems. However, during high-frequency rolling optimization, each wheel subsystem still needs to repeatedly solve the local optimization problem in each sampling period, while information exchange between subsystems is required to satisfy the overall optimization problem. In distributed model predictive control architectures, computational complexity increases exponentially with the prediction time domain and control dimension. Especially under complex conditions such as emergency braking, low adhesion, or sudden adhesion changes, the optimal control solution may experience significant jumps between adjacent time points, making the numerical solver sensitive to initial values. This leads to an increase in the number of online optimization iterations, a longer solution time, and a decrease in control real-time performance.

[0070] To improve the computational efficiency of the control system, this invention further introduces deep reinforcement learning to optimize the fast solution efficiency of distributed predictive control. Deep Deterministic Policy Gradient (DDPG), as a deep reinforcement learning algorithm, possesses powerful nonlinear function approximation and multi-dimensional feature extraction capabilities. It can deeply mine real-time vehicle state information, construct an implicit mapping from the state space to the optimal control space, and thus accurately predict and approximate the globally optimal control sequence as the initial value for the DMPC iterative solution, significantly improving the convergence efficiency of the solution.

[0071] The state space consists of the vehicle dynamics deviation state, driving commands, and road environment, as shown in Equation 28.

[0072] (28) In the formula, The yaw rate is angular velocity. For lateral velocity, Side slip angle, Longitudinal velocity, For the slip ratio of each wheel end, Road surface adhesion coefficient, This refers to the brake pedal opening. For deep reinforcement learning agents in the first The state vector received at each control moment is used to characterize the current vehicle dynamics deviation state, driving commands, and road environment.

[0073] The action space is the braking torque at each wheel end, as shown in Formula 29.

[0074] (29) For deep reinforcement learning agents in the first The action vector output at each control moment is used to characterize the predicted four-wheel braking torque control quantity under the current state.

[0075] To ensure consistency between the learning strategy and the optimization objective of the distributed model predictive control, the reward function... Designed as a distributed model predictive control cost function The negative mapping, i.e. ,in As an additional incentive.

[0076] The DDPG algorithm employs an Actor-Critic dual-network architecture for learning. Critic networks are used to fit deterministic policies from the state space to the optimal control space. A value function used to evaluate state-action pairs.

[0077] Let be the deterministic policy function represented by the Actor network, used to establish the nonlinear mapping relationship between the state vector s and the action vector a. The mapping from state vector s to action vector a is defined in equations 28 and 29. and . For the parameters of the Actor network, This represents a nonlinear mapping relationship from the state space to the action space.

[0078] The state-action value function represented by the Critic network is used to evaluate the long-term cumulative reward that can be obtained after performing action a in state s. These are the parameters for the Critic network.

[0079] During the offline training phase, the Critic network parameters are updated by minimizing the Bellman residuals. Its loss function As shown in Formula 30.

[0080] (30) In the formula, For experience replay pool, The temporal difference (TD) target value is calculated based on the target network; the target network is part of DDPG and includes a target Actor network and a target Critic network, used to stabilize the TD target value; the temporal difference (TD) target value is the temporal difference target value of the current Critic network, used as... The training labels enable the current Critic network output to gradually approach the target value, which is composed of immediate rewards and future discounted returns.

[0081] To perform the action The instant reward obtained afterward Indicates the replay of experiences from the pool The expected value is calculated from the training samples obtained by sampling. In actual training, this expected value is usually approximated by the mean of mini-batch samples. The role of the experience replay pool is to store historical interaction samples, break the temporal correlation between consecutive samples, improve sample utilization, and enhance the stability of the network training process; This is a discount factor used to adjust the proportion of future rewards in the current value assessment; The target Actor network is determined based on the state at the next time step. The output predicts the action at the next moment. For the target Actor network parameters; For the target Critic network, the value estimate of the state-action pair for the next time step; The parameters of the target Critic network.

[0082] For Actor networks, based on the Deterministic Policy Gradient Theorem, parameters are updated along the ascending direction of the Q-value gradient using a chain rule. To maximize the cumulative expected return As shown in Formula 31.

[0083] (31) In the formula, Let be the cumulative expected return objective function corresponding to the Actor network; This indicates that the objective function is relative to the Actor network parameters. The gradient; This represents the expectation of the sampled states in the experience replay pool; This indicates that the Critic network output value is relative to the action. The gradient; This indicates that the Actor network outputs actions relative to its network parameters. The gradient.

[0084] To improve the stability of the training process, this invention employs a target network soft update mechanism to recursively update the parameters of the target Actor network and the target Critic network. , The Actor network will eventually converge to an approximate function of the DMPC optimal control law, i.e. ; For calibration coefficients, .

[0085] Therefore, the Actor network trained using the DDPG algorithm for deep fitting of the globally optimal control law in the offline phase is deployed in the online controller. In each control cycle, the network rapidly generates a suboptimal control sequence based on real-time acquired vehicle state parameters. And directly map it to the initial value of the distributed model predictive control numerical solver, that is This strategy significantly reduces the number of iterations required for convergence of nonlinear optimization problems by utilizing high-precision initial values, thereby improving the online computational efficiency and real-time response capability of distributed model predictive control systems while ensuring control accuracy.

[0086] To effectively verify the validity of this invention, a hardware-in-the-loop platform based on the dSPACE real-time controller was built, such as... Figure 3 As shown. Compared with the benchmark strategy, this invention significantly reduces the efficiency of online optimization while effectively suppressing slip rate fluctuations and system oscillations caused by actuator dynamic time delays, greatly improving the steady-state accuracy of multi-wheel cooperative control and the safety performance of the vehicle during emergency braking.

[0087] The four-wheel slip ratio data of the method of the present invention are shown in the figure below. Figure 4 As shown in the figure, the four-wheel slip ratio data of the benchmark method are plotted as follows: Figure 5 As shown. By Figure 4 and Figure 5 The comparison results show that the baseline method exhibits significant slip ratio fluctuations. Due to the hysteresis characteristics of the EMB actuator, the control system model cannot accurately reflect actual state changes, leading to poor control performance and a persistently high slip ratio. However, the learning-based compensation method of this invention demonstrates better slip ratio performance and can fully utilize the vehicle's braking performance. The computation time data for the proposed method and the baseline method are shown in Table 1. Table 1:

[0088] As shown in Table 1, by comparing the results of the method of the present invention with those of the benchmark method, it can be seen that the present invention has a shorter computation time and significantly improved computational efficiency.

Claims

1. A learning-reinforcement-based four-wheel distributed brake-by-wire cooperative control method, characterized in that, The method is specifically as follows: S1. Steps for characterizing the dynamic hysteresis characteristics of the actuator in a brake-by-wire system: including EMB dynamic hysteresis characteristic testing and data processing, and EMB hysteresis characteristic characterization based on a linear attention recurrent neural network. S2. Steps for constructing the control framework of the brake-by-wire system: including explicit integration of actuator hysteresis characteristics and obtaining the optimal braking torque at each wheel end by establishing a distributed model predictive robust control method; S3. Steps to accelerate the solution of distributed model prediction robust control: Provide optimized initial values ​​for distributed model prediction by using the deep deterministic policy gradient algorithm, improve the efficiency of distributed model prediction, and thus improve the real-time response capability of the linear braking system.

2. The four-wheel distributed drive-by-wire braking cooperative control method based on learning reinforcement according to claim 1, characterized in that, The EMB dynamic hysteresis characteristic test and data processing are as follows: The actuator hysteresis behavior is quantitatively tested through the EMB dynamic characteristic test platform, and key characterization parameters that can reflect the EMB hysteresis characteristics are extracted. These parameters include response delay time, single input increment, and output deviation during loading and unloading processes. The raw data collected from the experiment are preprocessed, windowed, and feature-organized to construct an EMB hysteresis characteristic sample set for time-series modeling.

3. The four-wheel distributed drive-by-wire braking cooperative control method based on learning reinforcement according to claim 2, characterized in that, The specific method for characterizing EMB hysteresis characteristics based on a linear attention recurrent neural network involves selecting multidimensional variables representing input path information, internal electromechanical state information, and external operating condition information as the input feature vector of the linear attention recurrent neural network. The output vector is , Indicates time For a moment Input features First, the query vector, key vector, and value vector are obtained through linear mapping; then, a kernel function mapping is used. Linearize the attention calculation to obtain the result at time 1000. The attention state update quantity is obtained, and thus the linear attention output is obtained; Furthermore, a cyclical update is introduced to recursively memorize historical dynamic information, thereby obtaining the hidden state at the current moment. The hidden state is processed by the output layer of a linear attention recurrent neural network. Hysteresis characteristics characterization results mapped to EMB When training a linear attention recurrent neural network, physical constraints, output rate of change constraints, and output magnitude constraints are introduced.

4. The four-wheel distributed drive-by-wire braking cooperative control method based on learning reinforcement according to claim 3, characterized in that, The explicit integration of actuator hysteresis characteristics specifically involves: in the EMB braking system control, selecting the state matrix and control matrix, and establishing corresponding state-space discrete expressions using a seven-DOF vehicle dynamics model; constructing a constrained optimal braking torque distribution model from these state-space discrete expressions; and characterizing the hysteresis characteristics of the EMB. Explicitly integrated into the controller, at any time The equivalent hysteresis compensation amount of EMB is output by a linear attention recurrent neural network based on the current control input, actuator state, and historical response information. ,according to Calculate the equivalent braking torque actually applied to the vehicle at each wheel end. ;Will Substituting the discrete state-space expression, we construct an optimal braking torque allocation model that considers the effect of actuator hysteresis.

5. The four-wheel distributed drive-by-wire braking cooperative control method based on learning reinforcement according to claim 4, characterized in that, The specific steps of establishing a distributed model predictive robust control method are as follows: Based on the optimal braking torque allocation model that considers the effect of actuator hysteresis, a distributed model predictive control architecture is adopted to divide the vehicle braking control into multiple mutually coupled wheel-end subsystems. The optimal braking torque is solved for each subsystem, while the solution objective of each subsystem is limited to maintain the optimality of the overall objective, thereby obtaining the optimal braking torque of each wheel end.

6. The four-wheel distributed drive-by-wire braking cooperative control method based on learning reinforcement according to claim 5, characterized in that, Providing initial values ​​for distributed model prediction by using a deep deterministic policy gradient algorithm specifically involves: constructing the state space of the deep deterministic policy gradient algorithm. and action space The state space consists of vehicle dynamics deviation states, driving commands, and road environment, while the action space consists of the braking torque at each wheel end. A reward function for the deep deterministic policy gradient algorithm is constructed. , ,in This represents the predictive control cost function of the distributed model. This represents additional incentives; the deep deterministic policy gradient algorithm includes an Actor network and a Critic network, employing a target network soft update mechanism to recursively update the parameters of both the Actor and Critic networks; in each control cycle, the Actor network generates a suboptimal control sequence based on the real-time acquired vehicle state parameters. And directly map it to the initial value of the distributed model predictive control numerical solver.

7. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-6.

8. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Complete vehicle stability control method for composite braking system

    CN120863574A

  • Distributed driving vehicle stability control method and system and medium

    CN121989915A