A front-rear axle dual-motor driving and braking control method for a pure electric commercial vehicle
By combining learning-based model predictive control and deep reinforcement learning, a high-precision vehicle speed prediction model was established to optimize the energy distribution of dual-motor pure electric vehicles. This solved the problems of real-time energy management and battery health monitoring under complex operating conditions, achieving efficient energy distribution and extended battery life.
Patent Information
- Application Number
- CN202510048155.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-01-13
AI Technical Summary
Existing technologies in energy management of dual-motor driven pure electric vehicles struggle to achieve efficient energy distribution under complex and variable driving conditions, and traditional methods have shortcomings in real-time response and battery health status monitoring.
By combining the learning-based model predictive control (L-MPC) method with high-precision vehicle speed prediction algorithm and deep reinforcement learning (DDPG), a vehicle speed prediction model based on BiTCN-BiGRU is established to optimize energy distribution strategy and achieve adaptive adjustment to future operating conditions and dynamic monitoring of battery health status.
It improves the system's adaptability and real-time response capability under complex operating conditions, extends battery life, and enhances energy utilization efficiency and driving stability.
Smart Images

Figure CN119502719B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of new energy vehicle technology, and in particular to a method for controlling the braking of a pure electric commercial vehicle driven by dual motors on the front and rear axles. Background Technology
[0002] With increasing global focus on energy efficiency and environmental protection, new energy vehicles, especially pure electric vehicles, have gradually become a core research direction in the transportation sector. Against this backdrop, dual-motor driven pure electric commercial vehicles have garnered widespread attention due to their advantages in improving vehicle power performance, driving stability, and energy efficiency management. Compared to traditional single-motor drive systems, dual-motor drive systems on both the front and rear axles enable more flexible torque distribution, significantly improving the driving experience and effectively enhancing energy utilization. However, the complexity of dual-motor systems also places higher demands on the design of efficient drive and braking control strategies.
[0003] Energy management strategy is one of the key technologies for achieving efficient and energy-saving drive and braking control in pure electric vehicles. Traditional energy management methods, such as rule-based control, model predictive control (MPC), and deep reinforcement learning (DRL), have all achieved certain results under specific conditions. However, rule-based control strategies exhibit insufficient flexibility in actual driving, making it difficult to maintain the vehicle's optimal operating state under complex and changing conditions. While traditional model predictive control methods can predict and optimize future operating conditions, their computational complexity is high when facing long prediction periods, making it difficult to respond in real time to the rapid dynamic changes of multi-motor systems. Furthermore, although various speed prediction algorithms have been studied for vehicle speed prediction, there is still significant room for improvement in prediction accuracy. In addition, deep reinforcement learning methods based on artificial intelligence have shown superiority in handling nonlinear and multivariable problems, but due to their lack of systematicity in model predictive control, they often exhibit poor adaptability to changes in actual operating conditions.
[0004] Qi Xing et al. proposed an optimization strategy for torque distribution in dual-motor electric vehicles with front and rear axles. The paper outlines two torque distribution strategies: one based on motor efficiency optimization, which optimizes the torque distribution between the front and rear axle motors to ensure they operate within their optimal efficiency range, thereby improving the overall efficiency of the motor system; and the other based on battery efficiency optimization, which focuses on optimizing battery energy utilization and reducing unnecessary battery losses to increase driving range. However, these two strategies conflict, as optimal motor efficiency may lead to decreased battery efficiency, and vice versa. To balance these two factors, the researchers employed a multi-objective particle swarm optimization (MOPSO) algorithm, which can find a suitable torque distribution strategy while simultaneously considering both motor and battery efficiency.
[0005] First, the multi-objective particle swarm optimization (MOPSO) algorithm has high computational complexity. Since the algorithm needs to optimize multiple objectives simultaneously, real-time calculations may increase the computational burden on the system, leading to slower processing speeds. This can cause delays in real-world driving environments requiring rapid response, impacting the vehicle's real-time performance. Second, this technology only allocates torque based on instantaneous energy consumption, lacking the ability to predict future vehicle operating states, which can easily lead to efficiency degradation. Finally, this technology does not adequately monitor battery health and cannot dynamically adjust the energy allocation strategy according to the system's actual capabilities. Under high power demands or complex driving scenarios, this may result in excessively rapid battery health degradation. This application, however, utilizes a predictive model to predict future vehicle speed changes and optimizes energy allocation based on these future operating conditions. Furthermore, this application considers battery life degradation in the reinforcement learning reward function, achieving multi-objective optimization of both vehicle energy consumption and battery health.
[0006] Wu, Changcheng, and others proposed an energy management strategy for dual-motor pure electric vehicles based on self-attention deep reinforcement learning. To address the challenges of multiple operating conditions and complex road conditions, this technical solution employs a Deep Deterministic Policy Gradient (DDPG) algorithm, combined with a Gated Recurrent Unit (GRU) and a self-attention (SA) mechanism to optimize the energy management strategy. The GRU helps the controller acquire historical state information, improving the training effect of the strategy; the self-attention mechanism enhances the adaptability of the control system to uncertain environments, ensuring optimal energy allocation under varying operating conditions. The main optimization objectives of this strategy include reducing the battery aging rate and reducing overall vehicle energy consumption. By setting multi-objective optimization indices, the controller can adjust the torque output and energy consumption of the front and rear axle motors in real time, ensuring that overall energy consumption is reduced while extending battery life.
[0007] While some progress has been made in energy management of dual-motor driven pure electric vehicles, this deep reinforcement learning-based energy management strategy struggles to efficiently handle complex changes in real-world driving conditions. Although the strategy performs well under specific conditions, it cannot dynamically adjust to the optimal energy management strategy when dealing with complex and changing conditions, leading to decreased energy allocation efficiency and hindering system efficiency improvement and lifespan extension. This application, however, combines the L-MPC method with a high-precision vehicle speed prediction algorithm within the MPC framework, employing the real-time decision-making capabilities of reinforcement learning to achieve intelligent energy allocation for electric vehicles.
[0008] To overcome the above problems, this application proposes a learning-based model predictive control (L-MPC) method, which combines reinforcement learning with model predictive control to achieve the complementary advantages of the two methods. Summary of the Invention
[0009] The purpose of this invention is to address the problems existing in the background technology by proposing a drive and braking control method for a pure electric commercial vehicle driven by dual motors on both the front and rear axles. The invention also aims to provide an energy management strategy combining learning model predictive control (L-MPC) to achieve intelligent energy allocation for pure electric vehicles driven by dual motors on both the front and rear axles, thereby effectively improving the system's energy utilization efficiency, extending battery life, and ultimately enhancing the vehicle's overall economy. This invention aims to overcome the shortcomings of existing technologies in terms of adaptability, real-time performance, and multi-objective optimization, enabling pure electric vehicles to maintain efficient and energy-saving drive and braking control even under complex operating conditions.
[0010] Specifically, MPC struggles with complex systems and long-term optimization problems, while DRL can handle the uncertainties of complex systems and long-term scale optimization problems. On the other hand, DRL lacks adaptability to different driving conditions and has poor interpretability, leading to a significant decline in control performance when faced with new driving conditions. In contrast, the MPC framework is highly interpretable and can obtain local optima using speed predictors. Therefore, combining the two aims to achieve an efficient energy management strategy in complex driving environments. Based on this background, this application proposes an energy management strategy for a dual-motor driven pure electric commercial vehicle, combining the L-MPC method to achieve intelligent energy distribution for the electric vehicle. This strategy not only improves the energy efficiency of the dual-motor system but also significantly enhances the vehicle's adaptability and operational stability under multiple operating conditions, which is of great significance for extending the lifespan of the powertrain.
[0011] The following is a specific embodiment of the present invention:
[0012] A method for controlling the braking of a dual-motor driven pure electric commercial vehicle on both front and rear axles includes the following specific steps:
[0013] S1. Obtain historical operating condition data of dual-motor pure electric vehicles;
[0014] S2. Establish a vehicle speed prediction model based on BiTCN-BiGRU;
[0015] S3. Set the network parameters of the vehicle speed prediction model and train the prediction model;
[0016] S4. Adjust the network parameters based on the prediction accuracy to obtain a high-precision vehicle speed prediction model.
[0017] S5. Establish a model of a dual-motor pure electric vehicle, including a longitudinal dynamics model, a battery model, and a motor model.
[0018] S6. Establish an improved DDPG agent model based on truncated triple Q network and learning rate scheduling;
[0019] S7. Set up the state, action, and reward functions for the improved DDPG agent model;
[0020] S8. Based on the training dataset, a learning model predictive control energy management model combining a vehicle speed prediction model and an improved DDPG agent model is trained.
[0021] S9. Validate the energy management model based on learning model predictive control using a new operating condition dataset;
[0022] S10. Embed the model trained in step S9 into the vehicle controller, and calculate the distribution coefficient of the front and rear axle torque in real time to control the vehicle operation based on the current vehicle speed, battery charge SOC, and battery health status SOH.
[0023] Preferably, the historical vehicle operating condition data obtained in step S1 is the historical vehicle speed.
[0024] Preferably, in the vehicle speed prediction model trained in step S4, the input data is current and historical vehicle speed information, and BiTCN performs bidirectional processing on the input data; BiTCN captures the global dependencies in the sequence through bidirectional information flow;
[0025] BiGRU processes the features extracted by BiTCN and combines forward and backward information flows to model bidirectional dependencies of sequences.
[0026] Preferably, the power system of the dual-motor pure electric vehicle model in step S5 includes a first motor, a second motor, a reducer, a clutch, a power battery, a battery management system (BMS), a motor control unit, and a DC-DC converter.
[0027] Among them, the power battery is the power source of the entire system, providing electric energy to drive the vehicle;
[0028] The Battery Management System (BMS) is used to monitor battery status.
[0029] DC-DC converters are used to convert the high-voltage electrical energy from the battery into low-voltage electrical energy suitable for other systems in the vehicle.
[0030] The first motor and the second motor drive the tires on the front axle and the rear axle, respectively; the two motors are connected to the battery and the motor controller, respectively.
[0031] The motor control unit is responsible for controlling the operation of the motor and adjusting its speed and output torque.
[0032] A speed reducer is used to reduce the output speed of a motor while increasing the output torque to meet the rotational requirements of the wheels.
[0033] The clutch is used to disconnect the motor from the tire drive.
[0034] Preferably, the required driving torque for the entire vehicle in step S5 is calculated using longitudinal dynamics:
[0035]
[0036] in The conversion factor for rotating mass; For vehicle quality; For driving acceleration; It is the slope angle, and the rolling resistance coefficient is... The acceleration due to gravity is air drag coefficient Indicates; wheel radius is The speed is The windward area is Without considering the effect of slope, therefore Then the above equation becomes:
[0037]
[0038] The two motors meet the vehicle's required torque through torque coupling, based on the torque distribution factor. The torque calculations for the front and rear motors are as follows:
[0039]
[0040] In the formula, This indicates the output torque of the front motor. This indicates the output torque of the rear motor; additionally, when Only the rear motor provides power, when Only the front motor provides power; when This enables front and rear drive, and the two motors rotate at the same speed.
[0041] Model the motor; the motor's efficiency differs between driving and regenerative braking states, and the motor's output power is calculated as follows:
[0042]
[0043] also, The efficiency of the motor is represented by an efficiency map, which is used for lookup. ;
[0044] Modeling of power batteries includes equivalent internal resistance model and battery aging model;
[0045] In the equivalent internal resistance model, the battery's state of charge (SOC) and power are derived from the following formula:
[0046]
[0047] in, It is the battery's output power. It is the current in the circuit. It is open-circuit voltage. It is the battery's internal resistance. It's the battery capacity. It is the initial SOC;
[0048] Secondly, an aging model for the battery is constructed; the battery health (SOH) formula is expressed as follows:
[0049]
[0050] in, It is the equivalent number of cycles before the battery pack reaches the end of its lifespan. This is the discharge rate. The above formula is modified into a discrete form, expressed as follows:
[0051]
[0052] in, This refers to the current duration; the influence of discharge rate and internal temperature is calculated using an empirical model of capacity loss based on the Arrhenius equation, expressed by the following formula:
[0053]
[0054] in, It is the percentage of capacity loss. Indicates pre-exponential factor, It is the ideal gas constant, which is equal to 8.314 J / (mol•K). It is a power-law factor equal to 0.55. Indicates the throughput per hour; The activation energy is expressed in J / mol, and the formula is as follows:
[0055]
[0056] when When the battery level drops to 20%, the on-board power battery pack reaches the end of its lifespan. The formulas for the ampere-hour throughput and the equivalent number of cycles before the power battery pack reaches the end of its lifespan are as follows:
[0057]
[0058]
[0059] The battery health is calculated by combining the given current, temperature, and battery dynamics; in the above formula... This represents the average temperature inside the battery.
[0060] At any given moment, the powertrain components satisfy the power balance equation:
[0061]
[0062] in, It is the battery's output power. This represents the output power of the first motor. This represents the output power of the second motor.
[0063] Preferably, the DDPG network structure in step S6 consists of four networks: the current Actor network, the target Actor network, the current Critic network, and the target Critic network.
[0064] The current Actor network is used to generate policies. The current Critic network generates Q-values to evaluate the current policy; the experience replay pool stores data samples generated from each interaction between the agent and the environment. And it learns by random small-batch sampling;
[0065] The current Actor network is based on its current state. Select the action at the current moment. It is used to interact with the environment to generate the state for the next moment. and reward value The target Actor network is responsible for determining the next state. Select next action ;
[0066] Because the weights of the target network are periodically updated from the current network, the target Q-value can be temporarily fixed during training, making learning more stable; the target value of Q-value:
[0067]
[0068] in, It is the reward value at the current time step t. , which is the discount factor, representing the impact of future rewards on long-term returns, and its value ranges from [0, 1]. Generated for the target Critic network value; The policy function of the target Actor network Its network weight; The weights of the target Critic network;
[0069] The current goal of the Critic network is to minimize the loss function, which is used during network training:
[0070]
[0071] In the formula, K: the number of samples in a batch (i.e., mini-batch size), It is the target Q-value calculated through the target Critic network; The Q-value generated by the current Critic network; These are the current weights of the Critic network;
[0072] The current goal of the Actor network is to maximize the expected return, i.e.
[0073]
[0074] In the formula, It is the policy function of the current Actor network;
[0075] The current weight gradient of the Actor network is:
[0076]
[0077] It is the gradient of the Critic network's evaluation of the value of an action, measuring the value of the action. How do small changes affect The value; It is an Actor network strategy. The gradient describes how the Actor network adjusts according to the current state. Adjust parameters The selected actions are optimized; through the chain rule, the gradient of the Actor network is ultimately determined by the product of the value gradient of the Critic network and the policy gradient of the Actor, which guides the Actor network to update its weights so that the selected actions can better improve the cumulative revenue.
[0078] Update the weights of the current Actor network using an optimizer. and the weights of the current evaluation network Then, the weights of the two target networks are updated using a soft update method, that is:
[0079]
[0080] In the formula, These are soft update parameters.
[0081] Preferably, in step S6, a truncated triple Q-learning strategy is adopted to obtain a more accurate Q-value estimate;
[0082] The objective function expression is as follows:
[0083]
[0084] In the formula, and Representing different network parameters and Q output below; The next state is...
[0085] It is the set of possible actions; mean represents the mean operation on the truncated Q value.
[0086] In step S7, the CyclicLR learning rate scheduler is introduced to adjust the learning rate and periodically decay it.
[0087] The defined reward function is as follows:
[0088]
[0089] Furthermore, the boundaries of the learning process are defined as follows:
[0090]
[0091] In the formula, and Indicates the upper and lower limits of battery SOC. and Indicates the upper and lower limits of battery health status. This indicates the upper limit of the first motor's speed. This indicates the upper limit of the second motor's speed. This indicates the upper limit of the torque of the first motor. This indicates the upper limit of the torque of the second motor.
[0092] Preferably, in step S8, assuming the vehicle speed prediction time domain is 3 seconds, and the vehicle speed is v at the current time k, the historical vehicle speeds at times k-3, k-2, and k-1, as well as the current vehicle speed at time k, are input into the vehicle speed prediction module. This module outputs the predicted vehicle speeds at times k+1, k+2, and k+3. Then... The vehicle speed information is input into the learning-based predictive control agent module for training, and the vehicle speed is used... Battery SOC and SOH are used to update the state of the dual-motor power system in the prediction time domain:
[0093]
[0094] Obtain the optimal action at time t Then, using the state variables at time k and optimal action Receive the reward at time k. and the next state variable Then the sequence Add to the experience sequence;
[0095] The improved DDPG algorithm is used to train the policy network and value network to obtain an energy management model for dual-motor pure electric vehicles based on learning model predictive control.
[0096] Preferably, in step S10, after training is completed, the energy management model of the learning-based predictive control is loaded into the vehicle controller, specifically as follows:
[0097] During actual vehicle operation, the vehicle's short-term historical speed and current speed are input into the speed prediction model to solve the vehicle speed sequence in the prediction time domain in real time.
[0098] Then, the predicted vehicle speed in the time domain is input into the trained DDPG agent. At each moment, the optimal control action is calculated, including the torque distribution coefficients of the front and rear motors, and then input into the motor controller to distribute the torque of the front and rear motors, thereby achieving real-time rolling optimization.
[0099] Compared with the prior art, the present invention has the following beneficial technical effects:
[0100] The technical solution of this invention introduces a dual-motor energy management strategy based on learning model predictive control (L-MPC), which enables intelligent and economical energy allocation and health status perception for pure electric vehicles. The specific beneficial effects include:
[0101] 1. Improve adaptability to operating conditions
[0102] This invention employs a vehicle speed prediction model to predict vehicle speed, enabling the energy management strategy to adaptively adjust based on future driving conditions. This solution demonstrates exceptional adaptability to complex and dynamically changing road conditions and driving demands, ensuring the system can adjust the power distribution between the front and rear axle motors according to different operating conditions, thereby achieving optimal energy efficiency and driving smoothness.
[0103] 2. Achieve real-time response and improve control stability.
[0104] This invention overcomes the computational complexity problem of traditional model predictive control under complex operating conditions by combining L-MPC and deep reinforcement learning (DDPG), achieving efficient real-time response over long prediction periods. The policy updates provided by the DDPG algorithm enable the system to respond quickly to changes in driving conditions, improving vehicle economy while ensuring the real-time performance and stability of energy management, resulting in a smoother driving experience. Attached Figure Description
[0105] Figure 1 A diagram of an energy management system for a dual-motor pure electric vehicle based on learning model predictive control;
[0106] Figure 2 This is a structural diagram of a pure electric dual-motor automobile chassis.
[0107] Figure 3 This is a diagram illustrating learning rate scheduling.
[0108] Figure 4 This is a diagram of a bidirectional causal dilated convolution structure;
[0109] Figure 5 Here is a diagram of the BiTCN algorithm structure;
[0110] Figure 6 Here is a diagram of the GRU architecture;
[0111] Figure 7 A framework diagram for short-term vehicle speed prediction;
[0112] Figure 8 This provides a training framework for energy management strategies in dual-motor pure electric vehicles based on learning-based predictive control. Detailed Implementation
[0113] Example 1
[0114] This invention provides a dual-motor energy management strategy based on Learning-based Model Predictive Control (L-MPC) for pure electric vehicles with dual-motor drive on both the front and rear axles. This scheme combines deep reinforcement learning algorithms, multi-objective optimization, and high-precision vehicle speed prediction methods to achieve intelligent energy allocation and health status perception of the electric vehicle, thereby improving the system's energy utilization efficiency and extending its lifespan. Specific technical solutions are as follows: Figure 1 As shown.
[0115] This embodiment proposes an energy management method for dual-motor pure electric vehicles based on learning-based model predictive control, optimizing it by combining vehicle speed prediction with reinforcement learning strategies. The solution includes a training phase and an application phase. During the training phase, a longitudinal dynamics model and a battery model are established, and an improved deep deterministic policy gradient (DDPG) surrogate model based on a truncated triple Q-network and learning rate scheduling is set. The state, actions, and reward functions of the DDPG surrogate model are also defined. Simultaneously, historical operating data of the dual-motor pure electric vehicle is acquired, and a high-precision vehicle speed prediction model is trained using a bidirectional convolutional temporal network-bidirectional gated recurrent neural network (BiTCN-BiGRU) architecture. Based on this, vehicle speed prediction is combined with reinforcement learning to construct a learning-based predictive control energy management model.
[0116] In the application phase, the trained model is embedded into the vehicle controller. Based on observations such as current vehicle speed, battery SOC, and battery health state SOH, the torque distribution coefficient between the front and rear axles of the vehicle is calculated in real time to achieve efficient energy distribution and adaptability to driving conditions.
[0117] During the training phase, the specific modeling process is as follows:
[0118] 1. Dynamical system modeling
[0119] The model constructed in this application is a pure electric dual-motor power system. Figure 2 The powertrain includes major components such as a first motor, a second motor, a reducer, a clutch, a power battery, a battery management system (BMS), a motor control unit, and a DC-DC converter. The power battery provides the power source for the entire system, powering the vehicle. The battery management system monitors the battery status, ensuring it operates safely and efficiently. The DC-DC converter converts the high-voltage energy from the battery into low-voltage energy suitable for other vehicle systems. The first and second motors drive the tires on the front and rear axles, respectively. Each motor is connected to the battery and the motor controller. The motor control unit controls the motor's operation, adjusting its speed and output torque. The reducer reduces the motor's output speed while increasing output torque to meet the wheel's rotational demands. The clutch disconnects the motor from the tires. During vehicle operation, the power battery provides energy to the motors, which is then converted from high speed to appropriate torque by the reducer, and finally transmitted to the drive wheels via the clutch, allowing the vehicle to move smoothly forward. Meanwhile, the motor control unit dynamically adjusts the torque distribution between the front and rear axle motors to adapt to different road conditions and driving needs, improving energy efficiency and enhancing handling. The regenerative braking system recovers energy during deceleration, further extending the driving range. This dual-motor system on both the front and rear axles allows for flexible power distribution, with the torque distribution between the front and rear axles adjustable according to road conditions, driving modes, and other requirements to improve energy efficiency.
[0120] Next, the system and components are modeled. The required driving torque for the entire vehicle is calculated using longitudinal dynamics:
[0121]
[0122] in The conversion factor for rotating mass; For vehicle quality; For driving acceleration; It is the slope angle, and the rolling resistance coefficient is... The acceleration due to gravity is air drag coefficient Indicates; wheel radius is The speed is The windward area is Without considering the effect of slope, therefore Then the above equation becomes:
[0123]
[0124] Furthermore, in the dual-motor system of this application, the two motors meet the vehicle's required torque through torque coupling, based on the torque distribution factor. The torque calculations for the front and rear motors are as follows:
[0125]
[0126] In the formula, This indicates the output torque of the front motor. This indicates the output torque of the rear motor; additionally, when Only the rear motor provides power, when Only the front motor provides power; when This enables front and rear drive, and the two motors rotate at the same speed.
[0127] Next, we will model the motor. The motor's efficiency differs between drive and regenerative braking states. The motor's output power is calculated as follows:
[0128]
[0129] also, The efficiency of the motor is represented by an efficiency map, which is used for lookup. .
[0130] Next, we will model the power battery. The charging and discharging process of a power battery is a highly complex electrochemical reaction process, making the establishment of an accurate battery model crucial. The power battery model established in this application includes two sub-models: an equivalent internal resistance model and a battery aging model.
[0131] First, the battery's state of charge (SOC) and power in the equivalent internal resistance model are derived from the following formula:
[0132]
[0133] in, It is the battery's output power. It is the current in the circuit. It is open-circuit voltage. It is the battery's internal resistance. It's the battery capacity. It is the initial SOC.
[0134] Secondly, we construct a battery aging model. First, the battery health formula is expressed as follows:
[0135]
[0136] in, This is the equivalent number of cycles before the battery pack reaches the end of its lifespan. The above formula can be modified into a discrete form, as shown below:
[0137]
[0138] in, This refers to the current duration. The effect of discharge rate and internal temperature is calculated using an empirical model of capacity loss based on the Arrhenius equation, expressed by the following formula:
[0139]
[0140] in, It is the percentage of capacity loss. Indicates pre-exponential factor, It is the ideal gas constant, which is equal to 8.314 J / (mol•K). It is a power-law factor equal to 0.55. Indicates the throughput per hour; The activation energy is expressed in J / mol, and the formula is as follows:
[0141]
[0142] when When the battery level drops to 20%, the on-board power battery pack reaches the end of its lifespan. The formulas for the ampere-hour throughput and the equivalent number of cycles before the power battery pack reaches the end of its lifespan are as follows:
[0143]
[0144] in, It is the current duration. The battery health is calculated using formula (5) in combination with the given current, temperature and battery dynamics.
[0145] At any given moment, the powertrain components satisfy the power balance equation:
[0146]
[0147] in, It is the battery's output power. This represents the output power of the first motor. This represents the output power of the second motor.
[0148] 2. Reinforcement Learning Algorithm
[0149] 2.1 DDPG Algorithm
[0150] The DDPG algorithm is a reinforcement learning algorithm for continuous action spaces, using the Actor-Critic framework as its basic structure. The DDPG network structure consists of four networks: the current Actor network, the target Actor network, the current Critic network, and the target Critic network.
[0151] The current Actor network is used to generate policies. The Critic network currently generates Q-scores to evaluate the current policy. The experience replay pool stores data samples generated from each interaction between the agent and the environment. And it learns through random mini-batch sampling. The current Actor network is based on the current state. Select the action at the current moment. It is used to interact with the environment to generate the state for the next moment. and reward value The target Actor network is responsible for determining the next state. Select next action Because the weights of the target network are periodically updated from the current network, the target Q-value can be temporarily fixed during training, making learning more stable. The target Q-value is:
[0152]
[0153] in, It is the reward value at the current time step t. , which is the discount factor, representing the impact of future rewards on long-term returns, and its value ranges from [0, 1]. Generated for the target Critic network value; The policy function of the target Actor network Its network weight; The weights are those of the target Critic network.
[0154] The current goal of the Critic network is to minimize the loss function, which is used during network training:
[0155]
[0156] In the formula, K: the number of samples in a batch (i.e., mini-batch size), It is the target Q-value calculated through the target Critic network; The Q-value generated by the current Critic network; These are the weights of the current Critic network.
[0157] The current goal of the Actor network is to maximize the expected return, i.e.
[0158]
[0159] In the formula, It is the policy function of the current Actor network.
[0160] The current weight gradient of the Actor network is:
[0161]
[0162] It is the gradient of the Critic network's evaluation of the value of an action, measuring the value of the action. How do small changes affect The value; It is an Actor network strategy. The gradient describes how the Actor network adjusts according to the current state. Adjust parameters This optimizes the selected actions. Through the chain rule, the gradient of the Actor network is ultimately determined by the product of the Critic network's value gradient and the Actor's policy gradient, guiding the Actor network to update its weights so that the selected actions better improve cumulative returns.
[0163] Update the weights of the current Actor network using an optimizer. and the weights of the current evaluation network Then, the weights of the two target networks are updated using a soft update method, that is:
[0164]
[0165] In the formula, These are soft update parameters.
[0166] 2.2 Truncated Triple Q Learning Mechanism
[0167] In the DDPG algorithm, the target Critic network obtains its Q-value by using the maximum value as its expected value, which generally leads to overestimation compared to the true expected Q-value. This overestimation of the value function in DDPG typically causes two problems: firstly, overestimation leads to significant bias after multiple updates; secondly, biased value estimation further contributes to inaccurate policy updates. Overestimation of suboptimal actions by the Critic network can guide the policy network to choose suboptimal actions during policy updates.
[0168] To mitigate the Q-value overestimation problem in the DDPG algorithm, some studies have designed a dual-evaluation network mechanism. The core of this mechanism is to reduce the risk of Q-value overestimation by mutually correcting each other using two independent Q-networks. However, while this method can reduce the risk of overestimation, it can sometimes lead to underestimation. Therefore, this application adopts a truncated triple Q-learning strategy. The main idea is to simultaneously combine the biases of overestimation and underestimation to obtain a more accurate Q-value estimate.
[0169] The objective function expression of this method is as follows:
[0170]
[0171] In the formula, and Representing different network parameters and Q output below; The next state is...
[0172] It is the set of possible actions; mean represents the mean operation on the truncated Q value.
[0173] This mechanism reduces the risk of overestimation and avoids the negative impact of underestimation on policy optimization by simultaneously considering the minimum value of multiple Q-network outputs and averaging the results. This truncated triple Q-learning strategy improves the accuracy of Q-value estimation while retaining the advantages of dual-evaluation networks, thus more effectively guiding policy network optimization.
[0174] 2.3 Learning Rate Scheduling
[0175] Traditional learning rates (i.e., fixed learning rates) have shown certain limitations when training deep reinforcement learning models. The main characteristic of a fixed learning rate is that the step size remains constant during training. This approach is intuitive and simple to implement, but it also suffers from problems such as limited convergence speed, local optima, and overfitting in complex learning tasks. Specifically, these problems manifest in the following ways:
[0176] First, a fixed learning rate makes it difficult to achieve fast convergence in the early stages of training. If the learning rate is too low, the model needs more time to explore the parameter space and reach the optimal solution; but if the learning rate is too high, the model may skip the optimal solution during updates, leading to unstable training.
[0177] Secondly, in deep reinforcement learning, especially in policy network optimization, a fixed learning rate can easily cause the model to get stuck in local optima and find it difficult to escape.
[0178] Finally, the fixed learning rate did not decay in the later stages of training, resulting in the model fitting the training data too well and failing to generalize well to new data.
[0179] To address these issues, this application introduces learning rate schedulers into the reinforcement learning algorithm. Learning rate schedulers can dynamically adjust the learning rate during training, making the learning process more stable and effective, and facilitating the finding of the global optimum. CyclicLR, in particular, is a scheduler that periodically adjusts the learning rate, allowing it to fluctuate between a minimum and a maximum value. This helps the model to explore rapidly in the early stages of training and gradually converge in later stages.
[0180] Specifically, this application employs a learning rate scheduler with periodic fluctuations and peak decay, as follows: Figure 3 As shown in the diagram, the learning rate in each cycle gradually increases from its lowest value to its highest value, then decreases back to its lowest value, forming a triangular fluctuation pattern. At the beginning of each new cycle, the highest learning rate is halved compared to the previous cycle. An example of this process is as follows:
[0181] The first cycle: the learning rate gradually rises from the lowest value to the highest value, and then falls back to the lowest value, forming a complete triangular waveform.
[0182] The second cycle: the learning rate rises from its lowest value to half of the maximum value of the first cycle, and then falls back to its lowest value.
[0183] The third cycle: The highest learning rate is halved again, the learning rate rises from the lowest value to this new highest value, and then falls back to the lowest value.
[0184] As the training cycle progresses, the fluctuation range of the learning rate gradually decreases, eventually stabilizing within a low range. This pattern effectively allows for greater exploration in the early stages of training, while gradually reducing the fluctuation of the learning rate in the later stages, enabling the model to converge more stably.
[0185] 3. BiTCN-BiGRU vehicle speed prediction algorithm
[0186] 3.1 Bidirectional Convolutional Temporal Networks
[0187] Bidirectional Convolutional Temporal Network (BiTCN) is a deep learning model specifically designed for processing sequential data. By combining the characteristics of temporal convolutional networks with a bidirectional information flow mechanism, BiTCN effectively improves the prediction performance of sequential data. Specifically, temporal convolutional networks use causal convolution to ensure that each element of the output sequence depends only on the current and previous inputs, avoiding the influence of future information on current decisions. This causal constraint guarantees the reasonableness of the prediction and, to some extent, simulates the actual characteristics of time series. Furthermore, BiTCN uses dilated convolution, enabling the model to capture longer historical information, thereby improving performance when dealing with long-term dependencies.
[0188] like Figure 4 The structure diagram is shown below for bidirectional causal dilated convolution. By introducing a bidirectional information flow mechanism, the model can simultaneously consider historical information. Input and future information This bidirectional design allows the model to learn not only from current and past data but also from future information, thus enabling a more comprehensive modeling of sequential data. Furthermore, as shown in the figure, dilated convolutions significantly increase their receptive field through interval sampling, thereby achieving a larger receptive field with fewer network layers.
[0189] The output of BiTCN at time t, which combines bidirectional information flow and causal convolution, is calculated as follows:
[0190]
[0191] in, The weights of the convolution kernel, This represents the input item after interval sampling. d is the size of the convolution kernel, and d is the dilation coefficient.
[0192] The BiTCN algorithm mainly consists of causal convolutional layers, activation functions, weight normalization layers, and spatial dropout layers, enabling prediction of time series data.
[0193] 3.2 Bidirectional Gated Recurrent Neural Network
[0194] A Gated Recurrent Unit (GRU) consists of reset gates and update gates, used to control the flow of information. The reset gate determines how much past information is forgotten in the current state. The update gate determines how much information from the current state needs to be retained and passed to the next time step. The structure of a GRU unit is as follows: Figure 6 As shown.
[0195] The formulas for calculating its state and output are as follows:
[0196]
[0197] In the formula, It is the input vector at the current moment. It is the hidden state from the previous moment. It is the hidden state at the current moment. It's an update gate. It's a door reset. This refers to the new information generated at time t. Represents the input weight matrix.
[0198] The weight matrix represents the hidden state.
[0199] The Bidirectional Gated Recurrent Unit (BiGRU) combines the features of bidirectional recurrent neural networks and gated recurrent units, enabling it to consider both forward and backward information flows of data during processing. This enhances its ability to capture global dependencies and allows it to perform well in time series learning tasks.
[0200] 3.3 Short-term vehicle speed prediction framework
[0201] This application combines the advantages of BiTCN and BiGRU to achieve accurate prediction of future vehicle speed sequences. The input data consists of current and historical vehicle speed information. BiTCN performs bidirectional processing on the input data, utilizing its causal convolution and dilated convolution characteristics to expand the receptive field, ensuring that the model does not lose temporal order when capturing historical information. Through bidirectional information flow, BiTCN can capture global dependencies in the sequence without increasing computational complexity, helping the model understand the overall trend and change patterns of the sequence. Subsequently, BiGRU further processes the features extracted by BiTCN, combining forward and backward information flows to model the bidirectional dependencies of the sequence. Due to its fewer parameters and higher computational efficiency, BiGRU has advantages in vehicle speed prediction tasks with high real-time requirements. The bidirectional structure allows the model to better understand vehicle speed change trends, and is particularly suitable for capturing complex temporal dependencies.
[0202] Energy Management Strategy for Dual-Motor Pure Electric Vehicles Based on Learning-Based Predictive Control
[0203] 4.1 Reinforcement Learning Settings
[0204] In deep reinforcement learning algorithms, an agent observes the environmental state, takes corresponding actions, receives immediate rewards, and transitions to a new state. The design of the state space, action space, and reward function is crucial in this process. In the energy management of a dual-motor pure electric vehicle, the agent needs to acquire the vehicle's motion state and the battery's state. Specifically, the vehicle's motion state includes vehicle speed (v), acceleration (a), and the required torque; the battery's state includes the state of charge (SOC) and state of health (SOH), representing the battery's current charge and health status. Therefore, the state space is defined as follows: .
[0205] The action space of reinforcement learning defines the actions taken by the agent at each time step. This application uses the torque allocation factor. Defined as the motion space, it is used to control the torque distribution between the first motor and the second motor.
[0206] Furthermore, the core idea of reinforcement learning is that an agent learns a strategy that maximizes expected reward through interaction with the environment. Therefore, designing a reasonable reward function is crucial for training the agent. In this application, the optimization objectives include two aspects: first, reducing battery power consumption by minimizing the power consumption of the power battery to improve energy utilization efficiency; second, reducing battery health degradation by optimizing energy allocation to extend battery life. Therefore, the defined reward function is as follows:
[0207]
[0208] Furthermore, the boundaries of the learning process are defined as follows:
[0209]
[0210] 4.2 Energy Management Training Process for Learning-Based Predictive Control
[0211] Regarding the control strategy proposed in this application, the specific training framework is as follows: Figure 8 As shown:
[0212] Assuming the vehicle speed prediction time domain is 3 seconds, and the vehicle speed is v at the current time k, the historical vehicle speeds at times k-3, k-2, and k-1, along with the current vehicle speed at time k, are input into the vehicle speed prediction module. This module outputs the predicted vehicle speeds at times k+1, k+2, and k+3. Then... The vehicle speed information is input into the reinforcement learning agent module, and then the following steps are performed:
[0213] Table 1. Algorithm Execution Flow
[0214]
[0215] Utilizing vehicle speed Calculate and predict the state of the dual-motor power system in the time domain based on battery SOC and SOH:
[0216]
[0217] Obtain the optimal action at time t Then, using the state variables at time k and optimal action Receive the reward at time k. and the next state variable Then the sequence Add to the experience sequence.
[0218] Then, the improved DDPG algorithm's policy network and value network are trained using the above (10)~(14) to obtain the energy management strategy for dual-motor pure electric vehicles based on learning model predictive control.
[0219] 4.3 Energy Management Execution Process of Learning-Based Predictive Control
[0220] After training is complete, the energy management strategy of the learning-based predictive control is loaded into the vehicle controller, and the specific execution is as follows:
[0221] During actual vehicle operation, the vehicle's short-term historical speed and current speed are input into the speed prediction model to solve the vehicle speed sequence in the prediction time domain in real time.
[0222] Then, the predicted vehicle speed in the time domain is input into the trained DDPG agent. At each moment, the optimal control action (torque distribution coefficient of the front and rear motors) is obtained through the algorithm process in Table 1 and input into the motor controller to distribute the torque of the front and rear motors, thereby realizing real-time rolling optimization.
[0223] In this embodiment, 1. The traditional DDPG reinforcement learning algorithm is susceptible to overestimation bias, and a common improvement is to use two independent Q-networks to mutually correct each other, reducing the risk of Q-value overestimation. However, while this improvement method reduces the risk of overestimation, it sometimes leads to underestimation. This application adopts a truncated triple Q-learning strategy. The main idea is to combine the biases of overestimation and underestimation simultaneously, introducing three Q-networks to estimate action value, and using a truncation mechanism to obtain a more accurate Q-value estimate. Furthermore, the learning rate scheduling proposed in this application allows for greater exploration in the early stages of training, while gradually reducing the fluctuation of the learning rate in the later stages of training, enabling the model to converge more stably and improving algorithm performance.
[0224] 2. The vehicle speed prediction model based on BiTCN and BiGRU combines the advantages of convolutional neural networks and recurrent neural networks, which can efficiently capture long-term dependencies and local features in time series data, thereby improving the accuracy and robustness of vehicle speed prediction.
[0225] 3. This invention utilizes the framework of model predictive control and introduces a vehicle speed prediction algorithm to accurately predict future driving conditions. Combined with the decision-making ability of reinforcement learning, the system can adaptively adjust the torque output of the front and rear axle dual motors according to different driving conditions, thereby ensuring that the vehicle maintains efficient energy distribution in complex driving environments and achieves better energy consumption performance.
[0226] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.
Claims
1. A method for controlling the drive and braking of a pure electric commercial vehicle driven by dual motors on the front and rear axles, characterized in that, The specific steps include the following: S1. Obtain historical operating condition data of dual-motor pure electric vehicles; S2. Establish a vehicle speed prediction model based on BiTCN-BiGRU; S3. Set the network parameters of the vehicle speed prediction model and train the prediction model; S4. Adjust the network parameters according to the prediction accuracy to obtain a high-precision vehicle speed prediction model; In the vehicle speed prediction model trained in step S4, the input data is the current and historical vehicle speed information, and BiTCN performs bidirectional processing on the input data; BiTCN captures the global dependencies in the sequence through bidirectional information flow. BiGRU processes the features extracted by BiTCN and combines forward and backward information flows to model bidirectional dependencies of sequences. S5. Establish a dual-motor pure electric vehicle model, including a longitudinal dynamics model, a battery model, and a motor model. The power system of the dual-motor pure electric vehicle model in step S5 includes a first motor, a second motor, a reducer, a clutch, a power battery, a battery management system (BMS), a motor control unit, and a DC-DC converter. The power battery is the power source for the entire system, providing electrical energy to drive the vehicle. The BMS monitors the battery status. The DC-DC converter converts the high-voltage energy from the battery into low-voltage energy suitable for other vehicle systems. The first and second motors drive the tires on the front and rear axles, respectively. The two motors are connected to the battery and the motor controller, respectively. The motor control unit controls the operation of the motors, adjusting their speed and output torque. The reducer reduces the output speed of the motors while increasing the output torque to meet the rotational demands of the wheels. The clutch disconnects the motor from the tires. S6. Establish an improved DDPG agent model based on truncated triple Q network and learning rate scheduling; S7. Set up the state, action, and reward functions for the improved DDPG agent model; S8. Based on the training dataset, a learning model predictive control energy management model combining a vehicle speed prediction model and an improved DDPG agent model is trained. S9. Validate the energy management model based on learning model predictive control using a new operating condition dataset; S10. Embed the trained model from step S9 into the vehicle controller. Based on the current vehicle speed, battery SOC, and battery state of health (SOH), calculate the torque distribution coefficient between the front and rear axles in real time to control vehicle operation. In step S10, after training is completed, load the learning-based predictive control energy management model into the vehicle controller. The specific execution is as follows: During actual vehicle operation, the vehicle's short-term historical speed and current speed are input into the speed prediction model to solve the vehicle speed sequence in the prediction time domain in real time. Then, the predicted vehicle speed in the time domain is input into the trained DDPG agent. At each moment, the optimal control action, namely the torque distribution coefficient of the front and rear motors, is calculated and input into the motor controller to distribute the torque of the front and rear motors, thereby achieving real-time rolling optimization.
2. The driving and braking control method for a pure electric commercial vehicle driven by dual motors on the front and rear axles according to claim 1, characterized in that, The historical vehicle operating condition data obtained in step S1 is the historical vehicle speed.
3. The braking control method for a dual-motor driven pure electric commercial vehicle according to claim 1, characterized in that, In step S5, the required driving torque for the entire vehicle is calculated using longitudinal dynamics: in The conversion factor for the rotating mass; For vehicle quality; For driving acceleration; It is the slope angle, and the rolling resistance coefficient is... The acceleration due to gravity is air drag coefficient Indicates; wheel radius is The speed is The windward area is Without considering the effect of slope, therefore Then the above equation becomes: The two motors meet the vehicle's required torque through torque coupling, based on the torque distribution factor. The torque calculations for the front and rear motors are as follows: In the formula, This indicates the output torque of the front motor. This indicates the output torque of the rear motor; additionally, when Only the rear motor provides power, when Only the front motor provides power; when This enables front and rear drive, and the two motors rotate at the same speed. Model the motor; the motor's efficiency differs between driving and regenerative braking states, and the motor's output power is calculated as follows: also, The efficiency of the motor is represented by an efficiency map, which is used for lookup. ; Modeling of power batteries includes equivalent internal resistance model and battery aging model; In the equivalent internal resistance model, the battery's state of charge (SOC) and power are derived from the following formula: in, It is the battery's output power. It is the current in the circuit. It is open-circuit voltage. It is the battery's internal resistance. It's the battery capacity. It is the initial SOC; Secondly, an aging model for the battery is constructed; the battery health status (SOH) formula is expressed as follows: in, It is the equivalent number of cycles before the battery pack reaches the end of its lifespan. This is the discharge rate. The above formula is modified into a discrete form, expressed as follows: in, This refers to the current duration; the effect of discharge rate and internal temperature is calculated using an empirical model of capacity loss based on the Arrhenius equation, expressed by the following formula: in, It is the percentage of capacity loss. Indicates pre-exponential factor, It is the ideal gas constant, which is equal to 8.314 J / (mol•K). It is a power-law factor equal to 0.
55. Indicates the throughput per hour; The activation energy is expressed in J / mol, and the formula is as follows: when When the battery level drops to 20%, the on-board power battery pack reaches the end of its lifespan. The formulas for the ampere-hour throughput and the equivalent number of cycles before the power battery pack reaches the end of its lifespan are as follows: The battery health is calculated by combining the given current, temperature, and battery dynamics; in the above formula... The average temperature inside the battery; at the same moment, the powertrain components satisfy the power balance equation: in, It is the battery's output power. This represents the output power of the first motor. This represents the output power of the second motor.
4. The braking control method for a dual-motor driven pure electric commercial vehicle on the front and rear axles according to claim 1, characterized in that, In step S6, the DDPG network structure consists of four networks: the current Actor network, the target Actor network, the current Critic network, and the target Critic network. The current Actor network is used to generate policies. The current Critic network generates Q-values to evaluate the current policy; the experience replay pool stores data samples generated from each interaction between the agent and the environment. And it learns by random small-batch sampling; The current Actor network is based on its current state. Select the action at the current moment. It is used to interact with the environment to generate the state for the next moment. and reward value ; The target actor network is responsible for determining the next state. Select next action ; Because the weights of the target network are periodically updated from the current network, the target Q-value is temporarily fixed during training, making learning more stable; the target value of the Q-value is: in, It is the reward value at the current time step t. , which is the discount factor, representing the impact of future rewards on long-term returns, and its value ranges from [0, 1]. Generated for the target Critic network value; The policy function of the target Actor network Its network weight; The weights of the target Critic network; The current goal of the Critic network is to minimize the loss function, which is used during network training: In the formula, K: the number of samples in a batch, It is the target Q-value calculated through the target Critic network; The Q-value generated by the current Critic network; These are the current weights of the Critic network; The current goal of the Actor network is to maximize the expected return, i.e. In the formula, It is the policy function of the current Actor network; The current weight gradient of the Actor network is: It is the gradient of the Critic network's evaluation of the value of an action, measuring the value of the action. How do small changes affect The value; It is an Actor network strategy. The gradient describes how the Actor network adjusts according to the current state. Adjust parameters To optimize the selected actions; Update the weights of the current Actor network using an optimizer. and the weights of the current evaluation network Then, the weights of the two target networks are updated using a soft update method, that is: In the formula, These are soft update parameters.
5. The braking control method for a dual-motor driven pure electric commercial vehicle according to claim 1, characterized in that, In step S6, a truncated triple Q-learning strategy is adopted to obtain a more accurate Q-value estimate. The objective function expression is as follows: In the formula, and Representing different network parameters and Q output below; The next state is... It is the set of possible actions; mean represents the mean operation on the truncated Q value.
6. The braking control method for a pure electric commercial vehicle driven by dual motors on the front and rear axles according to claim 1, characterized in that, In step S7, the CyclicLR learning rate scheduler is introduced to adjust the learning rate and periodically decay it. The defined reward function is as follows: Furthermore, the boundaries of the learning process are defined as follows: In the formula, and Indicates the upper and lower limits of battery SOC. and Indicates the upper and lower limits of battery health status. This indicates the upper limit of the first motor's speed. This indicates the upper limit of the second motor's speed. This indicates the upper limit of the torque of the first motor. This indicates the upper limit of the torque of the second motor.
7. The driving and braking control method for a pure electric commercial vehicle driven by dual motors on the front and rear axles according to claim 1, characterized in that, In step S8, assuming the vehicle speed prediction time domain is 3 seconds, and the vehicle speed is v at the current time k, the historical vehicle speeds at times k-3, k-2, and k-1, along with the current vehicle speed at time k, are input into the vehicle speed prediction module. This module outputs the predicted vehicle speeds at times k+1, k+2, and k+3. Then... The vehicle speed information is input into the learning-based predictive control agent module for training, and the vehicle speed is used... Battery SOC and SOH are used to update the state of the dual-motor power system in the prediction time domain. Obtain the optimal action at time t Then, using the state variables at time k and optimal action Receive the reward at time k. and the next state variable Then the sequence Add to the experience sequence; The improved DDPG algorithm is used to train the policy network and value network to obtain an energy management model for dual-motor pure electric vehicles based on learning model predictive control.
Citation Information
Patent Citations
Reconstruction control method for distributed power driving system of electrically-driven vehicle and vehicle
CN110717218A
Intelligent variable time domain model prediction energy management method for hybrid power vehicle
CN111267831A