Whole vehicle system comprehensive inertia evaluation and transient response optimization control method
Patent Information
- Application Number
- CN202610975056.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-08-21
AI Technical Summary
此类策略存在固有滞后性,补偿动作在恶化发生后才能触发,无法保持系统最佳响应状态
[0017]This invention proposes a method for comprehensive inertia assessment and transient response optimization control of range-extended electric vehicle systems. It defines the response time constant of each subsystem to power commands as the equivalent inertia of that subsystem, synthesizes the equivalent inertia of heterogeneous subsystems into a unified scalar index—the comprehensive system inertia J_sys—and uses continuous minimization of J_sys as the control objective. Reinforcement learning is employed to coordinate the power distribution and operating states of the motor, battery, engine, and thermal management systems. Compared with existing technologies, this method has the following advantages: 1. Constructing equivalent inertia: It unifies and quantifies the three subsystems—motor, battery, and engine—with different physical responses and response time constants differing by more than 100 times, into a comparable and synthesizable equivalent inertia index with the same dimension (seconds). 、
,
Furthermore, based on this, the system's integrated inertia J_sys was defined as a single quantitative evaluation index for the vehicle's transient response capability, solving the technical problem of the inability to uniformly evaluate the response capability of heterogeneous subsystems. 2. The transient correction term in the J_sys synthesis formula is obtained through the battery power deficit ratio factor.
and the rate of change in power demand |
This precisely describes the physical phenomenon where the inertia difference is amplified when battery power is insufficient and driver demands change rapidly; when battery power is sufficient (
= 0) This item automatically resets to zero, achieving full-condition self-adaptation without the need for manual threshold setting. 3. The control concept shifts from threshold-triggered compensation to continuous minimization. The control objective is set as solving for the combination of control variables that minimizes J_sys in each control cycle. The system is always in the optimal response state under the current constraints, avoiding the passive lag of traditional strategies that only trigger after deterioration occurs. 4. Thermal inertia is incorporated into a unified inertia optimization framework. Thermal-electric-mechanical coupling optimization is achieved using existing vehicle components (four-way valve, electric water pump) without increasing hardware costs. The strategy network automatically determines when to intervene in thermal management to improve response capabilities. 5. Safety constraint layer design: A deterministic physical boundary safety check is set at the output of the strategy network, completely decoupling the exploratory nature of reinforcement learning from execution safety, ensuring that the final command does not violate the hardware physical boundaries of the battery and engine.
Smart Images

Figure CN122607298A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of new energy vehicle control technology, specifically relating to a comprehensive inertia assessment and transient response optimization control method for the whole vehicle system, based on equivalent inertia index and reinforcement learning, which is particularly suitable for range-extended electric vehicles under low battery charge conditions. Technical Background
[0002] The powertrain of a range-extended electric vehicle consists of an engine, a generator, a battery, and a drive motor connected in series. In this configuration, there is no mechanical connection between the engine and the wheels; the engine outputs electrical power to the bus solely through the generator, and the drive motor is the only power source for the wheels.
[0003] The power response time constant of the drive motor and power battery is approximately on the order of 5 to 10 milliseconds, while the engine's power response time constant is on the order of 200 milliseconds to 2 seconds, limited by the mechanical rotational inertia of the crankshaft-flywheel-generator set and the gas circulation delays such as intake manifold filling, in-cylinder combustion, and torque build-up. The difference between these three can be tens of times. Under high SOC conditions, the battery responds to transient power demands at the millisecond level, masking the engine's slow response and resulting in agile electric drive performance. However, under low SOC conditions, the Battery Management System (BMS) limits the discharge power to protect the battery, significantly weakening its rapid response capability. The vehicle shifts from an electric drive feel to an engine-driven feel, exhibiting noticeable acceleration lag and power lag. Furthermore, the lag in thermal parameters such as battery temperature, engine coolant temperature, and motor temperature further exacerbates the aforementioned response lag problem.
[0004] It is evident that there are significant differences in the transient response time constants of the motor, battery, and engine. The lack of a unified quantitative indicator and evaluation method for their transient responses during vehicle operation makes it difficult to eliminate response lag between the various subsystems. Existing technologies typically employ threshold-based trigger-based compensation strategies, passively adjusting power distribution or engine operating status only after a deterioration in response is detected. Such strategies inherently suffer from lag; compensation is only triggered after deterioration occurs, failing to maintain optimal system response. Thermal lag in the engine, battery, and motor leads to insufficient cooling capacity during power increases, resulting in system overheating and reduced system lifespan. Limiting the power increase rate, however, results in slow transient power response. In existing range-extended electric vehicles, the battery, motor, and engine each have independent thermal management systems and power control strategies, lacking a unified thermal-power coordination mechanism. The conflict between the thermal and power inertia of the battery, motor, and engine causes inconsistent response speeds among subsystems during sudden power demand changes, leading to power lag or thermal runaway risks, making it difficult to achieve optimal overall vehicle efficiency. The lack of effective quantitative indicators and evaluation methods makes it difficult to quantify the correlation between the thermal management of the motor, battery, and engine subsystems and the agile response to the transient power demand of the vehicle, which is a key technical bottleneck in the transient response optimization control of range-extended electric vehicles. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention proposes a method for comprehensive inertia assessment and transient response optimization control of the entire vehicle system. This method unifies the three subsystems (motor, battery, and engine) with different response time constants into comparable and synthesizable equivalent inertia indices. It also achieves continuous optimization of the vehicle's transient response through reinforcement learning, while taking into account SOC maintenance, fuel economy, and engine operating efficiency.
[0006] This invention achieves its purpose through the following measures: A method for comprehensive inertia assessment and transient response optimization control of a vehicle system, characterized by comprising the following steps: Step 1: Obtain the real-time status parameters of the motor, battery and engine, calculate the equivalent inertia of the motor, battery and engine respectively, and quantify the response capability of the heterogeneous subsystem into an equivalent inertia index with time dimension. Step 2: Calculate the overall inertia of the system based on the power allocation weights of each subsystem, the basic motor term, and the transient correction term driven by the battery power deficit ratio and the rate of change of power demand. Step 3: With minimizing the overall system inertia as the optimization objective, generate coordinated control commands including power distribution, engine speed, and thermal management actuators through a reinforcement learning policy network; Step 4: After performing a safety boundary check on the cooperative control command, it is output to the vehicle actuator to achieve continuous adaptive optimization of the vehicle's transient response capability.
[0007] In this invention, the calculation of the equivalent inertia of the motor is based on the electromagnetic time constant of the motor, and a temperature correction coefficient is introduced. When the motor temperature exceeds a preset threshold, the temperature correction coefficient increases to reflect the slowed response characteristics caused by motor overheating derating. Specifically, the equivalent inertia of the motor is calculated using the following formula. , (1), in The electromagnetic time constant of the motor (approximately 5~10ms) is determined by the inductance and resistance of the motor windings. The temperature correction factor is obtained by looking up a table using the preset motor temperature-derating curve within the normal operating temperature range. ≈ 1, when approaching the demagnetization threshold of permanent magnets >1; Motor temperature.
[0008] The calculation of the battery equivalent inertia in this invention is based on the battery's inherent electrochemical delay time constant, combined with the State of Charge (SOC) power limitation factor and the temperature penalty factor. The SOC power limitation factor increases in the low SOC range, and the temperature penalty factor increases under extreme temperature conditions to characterize the increase in response delay when the battery's available power is limited. Specifically, the battery equivalent inertia is calculated according to the following formula. , (2), This includes the inherent electrochemical response delay of the battery (approximately 5 ms). The critical SOC value for BMS power limitation (e.g., 30%). For SOC-power limiting sensitivity coefficient; , [This refers to the optimal operating temperature range of the battery (e.g., 25~35℃)]. , These are the penalty coefficients for low temperature and high temperature, respectively. The core physical meaning is: when the State of Charge (SOC) is lower than the maximum discharge power limit set by the rear BMS, the battery's power delivery capability decreases—although the battery's own electrochemical reaction rate remains unchanged, from the perspective of the entire vehicle, this is equivalent to a worse response. The lower the SOC, the further the temperature deviates from the optimal range. The larger the value, the more directional the temperature correction term becomes. exist[ , Within the interval, both max functions are zero; below... The low temperature item takes effect when the temperature is higher than 10℃. The high temperature item is in effect.
[0009] The equivalent inertia of the engine described in this invention is composed of the superposition of the rotational speed response time constant determined by the mechanical rotational inertia and the gas delay time constant; the gas delay time constant increases as the coolant temperature decreases to reflect the delay characteristics of the intake, charging, and combustion build-up processes in a cold engine state; specifically, the equivalent inertia of the engine is calculated according to the following formula. , (3), in = J_mech·Δω / T_max, which is the rotational speed response time constant determined by the mechanical moment of inertia (J_mech is the combined moment of inertia of crankshaft-flywheel-generator, and T_max is the maximum net torque). This is the gas delay time constant, which includes intake manifold filling, in-cylinder combustion, and torque build-up delays, and depends on engine coolant temperature. When the machine is cold The value increases significantly (up to 1~2 seconds), and tends to reach its minimum value (approximately 100~300ms) after warm-up. It is obtained by looking up the preset coolant temperature-gas delay mapping table.
[0010] Step 1 of this invention also includes thermal equivalent inertia. As a constraint input, The system thermal time constant does not directly enter the power response chain, but is constrained. and This can improve the overall vehicle responsiveness.
[0011] The calculation formula for the overall system inertia in step 2 of this invention includes a transient correction term. The value of the transient correction term is determined by the product or functional relationship between the battery power shortage ratio and the power demand change rate. When the battery power is sufficient, the transient correction term automatically returns to zero. When the battery power is insufficient and the power demand change rate is greater than a preset value, the transient correction term is adaptively amplified. Specifically, the synthesis of the overall system inertia J_sys is as shown in equation (5), which synthesizes the equivalent inertia of the heterogeneous subsystem into a unified scalar index J_sys. (5), in This represents the battery's current actual output power. = · This refers to the equivalent electrical power of the engine-generator set. This refers to the mechanical power output from the engine crankshaft end. This is the generator efficiency (approximately 0.90~0.95). To prevent the use of a pre-defined small constant with a denominator of zero; For transient correction of gain; is the battery power deficit ratio, is the maximum allowable discharge power of the battery obtained by the BMS based on the current SOC and a lookup table, is the driver's power demand, and clamp(x, a, b) means limiting x to the closed interval [a, b]. The rate of change of power demand; Formula (5) consists of three terms: (1) Power-weighted basic terms The actual output power of the battery and engine is used as the weight for the... and Weighted average, under high SOC conditions, the battery dominates the power output, and the basic terms are close to... (Approximately 5~10ms), reflecting the rapid response of the electric drive quality; under low SOC conditions, the engine dominates, and basic parameters tend to be close to... (Approximately 200ms to 2s), reflecting the engine's slow response characteristics; (2) Motor basic items The constant count is included in J_sys, representing the non-eliminable physical lower bound; (3) Transient correction term . A constant positive value indicates the inherent response difference between the engine and the battery. The value is large when the accelerator is pressed hard, and tends to zero during steady-state cruising. With sufficient battery power When the transient response is zero, the transient correction term is set to zero, the battery alone bears all transient demands, and the engine's slow response is not detected; when the battery power is insufficient... The value >0 indicates a larger gap that is closer to 1, and the degree of exposure to engine slow response is accurately reflected in J_sys. This design achieves full-condition adaptation without the need for manual threshold setting.
[0012] In step 3 of this invention, the reinforcement learning policy network adopts the PPO algorithm, whose reward function has the negative value of the system's overall inertia as the main component, and includes energy consumption penalty terms and safety constraint violation penalty terms; the input state space of the policy network includes driver pedal depth, pedal depth change rate, vehicle speed, temperature of each subsystem, SOC, and current power allocation state; specifically, it includes the following: Step 3-1: Construct a 14-dimensional state space. The policy network input s includes: pedal depth. Pedal depth change rate Vehicle speed v, acceleration a, power requirement SOC, battery temperature Battery equivalent inertia, engine speed ωe, coolant temperature Engine equivalent inertia , motor temperature \(T_{motor}\), motor equivalent inertia , ambient temperature \(T_{ambient}\), and , , are explicitly input as equivalent inertia information; Step 3-2: Design a PPO policy network to output a 4D control action (6). The policy network adopts a three-layer fully connected structure (128 neurons in each layer, ReLU activation). The output layer limits the output to [-1, 1] using tanh and then linearly maps it to the physical domain of each action. During operation, a deterministic forward propagation is performed once every control cycle (10 ms), without random sampling and parameter update. The 4D action includes: Battery-engine power distribution ratio \(\alpha\in[-\alpha_{chg\_max}, 1]\) (\(\alpha > 0\) for discharging, \(\alpha < 0\) for charging); Engine target speed \(\omega_{e\_target}\in ; Four-way valve opening \(u_{valve}\in[0, 1]\) (0 fully passes through the radiator, 1 fully passes through the battery heater); Water pump duty cycle \(u_{pump}\in[0, 1]\); Step 3-3: Construct a reward function with minimizing \(J_{sys}\) as the main objective: \(R=-w1\cdot J_{sys}\) \(-w2\cdot[\max(SOC_{min}-SOC, 0)]^2\) \(-w3\cdot[\max(SOC - SOC_{max}, 0)]^2\) \(-w4\cdot fuel\_rate\) \(-w5\cdot|\omega_e-\omega_{e\_opt}(P_{eng})|\) \(+w6\cdot\exp(- \cdot J_{sys})\) (7), where \(w1\sim w6\) are preset weight coefficients, \(fuel\_rate\) is the engine instantaneous fuel consumption rate, \(SOC_{min}\) and \(SOC_{max}\) are preset SOC safety intervals, and \(\omega_{e\_opt}(P_{eng})\) is the optimal efficiency speed.
[0013] The functions of the six terms in Step 3-3 of the present invention are as follows: ① Minimize \(J_{sys}\) (main optimization objective); ② Hard constraint on the lower limit of SOC (\(w2\) takes a large value, square penalty, to prevent over-discharging of the battery); ③ Soft constraint on the upper limit of SOC (\(w3\ll w2\)); ④ Minimize fuel consumption; ⑤ Penalty for engine efficiency deviation (complementary to ④); ⑥ Reward for transient response quality - sudden throttle pedal depression (|\ | Large) and J_sys receives a positive reward every hour, which complements item ① and jointly drives the policy network to learn engine pre-positioning and other behaviors that actively improve transient response.
[0014] In step 4 of this invention, the original output of the policy network... Issued after the following definitive security checks: (1) , where is the battery's current maximum allowable charging power; (2) ωe_target = ; (3) u_valve = u_pump = ; The security constraint layer completely decouples the learning exploration from the execution security, ensuring that the final instruction does not violate the physical boundaries of the hardware.
[0015] The present invention also includes engine pre-positioning and active thermal management control, wherein engine pre-positioning is based on... Signals to anticipate driver intent, in Before the pedal MAP filter delay is established, the engine speed is increased to the target range in advance, skipping the slowest speed climb phase (about 200~1000ms); the policy network automatically learns this behavior through the joint guidance of reward function term ⑥ and term ①, without the need for manual rule preset.
[0016] This invention employs active thermal management control using existing components such as a four-way valve and an electric water pump. It improves temperature conditions and reduces equivalent inertia through methods like engine waste heat heating the battery and motor pre-cooling. A strategy network automatically determines when energy is worth consuming for thermal management intervention to improve responsiveness. Specific control methods include engine coolant waste heat heating the battery (reducing...). ), engine delayed shutdown to maintain warm-up (reduce) ), motor pre-cooling (to prevent (Increases due to overheating and deflator).
[0017] This invention proposes a method for comprehensive inertia assessment and transient response optimization control of range-extended electric vehicle systems. It defines the response time constant of each subsystem to power commands as the equivalent inertia of that subsystem, synthesizes the equivalent inertia of heterogeneous subsystems into a unified scalar index—the comprehensive system inertia J_sys—and uses continuous minimization of J_sys as the control objective. Reinforcement learning is employed to coordinate the power distribution and operating states of the motor, battery, engine, and thermal management systems. Compared with existing technologies, this method has the following advantages: 1. Constructing equivalent inertia: It unifies and quantifies the three subsystems—motor, battery, and engine—with different physical responses and response time constants differing by more than 100 times, into a comparable and synthesizable equivalent inertia index with the same dimension (seconds). 、 , Furthermore, based on this, the system's integrated inertia J_sys was defined as a single quantitative evaluation index for the vehicle's transient response capability, solving the technical problem of the inability to uniformly evaluate the response capability of heterogeneous subsystems. 2. The transient correction term in the J_sys synthesis formula is obtained through the battery power deficit ratio factor. and the rate of change in power demand | This precisely describes the physical phenomenon where the inertia difference is amplified when battery power is insufficient and driver demands change rapidly; when battery power is sufficient ( = 0) This item automatically resets to zero, achieving full-condition self-adaptation without the need for manual threshold setting. 3. The control concept shifts from threshold-triggered compensation to continuous minimization. The control objective is set as solving for the combination of control variables that minimizes J_sys in each control cycle. The system is always in the optimal response state under the current constraints, avoiding the passive lag of traditional strategies that only trigger after deterioration occurs. 4. Thermal inertia is incorporated into a unified inertia optimization framework. Thermal-electric-mechanical coupling optimization is achieved using existing vehicle components (four-way valve, electric water pump) without increasing hardware costs. The strategy network automatically determines when to intervene in thermal management to improve response capabilities. 5. Safety constraint layer design: A deterministic physical boundary safety check is set at the output of the strategy network, completely decoupling the exploratory nature of reinforcement learning from execution safety, ensuring that the final command does not violate the hardware physical boundaries of the battery and engine. Attached Figure Description
[0018] Appendix Figure 1 This is a schematic diagram of the system topology in this invention.
[0019] Appendix Figure 2 This is a flowchart illustrating the overall architecture of the present invention. Detailed Implementation
[0020] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0021] This invention proposes a comprehensive inertia assessment and transient response optimization control method for range-extended electric vehicle systems. This method defines the power response time constants of the motor, battery, and engine as equivalent inertia, respectively. , The system inertia is synthesized into a comprehensive inertia J_sys. The PPO reinforcement learning policy network continuously minimizes J_sys, coordinates the battery-engine power distribution, engine target speed and thermal management actuator control, realizes the optimization of the vehicle's transient response, and applies physical boundary safety constraints to the policy network output before sending it to the controllers of each subsystem for execution. The equivalent inertia , , These three metrics respectively characterize the response speed of the motor, battery, and engine to power commands, with the unit being seconds. The smaller the τ, the stronger the transient response capability. These three metrics unify the response capabilities of subsystems with different physical natures into comparable indicators of the same dimension. The system's overall inertia J_sys is based on , , The actual power output status of each subsystem is obtained by power weighting and introducing an adaptive transient correction driven by the battery power gap and the rate of change of driver power demand. This is used as a single quantitative evaluation index of the vehicle's transient response capability. The transient correction term is automatically zeroed when the battery power is sufficient and adaptively amplified when the battery power is insufficient. The PPO reinforcement learning policy network includes the above, , The state vector is the input, and the output includes the control action of the battery-engine power allocation ratio α, the engine target speed ωe_target, and the thermal management actuator control quantity. The physical boundary security constraints limit the original control output of the policy network to within the operating range allowed by the hardware of each subsystem; including the following steps: Step S1: Obtain the operating status parameters of each subsystem in real time; Step S2: Calculate the equivalent inertia of the motor based on the operating state parameters. Battery equivalent inertia Equivalent inertia of the engine ; Step S3: Based on the above , , and power parameters, synthesized system inertia J_sys; Step S4: Perform forward inference using the PPO policy network and output control actions; Step S5: Perform safety constraint processing on the raw control quantity output by the policy network to obtain the final control quantity; Step S6: Convert the final control quantity into control commands for each subsystem and issue them, then return to step S1 to form a closed loop.
[0022] In step S2: the equivalent inertia of the motor, wherein The electromagnetic time constant of the motor is determined by the inductance and resistance of the motor windings. The temperature correction factor is obtained by looking up a table using a preset motor temperature-derating mapping relationship, within the normal operating temperature range of the motor. The value is 1, when the motor temperature approaches the demagnetization temperature threshold. The value is greater than 1; The battery equivalent inertia ,in Due to the inherent response delay of the battery, The preset battery power limit threshold SOC value, [ , [This refers to the preset optimal operating temperature range for the battery.] For SOC-power limiting sensitivity coefficient, , These are the penalty coefficients for low and high temperatures, respectively; the temperature correction term is... lie in[ , The value is zero within the interval, and below that. The low temperature correction takes effect when the temperature is higher than 10℃. The high-temperature correction item is now in effect. The engine equivalent inertia ,in The rotational speed response time constant, determined by the mechanical moment of inertia, is composed of the combined moment of inertia of the engine crankshaft, flywheel, and generator, as well as the engine's maximum net torque. This is the gas delay time constant, the value of which depends on the engine coolant temperature. The temperature is obtained by looking up a table based on a preset coolant temperature-gas delay mapping relationship, under cold conditions. The value is greater than the value during the warm-up state; The system inertia J_sys mentioned in step S3 is calculated according to the following formula:
[0023] in, This represents the battery's current actual output power. This refers to the equivalent electrical power of the engine-generator set. To prevent the use of a pre-defined small constant with a denominator of zero, For transient correction of gain coefficient; This represents the battery power deficit ratio. This represents the battery's current maximum allowable discharge power. For the driver's power requirements, The rate of change of power demand; In the formula, the first term As a power-weighted basic term, it uses the actual output power of the battery and engine as the weights. and Calculate the weighted average; the second term The third term is a fundamental term for the motor, representing the lower limit of the non-eliminable physical response; This is a transient correction term, applied when the battery power is sufficient. Reset to zero; this item will automatically reset when battery power is low. Adaptive amplification to reflect the degree of exposure to engine slow response; The PPO policy network described in step S4: The state vector is 14-dimensional and includes: driver pedal depth. Pedal depth change rate Vehicle speed v, acceleration a, driver power requirements, battery state of charge (SOC), battery temperature Battery equivalent inertia Current engine speed ωe, engine coolant temperature Engine equivalent inertia Motor temperature T_motor, motor equivalent inertia And the ambient temperature T_ambient; The control action is 4-dimensional, including: battery-engine power distribution ratio α, engine target speed ωe_target, four-way valve opening u_valve, and water pump duty cycle u_pump; The policy network takes a multi-objective reward function that includes minimizing J_sys, maintaining the SOC range, minimizing fuel consumption, and optimizing engine efficiency as its optimization objective. It adopts a three-layer fully connected forward propagation structure and performs deterministic forward inference once per control cycle during runtime. The expression for the reward function R is: R = -w1 · J_sys - w2 · [max(SOC_min - SOC, 0)]² - w3 · [max(SOC - SOC_max, 0)]² - w4 · fuel_rate - w5 · |ω_e - ω_e_opt(P_eng)| + w6 · exp( - · J_sys) Where w1~w6 are preset weighting coefficients, SOC_min and SOC_max are the preset lower and upper bounds of the SOC safety range, respectively, fuel_rate is the engine instantaneous fuel consumption rate, and ω_e_opt(P_eng) is the efficiency-optimal speed obtained by looking up the preset optimal efficiency curve with engine output power P_eng. The six items correspond to the following in order: minimizing the main optimization objective of J_sys, second-order penalty for lower limit of SOC, second-order penalty for upper limit of SOC, minimizing fuel consumption rate, penalty for engine efficiency deviation, and positive transient quality reward when the system responds quickly to sudden acceleration. The physical domain of the 4D control action is: α ∈ [-α_chg_max, 1], where α>0 represents battery discharge and α<0 represents battery charging; ωe_target ∈ ,in and These represent the minimum and maximum allowable engine speeds, respectively; u_valve ∈ [0, 1], where 0 indicates that the coolant is fully connected to the radiator circuit and 1 indicates that the coolant is fully connected to the battery heating circuit; u_pump ∈ [0, 1] is used to control the coolant circulation flow rate; The policy network has a three-layer fully connected structure: each layer has 128 neurons, the hidden layer uses the ReLU function as the activation function, and the output layer uses the tanh function to restrict the output to [-1, 1] and then linearly maps it to the physical domain of each action; deterministic reasoning is used at runtime, without random sampling and parameter updates.
[0024] During inference, the policy network uses the rate of change of pedal depth in the state vector as a basis. Anticipate the driver's acceleration intention and adjust accordingly to the driver's power requirements. Before the filter delay is fully established, the increased engine target speed ωe_target is output in advance, pre-raising the engine speed to the speed range corresponding to the expected power demand, so that... When the engine is actually set up, it is already at or close to the target speed. Only the torque build-up stage needs to be completed to output power, thereby shortening the effective response delay. The policy network actively adjusts the thermal management actuator state by outputting u_valve and u_pump, including: when the battery temperature... The engine coolant temperature is below the lower limit of the preset optimal temperature range. Higher than At this time, the four-way valve is switched to the battery heating circuit, using the waste heat of the engine coolant to heat the battery and reduce its temperature. When the motor temperature T_motor approaches the demagnetization derating threshold, increase the water pump duty cycle u_pump to increase the coolant circulation flow rate for pre-cooling the motor and preventing... Increased due to heat derating; and reduced by maintaining engine warm-up. ; The safety constraint processing described in step S5 is as follows: ,in This represents the battery's current maximum allowable charging power. This is the preset maximum charging rate limit; ωe_target = ; u_valve = ; u_pump = .
[0025] Step S1: Sensor data acquisition: The VCU collects the following real-time data via the CAN bus during each control cycle: pedal depth from the accelerator pedal sensor. (Normalized to 0~1); SOC (0~1) from BMS, battery temperature (°C), Current Actual Output Power of the Battery (W), maximum allowable discharge power of the battery (W) and maximum allowable charging power (W); Engine speed ωe (rpm) from the ECU, and mechanical output power at the engine crankshaft end. (W) Coolant temperature (°C), instantaneous fuel consumption rate fuel_rate (g / s); motor temperature T_motor (°C) from MCU; vehicle speed v (km / h), acceleration a (m / s²), and ambient temperature T_ambient (°C) from vehicle body sensors; After data acquisition, the VCU performs the following internal calculations: The driver's power requirement is obtained by performing pedal MAP mapping and low-pass filtering. (W); for Time difference yields the rate of change of pedal depth (1 / s); for Time difference yields the rate of change of power demand (W / s); based on generator efficiency Calculate the equivalent electrical power of the engine-generator set. = · (W); Step S2: Subsystem equivalent inertia calculation: Calculate the equivalent inertia of each subsystem according to (1) to (3): Step S21: Calculate the equivalent inertia of the motor according to Equation 1. (1), Obtain from motor specifications (approximately 5~10ms); Obtained by looking up a preset motor temperature-derating factor mapping table; Step S22: Calculate the battery's equivalent inertia according to formula (2). , (2), Retrieve preset value (approximately 5ms); , , , and[ , All of these are preset calibration parameters. lie in[ , The temperature correction term is zero within the specified interval. Step S23: Calculate the engine's equivalent inertia according to formula (3). (3), = J_mech·Δωref / T_max_ref (4), where J_mech (combined rotational inertia of crankshaft-flywheel-generator rotor), Δωref (reference speed span), and T_max_ref (reference operating condition maximum net torque) are all preset constants; The time is obtained by looking up the preset coolant temperature-gas delay mapping table. It can reach 1~2 seconds when the machine is cold and tends to be 100~300ms after the machine is warmed up. Step S24: Introduce the equivalent inertia of the thermal system Equivalent inertia of thermal system It is related to the temperature of the motor, battery, and coolant, and is coupled to the inertia of the motor, battery, and engine, serving as a constraint input to improve the overall vehicle response capability.
[0026] Step S3: Synthesis of system inertia J_sys: Calculate J_sys in three steps according to formula (5): Step S31: Calculate the power-weighted basic terms ,in Take a preset small positive number (such as 1W).
[0027] Step S32: It is directly included as a basic item for motors.
[0028] Step S33: Calculation and transient correction term ( The preset gain coefficient (dimensions s / W) is used to sum the three terms to obtain J_sys.
[0029] Step S4: PPO Policy Network Inference: Step S41: Construct a 14-dimensional state vector s, with each dimension as shown in Table 1.
[0030] Table 1 14-dimensional state vector
[0031] Step S42: Perform a three-layer fully connected forward propagation. The VCU loads the offline-trained PPO policy network parameters θ = {W1, b1, W2, b2, W3, b3, Wμ, bμ} and calculates them sequentially: First layer (14D → 128D) h1 = ReLU(W1 · s + b1); Second layer (128D → 128D) h2 = ReLU(W2 · h1 + b2); Third layer (128D → 128D) h3 = ReLU(W3 · h2 + b3); Output layer μ = tanh(Wμ · h3 + bμ); where W1 is a 128×14D matrix, b1 is a 128D vector, W2 and W3 are 128×128D matrices, b2 and b3 are 128D vectors, Wμ is a 128×4D matrix, bμ is a 4D vector, and ReLU(x) = max(0, x). Deterministic reasoning is used at runtime, and the four components of μ are directly taken as the original control variables: [αraw,ωe_raw, u_valve_raw, u_pump_raw] = [μ1, μ2, μ3, μ4]; Step S43: Map the original control quantity to the physical domain to obtain the 4-dimensional control action as shown in Table 2.
[0032] Table 2 4D Control Actions
[0033] Where α_chg_max is the preset maximum charging rate limit (e.g., 0.5). and These are the engine's minimum and maximum permissible idle speeds, respectively. Step S44: Engine pre-positioning: Policy network based on ,exist Before the pedal MAP filter delay (hundreds of milliseconds) is established, ωe_target is increased in advance to pre-pull the engine speed to the speed range corresponding to the expected power demand. In actual setup, the engine speed is already in place or close to in place. Only the torque setup phase (200~500ms) needs to be completed to output power, skipping the slowest speed ramp-up phase (200~1000ms). The effective response delay can be shortened by 30%~50%. This behavior is automatically learned by the policy network in offline training through the joint guidance of terms ⑥ and ① of the reward function (Equation 7). The policy network automatically finds the optimal balance between the response improvement benefits of pre-positioning and the costs of fuel consumption and efficiency deviation, without the need for manual pre-setting of rules. Step S45: Active Thermal Management Control. The policy network actively adjusts the thermal management actuators through u_valve and u_pump, utilizing existing vehicle components to improve temperature conditions and reduce equivalent inertia. Specific control methods include:
[0034] The above methods are all based on the vehicle's existing hardware, without the need for additional sensors or actuators. The policy network automatically learns the timing and intensity of thermal management interventions through a reward function. Step S5: Security constraint layer processing: Perform deterministic safety checks sequentially on the raw control values output by S4: (1)
[0035] (2) ωe_target = ; (3) u_valve = u_pump =
[0036] Step S6: Issuance of control commands: The VCU converts the processed final control quantity into control commands for each subsystem and sends them via the CAN bus: Sending the target battery discharge power to the BMS. Send the engine target mechanical power to the ECU. The engine target speed ωe_target; send the four-way valve position command u_valve and the water pump duty cycle command u_pump to the thermal management actuator.
[0037] When α>0, the battery discharges at the α ratio and the engine replenishes the remaining power; when α=0, the engine alone meets the full power demand; when α<0, the battery charges at the |α| ratio.
[0038] Each subsystem controller executes the underlying closed-loop control based on the received instructions. The next cycle returns to S1, forming a continuous closed loop.
[0039] Compared with the prior art, the present invention has the following advantages: 1. Constructing equivalent inertia, the invention unifies and quantifies the three subsystems—motor, battery, and engine—which have different physical responses and response time constants that differ by more than 100 times, into comparable and synthesizable equivalent inertia indices with the same dimension (seconds). 、 , Furthermore, based on this, the system's integrated inertia J_sys was defined as a single quantitative evaluation index for the vehicle's transient response capability, solving the technical problem of the inability to uniformly evaluate the response capability of heterogeneous subsystems. 2. The transient correction term in the J_sys synthesis formula is obtained through the battery power deficit ratio factor. and the rate of change in power demand | | It accurately describes the physical phenomenon that the inertia difference is amplified when battery power is insufficient and driver demand changes rapidly; when battery power is sufficient (= 0), this item automatically returns to zero, achieving full-condition adaptation without the need for manual threshold setting. 3. The control concept shifts from threshold-triggered compensation to continuous minimization, setting the control objective as solving for the combination of control variables that minimizes J_sys in each control cycle. The system is always in the optimal response state under the current constraints, avoiding the passive lag of traditional strategies that only trigger after deterioration occurs. 4. Thermal inertia is incorporated into a unified inertia optimization framework, using existing vehicle components (four-way valve, electric water pump) to achieve thermal-electric-mechanical coupling optimization without increasing hardware costs. The strategy network automatically determines when to intervene in thermal management to improve response capability. 5. The safety constraint layer design sets a deterministic physical boundary safety check at the output of the strategy network, completely decoupling the exploratory nature of reinforcement learning from execution safety, ensuring that the final command does not violate the hardware physical boundaries of the battery and engine.
Claims
1. A method for comprehensive inertia assessment and transient response optimization control of a vehicle system, characterized in that, Includes the following steps: Step 1: Obtain the real-time status parameters of the motor, battery and engine, calculate the equivalent inertia of the motor, battery and engine respectively, and quantify the response capability of the heterogeneous subsystem into an equivalent inertia index with time dimension. Step 2: Calculate the overall inertia of the system based on the power allocation weights of each subsystem, the basic motor term, and the transient correction term driven by the battery power deficit ratio and the rate of change of power demand. Step 3: With minimizing the overall system inertia as the optimization objective, generate coordinated control commands including power distribution, engine speed, and thermal management actuators through a reinforcement learning policy network; Step 4: After performing a safety boundary check on the cooperative control command, it is output to the vehicle actuator to achieve continuous adaptive optimization of the vehicle's transient response capability.
2. The method for comprehensive inertia assessment and transient response optimization control of a vehicle system according to claim 1, characterized in that, The calculation of the motor's equivalent inertia is based on the motor's electromagnetic time constant, and a temperature correction coefficient is introduced. When the motor temperature exceeds a preset threshold, the temperature correction coefficient increases to reflect the slowed response characteristics caused by motor overheating derating. Specifically, the motor's equivalent inertia is calculated using the following formula. : (1), in The electromagnetic time constant of the motor is determined by the inductance and resistance of the motor windings. The temperature correction factor is obtained by looking up a table using the preset motor temperature-derating curve; Motor temperature.
3. The method for comprehensive inertia assessment and transient response optimization control of a vehicle system according to claim 1, characterized in that, The calculation of the battery's equivalent inertia is based on the battery's inherent electrochemical delay time constant, combined with the State of Charge (SOC) power limitation factor and the temperature penalty factor. The SOC power limitation factor increases in the low SOC range, and the temperature penalty factor increases under extreme temperature conditions to characterize the increase in response delay when the battery's available power is limited. Specifically, the battery's equivalent inertia is calculated according to the following formula. : (2), This includes the inherent electrochemical response delay of the battery; The critical SOC value for BMS power limitation; For SOC-power limiting sensitivity coefficient; , This refers to the optimal operating temperature range for the battery. , These are the penalty coefficients for low temperature and high temperature, respectively.
4. The method for comprehensive inertia assessment and transient response optimization control of a vehicle system according to claim 1, characterized in that, The engine's equivalent inertia is composed of the superposition of the rotational speed response time constant determined by the mechanical rotational inertia and the gas delay time constant; the gas delay time constant increases as the coolant temperature decreases to reflect the delay characteristics of the intake, charging, and combustion build-up processes in a cold engine state; specifically, the engine's equivalent inertia is calculated according to the following formula. : (3), of which = J_mech·Δω / T_max, which is the rotational speed response time constant determined by the mechanical moment of inertia, J_mech is the combined moment of inertia of crankshaft-flywheel-generator, and T_max is the maximum net torque; is the gas delay time constant.
5. The method for comprehensive inertia assessment and transient response optimization control of a vehicle system according to claim 1, characterized in that, Step 1 also includes thermal equivalent inertia. As a constraint input.
6. The method for comprehensive inertia assessment and transient response optimization control of a vehicle system according to claim 1, characterized in that, The formula for calculating the overall system inertia in step 2 includes a transient correction term. The value of the transient correction term is determined by the product or function of the battery power shortage ratio and the rate of change in power demand. When the battery power is sufficient, the transient correction term automatically returns to zero. When the battery power is insufficient and the rate of change in power demand is greater than a preset value, the transient correction term is adaptively amplified. Specifically, the synthesis of the overall system inertia J_sys is as follows: The equivalent inertia of heteroproton systems is synthesized into a unified scalar index J_sys: (5), of which: This represents the battery's current actual output power. = · This refers to the equivalent electrical power of the engine-generator set. This refers to the mechanical power output from the engine crankshaft end. For generator efficiency, To prevent the use of a pre-defined small constant with a denominator of zero, For transient correction gain, This represents the battery power deficit ratio. For BMS based on the current SOC and The maximum allowable discharge power of the battery can be obtained by referring to the table. For the driver's power requirements, clamp(x, a, b) means restricting x to the closed interval [a, b]. This represents the rate of change in power demand.
7. The method for comprehensive inertia assessment and transient response optimization control of a vehicle system according to claim 1, characterized in that, In step 3, the reinforcement learning policy network adopts the PPO algorithm, whose reward function has the negative value of the system's overall inertia as the main component, and includes energy consumption penalty terms and safety constraint violation penalty terms; the input state space of the policy network includes driver pedal depth, pedal depth change rate, vehicle speed, temperature of each subsystem, SOC, and current power allocation state; specifically, it includes the following: Step 3-1: Construct a 14-dimensional state space. The policy network input s includes: pedal depth. Pedal depth change rate Vehicle speed v, acceleration a, power requirement SOC, battery temperature Battery equivalent inertia, engine speed ωe, coolant temperature Engine equivalent inertia Motor temperature T_motor, motor equivalent inertia Ambient temperature T_ambient , , As an explicit input of equivalent inertia information; Step 3-2: Design the PPO strategy network to output 4-dimensional control actions (6) The policy network adopts a three-layer fully connected structure. The output layer limits the output to [-1, 1] with tanh and then linearly maps it to the physical domain of each action. During runtime, a deterministic forward propagation is performed once per control cycle without random sampling or parameter updates. The 4-dimensional actions include: The battery-engine power distribution ratio α ∈ [-α_chg_max, 1], where α>0 is for discharging and α<0 is for charging; Engine target speed ωe_target ∈ ; The opening degree of the four-way valve u_valve ∈ [0, 1], 0 for full-way radiator, 1 for full-way battery heating; The pump duty cycle u_pump ∈ [0, 1]; Step 3-3: Construct a reward function with minimizing J_sys as the primary objective: R = -w1 · J_sys - w2 · [max(SOC_min - SOC, 0)]² - w3 · [max(SOC - SOC_max, 0)]² - w4 · fuel_rate - w5 · |ω_e - ω_e_opt(P_eng)| + w6 · exp( - · J_sys ) (7), Where w1~w6 are preset weighting coefficients, fuel_rate is the engine instantaneous fuel consumption rate, SOC_min and SOC_max are preset SOC safety ranges, and ωe_opt(P_eng) is the optimal speed for efficiency.
8. The method for comprehensive inertia assessment and transient response optimization control of a vehicle system according to claim 7, characterized in that, The six functions in step 3-3 are as follows: ① Minimize J_sys; ② Hard constraint on the lower limit of SOC (w2 takes the largest value, square penalty, to prevent battery over-discharge); ③ Soft constraint on the upper limit of SOC (w3 << w2); ④ Minimize fuel consumption; ⑤ Engine efficiency deviation penalty; ⑥ Transient response quality reward - sudden acceleration (| | Large) and J_sys receives a positive reward every hour, which complements item ① and jointly drives the policy network to learn engine pre-positioning and other behaviors that actively improve transient response.
9. The method for comprehensive inertia assessment and transient response optimization control of a vehicle system according to claim 1, characterized in that, In step 4, the original output of the policy network Issued after the following definitive security checks: (1) ,in This represents the battery's current maximum allowable charging power. (2)ωe_target = ; (3)u_valve = ,u_pump = ; The security constraint layer completely decouples the learning exploration from the execution security, ensuring that the final instruction does not violate the physical boundaries of the hardware.
10. The method for comprehensive inertia assessment and transient response optimization control of a vehicle system according to claim 1, characterized in that, It also includes engine pre-positioning, wherein engine pre-positioning is based on Signals to anticipate driver intent, in Before the pedal MAP filter delay is established, the engine speed is increased to the target range in advance, skipping the slowest speed increase phase.