Multi-strategy cooperative self-optimal hybrid power energy management control method and system
Through the collaborative framework of deep reinforcement learning and expert knowledge, a multi-dimensional state space and multi-objective reward function are constructed to generate real-time optimization control instructions, which solves the fuel economy, power and control smoothness problems of hybrid vehicle energy management strategies under different working conditions, and realizes efficient energy saving and real-time responsive energy management.
Patent Information
- Application Number
- CN202511023767.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-09-19
AI Technical Summary
Existing hybrid electric vehicle energy management control strategies are difficult to simultaneously take into account fuel economy, power and control smoothness under different operating conditions, and are computationally intensive, making real-time optimization impossible.
Adopting the collaborative framework of deep reinforcement learning and expert knowledge, by building a multi-dimensional state space representation system, defining the hybrid action space, constructing a multi-objective reward function, and using the deep deterministic policy gradient algorithm and model predictive control, real-time optimized control instructions are generated.
It realizes adaptive optimization under different working conditions, improves energy saving rate and real-time performance, enhances working condition adaptability and driving comfort, and solves the limitations of traditional strategies.
Smart Images

Figure CN120663900A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of energy-saving control technology for new energy vehicles, and specifically to a hybrid energy management control method and system with multi-strategy collaborative self-optimization, and more particularly to an energy management control method for hybrid vehicles based on deep reinforcement learning that integrates expert prior knowledge into the training and reasoning processes. Background Art
[0002] As the number of cars on the road increases, the non-renewable nature of fossil fuels is limiting the long-term development of the traditional automotive industry. Furthermore, high energy consumption in the transportation sector is leading to severe environmental problems. Therefore, to further energy conservation and emission reduction and achieve the "dual carbon" goals, developing new energy vehicles to gradually replace traditional vehicles has become a major development direction in academia and industry.
[0003] Vehicle energy management and control strategies are at the core of new energy vehicles, particularly hybrid vehicles and other multi-powered vehicles. The key challenge is to rationally distribute torque between these power sources. Researchers at home and abroad have proposed various energy management strategies for hybrid vehicle optimization to reduce vehicle energy consumption under different operating conditions. These strategies can generally be categorized as heuristic strategies based on prior knowledge and optimization-based strategies.
[0004] Heuristic strategies based on prior knowledge generally refer to deterministic rule-based control strategies and fuzzy rule-based control strategies. Rule-based strategies primarily rely on expert experience or offline optimization rules and primarily optimize vehicle energy distribution for the instantaneous operating conditions of new energy vehicles. Because the rules are predefined, they lack adaptability to changes in component models and operating conditions. Optimization-based strategies primarily utilize numerical algorithms for optimization and can be broadly categorized into offline and online optimization. Offline optimization strategies primarily serve as global optimization methods, aiming to minimize energy consumption across the entire operating condition. These strategies typically optimize a given operating condition using methods such as dynamic programming, stochastic dynamic programming, and the Pontryagin minimum principle. These offline optimization strategies are often used as benchmark strategies. However, these strategies require complete information about the entire vehicle speed or speed probability, and are computationally intensive. Therefore, they are not suitable for real-time control and are used only to evaluate other control strategies or as a design reference. Online optimization primarily includes model predictive control and equivalent fuel consumption minimization strategies. Existing methods rely heavily on precise mathematical solution algorithms and are computationally intensive, hindering their practical application in hybrid vehicles.
[0005] Patent document CN119078785A discloses a method for constructing a multi-cost energy management strategy for hybrid electric vehicles (HEVs), focusing on the intelligent utilization of new energy sources. This method constructs a driver module, a hierarchical cost function module, an ECMS regulation module, a hybrid power source module, and a SOC prediction module. The hierarchical cost function is then used to coordinate driving comfort, fuel economy, and mode switching costs. The driving comfort cost optimizes the discomfort losses caused by speed chain transitions during driving, creating a smoother speed transition sequence. The fuel economy cost utilizes the basic principles of ECMS to optimize equivalent fuel consumption based on real-time dual energy source output power. The mode switching cost converts the energy losses during gear shifts into a mode switching cost model. Finally, the PMP principle is used to solve the optimal total cost problem, effectively improving the driving comfort and fuel economy of the HEV. However, this method lacks a self-learning mechanism, making it difficult to adapt to changing operating conditions and driving environments. Summary of the Invention
[0006] In view of the defects in the prior art, the purpose of the present invention is to provide a hybrid energy management control method and system with multi-strategy coordinated self-optimization.
[0007] A multi-strategy coordinated self-optimization hybrid energy management control method provided by the present invention includes:
[0008] Step S1: collecting vehicle driving data to construct a multi-dimensional state space representation system;
[0009] The vehicle driving data includes vehicle speed, vehicle acceleration, power battery SOC, road slope and environmental parameters;
[0010] Step S2: defining a hybrid action space of engine torque and motor torque based on the multi-dimensional state space characterization system, and constraining the action feasible domain through real-time vehicle operating condition information;
[0011] Step S3: Construct a multi-objective reward function integrating fuel consumption rate, power balance and demand torque following;
[0012] Step S4: within the constrained action feasible domain, a deep deterministic policy gradient algorithm is used to construct a dual network architecture, and the multi-objective reward function is used as the optimization target. The output action and the action search space are constrained using expert prior knowledge, and benchmark control instructions are generated through priority experience replay;
[0013] Step S5: Use model predictive control to perform rolling optimization on the benchmark control instruction and output it.
[0014] Preferably, step S1 includes the following sub-steps:
[0015] Step S1.1: Obtaining vehicle operating condition information based on onboard sensors; the vehicle operating condition information includes current vehicle speed, vehicle longitudinal acceleration, battery SOC value, and environmental information;
[0016] Step S1.2: combining the vehicle operating condition information into a multi-dimensional state space of the vehicle at the current moment through timestamp synchronization;
[0017] Step S1.3: Normalize each dimension of the state in the established multi-dimensional state space.
[0018] Preferably, step S2 includes the following sub-steps:
[0019] Step S2.1: Analyze the powertrain configuration of the hybrid vehicle under study and model the power source motion space; define the engine torque T eng and the motor torque T mot A two-dimensional continuous action space composed of
[0020] Among them, T eng The value range is [0,F T_eng (ω eng )],T mot The value range is
[0021] [F T_mot_min (ω mot ),F T_mot_max (ω mot )];
[0022] F T_eng (·) represents the engine external characteristic curve function, ω eng Indicates engine speed, F T_mot_min (·) represents the maximum negative torque curve function of the motor, F T_mot_max (·) represents the maximum positive torque curve function of the motor, ω mot Indicates the speed of the motor;
[0023] Step S2.2: Analyze the torque boundaries of the engine and the electric motor in each mode of the hybrid vehicle based on the current vehicle operating condition information, and dynamically adjust the torque action boundaries of the dual power sources respectively.
[0024] Preferably, step S3 includes the following sub-steps:
[0025] Step S3.1: Construct the fuel consumption penalty term r fuel ; Calculate the engine instantaneous fuel consumption cost by looking up the engine speed and torque table fuel , and the negative value of fuel consumption is used as a penalty term. The formula is as follows:
[0026] r fuel =-costfuel
[0027] Step S3.2: Establish the energy balance penalty term r batt Based on expert prior knowledge and analysis of the internal resistance characteristics of vehicle batteries during charge and discharge, the following battery SOC maintenance guidance mechanism is established when the SOC is within a certain range:
[0028]
[0029] Among them, SOC lower_bound Indicates the battery SOC lower limit, SOC upper_bound Indicates the battery SOC upper limit.
[0030] Step S3.3: Establish the hybrid vehicle demand torque tracking penalty term r trq ; According to the current vehicle operating condition information, calculate the current required torque Trq req , according to the engine torque and motor torque in the two-dimensional continuous action space, the actual output torque Trq is obtained out , establish the following demand torque tracking penalty term:
[0031] r trq =-abs(Trq req -Trq out )
[0032] Step S3.4: Normalize the penalty term to the interval [0, 1] to eliminate the interference of dimensional differences on multi-objective optimization;
[0033] Step S3.5: weights are set for the fuel consumption penalty item, the power balance penalty item, and the vehicle demand tracking penalty item respectively to form a multi-objective combination optimization strategy.
[0034] Preferably, step S4 includes the following sub-steps:
[0035] Step S4.1: Establish a vehicle longitudinal dynamics model based on automobile theory, and establish simplified engine models, motor models, and battery models to calculate engine fuel consumption, motor power, battery power, battery current and voltage, and output torque, respectively;
[0036] Step S4.2: Based on the established multi-dimensional state space representation system and the two-dimensional continuous state space, an Actor-Critic dual network structure based on the DDPG algorithm is constructed, and its training stability is improved through a delayed update mechanism;
[0037] Step S4.3: Establishing expert experience action constraints and strategy guidance; establishing action constraints and guidance for the engine torque and motor torque output by the Actor network in the Actor-Critic dual network structure;
[0038] Step S4.4: Dynamically adjust the experience sampling weights in the experience replay pool based on the temporal difference error in reinforcement learning to accelerate key sample learning.
[0039] Preferably, the step S4.1 includes the following sub-steps:
[0040] Step S4.1.1: Establish the vehicle longitudinal dynamics model based on automobile theory. The formula is as follows:
[0041]
[0042] Among them, F t is the total driving force of the vehicle during operation, F f is the rolling resistance, F j is the acceleration resistance, F w is wind resistance, F i is the slope resistance, m is the weight of the car, g is the acceleration due to gravity, f is the rolling resistance coefficient, α is the slope, δ is the rotational mass coefficient, a is the acceleration of the car, and C d is the air resistance coefficient, A d is the frontal area, v is the acceleration of the car;
[0043] Step S4.1.2: Establish an engine model using the following formula:
[0044]
[0045] in, is the engine fuel consumption rate, f f (·) is the engine fuel consumption map, ω eng is the engine speed, T eng is the engine torque;
[0046] Step S4.1.3: Build the motor model using the following formula:
[0047]
[0048] Among them, P mg is the motor mechanical power, ω mg is the motor speed, T mg is the motor torque, η mg is the motor efficiency, f mg (·) is the motor efficiency Map, P mg,e is the electric power of the motor;
[0049] Step S4.1.4: Establish a battery model with the following formula:
[0050]
[0051] Among them, P bp is the battery power, P acs is the power of electric accessories, R bp is the internal resistance of the battery, V ocv is the battery open-loop voltage, f r (·) and f v (·) are the relationship between battery internal resistance and battery open-loop voltage with respect to SOC, I bp is the battery current, ΔSOC is the change in SOC, η c is the battery coulombic efficiency, Q bp is the battery capacity.
[0052] Preferably, the step S4.2 includes the following sub-steps:
[0053] Step S4.2.1: Construct a 5-layer online actor network; the first layer is the input layer, with the number of neurons equal to the dimension of the multidimensional state space; the second to fourth layers are hidden layers, which are 3-layer fully connected networks with the number of neurons [256, 256, 256] and the activation function is ReLU; the last layer is the output layer, with the number of neurons equal to the dimension of the two-dimensional continuous action space and the activation function is tanh, mapping the value to [-1, 1];
[0054] Step S422: Construct a 5-layer online critic network; the first layer is the input layer, and the number of neurons is the concatenation dimension of the multi-dimensional state space and the two-dimensional action space; the second to fourth layers are hidden layers, which are 3-layer fully connected networks, with the number of neurons being [256, 256, 256], and the activation function is ReLU; the last layer is the output layer, with the number of neurons being 1;
[0055] Step S423: Set the network delay update rate τ = 0.001, and synchronously update the target network parameters after each training step. The update formula is as follows:
[0056] θ target =τθ online +(1-τ)θ target
[0057] where θ target represents the target network parameters, θ online Indicates online update of network parameters.
[0058] Preferably, the step S4.3 includes the following sub-steps:
[0059] Step S4.3.1: The output ranges of the Actor network output action are: engine torque [0, 1]; motor torque [-1, 1]. Based on the combination of engine torque being 0, not 0, motor torque being less than 0, equal to 0, and greater than 0, the Actor network analyzes the current mode, calculates the engine speed and motor speed based on the vehicle speed, and dynamically adjusts the torque range.
[0060]
[0061] where π θ (·) represents the Actor network, s represents the current state, W eng Indicates engine speed, W motor Indicates the motor speed, f Mode (·) represents the speed calculation function, v represents the current vehicle speed;
[0062] Step S4.3.2: Based on the engine universal characteristic diagram, extract the optimal fuel consumption rate curve and construct the engine action search space.
[0063] Preferably, the step S4.3.2 further includes:
[0064] According to the test data of the battery internal resistance characteristics, the power balance penalty item is set to guide the engine and motor working points to balance the power release.
[0065] According to the present invention, a multi-strategy coordinated self-optimizing hybrid energy management and control system is provided, comprising:
[0066] Module M1: Collect vehicle driving data to build a multi-dimensional state space representation system;
[0067] The vehicle driving data includes vehicle speed, vehicle acceleration, power battery SOC, road slope and environmental parameters;
[0068] Module M2: defining a hybrid action space of engine torque and motor torque based on the multi-dimensional state space representation system, and constraining the action feasible domain through real-time vehicle operating condition information;
[0069] Module M3: Constructs a multi-objective reward function integrating fuel consumption rate, power balance and demand torque following;
[0070] Module M4: Within the constrained action feasible domain, a deep deterministic policy gradient algorithm is used to construct a dual-network architecture. The multi-objective reward function is optimized, and expert prior knowledge is used to constrain the output action and the action search space. Benchmark control instructions are generated through priority experience replay.
[0071] Module M5: Use model predictive control to perform rolling optimization on the benchmark control instruction and output it.
[0072] Preferably, the module M1 includes the following submodules:
[0073] Module M1.1: Acquires vehicle operating condition information based on onboard sensors; the vehicle operating condition information includes current vehicle speed, vehicle longitudinal acceleration, battery SOC value, and environmental information;
[0074] Module M1.2: combining the vehicle operating condition information into a multi-dimensional state space of the vehicle at the current moment through timestamp synchronization;
[0075] Module M1.3: Normalize each dimension of the state in the established multidimensional state space.
[0076] Preferably, the module M2 includes the following submodules:
[0077] Module M2.1: Analyze the powertrain configuration of the hybrid vehicle under study and model the power source motion space; define the engine torque T eng and the motor torque T mot A two-dimensional continuous action space composed of
[0078] Among them, T eng The value range is [0,F T_eng (ω eng )],T mot The value range is
[0079] [F T_mot_min (ω mot ),F T_mot_max (ω mot )];
[0080] F T_eng (·) represents the engine external characteristic curve function, ω eng Indicates engine speed, F T_mot_min (·) represents the maximum negative torque curve function of the motor, F T_mot_max (·) represents the maximum positive torque curve function of the motor, ω mot Indicates the speed of the motor;
[0081] Module M2.2: Analyzes the torque boundaries of the engine and electric motor in each mode of the hybrid vehicle based on the current vehicle operating condition information, and dynamically adjusts the torque action boundaries of the dual power sources.
[0082] Preferably, the module M3 includes the following submodules:
[0083] Module M3.1: Constructing the fuel consumption penalty term r fuel ; Calculate the engine instantaneous fuel consumption cost by looking up the engine speed and torque table fuel, and the negative value of fuel consumption is used as a penalty term. The formula is as follows:
[0084] r fuel =-cost fuel
[0085] Module M3.2: Establishing the energy balance penalty term r batt Based on expert prior knowledge and analysis of the internal resistance characteristics of vehicle batteries during charge and discharge, the following battery SOC maintenance guidance mechanism is established when the SOC is within a certain range:
[0086]
[0087] Among them, SOC lower_bound Indicates the battery SOC lower limit, SOC upper_bound Indicates the upper limit of battery SOC;
[0088] Module M3.3: Establishing the hybrid vehicle demand torque tracking penalty term r trq ; According to the current vehicle operating condition information, calculate the current required torque Trq req , according to the engine torque and motor torque in the two-dimensional continuous action space, the actual output torque Trq is obtained out , establish the following demand torque tracking penalty term:
[0089] r trq =-abs(Trq req -Trq out )
[0090] Module M3.4: Normalize the penalty term to the interval [0,1] to eliminate the interference of dimensional differences on multi-objective optimization;
[0091] Module M3.5: Set weights for the fuel consumption penalty, power balance penalty, and vehicle demand tracking penalty to form a multi-objective combination optimization strategy.
[0092] Preferably, the module M4 includes the following submodules:
[0093] Module M4.1: Establish a vehicle longitudinal dynamics model based on automotive theory, and build simplified engine, motor, and battery models to calculate engine fuel consumption, motor power, battery power, battery current and voltage, and output torque, respectively.
[0094] Module M4.2: Construct an Actor-Critic dual network structure based on the DDPG algorithm based on the established multi-dimensional state space representation system and two-dimensional continuous state space, and improve its training stability through a delayed update mechanism;
[0095] Module M4.3: Establishing expert experience action constraints and policy guidance; establishing action constraints and guidance for the engine torque and motor torque output by the actor network in the actor-critic dual network structure;
[0096] Module M4.4: Dynamically adjust the experience sampling weights in the experience replay pool based on the temporal difference error in reinforcement learning to accelerate key sample learning.
[0097] Preferably, the module M4.1 includes the following submodules:
[0098] Module M4.1.1: Establish a vehicle longitudinal dynamics model based on automobile theory. The formula is as follows:
[0099]
[0100] Among them, F t is the total driving force of the vehicle during operation, F f is the rolling resistance, F j is the acceleration resistance, F w is wind resistance, F i is the slope resistance, m is the weight of the car, g is the acceleration due to gravity, f is the rolling resistance coefficient, α is the slope, δ is the rotational mass coefficient, a is the acceleration of the car, and C d is the air resistance coefficient, A d is the frontal area, v is the acceleration of the car;
[0101] Module M4.1.2: Build an engine model with the following formula:
[0102]
[0103] in, is the engine fuel consumption rate, f f (·) is the engine fuel consumption map, ω eng is the engine speed, T eng is the engine torque;
[0104] Module M4.1.3: Build the electric motor model. The formula is as follows:
[0105]
[0106] Among them, P mg is the motor mechanical power, ω mg is the motor speed, T mg is the motor torque, η mg is the motor efficiency, f mg (·) is the motor efficiency Map, P mg,e is the electric power of the motor;
[0107] Module M4.1.4: Establish a battery model with the following formula:
[0108]
[0109] Among them, P bp is the battery power, P acs is the power of electric accessories, R bp is the internal resistance of the battery, V ocv is the battery open-loop voltage, f r (·) and f v (·) are the relationship between battery internal resistance and battery open-loop voltage with respect to SOC, I bp is the battery current, ΔSOC is the change in SOC, η c is the battery coulombic efficiency, Q bp is the battery capacity.
[0110] Preferably, the module M4.2 includes the following submodules:
[0111] Module M4.2.1: Construct a 5-layer online actor network; the first layer is the input layer, with the number of neurons equal to the dimensionality of the multidimensional state space; the second to fourth layers are hidden layers, which are 3-layer fully connected networks with the number of neurons [256, 256, 256] and the activation function is ReLU; the last layer is the output layer, with the number of neurons equal to the dimensionality of the two-dimensional continuous action space and the activation function is tanh, mapping values to [-1, 1];
[0112] Module M422: Constructs a 5-layer online critic network; the first layer is the input layer, with the number of neurons equal to the concatenation dimension of the multidimensional state space and the two-dimensional action space; the second to fourth layers are hidden layers, which are 3-layer fully connected networks with the number of neurons [256, 256, 256] and the activation function is ReLU; the last layer is the output layer, with the number of neurons being 1;
[0113] Module M423: Set the network delay update rate τ = 0.001, and update the target network parameters synchronously after each training step. The update formula is as follows:
[0114] θ target =τθ online +(1-τ)θ target
[0115] where θ target represents the target network parameters, θ online Indicates online update of network parameters.
[0116] Preferably, the module M4.3 includes the following submodules:
[0117] Module M4.3.1: The output ranges of the Actor network output action are: engine torque [0, 1]; motor torque [-1, 1]. Based on the combination of engine torque being 0, not 0, motor torque being less than 0, equal to 0, and greater than 0, the mode is analyzed, and the engine speed and motor speed are calculated based on the vehicle speed, dynamically adjusting the torque range.
[0118]
[0119] where π θ (·) represents the Actor network, s represents the current state, W eng Indicates engine speed, W motor Indicates the motor speed, f Mode (·) represents the speed calculation function, and v represents the current vehicle speed.
[0120] Module M4.3.2: Based on the engine universal characteristic diagram, extract the optimal fuel consumption rate curve and construct the engine action search space.
[0121] Preferably, the module M4.3.2 further includes:
[0122] According to the test data of the battery internal resistance characteristics, the power balance penalty item is set to guide the engine and motor working points to balance the power release.
[0123] Compared with the prior art, the present invention has the following beneficial effects:
[0124] 1. This invention provides a multi-strategy, self-optimizing hybrid energy management method, successfully resolving the core issue of traditional control strategies, which struggle to balance fuel economy, power, and control smoothness. By leveraging a framework of deep reinforcement learning and expert knowledge, this method achieves adaptive optimization of energy distribution, achieving excellent performance under various operating conditions, with high energy savings, excellent real-time performance, and strong adaptability to various operating conditions.
[0125] 2. This invention constructs a heuristic reward function and action space constraint model based on expert prior knowledge, achieving efficient sample utilization and significantly improving convergence speed during training. Pre-screening the state-action space and limiting the feasible region effectively avoids illegal actions during exploration, improving training stability and security.
[0126] 3. This invention integrates real-time environmental and operating condition information with vehicle state parameters into the control process, constructing a multimodal sensor data-fused operating condition perception model that effectively adapts to changing operating conditions. In particular, the integration of a model predictive control framework enables smooth control command output, significantly improving driving comfort while ensuring rapid powertrain response. BRIEF DESCRIPTION OF THE DRAWINGS
[0127] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0128] Figure 1 Schematic diagram of the flow of the multi-strategy coordinated self-optimal hybrid energy management control method of the present invention.
[0129] Figure 2 This is a schematic diagram of the basic process of deep reinforcement learning based on the Actor-Critic dual network structure in the present invention.
[0130] Figure 3 Schematic diagram of the vehicle longitudinal dynamics model. DETAILED DESCRIPTION
[0131] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0132] Reference Figure 1 As shown in FIG, a multi-strategy coordinated self-optimal hybrid energy management control method includes:
[0133] Step S1: Using sensors to collect real-time vehicle driving data, including vehicle speed, vehicle acceleration, power battery SOC, road slope, and environmental parameters, to construct a multi-dimensional state space representation system;
[0134] The specific implementation process of constructing the multidimensional state space representation system in step S1 is:
[0135] Step S11: Acquire vehicle operating condition information based on onboard sensors, the vehicle operating condition information including vehicle information such as current vehicle speed, vehicle longitudinal acceleration, battery SOC value, and environmental information such as road slope and type;
[0136] Step S12: combining the vehicle operating condition information into a multi-dimensional state space of the vehicle at the current moment through timestamp synchronization;
[0137] Step S13: normalize each dimension of the state in the established multi-dimensional state space.
[0138] Step S2: define a hybrid action space including engine torque and motor torque, and constrain the action feasible domain through real-time vehicle operating condition information;
[0139] The specific implementation process is:
[0140] Step S21: Analyze the powertrain configuration of the hybrid power system under study and perform power source motion space modeling. Define the engine torque T eng and the motor torque T mot The two-dimensional continuous action space composed of T eng The value range is [0,F T_eng (ω eng )],T mot The value range is [F T_mot_min (ω mot ),F T_mot_max (ω mot )]. Where F T_eng (·) represents the engine external characteristic curve function, ω eng Indicates engine speed, F T_mot_min (·) represents the maximum negative torque curve function of the motor, F T_mot_max (·) represents the maximum positive torque curve function of the motor, ω mot Indicates the speed of the motor.
[0141] Step S22: analyzing the torque boundaries of the engine and the electric motor in each mode of the hybrid vehicle according to the current vehicle operating condition information, and dynamically adjusting the torque action boundaries of the dual power sources respectively;
[0142] Step S3: Design a multi-objective reward function including fuel consumption rate, power balance and demand torque following;
[0143] The specific implementation process is:
[0144] Step S31: Constructing the fuel consumption penalty term r fuel Calculate the instantaneous fuel consumption cost of the engine by looking up the engine speed and torque table. fuel , and the negative value of fuel consumption is used as a penalty term. The formula is as follows;
[0145] r fuel =-cost fuel
[0146] Step S32: Establishing the power balance penalty term r batt Based on expert prior knowledge and analysis of the internal resistance characteristics of vehicle batteries during charge and discharge, it is found that when the SOC is within a certain range, the internal resistance is small. Therefore, the following battery SOC maintenance guidance mechanism is established:
[0147]
[0148] Among them, SOC lower_bound Indicates the battery SOC lower limit, SOC upper_bound Indicates the battery SOC upper limit.
[0149] Step S33: Establishing a hybrid vehicle demand torque tracking penalty term r trq According to the current vehicle operating condition information, the current required torque Trq can be calculated req , then according to the engine torque and motor torque in the two-dimensional continuous action space, the actual output torque Trq can be obtained out , so the following demand torque tracking penalty term is established:
[0150] r trq =-abs(Trq req -Trq out )
[0151] Step S34: To address the dimensional differences among the three indicators of fuel consumption, power balance, and torque tracking, these penalty terms need to be normalized to the interval [0, 1] to eliminate the interference of dimensional differences on the multi-objective optimization.
[0152] Step S35: setting weights for the fuel consumption penalty item, the power balance penalty item, and the vehicle demand tracking penalty item respectively to form a multi-objective combination optimization strategy.
[0153] Step S4: Build a dual-network architecture based on the deep deterministic policy gradient algorithm to generate control instructions. During the training process, use expert prior knowledge to constrain the output actions and action search space, and use priority experience replay to improve sample utilization.
[0154] The specific implementation process is:
[0155] Step S41: Establishing a vehicle longitudinal dynamics model based on automobile theory, and establishing simplified engine models, motor models, and battery models to calculate necessary parameters such as engine fuel consumption, motor power, battery power, battery current and voltage, and output torque;
[0156] The specific implementation process is:
[0157] Step S411: Establish a vehicle longitudinal dynamics model based on automobile theory. The formula is as follows:
[0158]
[0159] Among them F t is the total driving force of the vehicle during operation, F f is the rolling resistance, F j is the acceleration resistance, F w is wind resistance, F i is the slope resistance, m is the weight of the car, g is the acceleration due to gravity, f is the rolling resistance coefficient, α is the slope, δ is the rotational mass coefficient, a is the acceleration of the car, and C d is the air resistance coefficient, A dis the frontal area, and v is the acceleration of the car.
[0160] Step S412: Establish an engine model, the formula is as follows:
[0161]
[0162] in, is the engine fuel consumption rate, f f (·) is the engine fuel consumption map, ω eng is the engine speed, T eng is the engine torque.
[0163] Step S413: Establish a motor model. The formula is as follows:
[0164]
[0165] Among them, P mg is the motor mechanical power, ω mg is the motor speed, T mg is the motor torque, η mg is the motor efficiency, f mg (·) is the motor efficiency Map, P mg,e is the electric power of the motor.
[0166] Step S414: Establish a battery model, the formula is as follows:
[0167]
[0168] Among them, P bp is the battery power, P acs is the electric accessory power, R bp is the internal resistance of the battery, V ocv is the battery open-loop voltage, f r (·) and f v (·) are the relationship between battery internal resistance and battery open-loop voltage with respect to SOC, I bp is the battery current, ΔSOC is the change in SOC, η c is the battery coulombic efficiency, Q bp is the battery capacity.
[0169] Step S42: constructing an Actor-Critic dual network structure based on the DDPG algorithm based on the established multi-dimensional state representation system and the two-dimensional continuous state space, and improving its training stability through a delayed update mechanism;
[0170] The specific implementation process of step S42 is:
[0171] Step S421: The constructed online Actor network mainly has 5 layers. The first layer is the input layer, and the number of neurons is the dimension of the multidimensional state space; the second to fourth layers are hidden layers, specifically a 3-layer fully connected network, the number of neurons is [256, 256, 256], and the activation function is ReLU; the last layer is the output layer, the number of neurons is the dimension of the two-dimensional continuous action space, and the activation function is tanh, which maps the value to [-1, 1];
[0172] Step S422: The constructed online critic network mainly has 5 layers. The first layer is the input layer, and the number of neurons is the concatenation dimension of the multi-dimensional state space and the two-dimensional action space; the second to fourth layers are hidden layers, specifically a 3-layer fully connected network, the number of neurons is [256, 256, 256], and the activation function is ReLU; the last layer is the output layer, and the number of neurons is 1;
[0173] Step S423: Set the network delay update rate τ = 0.001, and synchronously update the target network parameters after each training step to ensure a smooth transition of the strategy. The update formula is as follows:
[0174] θ target =τθ online +(1-τ)θ target
[0175] where θ target represents the target network parameters, θ online Indicates online update of network parameters.
[0176] Step S43: Establishing expert experience action constraints and strategy guidance. Establish action constraints and guidance for the engine torque and motor torque output by the Actor network in the Actor-Critic dual network structure;
[0177] The specific implementation process of step S43 is:
[0178] Step S431: The output intervals of the Actor network output action are: engine torque T eng [0,1]; motor torque T motor [-1,1]. Analyze the current mode based on the combination of engine torque being 0, not 0, and motor torque being less than 0, equal to 0, and greater than 0. Calculate the engine speed and motor speed based on vehicle speed to dynamically adjust the torque range:
[0179]
[0180] where π θ (·) represents the Actor network, s represents the current state, W eng Indicates engine speed, W motorIndicates the motor speed, f Mode (·) represents the speed calculation function, and v represents the current vehicle speed.
[0181] Step S432: Based on the engine universal characteristic diagram, the optimal fuel consumption rate curve is extracted to construct the engine action search space. According to the battery internal resistance characteristic test data, the power balance penalty item is set to guide the engine and motor operating points to balance power release.
[0182] Step S44: Dynamically adjust the experience sampling weights in the experience replay pool based on the temporal difference error in reinforcement learning to accelerate key sample learning.
[0183] Step S5: After the training is completed, the integrated model predictive control framework performs rolling optimization on the benchmark control instructions output by deep reinforcement learning to achieve smooth control.
[0184] The specific implementation process of step S5 is:
[0185] Step S51: Save the trained Actor network parameters as a binary weight file and load it into the Actor network. Use vehicle sensors to obtain real-time vehicle operating condition information to form a multi-dimensional state space. Then, use the Actor network to obtain deterministic actions based on the state parameters.
[0186] Step S52: constructing a hierarchical control architecture, using the benchmark command output by the DRL as the reference trajectory of the MPC, and achieving dynamic correction and smoothness enhancement through rolling horizon optimization;
[0187] Step S53: Suppress high-frequency control fluctuations by combining differential constraints with low-pass filtering.
[0188] The present invention realizes the energy optimization control of hybrid electric vehicles based on multi-dimensional state perception. By integrating multi-source data such as vehicle speed, acceleration, battery SOC, road slope, etc. in real time, a complete system state representation is constructed, thereby realizing the full-operating-condition dynamic energy distribution of the hybrid system. In addition, the present invention realizes the unification of optimal and local smoothness of control strategy based on the hybrid decision-making architecture of deep deterministic policy gradient algorithm and expert knowledge collaboration, fully considers the energy conversion efficiency and power performance requirements under complex driving scenarios, and breaks through the limitations of traditional rule control and equivalent fuel consumption minimization strategy. The present invention significantly improves the algorithm convergence speed and training efficiency through the search space optimization and priority experience replay mechanism constrained by expert knowledge, and realizes the smooth transition and real-time adjustment of control instructions through the rolling optimization framework of model predictive control. In summary, the present invention simultaneously realizes the optimal energy distribution and real-time dynamic adaptive control of the hybrid system, and has the characteristics of high energy utilization efficiency, good control stability and rapid system response.
[0189] The present invention also provides a multi-strategy collaborative self-optimizing hybrid energy management and control system, which can be implemented by executing the process steps of the multi-strategy collaborative self-optimizing hybrid energy management control method, that is, those skilled in the art can understand the multi-strategy collaborative self-optimizing hybrid energy management and control method as a preferred implementation method of the multi-strategy collaborative self-optimizing hybrid energy management and control system.
[0190] Specifically, a multi-strategy coordinated self-optimizing hybrid energy management and control system includes:
[0191] Module M1: Collect vehicle driving data to build a multi-dimensional state space representation system;
[0192] The vehicle driving data includes vehicle speed, vehicle acceleration, power battery SOC, road slope and environmental parameters;
[0193] Module M2: defining a hybrid action space of engine torque and motor torque based on the multi-dimensional state space representation system, and constraining the action feasible domain through real-time vehicle operating condition information;
[0194] Module M3: Constructs a multi-objective reward function integrating fuel consumption rate, power balance and demand torque following;
[0195] Module M4: Within the constrained action feasible domain, a deep deterministic policy gradient algorithm is used to construct a dual-network architecture. The multi-objective reward function is optimized, and expert prior knowledge is used to constrain the output action and the action search space. Benchmark control instructions are generated through priority experience replay.
[0196] Module M5: Use model predictive control to perform rolling optimization on the benchmark control instruction and output it.
[0197] The module M1 includes the following submodules:
[0198] Module M1.1: Acquires vehicle operating condition information based on onboard sensors; the vehicle operating condition information includes current vehicle speed, vehicle longitudinal acceleration, battery SOC value, and environmental information;
[0199] Module M1.2: combining the vehicle operating condition information into a multi-dimensional state space of the vehicle at the current moment through timestamp synchronization;
[0200] Module M1.3: Normalize each dimension of the state in the established multidimensional state space.
[0201] The module M2 includes the following submodules:
[0202] Module M2.1: Analyze the powertrain configuration of the hybrid vehicle under study and model the power source motion space; define the engine torque T eng and the motor torque T mot A two-dimensional continuous action space composed of
[0203] Among them, T eng The value range is [0,F T_eng (ω eng )],T mot The value range is
[0204] [F T_mot_min (ω mot ),F T_mot_max (ω mot )];
[0205] F T_eng (·) represents the engine external characteristic curve function, ω eng Indicates engine speed, F T_mot_min (·) represents the maximum negative torque curve function of the motor, F T_mot_max (·) represents the maximum positive torque curve function of the motor, ω mot Indicates the speed of the motor;
[0206] Module M2.2: Analyzes the torque boundaries of the engine and electric motor in each mode of the hybrid vehicle based on the current vehicle operating condition information, and dynamically adjusts the torque action boundaries of the dual power sources.
[0207] The module M3 includes the following submodules:
[0208] Module M3.1: Constructing the fuel consumption penalty term r fuel ; Calculate the engine instantaneous fuel consumption cost by looking up the engine speed and torque table fuel , and the negative value of fuel consumption is used as a penalty term. The formula is as follows:
[0209] r fuel =-cost fuel
[0210] Module M3.2: Establishing the energy balance penalty term r batt Based on expert prior knowledge and analysis of the internal resistance characteristics of vehicle batteries during charge and discharge, the following battery SOC maintenance guidance mechanism is established when the SOC is within a certain range:
[0211]
[0212] Among them, SOC lower_bound Indicates the battery SOC lower limit, SOC upper_bound Indicates the battery SOC upper limit.
[0213] Module M3.3: Establishing the hybrid vehicle demand torque tracking penalty term r trq ; According to the current vehicle operating condition information, calculate the current required torque Trq req , according to the engine torque and motor torque in the two-dimensional continuous action space, the actual output torque Trq is obtained out , establish the following demand torque tracking penalty term:
[0214] r trq =-abs(Trq req -Trq out )
[0215] Module M3.4: Normalize the penalty term to the interval [0,1] to eliminate the interference of dimensional differences on multi-objective optimization;
[0216] Module M3.5: Set weights for the fuel consumption penalty, power balance penalty, and vehicle demand tracking penalty to form a multi-objective combination optimization strategy.
[0217] The module M4 includes the following submodules:
[0218] Module M4.1: Establish a vehicle longitudinal dynamics model based on automotive theory, and build simplified engine, motor, and battery models to calculate engine fuel consumption, motor power, battery power, battery current and voltage, and output torque, respectively.
[0219] Module M4.2: Construct an Actor-Critic dual network structure based on the DDPG algorithm based on the established multi-dimensional state space representation system and two-dimensional continuous state space, and improve its training stability through a delayed update mechanism;
[0220] Module M4.3: Establishing expert experience action constraints and policy guidance; establishing action constraints and guidance for the engine torque and motor torque output by the actor network in the actor-critic dual network structure;
[0221] Module M4.4: Dynamically adjust the experience sampling weights in the experience replay pool based on the temporal difference error in reinforcement learning to accelerate key sample learning.
[0222] The module M4.1 includes the following submodules:
[0223] Module M4.1.1: Establish a vehicle longitudinal dynamics model based on automobile theory. The formula is as follows:
[0224]
[0225] Among them, F t is the total driving force of the vehicle during operation, Ff is the rolling resistance, F j is the acceleration resistance, F w is wind resistance, F i is the slope resistance, m is the weight of the car, g is the acceleration due to gravity, f is the rolling resistance coefficient, α is the slope, δ is the rotational mass coefficient, a is the acceleration of the car, and C d is the air resistance coefficient, A d is the frontal area, v is the acceleration of the car;
[0226] Module M4.1.2: Build an engine model with the following formula:
[0227]
[0228] in, is the engine fuel consumption rate, f f (·) is the engine fuel consumption map, ω eng is the engine speed, T eng is the engine torque;
[0229] Module M4.1.3: Build the electric motor model. The formula is as follows:
[0230]
[0231] Among them, P mg is the motor mechanical power, ω mg is the motor speed, T mg is the motor torque, η mg is the motor efficiency, f mg (·) is the motor efficiency Map, P mg,e is the electric power of the motor;
[0232] Module M4.1.4: Establish a battery model with the following formula:
[0233]
[0234] Among them, P bp is the battery power, P acs is the electric accessory power, R bp is the internal resistance of the battery, V ocv is the battery open-loop voltage, f r (·) and f v (·) are the relationship between battery internal resistance and battery open-loop voltage with respect to SOC, I bp is the battery current, ΔSOC is the change in SOC, η c is the battery coulombic efficiency, Q bp is the battery capacity.
[0235] The module M4.2 includes the following submodules:
[0236] Module M4.2.1: Construct a 5-layer online actor network; the first layer is the input layer, with the number of neurons equal to the dimensionality of the multidimensional state space; the second to fourth layers are hidden layers, which are 3-layer fully connected networks with the number of neurons [256, 256, 256] and the activation function is ReLU; the last layer is the output layer, with the number of neurons equal to the dimensionality of the two-dimensional continuous action space and the activation function is tanh, mapping values to [-1, 1];
[0237] Module M422: Constructs a 5-layer online critic network; the first layer is the input layer, with the number of neurons equal to the concatenation dimension of the multidimensional state space and the two-dimensional action space; the second to fourth layers are hidden layers, which are 3-layer fully connected networks with the number of neurons [256, 256, 256] and the activation function is ReLU; the last layer is the output layer, with the number of neurons being 1;
[0238] Module M423: Set the network delay update rate τ = 0.001, and update the target network parameters synchronously after each training step. The update formula is as follows:
[0239] θ target =τθ online +(1-τ)θ target
[0240] where θ target represents the target network parameters, θ online Indicates online update of network parameters.
[0241] The module M4.3 includes the following submodules:
[0242] Module M4.3.1: The output ranges of the Actor network output action are: engine torque [0, 1]; motor torque [-1, 1]. Based on the combination of engine torque being 0, not 0, motor torque being less than 0, equal to 0, and greater than 0, the mode is analyzed, and the engine speed and motor speed are calculated based on the vehicle speed, dynamically adjusting the torque range.
[0243]
[0244] where π θ (·) represents the Actor network, s represents the current state, W eng Indicates engine speed, W motor Indicates the motor speed, f Mode (·) represents the speed calculation function, and v represents the current vehicle speed.
[0245] Module M4.3.2: Based on the engine universal characteristic diagram, extract the optimal fuel consumption rate curve and construct the engine action search space.
[0246] The module M4.3.2 also includes:
[0247] According to the test data of the battery internal resistance characteristics, the power balance penalty item is set to guide the engine and motor working points to balance the power release.
[0248] Those skilled in the art will appreciate that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, and units for implementing various functions can also be considered as both software modules implementing the method and structures within the hardware component.
[0249] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.
Claims
1. A multi-strategy coordinated self-optimization hybrid energy management control method, characterized by: include: Step S1: collecting vehicle driving data to construct a multi-dimensional state space representation system; The vehicle driving data includes vehicle speed, vehicle acceleration, power battery SOC, road slope and environmental parameters; Step S2: defining a hybrid action space of engine torque and motor torque based on the multi-dimensional state space characterization system, and constraining the action feasible domain through real-time vehicle operating condition information; Step S3: Construct a multi-objective reward function integrating fuel consumption rate, power balance and demand torque following; Step S4: within the constrained action feasible domain, a deep deterministic policy gradient algorithm is used to construct a dual network architecture, and the multi-objective reward function is used as the optimization target. The output action and the action search space are constrained using expert prior knowledge, and benchmark control instructions are generated through priority experience replay; Step S5: Use model predictive control to perform rolling optimization on the benchmark control instruction and output it.
2. The multi-strategy coordinated self-optimization hybrid energy management control method according to claim 1 is characterized in that: The step S1 includes the following sub-steps: Step S1.1: Obtaining vehicle operating condition information based on onboard sensors; the vehicle operating condition information includes current vehicle speed, vehicle longitudinal acceleration, battery SOC value, and environmental information; Step S1.2: combining the vehicle operating condition information into a multi-dimensional state space of the vehicle at the current moment through timestamp synchronization; Step S1.3: Normalize each dimension of the state in the established multi-dimensional state space.
3. The multi-strategy coordinated self-optimization hybrid energy management control method according to claim 2 is characterized in that: The step S2 includes the following sub-steps: Step S2.1: Analyze the powertrain configuration of the hybrid vehicle under study and model the power source motion space; define the engine torque T eng and the motor torque T mot A two-dimensional continuous action space composed of Among them, T eng The value range is [0,F T_eng (ω eng )],T mot The value range is [F T_mot_min (ω mot ),F T_mot_max (ω mot )]; F T_eng (·) represents the engine external characteristic curve function, ω eng Indicates engine speed, F T_mot_min (·) represents the maximum negative torque curve function of the motor, F T_mot_max (·) represents the maximum positive torque curve function of the motor, ω mot Indicates the speed of the motor; Step S2.2: Analyze the torque boundaries of the engine and the electric motor in each mode of the hybrid vehicle based on the current vehicle operating condition information, and dynamically adjust the torque action boundaries of the dual power sources respectively.
4. The multi-strategy coordinated self-optimization hybrid energy management control method according to claim 3 is characterized in that: The step S3 includes the following sub-steps: Step S3.1: Construct the fuel consumption penalty term r fuel ; Calculate the engine instantaneous fuel consumption cost by looking up the engine speed and torque table fuel , and the negative value of fuel consumption is used as a penalty term. The formula is as follows: r fuel =-cost fuel Step S3.2: Establish the energy balance penalty term r batt Based on expert prior knowledge and analysis of the internal resistance characteristics of vehicle batteries during charge and discharge, the following battery SOC maintenance guidance mechanism is established when the SOC is within a certain range: Among them, SOC lower_bound Indicates the battery SOC lower limit, SOC upper_bound Indicates the upper limit of battery SOC; Step S3.3: Establish the hybrid vehicle demand torque tracking penalty term r trq ; According to the current vehicle operating condition information, calculate the current required torque Trq req , according to the engine torque and motor torque in the two-dimensional continuous action space, the actual output torque Trq is obtained out , establish the following demand torque tracking penalty term: r trq =-abs(Trq req -Trq out ) Step S3.4: Normalize the penalty term to the interval [0, 1] to eliminate the interference of dimensional differences on multi-objective optimization; Step S3.5: weights are set for the fuel consumption penalty item, the power balance penalty item, and the vehicle demand tracking penalty item respectively to form a multi-objective combination optimization strategy.
5. The multi-strategy coordinated self-optimization hybrid energy management control method according to claim 4 is characterized in that: The step S4 includes the following sub-steps: Step S4.1: Establish a vehicle longitudinal dynamics model based on automobile theory, and establish simplified engine models, motor models, and battery models to calculate engine fuel consumption, motor power, battery power, battery current and voltage, and output torque, respectively; Step S4.2: Based on the established multi-dimensional state space representation system and the two-dimensional continuous state space, an Actor-Critic dual network structure based on the DDPG algorithm is constructed, and its training stability is improved through a delayed update mechanism; Step S4.3: Establishing expert experience action constraints and strategy guidance; establishing action constraints and guidance for the engine torque and motor torque output by the Actor network in the Actor-Critic dual network structure; Step S4.4: Dynamically adjust the experience sampling weights in the experience replay pool based on the temporal difference error in reinforcement learning to accelerate key sample learning.
6. The multi-strategy coordinated self-optimization hybrid energy management control method according to claim 5 is characterized in that: The step S4.1 includes the following sub-steps: Step S4.1.1: Establish the vehicle longitudinal dynamics model based on automobile theory. The formula is as follows: Among them, F t is the total driving force of the vehicle during operation, F f is the rolling resistance, F j is the acceleration resistance, F w is wind resistance, F i is the slope resistance, m is the weight of the car, g is the acceleration due to gravity, f is the rolling resistance coefficient, α is the slope, δ is the rotational mass coefficient, a is the acceleration of the car, and C d is the air resistance coefficient, A d is the frontal area, v is the acceleration of the car; Step S4.1.2: Establish an engine model using the following formula: in, is the engine fuel consumption rate, f f (·) is the engine fuel consumption map, ω eng is the engine speed, T eng is the engine torque; Step S4.1.3: Build the motor model using the following formula: Among them, P mg is the motor mechanical power, ω mg is the motor speed, T mg is the motor torque, η mg is the motor efficiency, f mg (·) is the motor efficiency Map, P mg,e is the electric power of the motor; Step S4.1.4: Establish a battery model with the following formula: Among them, P bp is the battery power, P acs is the power of electric accessories, R bp is the internal resistance of the battery, V ocv is the battery open-loop voltage, f r (·) and f v (·) are the relationship between battery internal resistance and battery open-loop voltage with respect to SOC, I bp is the battery current, ΔSOC is the change in SOC, η c is the battery coulombic efficiency, Q bp is the battery capacity.
7. The multi-strategy coordinated self-optimization hybrid energy management control method according to claim 5 is characterized in that: The step S4.2 includes the following sub-steps: Step S4.2.1: Construct a 5-layer online Actor network; the first layer is the input layer, and the number of neurons is the dimension of the multidimensional state space; The second to fourth layers are hidden layers, which are three-layer fully connected networks with the number of neurons being [256, 256, 256] and the activation function being ReLU. The last layer is the output layer with the number of neurons being the dimension of the two-dimensional continuous action space and the activation function being tanh, which maps the value to [-1, 1]. Step S422: Construct a 5-layer online critic network; the first layer is the input layer, and the number of neurons is the concatenation dimension of the multi-dimensional state space and the two-dimensional action space; The second to fourth layers are hidden layers, which are 3-layer fully connected networks with the number of neurons being [256, 256, 256] and the activation function being ReLU; the last layer is the output layer with 1 neuron. Step S423: Set the network delay update rate τ = 0.001, and synchronously update the target network parameters after each training step. The update formula is as follows: i target =tθ online +(1-τ)θ target Among them, θ target represents the target network parameters, θ online Indicates online update of network parameters.
8. The multi-strategy coordinated self-optimization hybrid energy management control method according to claim 5 is characterized in that: The step S4.3 includes the following sub-steps: Step S4.3.1: The output intervals of the Actor network output action are: engine torque [0,1]; Motor torque [-1, 1]; based on the combination of engine torque being 0, not 0, motor torque being less than 0, equal to 0, and greater than 0, analyze which mode it is in, calculate the engine speed and motor speed based on the vehicle speed, and dynamically adjust the torque range: Among them, π θ (·) represents the Actor network, s represents the current state, W eng Indicates engine speed, W motor Indicates the motor speed, f Mode (·) represents the speed calculation function, v represents the current vehicle speed; Step S4.3.2: Based on the engine universal characteristic diagram, extract the optimal fuel consumption rate curve and construct the engine action search space.
9. The multi-strategy coordinated self-optimization hybrid energy management control method according to claim 8, characterized in that: The step S4.3.2 further includes: According to the test data of the battery internal resistance characteristics, the power balance penalty item is set to guide the engine and motor working points to balance the power release.
10. A multi-strategy coordinated self-optimizing hybrid energy management and control system, characterized by: include: Module M1: Collect vehicle driving data to build a multi-dimensional state space representation system; The vehicle driving data includes vehicle speed, vehicle acceleration, power battery SOC, road slope and environmental parameters; Module M2: defining a hybrid action space of engine torque and motor torque based on the multi-dimensional state space representation system, and constraining the action feasible domain through real-time vehicle operating condition information; Module M3: Constructs a multi-objective reward function integrating fuel consumption rate, power balance and demand torque following; Module M4: Within the constrained action feasible domain, a deep deterministic policy gradient algorithm is used to construct a dual-network architecture. The multi-objective reward function is optimized, and expert prior knowledge is used to constrain the output action and the action search space. Benchmark control instructions are generated through priority experience replay. Module M5: Use model predictive control to perform rolling optimization on the benchmark control instruction and output it.
Citation Information
Patent Citations
Multi-cost energy management strategy construction method for hybrid vehicle
CN119078785A
Cited By
Hybrid electric vehicle energy management method based on reinforcement learning
CN121425180A