A hybrid vehicle thermal management strategy generation method based on deep reinforcement learning
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-22
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本发明的目的是克服现有技术中存在的热管理策略制定周期长、成本高、精确性差的问题,提供了一种通过建模方式生成热管理策略、提高策略精准性的基于深度强化学习的混动汽车热管理策略生成方法
[0064] 1. This invention discloses a method for generating thermal management strategies for hybrid vehicles based on deep reinforcement learning. First, relevant experimental and informational data about the vehicle are collected. Then, using this data, a model is created to simulate the vehicle's powertrain and thermal management system (i.e., the heat generation system). The powertrain model and the thermal management system model are then coupled to obtain the vehicle's energy model. This energy model can calculate the heat generation of heat-generating components based on the vehicle's state, and then calculate the temperature of each heat-generating and heat-dissipating component on the vehicle, using this as the basis for the next round of calculations. The constructed model can simulate the heat generation and dissipation processes on the vehicle, thereby simulating temperature changes and providing a foundation for the simulation of automotive thermal management strategy generation. Therefore, this design constructs a vehicle energy model to simulate the temperature change process of the vehicle.
Smart Images

Figure CN115840987B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for generating thermal management strategies for hybrid vehicles based on deep reinforcement learning, specifically applicable to improving the adaptability of vehicle thermal management strategies. Background Technology
[0002] In recent years, with the increasing popularity of automobiles globally, vehicle thermal management has received growing attention from major automakers and is seen as a promising emerging field awaiting development. To ensure that the power components of a vehicle operate within a reasonable temperature range, the rational regulation of the temperature requirements of various vehicle systems has become an important research and development direction in the field of automotive thermal management technology. Traditional automotive thermal management research mainly focuses on the engine cooling system, while electric vehicle research primarily concentrates on the technical research of the temperature field of the power battery.
[0003] Currently, most thermal management strategies are rule-based or fuzzy control algorithm-based. While these methods are intuitive and effective for specific vehicle models, they require calibration based on engineering experience, and different models need to be reconfigured, which consumes a lot of human resources and time. In addition, the calibration results are also highly subjective and have poor accuracy. Summary of the Invention
[0004] The purpose of this invention is to overcome the problems of long development cycle, high cost and poor accuracy of thermal management strategies in the prior art, and to provide a method for generating thermal management strategies for hybrid vehicles based on deep reinforcement learning by generating thermal management strategies through modeling and improving the accuracy of the strategies.
[0005] To achieve the above objectives, the technical solution of the present invention is:
[0006] A method for generating thermal management strategies for hybrid vehicles based on deep reinforcement learning includes the following steps:
[0007] S1: Obtain vehicle information and status information for hybrid vehicles:
[0008] Obtain vehicle information for hybrid vehicles: Collect vehicle information data for the models whose strategies are to be generated.
[0009] The vehicle information data includes: vehicle weight. Vehicle frontal area and battery nominal capacity Engine heat output map; Motor efficiency map;
[0010] Obtain hybrid vehicle status information: Collect relevant real-vehicle status data, battery status data, and environmental status data of the model to be used for strategy generation;
[0011] The vehicle status information includes: vehicle speed. Engine speed Engine output torque Air conditioner compressor speed Fan speed and the open / closed state of the solenoid valve ;
[0012] The battery status information includes: battery current. Voltage V, internal resistance and battery temperature ;
[0013] The environmental status information includes: ambient temperature. ;
[0014] S2: Building a Hybrid Electric Vehicle Simulation Model: Establishing a vehicle energy model; establishing a vehicle power model in Simulink, establishing a thermal management system model in GT-SUITE, and coupling the vehicle power system model and the thermal management system model in Simulink to obtain the vehicle energy model;
[0015] S3: Utilize deep reinforcement learning algorithms to construct a thermal management strategy for hybrid vehicles, solve a multi-objective optimization problem including fuel economy, battery efficiency, and battery heat dissipation, and thus obtain the optimal thermal management strategy;
[0016] First, define the reward function and simulate the vehicle energy model obtained in S2, obtaining the current state S at each simulation step. t And the reward information R t Make a decision and take action A t In the next time step, obtain the new environmental state S. t+1 And reward information R t+1 The reinforcement learning strategy is learned and updated through this process, with the goal of improving system performance through trial and error, and maximizing the cumulative value of reward information. As training progresses, i.e. the loss converges, the output state-action set becomes the optimal control strategy. At this point, the hybrid vehicle thermal management strategy is generated.
[0017] S2: Building a hybrid vehicle simulation model:
[0018] S2.1 Build a vehicle powertrain model in Simulink. The vehicle's drive power is:
[0019]
[0020] in, For drive power, For engine output power, For battery power, For motor efficiency, For the efficiency of the transmission and axles;
[0021]
[0022] in, For engine output power, For the weight of the vehicle, For rolling resistance coefficient, For air drag coefficient, For windward area, For slope, For the rotational mass conversion factor, For vehicle speed, For the efficiency of the transmission and axles;
[0023] Establish an engine model:
[0024]
[0025] in, For engine output torque, For engine speed, This refers to the engine's output power. and fuel consumption rate If they are directly proportional, then the fuel consumption rate With engine output torque Engine speed It is a functional relationship, that is:
[0026]
[0027] Establish a power battery model:
[0028]
[0029] in, For battery current, For battery open circuit voltage, For battery internal resistance, Battery power;
[0030]
[0031] in, The derivative of the battery's state of charge with respect to time. This refers to the battery's nominal capacity.
[0032]
[0033] in, For battery temperature, This represents the change in battery temperature. yes and Functions of both
[0034] S2.2: Build a thermal management system model in GT-SUITE.
[0035] Establishment of thermal management model: In GT-SUITE, adjust the parameters in the thermal management model according to the parameters of the actual vehicle, model and calibrate the heat-generating components and radiators on the whole vehicle respectively, and then build the system model according to the actual thermal management system configuration.
[0036] S2.3: Couple the vehicle powertrain model and thermal management system model to the vehicle energy model in Smulink:
[0037] Based on the known vehicle state data, the heat generated by the engine, motor, and battery can be obtained in the vehicle power system model. These heats are then input into the thermal management system model. After simulation, the thermal management system model feeds back the engine temperature, motor temperature, battery temperature, and power consumption information of each energy-consuming component to the vehicle power system model. The vehicle power system model updates the relevant vehicle state data based on the feedback temperature information.
[0038] The vehicle powertrain model outputs the engine speed. Torque The engine heat output is obtained by looking up a table based on the engine heat output map;
[0039] The vehicle powertrain model outputs the motor output power, speed, and torque. Based on the motor efficiency map, the motor efficiency value for the corresponding state is obtained, and then the heat generated by the motor is calculated.
[0040] The vehicle powertrain model outputs battery output power and current. and internal resistance According to the formula The heat generated by the battery was calculated.
[0041] In S3: When constructing a hybrid vehicle thermal management strategy using a deep reinforcement learning algorithm, the reward function R in the thermal management strategy is defined as:
[0042]
[0043] Among them, respectively It is the weighting factor of fuel economy in the reward signal. It is the weighting factor for maintaining battery SOC in the reward signal. It is a weighting factor for maintaining battery temperature in the reward signal. The derivative of the battery's state of charge with respect to time. For battery temperature change, The set temperature difference limit is a constant.
[0044] The weighting factors are set with different values according to different strategy objectives.
[0045] The vehicle energy model obtained in S2 is simulated, and the current state S is obtained at each simulation step. t And the reward information R t Make a decision and take action A t In the next time step, obtain the new environmental state S. t+1 And reward information R t+1 The goal is to learn and update reinforcement learning strategies through this process, with the aim of improving system performance through trial and error, and maximizing the cumulative value of reward information.
[0046] according to The algorithm selects the action. The probability of random selection in the algorithm is In each state S t Based on previous training round selection experience, there are... Choose action A with the highest probability of obtaining the maximum reward. t ,have The probability of randomly selecting actions is intended to promote exploration. The initial value is large, but after each training round, Attenuation rate Attenuation, according to the formula right Update to make Gradually reduce random exploration as training progresses, eventually bringing the selection closer to the optimal action;
[0047] Current status ,in This represents a set of states that includes all states. For ambient temperature, For engine torque, Engine speed, For battery power, Battery temperature;
[0048] Based on the current state The algorithm selects actions, and the action set A = , , ,in This refers to the air conditioner compressor speed. This refers to the fan speed. This refers to the open / closed state of each solenoid valve. );
[0049] The Deep Q-Network (DQN) algorithm is employed, with the goal of maximizing the expected cumulative reward the agent receives from the environment. This can be calculated using the Bellman equation:
[0050]
[0051] in, A value function representing the current state-action pair; In order to achieve expectations; For the discount factor of the future value function, the update rule of Q-learning, for Assign a value:
[0052]
[0053] in, For learning rate, The value function for the current state-action pair;
[0054] As the algorithm iterates, the value function will gradually converge to the optimal value, thus achieving the optimal control strategy. That is, the sequence of actions to maximize the Q-value function:
[0055]
[0056] Using parameters A deep Q-network is used to fit the value function, avoiding state discretization:
[0057]
[0058] To improve algorithm performance, a target value network approach is adopted, designing two identical networks: an evaluation network and a target network. The evaluation network selects actions and updates parameters; periodically, the parameters are copied to the target network for delayed updates. This method reduces the correlation between the current Q-value and the target Q-value, improving the algorithm's stability. The algorithm's objective is to minimize the loss function. :
[0059]
[0060] The network parameters are continuously updated using the gradient descent algorithm. This continues until learning converges. To balance the relationship between "exploration" and "utilization" in the learning process, the following approach is adopted: During algorithm execution, the strategy has a small probability. Randomly selected action, with a relatively high probability of 1- Choose the action that maximizes the Q value; initial learning phase It is relatively large, which enhances the network's exploration capabilities, and as training progresses, Gradually decrease the decay rate to accelerate the learning process;
[0061] As training progresses, i.e. the loss function converges, the output state-action set becomes the optimal control policy.
[0062] At this point, the hybrid vehicle thermal management strategy has been generated.
[0063] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0064] 1. This invention discloses a method for generating thermal management strategies for hybrid vehicles based on deep reinforcement learning. First, relevant experimental and informational data about the vehicle are collected. Then, using this data, a model is created to simulate the vehicle's powertrain and thermal management system (i.e., the heat generation system). The powertrain model and the thermal management system model are then coupled to obtain the vehicle's energy model. This energy model can calculate the heat generation of heat-generating components based on the vehicle's state, and then calculate the temperature of each heat-generating and heat-dissipating component on the vehicle, using this as the basis for the next round of calculations. The constructed model can simulate the heat generation and dissipation processes on the vehicle, thereby simulating temperature changes and providing a foundation for the simulation of automotive thermal management strategy generation. Therefore, this design constructs a vehicle energy model to simulate the temperature change process of the vehicle.
[0065] 2. This invention, a method for generating thermal management strategies for hybrid vehicles based on deep reinforcement learning, utilizes reward functions to design different reward functions according to different needs, thereby generating different thermal management strategies to adapt to changes in the objective environment. Simultaneously, it employs deep reinforcement learning to generate the management strategies, trying as many different actions as possible during simulation for trial and error. Compared to traditional experience-based calibration methods, this provides the vehicle system with more choices and computational basis, making the strategy objectives clearer and achieving real-time and optimal thermal management strategies. Therefore, this design provides a more quantitative and explicit method for generating automotive thermal management strategies, achieving real-time and optimal thermal management strategies.
[0066] 3. This invention provides a method for generating thermal management strategies for hybrid vehicles based on deep reinforcement learning. This method is adaptive to different vehicle models, solving the problem that traditional thermal management strategies are not applicable to different vehicle models. This method only requires modeling the new vehicle model and then establishing different reward functions for deep reinforcement learning to obtain thermal management strategies based on different objectives. Therefore, this design is highly adaptable and can meet the strategy generation needs of different vehicle models.
[0067] 4. The hybrid vehicle thermal management strategy generation method based on deep reinforcement learning in this invention considers fuel economy while maximizing the efficiency of the power battery and maintaining the battery temperature near its optimal value. Therefore, the reward function objective of this invention is reasonably designed and meets the performance requirements of the vehicle. Attached Figure Description
[0068] Figure 1 This is a flowchart of the strategy generation process of the present invention.
[0069] Figure 2 This is a logic diagram of the deep reinforcement learning algorithm used in this invention.
[0070] Figure 3 This is a schematic diagram of the whole vehicle energy model obtained by coupling in Smulink in Example 4.
[0071] Figure 4 This is a line graph of the reward value during the training process in Example 4.
[0072] Figure 5 It is the average loss in each training cycle of Example 4.
[0073] Figure 6 This is a comparison curve of fuel consumption and battery SOC changes under training and verification conditions in Example 4.
[0074] Figure 7 This is a comparison chart of temperature change curves of the crew compartment, engine outlet, and battery under training and verification conditions in Example 4. Detailed Implementation
[0075] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0076] See Figures 1 to 2 A method for generating thermal management strategies for hybrid vehicles based on deep reinforcement learning, the method comprising the following steps:
[0077] S1: Obtain vehicle information and status information for hybrid vehicles:
[0078] Obtain vehicle information for hybrid vehicles: Collect vehicle information data for the models whose strategies are to be generated.
[0079] The vehicle information data includes: vehicle weight. Vehicle frontal area and battery nominal capacity Engine heat output map; Motor efficiency map;
[0080] Obtain hybrid vehicle status information: Collect relevant real-vehicle status data, battery status data, and environmental status data of the model to be used for strategy generation;
[0081] The vehicle status information includes: vehicle speed. Engine speed Engine output torque Air conditioner compressor speed Fan speed and the open / closed state of the solenoid valve ;
[0082] The battery status information includes: battery current. Voltage V, internal resistance and battery temperature ;
[0083] The environmental status information includes: ambient temperature. ;
[0084] S2: Building a Hybrid Electric Vehicle Simulation Model: Establishing a vehicle energy model; establishing a vehicle power model in Simulink, establishing a thermal management system model in GT-SUITE, and coupling the vehicle power system model and the thermal management system model in Simulink to obtain the vehicle energy model;
[0085] S3: Utilize deep reinforcement learning algorithms to construct a thermal management strategy for hybrid vehicles, solve a multi-objective optimization problem including fuel economy, battery efficiency, and battery heat dissipation, and thus obtain the optimal thermal management strategy;
[0086] First, define the reward function and simulate the vehicle energy model obtained in S2, obtaining the current state S at each simulation step. t And the reward information R t Make a decision and take action A t In the next time step, obtain the new environmental state S. t+1 And reward information R t+1 The reinforcement learning strategy is learned and updated through this process, with the goal of improving system performance through trial and error, and maximizing the cumulative value of reward information. As training progresses, i.e. the loss converges, the output state-action set becomes the optimal control strategy. At this point, the hybrid vehicle thermal management strategy is generated.
[0087] S2: Building a hybrid vehicle simulation model:
[0088] S2.1 Build a vehicle powertrain model in Simulink. The vehicle's drive power is:
[0089]
[0090] in, For drive power, For engine output power, For battery power, For motor efficiency, For the efficiency of the transmission and axles;
[0091]
[0092] in, For engine output power, For the weight of the vehicle, For rolling resistance coefficient, For air drag coefficient, For windward area, For slope, For the rotational mass conversion factor, For vehicle speed, For the efficiency of the transmission and axles;
[0093] Establish an engine model:
[0094]
[0095] in, For engine output torque, For engine speed, This refers to the engine's output power. and fuel consumption rate If they are directly proportional, then the fuel consumption rate With engine output torque Engine speed It is a functional relationship, that is:
[0096]
[0097] Establish a power battery model:
[0098]
[0099] in, For battery current, For battery open circuit voltage, For battery internal resistance, Battery power;
[0100]
[0101] in, The derivative of the battery's state of charge with respect to time. This refers to the battery's nominal capacity.
[0102]
[0103] in, For battery temperature, This represents the change in battery temperature. yes and Functions of both
[0104] S2.2: Build a thermal management system model in GT-SUITE.
[0105] Establishment of thermal management model: In GT-SUITE, adjust the parameters in the thermal management model according to the parameters of the actual vehicle, model and calibrate the heat-generating components and radiators on the whole vehicle respectively, and then build the system model according to the actual thermal management system configuration.
[0106] S2.3: Couple the vehicle powertrain model and thermal management system model to the vehicle energy model in Smulink:
[0107] Based on the known vehicle state data, the heat generated by the engine, motor, and battery can be obtained in the vehicle power system model. These heats are then input into the thermal management system model. After simulation, the thermal management system model feeds back the engine temperature, motor temperature, battery temperature, and power consumption information of each energy-consuming component to the vehicle power system model. The vehicle power system model updates the relevant vehicle state data based on the feedback temperature information.
[0108] The vehicle powertrain model outputs the engine speed. Torque The engine heat output is obtained by looking up a table based on the engine heat output map;
[0109] The vehicle powertrain model outputs the motor output power, speed, and torque. Based on the motor efficiency map, the motor efficiency value for the corresponding state is obtained, and then the heat generated by the motor is calculated.
[0110] The vehicle powertrain model outputs battery output power and current. and internal resistance According to the formula The heat generated by the battery was calculated.
[0111] In S3: When constructing a hybrid vehicle thermal management strategy using a deep reinforcement learning algorithm, the reward function R in the thermal management strategy is defined as:
[0112]
[0113] Among them, respectively It is the weighting factor of fuel economy in the reward signal. It is the weighting factor for maintaining battery SOC in the reward signal. It is a weighting factor for maintaining battery temperature in the reward signal. The derivative of the battery's state of charge with respect to time. For battery temperature change, The set temperature difference limit is a constant.
[0114] The weighting factors are set with different values according to different strategy objectives.
[0115] In S3: the definition of the reward function R is adjusted according to different usage environments and control requirements; the weighting factor is set with different values according to different strategy objectives; the actual value of the considered parameter is subtracted from the target value, the square is taken, and then multiplied by the set reward signal weighting factor; the energy consumption component parameters considered include: compressor power. Water pump power Motor inlet temperature Engine outlet temperature Crew cabin temperature .
[0116] The vehicle energy model obtained in S2 is simulated, and the current state S is obtained at each simulation step. t And the reward information R t Make a decision and take action A t In the next time step, obtain the new environmental state S. t+1 And reward information R t+1 The goal is to learn and update reinforcement learning strategies through this process, with the aim of improving system performance through trial and error, and maximizing the cumulative value of reward information.
[0117] according to The algorithm selects the action. The probability of random selection in the algorithm is In each state S t Based on previous training round selection experience, there are... Choose action A with the highest probability of obtaining the maximum reward. t ,have The probability of randomly selecting actions is intended to promote exploration. The initial value is large, but after each training round, Attenuation rate Attenuation, according to the formula right Update to make Gradually reduce random exploration as training progresses, eventually bringing the selection closer to the optimal action;
[0118] Current status ,in This represents a set of states that includes all states. For ambient temperature, For engine torque, Engine speed, For battery power, Battery temperature;
[0119] Based on the current state The algorithm selects actions, and the action set A = , , ,in This refers to the air conditioner compressor speed. This refers to the fan speed. This refers to the open / closed state of each solenoid valve. );
[0120] The Deep Q-Network (DQN) algorithm is employed, with the goal of maximizing the expected cumulative reward the agent receives from the environment. This can be calculated using the Bellman equation:
[0121]
[0122] in, A value function representing the current state-action pair; In order to achieve expectations; For the discount factor of the future value function, the update rule of Q-learning, for Assign a value:
[0123]
[0124] in, For learning rate, The value function for the current state-action pair;
[0125] As the algorithm iterates, the value function will gradually converge to the optimal value, thus achieving the optimal control strategy. That is, the sequence of actions to maximize the Q-value function:
[0126]
[0127] Using parameters A deep Q-network is used to fit the value function, avoiding state discretization:
[0128]
[0129] To improve algorithm performance, a target value network approach is adopted, designing two identical networks: an evaluation network and a target network. The evaluation network selects actions and updates parameters; periodically, the parameters are copied to the target network for delayed updates. This method reduces the correlation between the current Q-value and the target Q-value, improving the algorithm's stability. The algorithm's objective is to minimize the loss function. :
[0130]
[0131] The network parameters are continuously updated using the gradient descent algorithm. Until learning converges; in order to balance the relationship between "exploration" and "utilization" in the learning process, the following approach is adopted. During algorithm execution, the strategy has a small probability. Randomly selected action, with a relatively high probability of 1- Choose the action that maximizes the Q value; initial learning phase It is relatively large, which enhances the network's exploration capabilities, and as training progresses, Gradually decrease the decay rate to accelerate the learning process;
[0132] As training progresses, i.e. the loss function converges, the output state-action set becomes the optimal control policy.
[0133] At this point, the hybrid vehicle thermal management strategy has been generated.
[0134] The principle of this invention is explained as follows:
[0135] The vehicle energy model obtained in S2 is simulated, and the current state S is obtained at each simulation step. t And the reward information R t Make a decision and take action A t After this round of simulation is completed, the new environmental state S will be obtained in the vehicle energy model at the next time step. t+1 And reward information R t+1 This process is repeated repeatedly; through this process, the reinforcement learning strategy is learned and updated. The goal is to improve system performance through trial and error, so that the cumulative value of the reward information is maximized and the loss function converges.
[0136] Example 1:
[0137] A method for generating thermal management strategies for hybrid vehicles based on deep reinforcement learning, the method comprising the following steps:
[0138] S1: Obtain vehicle information and status information for hybrid vehicles:
[0139] Obtain vehicle information for hybrid vehicles: Collect vehicle information data for the models whose strategies are to be generated.
[0140] The vehicle information data includes: vehicle weight. Vehicle frontal area and battery nominal capacity Engine heat output map; Motor efficiency map;
[0141] Obtain hybrid vehicle status information: Collect relevant real-vehicle status data, battery status data, and environmental status data of the model to be used for strategy generation;
[0142] The vehicle status information includes: vehicle speed. Engine speed Engine output torque Air conditioner compressor speed Fan speed and the open / closed state of the solenoid valve ;
[0143] The battery status information includes: battery current. Voltage V, internal resistance and battery temperature ;
[0144] The environmental status information includes: ambient temperature. ;
[0145] S2: Building a Hybrid Electric Vehicle Simulation Model: Establishing a vehicle energy model; establishing a vehicle power model in Simulink, establishing a thermal management system model in GT-SUITE, and coupling the vehicle power system model and the thermal management system model in Simulink to obtain the vehicle energy model;
[0146] S3: Utilize deep reinforcement learning algorithms to construct a thermal management strategy for hybrid vehicles, solve a multi-objective optimization problem including fuel economy, battery efficiency, and battery heat dissipation, and thus obtain the optimal thermal management strategy;
[0147] First, define the reward function and simulate the vehicle energy model obtained in S2, obtaining the current state S at each simulation step. t And the reward information R t Make a decision and take action A t In the next time step, obtain the new environmental state S. t+1 And reward information R t+1The reinforcement learning strategy is learned and updated through this process, with the goal of improving system performance through trial and error, and maximizing the cumulative value of reward information. As training progresses, i.e. the loss converges, the output state-action set becomes the optimal control strategy. At this point, the hybrid vehicle thermal management strategy is generated.
[0148] Example 2:
[0149] Example 2 is basically the same as Example 1, except that:
[0150] S2: Building a hybrid vehicle simulation model:
[0151] S2.1 Build a vehicle powertrain model in Simulink. The vehicle's drive power is:
[0152]
[0153] in, For drive power, For engine output power, For battery power, For motor efficiency, For the efficiency of the transmission and axles;
[0154]
[0155] in, For engine output power, For the weight of the vehicle, For rolling resistance coefficient, For air drag coefficient, For windward area, For slope, For the rotational mass conversion factor, For vehicle speed, For the efficiency of the transmission and axles;
[0156] Establish an engine model:
[0157]
[0158] in, For engine output torque, For engine speed, This refers to the engine's output power. and fuel consumption rate If they are directly proportional, then the fuel consumption rate With engine output torque Engine speed It is a functional relationship, that is:
[0159]
[0160] Establish a power battery model:
[0161]
[0162] in, For battery current, For battery open circuit voltage, For battery internal resistance, Battery power;
[0163]
[0164] in, The derivative of the battery's state of charge with respect to time. This refers to the battery's nominal capacity.
[0165]
[0166] in, For battery temperature, This represents the change in battery temperature. yes and Functions of both
[0167] S2.2: Build a thermal management system model in GT-SUITE.
[0168] Establishment of thermal management model: In GT-SUITE, adjust the parameters in the thermal management model according to the parameters of the actual vehicle, model and calibrate the heat-generating components and radiators on the whole vehicle respectively, and then build the system model according to the actual thermal management system configuration.
[0169] S2.3: Couple the vehicle powertrain model and thermal management system model to the vehicle energy model in Smulink:
[0170] Based on the known vehicle state data, the heat generated by the engine, motor, and battery can be obtained in the vehicle power system model. These heats are then input into the thermal management system model. After simulation, the thermal management system model feeds back the engine temperature, motor temperature, battery temperature, and power consumption information of each energy-consuming component to the vehicle power system model. The vehicle power system model updates the relevant vehicle state data based on the feedback temperature information.
[0171] The vehicle powertrain model outputs the engine speed. Torque The engine heat output is obtained by looking up a table based on the engine heat output map;
[0172] The vehicle powertrain model outputs the motor output power, speed, and torque. Based on the motor efficiency map, the motor efficiency value for the corresponding state is obtained, and then the heat generated by the motor is calculated.
[0173] The vehicle powertrain model outputs battery output power and current. and internal resistance According to the formula The heat generated by the battery was calculated.
[0174] In S3: When constructing a hybrid vehicle thermal management strategy using a deep reinforcement learning algorithm, the reward function R in the thermal management strategy is defined as:
[0175]
[0176] Among them, respectively It is the weighting factor of fuel economy in the reward signal. It is the weighting factor for maintaining battery SOC in the reward signal. It is a weighting factor for maintaining battery temperature in the reward signal. The derivative of the battery's state of charge with respect to time. For battery temperature change, The set temperature difference limit is a constant.
[0177] The weighting factors are set with different values according to different strategy objectives.
[0178] The vehicle energy model obtained in S2 is simulated, and the current state S is obtained at each simulation step. t And the reward information R t Make a decision and take action A t In the next time step, obtain the new environmental state S. t+1 And reward information R t+1 The goal is to learn and update reinforcement learning strategies through this process, with the aim of improving system performance through trial and error, and maximizing the cumulative value of reward information.
[0179] according to The algorithm selects the action. The probability of random selection in the algorithm is In each state S t Based on previous training round selection experience, there are... Choose action A with the highest probability of obtaining the maximum reward. t ,have The probability of randomly selecting actions is intended to promote exploration. The initial value is large, but after each training round, Attenuation rate Attenuation, according to the formula right Update to make Gradually reduce random exploration as training progresses, eventually bringing the selection closer to the optimal action;
[0180] Current status ,in This represents a set of states that includes all states. For ambient temperature, For engine torque, Engine speed, For battery power, Battery temperature;
[0181] Based on the current state The algorithm selects actions, and the action set A = , , ,in This refers to the air conditioner compressor speed. This refers to the fan speed. This refers to the open / closed state of each solenoid valve. );
[0182] The Deep Q-Network (DQN) algorithm is employed, with the goal of maximizing the expected cumulative reward the agent receives from the environment. This can be calculated using the Bellman equation:
[0183]
[0184] in, A value function representing the current state-action pair; In order to achieve expectations; For the discount factor of the future value function, the update rule of Q-learning, for Assign a value:
[0185]
[0186] in, For learning rate, The value function for the current state-action pair;
[0187] As the algorithm iterates, the value function will gradually converge to the optimal value, thus achieving the optimal control strategy. That is, the sequence of actions to maximize the Q-value function:
[0188]
[0189] Using parameters A deep Q-network is used to fit the value function, avoiding state discretization:
[0190]
[0191] To improve algorithm performance, a target value network approach is adopted, designing two identical networks: an evaluation network and a target network. The evaluation network selects actions and updates parameters; periodically, the parameters are copied to the target network for delayed updates. This method reduces the correlation between the current Q-value and the target Q-value, improving the algorithm's stability. The algorithm's objective is to minimize the loss function. :
[0192]
[0193] The network parameters are continuously updated using the gradient descent algorithm. This continues until learning converges. To balance the relationship between "exploration" and "utilization" in the learning process, the following approach is adopted: During algorithm execution, the strategy has a small probability. Randomly selected action, with a relatively high probability of 1- Choose the action that maximizes the Q value; initial learning phase It is relatively large, which enhances the network's exploration capabilities, and as training progresses, Gradually decrease the decay rate to accelerate the learning process;
[0194] As training progresses, i.e. the loss function converges, the output state-action set becomes the optimal control policy.
[0195] At this point, the hybrid vehicle thermal management strategy has been generated.
[0196] Example 3:
[0197] Example 3 is basically the same as Example 2, except that:
[0198] In S3: the definition of the reward function R is adjusted according to different usage environments and control requirements; the weighting factor is set with different values according to different strategy objectives; the actual value of the considered parameter is subtracted from the target value, the square is taken, and then multiplied by the set reward signal weighting factor; the energy consumption component parameters considered include: compressor power. Water pump power Motor inlet temperature Engine outlet temperature Crew cabin temperature .
[0199] Example 4:
[0200] A multi-objective thermal management strategy for a hybrid electric vehicle was generated using a reinforcement learning algorithm, taking into account fuel consumption, battery SOC, engine outlet temperature, battery temperature, and passenger compartment temperature. The relationship between fuel consumption, battery SOC, and the thermal management system is as follows: the power of energy-consuming components such as the compressor and water pump is provided by the battery, and the engine can charge the battery. If the thermal management strategy is optimal, the power of energy-consuming components such as the compressor and water pump will be low, and the temperature of each component will be suitable, resulting in low battery output power and low SOC fluctuation. The charging power of the engine to the battery through the generator will be low, which will give the engine a better chance to operate in the high-efficiency range.
[0201] The vehicle model generated through step S3 is as follows: Figure 3 As shown.
[0202] At each training step, the agent acquires vehicle state information S. t Engine outlet temperature and crew cabin temperature And randomly output action set A t ={ , , },in For compressor speed, For engine water pump speed, The next state S fed back from the action set system after the action set system responds to the opening degree of the electronic thermostat in the engine cooling circuit (K=0: electronic thermostat closed, i.e., the engine is cooled through the small loop; K=100: electronic thermostat fully open, i.e., the engine is cooled entirely through the large loop). t+1 Through the formula:
[0203] Calculate the reward value R of the current action. t The significance of this reward function lies in minimizing fuel consumption, minimizing battery SOC fluctuations, controlling battery temperature between 25-40℃, controlling engine coolant outlet temperature between 95-115℃, and controlling passenger compartment temperature between 15-25℃. This can be achieved by adjusting the weighting factors of each item to prioritize control. And the above (S) t A t R t S t+1 The dataset is stored in the experience pool. As training progresses, when the size of the dataset in the experience pool reaches a preset value, at each subsequent training step, a certain amount of data is randomly drawn from the experience pool and stored in the current state S. t Below, some 1- The probability selects A, which has the highest reward value from the sampled dataset. t , there are also The probability continues to randomly select actions. As training continues, the newly generated dataset replaces the oldest dataset in the experience pool. In each training epoch... With attenuation rate The value is reduced initially, which encourages random action selection, i.e., promotes the agent's exploration. As training progresses, the value is gradually reduced, meaning that the agent tries to select the optimal action each time. By the end of training, the optimal action is output at each step, i.e., the optimal control strategy is generated. Figure 4 The figure shows the curve of reward values during the training process. Figure 5 The average loss in each training cycle is shown in the figure. When the reward no longer increases or the loss no longer decreases (fluctuating within an acceptable range), convergence is achieved, and training can be considered to have ended. Figure 6 At the end of the training, the fuel consumption curves and battery SOC change curves under the training and verification conditions are shown. Figure 7 This document presents the temperature variation curves of the crew compartment, engine outlet water temperature, and battery temperature under training and verification operating conditions. Figure 6 , Figure 7 It can be seen that the generated thermal management strategy still has a good control effect under different operating conditions.
Claims
1. A method for generating thermal management strategies for hybrid vehicles based on deep reinforcement learning, characterized in that: The strategy generation method includes the following steps: S1: Obtain vehicle information and status information for hybrid vehicles: Obtain vehicle information for hybrid vehicles: Collect vehicle information data for the models whose strategies are to be generated. The vehicle information data includes: vehicle weight. vehicle frontal area and battery nominal capacity Engine heat output map, motor efficiency map; Obtain hybrid vehicle status information: Collect relevant real-vehicle status data, battery status data, and environmental status data of the model to be generated from the actual vehicle test; The vehicle status information includes: vehicle speed. Engine speed Engine output torque Air conditioner compressor speed Fan speed and the open / closed state of the solenoid valve ; The battery status information includes: battery current. Voltage V, internal resistance and battery temperature ; The environmental status information includes: ambient temperature. ; S2: Building a Hybrid Electric Vehicle Simulation Model: Establishing a vehicle energy model; establishing a vehicle power model in Simulink, establishing a thermal management system model in GT-SUITE, and coupling the vehicle power system model and the thermal management system model in Simulink to obtain the vehicle energy model; Based on the known vehicle state data, the heat generated by the engine, motor, and battery can be obtained in the vehicle power system model. These heats are then input into the thermal management system model. After simulation, the thermal management system model feeds back the engine temperature, motor temperature, battery temperature, and power consumption information of each energy-consuming component to the vehicle power system model. The vehicle power system model updates the relevant vehicle state data based on the feedback temperature information. The vehicle powertrain model outputs the engine speed. Torque The engine heat output is obtained by looking up a table based on the engine heat output map; The vehicle powertrain model outputs the motor output power, speed, and torque. Based on the motor efficiency map, the motor efficiency value for the corresponding state is obtained, and then the heat generated by the motor is calculated. The vehicle powertrain model outputs battery output power and current. and internal resistance According to the formula The heat generated by the battery was calculated; S3: Utilize deep reinforcement learning algorithms to construct a thermal management strategy for hybrid vehicles, solve a multi-objective optimization problem that includes fuel economy, battery efficiency, and battery heat dissipation, and thus obtain the optimal thermal management strategy; First, define the reward function and simulate the vehicle energy model obtained in S2, obtaining the current state S at each simulation step. t and the reward information R t Make a decision and take action A t In the next time step, obtain the new environmental state S. t+1 And reward information R t+1 The reinforcement learning strategy is learned and updated through this process, with the goal of improving system performance through trial and error, and maximizing the cumulative value of reward information. As training progresses, i.e. the loss converges, the output state-action set becomes the optimal control strategy. At this point, the hybrid vehicle thermal management strategy is generated.
2. The method for generating a hybrid vehicle thermal management strategy based on deep reinforcement learning according to claim 1, characterized in that: S2: Building a hybrid vehicle simulation model: S2.1 Build a vehicle powertrain model in Simulink. The vehicle's drive power is: ; in, For drive power, For engine output power, For battery power, For motor efficiency, For the efficiency of the transmission and axles; ; in, For engine output power, For the weight of the vehicle, For rolling resistance coefficient, For air drag coefficient, For windward area, For slope, For the rotational mass conversion factor, For vehicle speed, For the efficiency of the transmission and axles; Establish an engine model: ; in, For engine output torque, For engine speed, This refers to the engine's output power. and fuel consumption rate If they are directly proportional, then the fuel consumption rate With engine output torque Engine speed It is a functional relationship, that is: ; Establish a power battery model: ; in, For battery current, For battery open circuit voltage, For battery internal resistance, Battery power; ; in, The derivative of the battery's state of charge with respect to time. This refers to the battery's nominal capacity. ; in, For battery temperature, This represents the change in battery temperature. yes and The functions of both.
3. The method for generating a hybrid vehicle thermal management strategy based on deep reinforcement learning according to claim 2, characterized in that: S2: Building a hybrid vehicle simulation model: S2.2: Build a thermal management system model in GT-SUITE. Establishment of thermal management model: In GT-SUITE, adjust the parameters in the thermal management model according to the actual vehicle parameters, model and calibrate the heat-generating components and radiators on the whole vehicle, and then build the system model according to the actual thermal management system configuration.
4. The method for generating a hybrid vehicle thermal management strategy based on deep reinforcement learning according to claim 3, characterized in that: In S3: When constructing a hybrid vehicle thermal management strategy using a deep reinforcement learning algorithm, the reward function R in the thermal management strategy is defined as: ; Among them, respectively It is the weighting factor of fuel economy in the reward signal. It is the weighting factor for maintaining battery SOC in the reward signal. It is a weighting factor for maintaining battery temperature in the reward signal. The derivative of the battery's state of charge with respect to time. For battery temperature change, The set temperature difference limit is a constant. The weighting factors are set with different values according to different strategy objectives.
5. The method for generating a hybrid vehicle thermal management strategy based on deep reinforcement learning according to claim 4, characterized in that: The vehicle energy model obtained in S2 is simulated, and the current state is obtained at each simulation step. and information on rewards received Make decisions and take action In the next time step, the new state of the environment will be obtained. and reward information The goal is to learn and update reinforcement learning strategies through this process, with the aim of improving system performance through trial and error, and maximizing the cumulative value of reward information. according to The algorithm selects the action. The probability of random selection in the algorithm is In every state Based on previous training round selection experience, there are... Choose the action that yields the greatest reward based on probability. ,have The probability of randomly selecting actions is intended to promote exploration. The initial value is large, but after each training round, Attenuation rate Attenuation, according to the formula right Update to make Gradually reduce random exploration as training progresses, eventually bringing the selection closer to the optimal action; Current status ,in This represents a set of states that includes all states. For ambient temperature, For engine torque, Engine speed, For battery power, Battery temperature; Based on the current state The algorithm selects actions, and the action set A = , , ,in This refers to the air conditioner compressor speed. This refers to the fan speed. This represents the open / closed state of each solenoid valve. The Deep Q-Network algorithm is employed, with the goal of maximizing the expected cumulative reward the agent receives from the environment. This can be calculated using the Bellman equation: ; in, A value function representing the current state-action pair; In order to achieve expectations; For the discount factor of the future value function, the update rule of Q-learning, for Assign a value: ; in, For learning rate, The value function for the current state-action pair; As the algorithm iterates, the value function will gradually converge to the optimal value, thus achieving the optimal control strategy. That is, the sequence of actions to maximize the Q-value function: ; Using parameters A deep Q-network is used to fit the value function, avoiding state discretization: ; To improve algorithm performance, a target value network approach is adopted, designing two identical networks: an evaluation network and a target network. The evaluation network selects actions and updates parameters; periodically, the parameters are copied to the target network for delayed updates. This method reduces the correlation between the current Q-value and the target Q-value, improving the algorithm's stability. The algorithm's objective is to minimize the loss function. : ; The network parameters are continuously updated using the gradient descent algorithm. Until learning converges; in order to balance the relationship between "exploration" and "utilization" in the learning process, the following approach is adopted. During algorithm execution, the strategy has a small probability. Randomly selected action, with a relatively high probability of 1- Choose the action that maximizes the Q value; initial learning phase It is relatively large, which enhances the network's exploration capabilities, and as training progresses, Gradually decrease the decay rate to accelerate the learning process; As training progresses, i.e. the loss function converges, the output state-action set becomes the optimal control policy. At this point, the hybrid vehicle thermal management strategy has been generated.
Citation Information
Patent Citations
Fuel cell vehicle energy management method based on deep reinforcement learning algorithm
CN112287463A