An energy efficiency optimization method for a heat pump type thermal management system
By introducing a deep Q network model with adaptive discount factors into the electric vehicle thermal management system, combining thermal state deviation and cyclic pressure ratio to calculate the system state index, and dynamically adjusting the compressor speed, the problem that traditional strategies cannot adapt to complex working conditions and achieve a significant improvement in energy efficiency.
Patent Information
- Application Number
- CN202510542581.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-28
AI Technical Summary
Traditional control strategies and deep Q networks with fixed discount factors are difficult to effectively cope with the complex and variable operating conditions of electric vehicle thermal management systems, and cannot achieve a dynamic balance between immediate demand and long-term goals, resulting in limited energy efficiency improvement.
The deep Q network model using an adaptive discount factor mechanism is used to calculate the system state index by combining the absolute value of thermal state deviation and the cyclic pressure ratio, and dynamically adjust the compressor speed to optimize the energy efficiency of the thermal management system.
It has achieved the adaptability and energy efficiency improvement of the electric vehicle thermal management system under complex and variable operating conditions, and can intelligently switch strategies based on the system's deviation from the target and operating intensity to meet the dynamic balance of comfort and endurance needs.
Smart Images

Figure CN120056691B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to an energy efficiency optimization method for a heat pump type thermal management system. Background Art
[0002] As a key component of sustainable transportation, the development of electric vehicles (EVs) is deeply influenced by factors such as driving range, charging convenience, and user experience. Among them, the heat pump type thermal management system has become an indispensable core component of modern electric vehicles due to its potential for efficient cooling and heating in different environments. This system undertakes two key tasks: one is to provide a comfortable temperature environment for the passenger compartment; the other is to precisely manage the working temperature of the power battery pack. However, the operation of the thermal management system is the main non-driving energy consumption source of electric vehicles.
[0003] The actual operating environment of electric vehicles is extremely complex and dynamically variable. Traditional control strategies, such as logic control based on preset rules or classical PID (Proportional-Integral-Derivative) controllers, are difficult to effectively handle systems with highly non-linear, multi-variable coupling, and strong time-varying characteristics. Rule-based methods often rely on the experience of engineers and a large number of experimental calibrations, and it is difficult to cover all working conditions and ensure optimality, and their control effects are often sub-optimal; while PID controllers are overwhelmed when dealing with systems with multiple inputs and outputs, strong non-linear coupling, difficult parameter tuning, and it is difficult to balance conflicting control objectives (such as rapid cooling and low energy consumption).
[0004] In recent years, with the development of artificial intelligence technology, reinforcement learning (RL), especially deep reinforcement learning methods such as Deep Q-Network (DQN), have shown strong self-learning and optimization capabilities in complex decision-making problems. However, the existing standard Deep Q-Network algorithms still have inherent defects when applied to scenarios such as electric vehicle thermal management that require balancing immediate needs and long-term goals. The fixed discount factor limits the flexibility of the DQN strategy and its adaptability to variable scenarios, and cannot achieve the optimal dynamic balance, thus limiting its potential in improving the real-world energy efficiency of electric vehicles. Summary of the Invention
[0005] In view of the problem that the above fixed discount factor limits the flexibility of the DQN strategy and its adaptability to variable scenarios, the present invention proposes an energy efficiency optimization method for a heat pump type thermal management system, including: obtaining historical state parameters of an electric vehicle heat pump type thermal management system, where the historical state parameters at least include: ambient temperature, target temperature, actual temperature, low-pressure side pressure, high-pressure side pressure, and compressor input power; training a deep Q network model based on the historical state parameters, using multiple state parameters at the same moment as the state vector of the deep Q network model; obtaining the real-time state vector of the heat pump type thermal management system, inputting the real-time state vector into the trained deep Q network model, and adjusting the operating speed of the compressor in the thermal management system according to the action instruction output by the model to optimize its energy efficiency; the deep Q network model also includes using an adaptive discount factor for training, and there is: ; where represents the maximum setting value of the discount factor; represents the minimum setting value of the discount factor; is a parameter tuning factor; represents the system state index; the system state index is the product of the absolute value of the thermal state deviation and the cycle pressure ratio; the absolute value of the thermal state deviation is the absolute value of the difference between the current actual temperature and the target temperature; the cycle pressure ratio is the ratio of the current high-pressure side pressure to the low-pressure side pressure.
[0006] The present invention dynamically adjusts the adaptive discount factor through a deep Q network combined with a system state index jointly determined by the thermal state deviation and the cycle pressure ratio, solving the problem in the prior art that fixed control logic or standard reinforcement learning cannot effectively adapt to the core contradiction of dynamically balancing heat demand and energy-saving endurance under the complex and variable working conditions of electric vehicles. This method enables the control strategy to intelligently switch the focus according to the degree of deviation of the system from the target and the operating degree, quickly respond when needed, and be energy-efficient in detail when stable. Compared with traditional methods and standard DQN, it significantly improves the adaptive ability and overall energy efficiency of the electric vehicle thermal management system.
[0007] Further, the state vector of the deep Q network model is specifically:
[0008] ;
[0009] where represents the state vector at time ; represents the ambient temperature at time ; represents the target temperature at time ; represents the actual temperature at time ; Indicates the moment The low - pressure side pressure; Indicates the moment The high - pressure side pressure; Indicates the moment The compressor input power at the moment.
[0010] Furthermore, the calculation method of the absolute value of the thermal state deviation is specifically as follows:
[0011] ;
[0012] Where Indicates the absolute value of the thermal state deviation at the moment ; Is the actual temperature at the moment ; Is the target temperature at the moment ;
[0013] Furthermore, the calculation method of the cycle pressure ratio is specifically as follows:
[0014] ;
[0015] Where Indicates the cycle pressure ratio at the moment ; Is the high - pressure side pressure at the moment ; Is the low - pressure side pressure at the moment ; Indicates the parameter - tuning factor.
[0016] The present invention reflects the current operating intensity or potential efficiency range of the heat pump cycle through the cycle pressure ratio. The pressure ratio is directly related to the compression work, temperature rise, and theoretical efficiency. Taking it as another key input for calculating the system state index enables subsequent adaptive adjustment to consider the current operating load of the system, rather than just the temperature deviation, improving the dimension and accuracy of adaptability.
[0017] Furthermore, the calculation method of the system state index is specifically as follows:
[0018] ;
[0019] Where Indicates the system state index; Indicates the absolute value of the thermal state deviation at the moment ; Indicates the cycle pressure ratio at the moment ;
[0020] The present invention realizes a non-linear combination method by constructing a system state index, which can effectively amplify the bad states that require priority attention, namely, large deviation and high operating intensity, while avoiding the introduction of weights that need to be manually adjusted. This construction method can better reflect the operating intensity of the system under the combined action of multiple factors than a simple linear combination, providing a more sensitive and reasonable input for the adjustment of the adaptive discount factor.
[0021] Further, the action instruction output by the deep Q-network model is defined as the adjustment amount of the compressor speed, specifically:
[0022] ;
[0023] where means reducing the operating speed of the compressor by a preset step; 0 means keeping the current speed of the compressor unchanged; means increasing the operating speed of the compressor by a preset step.
[0024] Further, the reward function used during the training of the deep Q-network model is specifically:
[0025] ;
[0026] where represents the reward function; represents the actual temperature at time ; represents the target temperature at time .
[0027] Further, the training process of the deep Q-network model further includes:
[0028] The hidden layer uses the ReLU activation function; the learning rate is adaptively adjusted using the Adam optimizer; the action decision adopts the ε-greedy strategy.
[0029] By applying the mature technologies in the current deep reinforcement learning field (ReLU, Adam, ε-greedy), the learning efficiency, stability of the model in dealing with the complex non-linear dynamics of the electric vehicle thermal management system, and the possibility of converging to a high-quality strategy are effectively improved.
[0030] Further, it also includes data cleaning and data standardization of the historical state parameters.
[0031] Further, the data cleaning is the median filtering algorithm; the data standardization is the Z-Score standardization algorithm
[0032] By preprocessing the input data, the quality and consistency of the data are improved, and the adverse effects of noise interference and different parameter scale differences on model training are reduced. This enhances the robustness of the trained model, improves the generalization ability of the model to real-world data, ensures the reliability and stability of the final control strategy in practical applications, and is superior to directly using the original data for training.
[0033] The technical effects of the present invention are as follows:
[0034] Regarding the problem of energy efficiency optimization of the electric vehicle thermal management system, the present invention proposes an intelligent control method based on the Deep Q-Network (DQN), and the key is to introduce an adaptive discount factor mechanism. Different from the prior art, this adaptive adjustment does not rely on complex external parameter estimation or artificial weights, but is driven by a system state index closely related to the physical logic of the scenario. This index is obtained by multiplying the absolute value of the real-time thermal state deviation by the pressure ratio reflecting the cycle operation intensity. The discount factor dynamically adjusted based on this index enables the DQN to intelligently switch the time-scale focus of its decision-making according to the current degree of deviation of the system from the target and the degree of "effort" in operation, so as to achieve a dynamic and effective balance between meeting immediate needs such as comfort and battery health and the long-term energy efficiency goal of maximizing the driving range. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0036] Figure 1 is a flowchart schematically showing an energy efficiency optimization method for a heat pump type thermal management system in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] The following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings.
[0039] An embodiment of an energy efficiency optimization method for a heat pump type thermal management system:
[0040] As Figure 1 shown, an energy efficiency optimization method for a heat pump type thermal management system of the present invention includes:
[0041] S1. Collect the core operating state parameters of the electric vehicle heat pump system and perform preprocessing.
[0042] The heat pump system of an electric vehicle usually includes components such as a compressor, a condenser, an evaporator, a throttling device, multiple heat exchangers, a coolant circulation circuit, and related valve groups. To achieve energy efficiency optimization control, the following parameters can be monitored in real time in this embodiment:
[0043] Collect the ambient temperature through a temperature sensor installed on the vehicle, which determines the basic condition of heat exchange between the vehicle and the outside world; synthesize the occupant compartment temperature set by the user and the target temperature of the BMS (Battery Management System), which represents the current control target; obtain the actual temperature through the occupant compartment air temperature sensor and the battery pack internal / coolant temperature sensor, which corresponds to the actual measured value of the target temperature; collect the low-pressure side pressure through the compressor suction side pressure sensor, which reflects the operating state of the evaporator; collect the high-pressure side pressure through the compressor discharge side pressure sensor, which reflects the operating state of the condenser; monitor the compressor input power of the compressor through a power sensor.
[0044] All the collected data above can be transmitted to the central control unit of the vehicle (such as the vehicle control unit VCU or a dedicated thermal management controller) in real time through the in-vehicle controller area network (CAN) bus. In this embodiment, the sampling frequency of all sensors can be set to 1Hz to capture the dynamic changes in the vehicle operating state.
[0045] Furthermore, all the collected raw data needs to be preprocessed, including:
[0046] Use methods such as median filtering to filter out the instantaneous noise or abnormal jumps in the sensor data to ensure the reliability of the data; then adopt Z-Score standardization to eliminate the influence of different physical dimensions and numerical range differences of temperature, pressure, power, etc. on the subsequent neural network model training, so that the model can fairly learn the importance of each input feature. The above preprocessing operations are well-known technologies, and the specific implementation methods will not be elaborated here.
[0047] S2. Define the state, action, and reward of the deep Q-network algorithm; construct the model of the deep Q-network algorithm.
[0048] Use the deep Q-network (DQN) as the core intelligent algorithm to learn the control strategy for optimizing the heat pump energy efficiency in a complex and dynamically changing electric vehicle operating environment.
[0049] The definitions of the key elements are as follows:
[0050] State vector ; This vector provides the set of information required for decision-making by DQN, including environmental conditions, thermal management requirements, demand satisfaction, key thermodynamic states of the system, and the current energy consumption cost. Among them represents the moment of the ambient temperature; represents the moment of the target temperature; represents the moment of the actual temperature; represents the moment of the low-pressure side pressure; represents the moment of the high-pressure side pressure; represents the moment of the compressor input power.
[0051] Action space is defined as the set of operations that the agent can choose to execute in state . In this embodiment, it is set as the adjustment amount of the compressor speed: ; Among them means reducing the compressor operating speed by a preset step; 0 means keeping the current compressor speed unchanged; means increasing the compressor operating speed by a preset step. In this embodiment, the preset step can be set as the empirical value 100, with the unit of RPM.
[0052] Reward function is used to guide the learning of DQN and is defined as the immediate scalar signal feedback by the environment after the agent executes action in state . It is used to evaluate the quality of this action. That is to say, it needs to reflect the core trade-off goal of electric vehicle thermal management. Then there is: ; Among them represents the moment of the actual temperature; represents the moment of the target temperature. When the actual temperature is closer to the target temperature, the absolute value of the deviation is smaller, and the reward value is larger (closer to 0), encouraging the agent to take actions that can minimize the temperature deviation.
[0053] Further construct the DQN model:
[0054] First, determine its network structure: adopt a multi-layer fully connected neural network with a ReLU (Rectified Linear Unit) activation function to fit the function of state to action value (Q value) . The input layer is used to receive the state vector , with a size equal to the dimension of the state vector, which is 6 in this embodiment; the output layer corresponds to the Q-values of the action set, and its size is the dimension of the action space, which is 3 in this embodiment.
[0055] Then, determine the loss function of the network, which is the mean square value of the temporal difference (TD) error based on the Bellman equation:
[0056] ;
[0057] where represents the value of the loss function; represents the expectation operation function; represents the immediate reward after executing the action; represents the discount factor; represents in the next state among all possible actions the maximum value of the corresponding Q-values; represents the current state under which the action is executed, and the expected cumulative reward after that. And the specific Q-value update formula is as follows:
[0058] ;
[0059] where represents the learning rate, which can be adaptively adjusted using the Adam optimizer in this embodiment; the meanings of the remaining parameters and symbols are the same as those in the above loss function. Finally, in this embodiment, the action decision can adopt the ε-greedy strategy, which is used to balance exploration (trying new actions) and exploitation (executing known optimal actions) during the training process.
[0060] S3. Calculate the system state index based on the operating state of the electric vehicle, and then calculate the adaptive discount factor.
[0061] In step S2, the model construction and training of the deep Q-network are completed. The core parameter of this algorithm includes the discount factor 𝛾, which determines the influence degree of the current action on future rewards. However, the operating scenarios of electric vehicles are diverse (such as urban congestion, highway cruising, fast charging, cold start, etc.), and the requirements for the focus of the thermal management strategy (quick response or extreme energy saving) are also different. A fixed discount factor cannot adapt to this dynamic demand. Therefore, in this embodiment, the discount factor is further improved later to enhance the adaptability and optimization effect of the control system.
[0062] First, quantify the most direct performance of the current system, that is, the actual temperature and the desired target The gap between them. This deviation is the fundamental reason for the action of the drive control system, and its magnitude directly reflects the completion or deficiency of the current thermal management task. There is:
[0063] ;
[0064] where (unit: °C) represents the absolute value of the thermal state deviation at time ; (unit: °C) is the actual measured temperature at time ; (unit: °C) is the target set temperature at time .
[0065] Calculate the absolute difference between the actual temperature and the desired target to obtain the absolute value of the thermal state deviation. The larger the value, the farther away from the target, and the greater the possibility or amplitude of adjustment required.
[0066] Then calculate the cycle pressure ratio, which is an index reflecting the current "working intensity" or "operating condition" of the heat pump cycle. The suction and discharge pressures of the compressor ( ) are the core thermodynamic state parameters. Their ratio, i.e., the cycle pressure ratio, is closely related to the theoretical work required by the compressor, the temperature rise capacity of the cycle, and the efficiency thermodynamically.
[0067] Generally, when a larger temperature rise is required (such as when the temperature difference between inside and outside is large) or the performance of some system components (such as heat exchangers) is poor, the pressure ratio will increase. A high pressure ratio often means that the compressor operates more laboriously and the potential efficiency of the entire cycle is lower. Therefore, the cycle pressure ratio is calculated as:
[0068] ;
[0069] where (dimensionless) represents the cycle pressure ratio at time ; (unit: Pa) is the high-side pressure at time ; (unit: Pa) is the low-side pressure at time ; represents the tuning parameter factor, which is used to avoid the denominator being zero and can be set to the empirical value 1e-5. The higher this ratio, the more the heat pump cycle operates in the range that requires more compression work and lower theoretical efficiency.
[0070] Then obtain the system state index. The adjustment requirement of the system not only depends on the degree of deviation from the target , it also depends on the current operating intensity of the system. When the system is far from the target and operating under high pressure ratio (high intensity / low potential efficiency) conditions, it indicates an unfavorable situation and urgent adjustment is needed. At this time, the algorithm needs to pay more attention. The system state index is:
[0071] ;
[0072] Among them, (unit: °C) is the calculated system state index; represents the absolute value of the thermal state deviation at time ; represents the cycle pressure ratio at time . When the temperature deviation is large and the cycle pressure ratio is high, the system state index will be amplified sharply.
[0073] Finally, determine the adaptive discount factor. According to the calculated system state index, dynamically adjust the discount factor of the DQN algorithm when evaluating future rewards . When has a high value, it indicates that the system is in a challenging state (large deviation or high operating intensity), and the algorithm should pay more attention to short-term effects and respond quickly; when has a low value, it indicates that the system state is good (close to the target and operating stably), and the algorithm should pay more attention to long-term benefits and finely optimize energy efficiency. Then the adaptive discount factor is:
[0074] ;
[0075] Among them (dimensionless) represents the adaptive discount factor; represents the preset maximum value of the discount factor, which can be set to 0.99 in this embodiment, representing the maximum degree of attention to the future; represents the preset minimum value of the discount factor, which can be set to 0.5 in this embodiment, representing the minimum degree of attention to the future; (unit: 1 / °C) is a positive sensitivity adjustment parameter used to control the sensitivity of the discount factor to the system state index, which can be set to the empirical value 0.5 in this embodiment; represents the system state index; is the base of the natural logarithm.
[0076] When tends to 0 (ideal state), tends to 1, tends to 0, tends to , at this time, the deep Q network balances the current and future rewards and maintains the stable control of the system; when increases, tends to 0, tends to 1, tends to , at this time, the deeper the Q - network strengthens the role of the current reward, the faster it adjusts the operating speed of the compressor.
[0077] S4. Integrate the deep Q - network model with an adaptive discount factor into the thermal management control unit of the electric vehicle to complete the intelligent optimization control of energy efficiency.
[0078] In step S3, the adaptive discount factor is obtained , then combining with the deep Q - network loss function determined in step S2, the final loss function is:
[0079] ;
[0080] where represents the value of the loss function; represents the expectation operation function; represents the immediate reward after executing an action; represents the adaptive discount factor; represents in the next state among all possible actions the maximum value of the corresponding Q - values; represents the current state under which, after executing the action the expected cumulative reward, and the improved Q - value update formula is as follows:
[0081] ;
[0082] where represents the learning rate, which can be adaptively adjusted using the Adam optimizer in this embodiment; the meanings of the remaining parameters and symbols are the same as those in the above - mentioned loss function.
[0083] Apply the improved DQN model containing the calculation logic to actual control:
[0084] Use the historical operation data of the electric vehicle containing various real - driving cycles, charging scenarios, and environmental conditions, that is, the operation data containing the parameters described in step S2. In one embodiment, it can be the time - series data of the parameters in step S2, to perform offline training on the model so that it learns the optimal action strategies under different and ; then deploy the trained model to the thermal management controller or the vehicle - control unit (VCU) of the electric vehicle.
[0085] The real-time control process is as follows: First, the controller collects sensor data in real time through the CAN bus to form the current state ; Then the controller according to Calculate the current system status index ; Then Input into the DQN model, the model outputs the Q value corresponding to each action (compressor speed adjustment); the controller further selects the optimal action based on the Q value ; The controller then sends control instructions to the compressor inverter and other actuators to execute the action ; Finally, the system state changes and enters the next control cycle, repeating the closed-loop control process.
Claims
1. An energy efficiency optimization method for a heat pump type thermal management system, characterized in that The method includes: Obtaining historical state parameters of an electric vehicle heat pump type thermal management system, where the historical state parameters at least include: ambient temperature, target temperature, actual temperature, low-pressure side pressure, high-pressure side pressure, and compressor input power; training a deep Q-network model based on the historical state parameters, and using multiple state parameters at the same moment as the state vector of the deep Q-network model; Obtaining the real-time state vector of the heat pump type thermal management system, inputting the real-time state vector into the trained deep Q-network model, and adjusting the operating speed of the compressor in the thermal management system according to the action instruction output by the model to optimize its energy efficiency; The deep Q-network model also includes using an adaptive discount factor for training, and there is: ; where represents the maximum set value of the discount factor; represents the minimum set value of the discount factor; is a parameter tuning factor; represents the system state index; The system state index is the product of the absolute value of the thermal state deviation and the cycle pressure ratio; the specific calculation method of the system state index is: ; wherein represents the system state index; represents at the moment the absolute value of the thermal state deviation; represents the moment the cycle pressure ratio; the absolute value of the thermal state deviation is the absolute value of the difference between the current actual temperature and the target temperature; the cycle pressure ratio is the ratio of the current high-side pressure to the low-side pressure.
2. The energy efficiency optimization method of a heat pump type thermal management system according to claim 1, characterized in that, The state vector of the deep Q-network model is specifically: ; wherein represents the state vector at time ; represents the ambient temperature at time ; represents the target temperature at time ; represents the actual temperature at time ; represents the low-pressure side pressure at time ; represents the high-pressure side pressure at time ; represents the compressor input power at time .
3. The energy efficiency optimization method of a heat pump type thermal management system according to claim 1, characterized in that, The specific calculation method of the absolute value of the thermal state deviation is: ; wherein represents the absolute value of the thermal state deviation at time ; is the actual temperature at time ; is the target temperature at time .
4. An energy efficiency optimization method for a heat pump type thermal management system according to claim 1, characterized in that, The specific calculation method of the cycle pressure ratio is: ; wherein represents the cycle pressure ratio at time ; is the high-pressure side pressure at time ; is the low-pressure side pressure at time ; represents the parameter adjustment factor.
5. The energy efficiency optimization method of a heat pump type thermal management system according to claim 1, characterized in that The action instruction output by the deep Q-network model is defined as the adjustment amount of the compressor speed, specifically: ; wherein means reducing the operating speed of the compressor by a preset step; 0 means keeping the current speed of the compressor unchanged; means increasing the operating speed of the compressor by a preset step.
6. The energy efficiency optimization method of a heat pump type thermal management system according to claim 1, characterized in that, The reward function used during the training of the deep Q-network model is specifically: ; Among them represents the reward function; represents the moment of the actual temperature; represents the moment of the target temperature.
7. An energy efficiency optimization method for a heat pump type thermal management system according to claim 1, characterized in that, The training process of the deep Q-network model further includes: The hidden layer uses the ReLU activation function; the learning rate is adaptively adjusted using the Adam optimizer; the action decision uses the ε-greedy strategy.
8. An energy efficiency optimization method for a heat pump type thermal management system according to claim 1, characterized in that, It also includes data cleaning and data standardization of the historical state parameters.
9. An energy efficiency optimization method for a heat pump type thermal management system according to claim 8, characterized in that, The data cleaning is the median filtering algorithm; the data standardization is the Z-Score standardization algorithm.
Citation Information
Patent Citations
Electric automobile air conditioner and passenger compartment heat management control method based on TD3 algorithm
CN116714411A
New energy vehicle thermal management system control method based on deep learning
CN118560223A
AI body measurement-based motion data management method and system
CN119580942A