Automotive thermal management system and method
A reinforcement learning-based thermal management system addresses the inefficiencies and safety concerns of current systems by using real-time sensor data and adaptive control methods, resulting in enhanced efficiency and extended vehicle range and lifetime.
Patent Information
- Application Number
- PCT/JP2024/033873
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-03
- Filing Date
- 2024-09-24
- Publication Date
- 2025-05-08
AI Technical Summary
Current automotive thermal management systems face challenges in achieving optimal control efficiency and safety, particularly during real driving conditions and system degradation, due to the complexity of calibrating multiple actuators and the lack of consideration for degradation mechanisms.
The implementation of a reinforcement learning-based thermal management system that utilizes sensor data and user requests to generate optimal control settings in real-time, while ensuring safety through the use of control barrier functions and adaptive control mechanisms.
This solution enables improved efficiency and reduced energy consumption, extends the range and lifetime of battery electric vehicles, and reduces calibration efforts by allowing for automatic online training of AI models.
Smart Images

Figure JP2024033873_08052025_PF_FP_ABST
Abstract
Description
AUTOMOTIVE THERMAL MANAGEMENT SYSTEM AND METHODCross Reference
[0001] This application claims the benefit of German Patent Application No. 102023130394.5 filed on November 3, 2023. The entire disclosures of the above application are incorporated herein by reference.
[0002] The present invention relates to an automotive thermal management system and method with which, for instance, a continuous high efficient and adaptive operation of a thermal heat pump system is possible.
[0003] A heat pump is a key thermal system to condition, i.e., heat or cool, the vehicle’s cabin to a comfortable temperature range and / or to condition, i.e., heat or cool, the high voltage battery to an optimal working condition. The coolant system is required to condition the electric powertrain and the battery as energy carrier connected to the refrigerant system.
[0004] Current state-of-the-art heat pumps systems contain many actuators that are controlled to realize a target temperature in the cabin, the battery and the electric powertrain. Traditionally, in order to realize the required heating or cooling power with minimum energy losses, a lot of testing is performed under many different conditions to calibrate the system, which increases the calibration efforts significantly. Calibrations, which are found as a result of the testing, are stored into lookup maps and a control is being configured. Despite all calibration efforts not always full optimal controls can be guaranteed in the complete operation domain. Additionally, degradation mechanisms are not considered throughout the vehicle system and component lifetime.
[0005] Existing protective rights on control of automotive HVAC (heating-cooling-air-conditioning) systems based on AI (artificial intelligence) focus on single operation mode of the heat-pump system. Moreover, control system safety is not guaranteed during the learning and implementation phase. Patent US 11002202 B2 identifies the benefit of Reinforcement Learning-based adaptive on-line control in reducing calibration efforts compared to map-based control. Document CN109193075 discloses a method for controlling a power battery cooling system / pump of an electric vehicle based on reinforcement learning. A policy gradient method is used and varying scenes are mentioned. In document US2022077810, a reinforced learning algorithm for thermal systems is used and reward functions based on modified states are mentioned.
[0006] The application DE 10 2022 115 096 disclosed a data driven supervised learning model unit which is configured to compute a cost function in an optimization domain for a control optimization unit and to generate calculated intermediate outputs.
[0007] However, in any of the mentioned documents, real and safe driving operation of the thermal system and practical implementation by considering real sensors (e.g., needed for safety) is not considered.
[0008] Automotive thermal management system and method are to be provided, for which a safety guarantee in on-line calibration is given and a control compensation for system aging is possible.
[0009] This object is solved by the subject matters of the independent claims. Further aspects are defined in the subclaims.
[0010] According to a first aspect of the present invention, a thermal management system for an electric vehicle with a cabin is provided, which comprises a thermal system with sensor means for detecting ambient parameters, and operating parameters and conditions of the thermal system of the electric vehicle, and reinforcement learning based control means configured to generate an output for operating the thermal management system based on a control means input including at least one user request and / or at least one parameter detected by the sensor means, wherein the thermal system includes cooling and / or heating components with and without electric driven components and electric driven auxiliary components, wherein the control means input is selected from a group of parameters defining target air conditions in the cabin, ambient conditions, thermal system conditions and vehicle states, wherein the control means are configured to capture the relation between the inputs, in the form of a state, and outputs, in the form of at least one action, of the control means as input to the thermal system, which comprise operating parameters of the thermal systems and / or temperature conditions, for implementing real-time operation by computing an optimization cost function for the control means and to receive outputs of the thermal system as input for the control means as reward, and wherein the outputs of thermal system comprise system performance parameters of the thermal system when exploring the thermal system for finding improved rewards, and wherein the control means are configured to determine control settings for real-time operation as output for operating the thermal system with improved rewards in view of the at least one user request while avoiding a risky operational area. Preferably, the system performance parameters might include heat pump efficiency characterized by any of a coefficient of performance, a driver's comfort and a safety term.
[0011] With the thermal management system according to the first aspect, an increased range of a battery electric vehicle (BEV) can be obtained during real driving conditions, but also throughout component testing and system testing, wherein, preferably, an optimal system performance is guaranteed. Furthermore, the calibration efforts can be reduced since the thermal management system according to the present invention replaces map-based calibrations according to the prior art. As a result, there are a reduced requirement with respect to testing time and reduced engineering efforts. Moreover, such a thermal management system might have increased robustness since an optimal operation is able to be guaranteed due to adaptive and safe control.
[0012] According to a second aspect of the present invention, in the thermal management system according to the first aspect, the system performance parameters include at least an indicator for energy consumption or energy efficiency, for instance a coefficient of performance (COP) of the thermal system, and a safety term for avoiding the risky operational area of the thermal system.
[0013] According to a third aspect of the present invention, in the thermal management system according to the first or second aspect, the system performance parameters include a driving comfort term, for instance to guarantee the cooling requested.
[0014] According to a fourth aspect of the present invention, in the thermal management system according to one of the preceding aspects, in the safety term for avoiding a risky operational area of the outputs of thermal system, a control barrier function, which approaches an infinite value close to a boundary of the predefined safe set of states, is reflected and causes the states of the thermal system to lie within a predefined safe set of states, preferably a reciprocal control barrier function.
[0015] According to a further aspect of the present invention, the thermal management system according to the first or second aspect additionally comprises a digital twin of the thermal system in simulation for initially calibrating the control means. By using such a digital twin of the vehicle in simulation environment an initial calibration can be carried out. Subsequently, further improvements can be implemented in real-time operation in an automatic manner.
[0016] According to a further aspect of the present invention, in the thermal management system according to the previous aspect, the at least one action defines a safe policy to collect data for initial calibration while a target policy is updated towards an optimal policy considering the safety term.
[0017] According to a fifth aspect of the present invention, in the thermal management system according to one of the preceding aspects, the control means is configured to use a model-free or a model based learning approach.
[0018] According to a sixth aspect of the present invention, in the thermal management system according to the fifth aspect, the control means uses tabular Q-learning or Deep-Q-learning, wherein the Q-value is an estimation of the quality of the interrelationship between the state as input and the action as output.
[0019] According to a seventh aspect of the present invention, in the thermal management system according to one of the preceding aspects, the user request of the control means input includes at least one of the requested cooling power of the evaporator or the temperature at its outlet, the requested heating power of the inner condenser or the temperature at its outlet, the requested cooling power of the chiller or the temperature at its outlet, the requested heating power of the condenser or the temperature at its outlet, the recycling ratio of thermal system, the HVAC blower duty of the thermal system.
[0020] According to an eighth aspect of the present invention, in the thermal management system according to one of the preceding aspects, the parameter detected by the sensor means of the control means input includes at least one of the ambient temperature, the vehicle speed, the ambient humidity, the cabin temperature, the cabin humidity.
[0021] According to a ninth aspect of the present invention, in the thermal management system according to one of the preceding aspects, the output of the control means for operating the thermal system includes, preferably with a higher priority, at least one of the compressor speed, a control signal for a first expansion valve for controlling the two phase flow for the cabin cooling by the evaporator, wherein optionally calculating reference setpoints are considered and / or wherein the output of the control means for operating the thermal system includes, preferably with a lower priority, at least one of a control signal for a second expansion valve for controlling the two phase flow for heating, a control signal for a third expansion valve for controlling the two phase flow for battery cooling by the chiller, a control signal for a first solenoid valve for realizing parallel dehumidification, a control signal for a second solenoid valve for disabling cooling, a control signal for the fan speed, a control signal for the air shutter.
[0022] According to an tenth aspect of the present invention, in the thermal management system according to one of the first to sixth or previous further aspects, the user request of the control means input includes the requested cooling power of the evaporator or the temperature at its outlet, the parameter detected by the sensor means of the control means input includes at least one of the ambient temperature and the vehicle speed, and the output of the control means for operating the thermal system includes at least one of the compressor speed and a control signal for a first expansion valve for controlling the two phase flow for the cabin cooling by the evaporator.
[0023] According to an eleventh aspect of the present invention, in the thermal management system according to one of the preceding aspects, the thermal system comprises the refrigerant system of a heat-pump system.
[0024] According to a twelfth aspect of the present invention, in the thermal management system according to one of the preceding aspects, the thermal system comprises an electric power train / battery coolant system.
[0025] According to a thirteenth aspect of the present invention, a thermal management method for an electric vehicle with a cabin is provided, which comprises providing a thermal system with sensor means for detecting ambient parameters, and operating parameters and conditions of the thermal system of the electric vehicle, wherein the thermal system includes cooling and / or heating components with and without electric driven components and electric driven auxiliary components, providing a reinforcement learning based control means configured to generate an output for operating the thermal management system based on a control means input including at least one user request and / or at least one parameter detected by the sensor means, providing the control means input which is selected from a group of parameters defining target air conditions in the cabin, ambient conditions, thermal system conditions and vehicle states, providing the control means so that it is able to capture the relation between the inputs, in the form of a state, and outputs, in the form of at least one action, of the control means as input to the thermal system, which comprise operating parameters of the thermal systems and / or temperature conditions, for implementing real-time operation by computing an optimization cost function for the control means and to receive outputs of the thermal system as input for the control means as reward, and providing the outputs of thermal system so that they comprise system performance parameters of the thermal system when exploring the thermal system for finding improved rewards, and wherein the thermal management method comprises the step of determining by the control means control settings for real-time operation as output for operating the thermal system with improved rewards in view of the at least one user request while avoiding a risky operational area.
[0026] According to a fourteenth aspect of the present invention, in the thermal management method according to the thirteenth aspect, the system performance parameters include at least a coefficient of performance of the thermal system and a safety term for avoiding the risky operational area of the thermal system and optionally a driving comfort term, for instance to guarantee the cooling requested.
[0027] According to a fifteenth aspect of the present invention, in the thermal management method according to the thirteenth or fourteenth aspect, in the safety term for avoiding a risky operational area of the outputs of thermal system, control barrier function, preferably a reciprocal control barrier function, which approaches an infinite value close to a boundary of the predefined safe set of states, is reflected and causes the states of the thermal system to lie within a predefined safe set of states.
[0028] According to a sixteenth aspect of the present invention, the thermal management method according to thirteenth, fourteenth or fifteenth aspect additionally comprises providing a digital twin of the thermal system in simulation for initially calibrating the control means prior to the step of determining by the control means control settings for real-time operation.
[0029] According to a seventeenth aspect of the present invention, in the thermal management method according to sixteenth aspect, the at least one action defines a safe policy to collect data for initial calibration while a target policy is updated towards an optimal policy considering the safety term.
[0030] According to an eighteenth aspect of the present invention, in the thermal management method according to one of the thirteenth through sixteenth aspects, the control means is provided so that is able to use a model-free or a model based learning approach.
[0031] According to a nineteenth aspect of the present invention, the thermal management method according to the eighteenth aspect, the control means uses tabular Q-learning or Deep-Q-learning, wherein the Q-value is an estimation of the quality of the interrelationship between the state as input and the action as output.
[0032] For the thermal management method, it is possible to use further restrictions mentioned in relation to the thermal management system for obtaining the corresponding benefits.
[0033] With reference to the drawings and the corresponding detailed description, the following object of the present invention is described more in detail together with its other objects, features and advantages.
[0034] Fig. 1 illustrates a heat pump system for heating and cooling the cabin air as part of the thermal system to be managed and controlled by the thermal management system according to the present invention;Fig. 2 illustrates an electric power train / battery coolant system for keeping electric powertrain and battery in desired temperature ranges;Fig. 3 illustrates control means 300 controlling the thermal system 100, 200 as shown in Fig. 1 and 2 according to a first embodiment of the present invention;Fig. 4 illustrates an example for the internal structure of control means 300 and heat pump system 100 according to the first embodiment of the present invention;Fig. 5 shows an aspect of the control means 300 controlling the thermal system 100, 200 according to the first embodiment of the present invention in which digital twins 100’, 200’ are used;Fig. 6 shows the result of the control with control means 300 according to the first embodiment of the present invention from an initial state to an equilibrium state;Fig. 7 shows details of operation of a further aspect of the control means 300 according to the first embodiment of the present invention;Fig. 8 shows details of the use of a reciprocal Control Barrier Function with the control means 300 according to the first embodiment of the present invention;Fig. 9A shows an example for tabular Q-Learning with the control means 300 according to the first embodiment of the present invention in form of a chart in tabular form;Fig. 9B shows an example for tabular Q-Learning with the control means 300 according to the first embodiment of the present invention in form of a chart in tabular form;Fig. 10 shows an example for Deep Q-Network with the control means 300 according to the first embodiment of the present invention;Fig. 11 shows a second embodiment of the present invention; andFig. 12 shows a third embodiment of the present invention.
[0035] With the automotive thermal management system and method according to the present invention, an optimal efficiency of, for instance, a refrigerant system is to be ensured automatically under all conditions. Therefore, it is preferable, to determine on-line optimal control signals for varying operating conditions and for a slowly changing system due to degradation while at the same time the safety of the systems is to be guaranteed. As a result, the range of the vehicle can be maximized and lifetime can be increased.
[0036] Furthermore, with the present invention, the calibration efforts can significantly be reduced, as it allows automatic training of AI models online under real time operation of the vehicle, e.g., by maps or neural networks. Since with the present invention, it is not necessary to have complete knowledge of the system dynamics, the present invention significantly reduces modelling efforts.
[0037] More specifically, with the present invention, it is preferable to introduce a control-oriented model that captures the relationship between all relevant inputs, i.e., states, and outputs, i.e., actions, of the actual thermal system that can realize optimal real time operation on existing hardware, which is reflected in a reward.
[0038] Moreover, with the present invention, it is preferable to introduce an AI-based adaptive controller that can safely determine optimal control settings for optimal real-time operation under varying operating conditions and under system degradation on existing hardware.
[0039] For implementing the above control-oriented model and the above an AI-based adaptive controller, it is preferable: -to introduce a reinforcement learning based control method, wherein preferable a model-free reinforcement learning based control method is introduced, -to define states, e.g., ambient temperature, vehicle speed, route information, -to define a reward function, like one of the performance indicators: the heat pump’s efficiency COP, -to select actions, like the control of the expansion valve position, and -to calibrate hyperparameters as specified below for convergence.
[0040] According to the present invention, it is possible to merge an AI-method with an engineering application, more specifically a heat pump system. Furthermore, it is possible to implement a suitable RL (reinforced learning)-controller developed specifically for a real, state-of-the-art heat pump system. With the present invention, a reward function R(x, a), states and actions are defined as the formula in Math. 1.
[0041]
[0042] The thermal system to which the present invention is applicable can be, as an example, a heat pump system 100 and / or an electric powertrain / battery coolant system 200. Fig. 1 shows a heat pump or heat pump system 100 for heating and cooling the cabin air as an example, Fig. 2 shows an electric powertrain / battery coolant system 200 for keeping electric powertrain and battery in desired temperature ranges as an example and Fig. 3 shows control means 300 for controlling the respective thermal management system 100, 200.
[0043] The heat pump system 100 comprises in this example a chiller 2, an inner condenser 102, an evaporator 104, a compressor 106, and an outer heat exchanger 108 with a fan 110 interconnected via a refrigerant loop 112 and an air blower 114 for air as cooling fluid to evaporator 104 and inner condenser 102. Furthermore, an accumulator 116 is provided for storing liquid refrigerant and for separating liquid and gaseous refrigerant. The outlet of compressor 106 is connected to the inlet of inner condenser 102. The outlet of inner condenser 102 is connected to an outer heat exchanger expansion valve 118, which in turn is connected to the inlet of outer heat exchanger 108. The outer heat exchanger expansion valve 118 is also referenced as EXVHEAT in the following. The outlet of outer heat exchanger 108 is connected via a first check valve 120 to a point between an evaporator expansion valve 122 and a chiller expansion valve 124. The evaporator expansion valve 122 is also referenced as EXVEVA. The chiller expansion valve 124 in turn is connected to the refrigerant inlet of chiller 2. The evaporator expansion valve 122 in turn is connected to the inlet of evaporator 104. The outlet of evaporator 104 is connected to the inlet of accumulator 116 via a pressure regulator valve 126. The refrigerant outlet of chiller 2 is also connected to the inlet of accumulator 116. The outlet of outer heat exchanger 108 is likewise connected via a dehumidification control valve 130 and a second check valve 132 to the inlet of accumulator 116. A point between the outlet of inner condenser 102 and the outer heat exchanger expansion valve 118 is connected via a heating control valve 134 to a point which is in connection with the first check valve 120, the chiller expansion valve 124 and the evaporator expansion valve 122. The inner condenser 102 and evaporator 104 are arranged in a heating-cooling-air-conditioning or HVAC channel 136 entering into the vehicle cabin.
[0044] The electric powertrain / battery coolant system 200 shown in the example of Fig. 2 comprises the coolant side of chiller 2, a battery 202, a battery heater 203, an electric powertrain 204 and a radiator 206 interconnected via a powertrain / battery coolant loop 208. The powertrain / battery coolant loop 208 connects the coolant outlet of chiller 2 with the coolant inlet of battery heater 203. The coolant outlet of the battery heater 203 is connected to the coolant inlet of battery 202. The coolant outlet of battery 202 is connected to the coolant inlet of chiller 2 via a two-way-valve 210 and a first pump 212. The coolant outlet of chiller 2 is likewise connected to the coolant inlet of electric power train 204 via a second pump 216. The coolant outlet of electric powertrain 204 is connected via a 3-way-valve 218 to the inlet of the first pump 212 and the coolant inlet of radiator 206. As shown in Fig. 1 and 2, the outer heat exchanger 108 and the radiator 206 having an air fan 110 and an active grill shutter 4 are arranged as a stack.
[0045] A heat pump system contains many actuators to make sure that the required performance is obtained. To simplify the controls of all actuators, many operation-modes might be introduced wherein these modes pre-define the range of usage per actuator per condition. For each mode, the activated actuators need to be calibrated to ensure that an acceptable efficiency is reached. The efficiency of a thermal system, more specifically a refrigerant system, can be described by the Coefficient Of Performance COP as already specified above. The COP is the ratio between the useful heating or cooling provided (QH) and work / energy put into the system (WIN) and is computed according to the formula in Math. 2.
[0046] In the cooling mode of the heat pump system 100, the COP can be expressed in the formula in Math. 3. In the formula in Math 3, Pevais the required evaporator power [W], Pcmpis the required compressor power [W], Pfanis required fan power [W].
[0047] The higher the COP, the better the efficiency. Heating or cooling can be maximized by optimizing the phase change between gas and liquid during condensation and evaporation, i.e., during enthalpy change. By controlling the actuators, the desired phase change for performance and efficiency is implemented.
[0048] A first embodiment of the present invention is described in the following with reference to Fig. 3.
[0049] Control means 300 inputs actions 304 into the thermal system 100, 200. The thermal system 100, 200 provide states 302 of the thermal system 100, 200 and rewards 306 obtained with the thermal system 100, 200 as input to control means 300.
[0050] Inputs 302 to the control means might contain user requests which includes at least one of - the requested cooling power of the evaporator 104 or the temperature at its outlet, -the requested heating power of the inner condenser 102 or the temperature at its outlet, - the requested cooling power of the chiller 2 or the temperature at its outlet, - the requested heating power of the condenser 102 or the temperature at its outlet, - the recycling ratio of thermal system 100, 200, - the HVAC blower duty of the thermal system 100,200.
[0051] The parameter detected by sensor means of the control means input 302 might include at least one of - the ambient temperature, - the vehicle speed, - the ambient humidity, - the cabin temperature.
[0052] The output 304 of the control means 300 for operating the thermal system 100, 200 includes with a higher priority at least one of - the compressor speed, - a control signal for a first expansion valve EXV1 for controlling the two phase flow for the cabin cooling by the evaporator, wherein optionally calculating reference setpoints are considered.
[0053] The output 304 of the control means 300 for operating the thermal system 100, 200 might include with a lower priority at least one of - a control signal for a second expansion valve EXV2 for controlling the two phase flow for heating, - a control signal for an expansion valve EXV3 for controlling the two phase flow for battery cooling by the chiller, - a control signal for a first solenoid valve SOV1 for realizing parallel dehumidification, - a control signal for a first solenoid valve SOV2 for disabling cooling, - a control signal for the fan speed, - a control signal for the air shutter.
[0054] Furthermore, the driver can also provide one of the following inputs 302 to the control means 300: - a control signal for the speed of the heating-cooling-air-conditioning (HVAC) blower, - a control signal for the HVAC recycling amount, - a control signal for the HVAC mixing flap.
[0055] In a preferred aspect of the first embodiment, the user request of the control means input 302 includes the requested cooling power of the evaporator or the temperature at its outlet, the parameter detected by the sensor means of the control means input 302 includes at least one of the ambient temperature and the vehicle speed, and the output 304 of the control means 300 for operating the thermal system 100, 200 includes at least one of the compressor speed and a control signal for a first expansion valve EXV for controlling the two phase flow for the cabin cooling by the evaporator.
[0056] Fig. 6 shows the result of the control with control means 300 of the present invention for state 1 and state 2 from an initial state to an equilibrium state. In order to be able to implement safe reinforcement learning via control means 300 according to the present invention, the following considerations are helpful: -safety considerations in relation to the control system are made in view of on-line calibration, -the safety considerations are helpful not only during the learning phase but also in the implementation phase, -with off-policy Reinforcement Learning with Control Barrier Functions, it is possible to avoid unsafe system states during learning, -with an off-policy algorithm, a safe policy is used to collect data for learning while a target policy is updated towards an optimal policy, -an optimization cost function (preferably by maximizing reward) is modified by including a reciprocal barrier function which approaches an infinite value close to the boundary of the safe set, i.e., the circle in Fig. 6, -the use of barrier functions causes that the system states lie within predefined safe set of states during the operation of the thermal system 100, 200.
[0057]
[0058] One example for implementing reinforced learning in the first embodiment of the present invention with thermal system 100 with heat pump system 170 and control means 300 is shown in Fig. 4. Control means 300 has here a replay memory 180, the learning algorithm and a policy, which is updated and can be stored in replay memory 180. The thermal system 100 has here an external controller 160. In control means 300 according to the present invention, the map-based controller according to the prior art, is replaced by a reinforced learning controller with the aim to improve system performance and reduce calibration efforts. In order to optimize control means 300, the following implementation steps can be applied: -analyzing the control problem, -designing the reward function, -defining the space of states and actions, -selecting an algorithm, -training the learning algorithm, -evaluating the results obtained by optimized control means 300.
[0059] In Fig. 4, external inputs w, actuator inputs u, disturbances d and outputs y are designated while a previous reward function Rtis updated by the next reward function Rt+1. The external input w and the output y of heat pump system 170 input to replay memory 180 in form of Siare constraints with respect to reinforced learning.
[0060] According to the first embodiment, it is preferable that, in the safety term for avoiding a risky operational area of the outputs 306 of thermal system 100, 200, a reciprocal barrier function, which approaches an infinite value close to a boundary of the predefined safe set of states, is reflected and causes the states of the thermal system 100, 200 to lie within a predefined safe set of states.
[0061]
[0062]
[0063] The reciprocal CBF term can indicate that the control means 300 is approaching the boundary of the safe set, as shown for instance in Fig. 6. This information can be used to avoid actions that lead to an exponential increase in reward.
[0064]
[0065]
[0066]
[0067] Control system 300 according to the first embodiment is a self-learning control system. Preferably, the control system 300 is configured to use a model-free learning approach. For this model-free learning approach, for instance, tabular Q-Learning or Deep-Q-learning can be used, wherein the Q-value is an estimation of the quality of the interrelationship between the state as input 302 and the action as output 304. Fig. 9A and 9B shows an example for tabular Q-Learning while, in Fig. 10, a Deep Q-Network is shown.
[0068] In the tabular Q-Learning of Fig. 9A and 9B, the following approach is applied: -the Q-table of Fig. 9B is initialized, -an action is chosen, -the action is performed, -the reward is measured, -the Q-table is updated and then it is jumped back to the step of choosing an action.
[0069] At the start, a random action can be selected and the action with the maximum Q is finally used. In the example of Fig. 9A and 9B, it is started with state (2,1) with Q=0. Then (1,1) is chosen with Q=0. Subsequently, (1,2) results in Q=-10, then with the state (2,2) Q=-10 is obtained. In the next self-learning, after the state (2,2) with Q=0, the state (2,3) results in Q=+5. A new or unexplored system usually starts with more exploration, wherein a random action selection takes place. In an already known system, it is tried to exploit the result with the highest reward, which then results in an optimal control.
[0070] In the case of Deep Q-Network of Fig. 10, value-based algorithms are selected for ease of implementation and of possible analysis. In this case, the Q-value directly shows the expected COP. The Deep Q-Network (DQN) is selected according to continuous state input, the generalizability over state-action combinations and the scalability for larger states.
[0071] In the following, the results, which can be obtained by the first embodiment of the present invention in comparison to a benchmark, in which a map-based controller is used, are illustrated. Here the output of control means 300 to optimal control setpoints are analyzed for showing the effect of discretization and the training success of control means 300. In this test, an evaluation of the control means 300 for a WLTC (Worldwide Harmonized Light Vehicles Test Cycle) driving cycle at 28 °C results in a 10.6 % lower energy consumption while at the same time increasing the COP.
[0072] From this test, it can be concluded that by using the RL-based controller in control system 300, the system efficiency is increased, and the calibration effort is decreased.
[0073] While Fig. 3 shows a first embodiment of the present invention, in which a single action 304 is provided from control means 300 to thermal system 100, 200, the present invention is not limited thereto.
[0074] Fig. 11 shows a second embodiment of the present invention in which a plurality of actions, for instance 304a-c, are output from control means 300 and are input into the thermal system 100, 200. In this way, one control agent is able to provide multiple outputs so that with one control means 300 the whole range of relationship between ambient temperature and target temperature in the case of using thermal system 100 in the cooling mode can be covered.
[0075] Fig. 12 shows a third embodiment of the present invention in which a plurality of actions, for instance 304a-c, are output from respective control means 300a-c and are respectively input into the thermal system 100, 200. In this way, individual control means 300a-c are able to provide respective outputs so that each of the control means 300a-c has specific ranges in the relationship between ambient temperature and target temperature in the case of using thermal system 100 in the cooling mode.
[0076] The other aspects of the first embodiment are equally applicable to the second and third embodiments, too.
Claims
1. A thermal management system for an electric vehicle with a cabin, comprising a thermal system (100, 200) with a sensor means for detecting ambient parameters, and operating parameters and conditions of the thermal system of the electric vehicle, and a reinforcement learning based control means (300) configured to generate an output (304) for operating the thermal system (100, 200) based on a control means input (302) including at least one user request and / or at least one parameter detected by the sensor means, wherein the thermal system (100, 200) includes cooling and / or heating components with and without electric driven components and electric driven auxiliary components, wherein the control means input (302) is selected from a group of parameters defining target air conditions in the cabin, ambient conditions, thermal system conditions, and vehicle states, wherein the reinforcement learning based control means (300) is configured to capture a relation between the control means input (302), in a form of a state, and the output (304), in a form of at least one action, of the reinforcement learning based control means (300) as input to the thermal system (100, 200), which comprise operating parameters of the thermal systems and / or temperature conditions, for implementing real-time operation by computing an optimization cost function for the reinforcement learning based control means (300) and to receive outputs (306) of the thermal system (100, 200) as the control means input for the reinforcement learning based control means (300) as reward, and wherein the outputs (306) of the thermal system (100, 200) comprise system performance parameters of the thermal system (100, 200) when exploring the thermal system (100, 200) for finding improved rewards, and wherein the reinforcement learning based control means (300) is configured to determine control settings for real-time operation as the output (304) for operating the thermal system (100, 200) with the improved rewards in view of the at least one user request while avoiding a risky operational area.
2. The thermal management system according to claim 1, wherein the system performance parameters include at least an indicator for energy consumption or energy efficiency of the thermal system (100, 200) and a safety term for avoiding the risky operational area of the thermal system (100, 200).
3. The thermal management system according to claim 1 or 2, wherein the system performance parameters include a driving comfort term.
4. The thermal management system according to one of the preceding claims, wherein, in a safety term for avoiding the risky operational area of the outputs (306) of the thermal system (100, 200), a control barrier function, which approaches an infinite value close to a boundary of a predefined safe set of states, is reflected and causes states of the thermal system (100, 200) to lie within the predefined safe set of states.
5. The thermal management system according to one of the preceding claims, wherein the reinforcement learning based control means (300) is configured to use a model-free learning approach.
6. The thermal management system according to claim 5, wherein the reinforcement learning based control means (300) uses tabular Q-learning or Deep-Q-learning, wherein a Q-value is an estimation of a quality of an interrelationship between the state as the input (302) and the action as the output (304).
7. The thermal management system according to one of the preceding claims, wherein the at least one user request of the control means input (302) includes at least one of: a requested cooling power of an evaporator or a temperature at an outlet of the evaporator; a requested heating power of an inner condenser or a temperature at an outlet of the inner condenser; a requested cooling power of a chiller or a temperature at an outlet of the chiller; a requested heating power of a condenser or a temperature at an outlet of the condenser; a recycling ratio of the thermal system (100, 200); and a HVAC blower duty of the thermal system (100, 200).
8. The thermal management system according to one of the preceding claims, wherein the at least one parameter detected by the sensor means of the control means input (302) includes at least one of: an ambient temperature; a vehicle speed; an ambient humidity; a cabin temperature; and a cabin humidity.
9. The thermal management system according to one of the preceding claims, wherein the output (304) of the reinforcement learning based control means (300) for operating the thermal system (100, 200) includes at least one of: a compressor speed; a control signal for a first expansion valve (EXV1) for controlling a two phase flow for cabin cooling by an evaporator; a control signal for a second expansion valve (EXV2) for controlling a two phase flow for heating; a control signal for a third expansion valve (EXV3) for controlling a two phase flow for battery cooling by a chiller; a control signal for a first solenoid valve (SOV1) for realizing parallel dehumidification; a control signal for a second solenoid valve (SOV2) for disabling cooling; a control signal for a fan speed; and a control signal for an air shutter, wherein optionally calculating reference setpoints are considered.
10. The thermal management system according to one of the claims 1 to 6, wherein the at least one user request of the control means input (302) includes a requested cooling power of an evaporator or a temperature at an outlet of the evaporator, wherein the at least one parameter detected by the sensor means of the control means input (302) includes at least one of an ambient temperature and a vehicle speed, and wherein the output (304) of the reinforcement learning based control means (300) for operating the thermal system (100, 200) includes at least one of a compressor speed and a control signal for a first expansion valve (EXV1) for controlling a two phase flow for cabin cooling by the evaporator.
11. The thermal management system according to one of the preceding claims, wherein the thermal system (100, 200) comprises a refrigerant system of a heat-pump system (100).
12. The thermal management system according to one of the preceding claims, wherein the thermal system (100, 200) comprises an electric power train / battery coolant system (200).
13. A thermal management method for an electric vehicle with a cabin, comprising: providing a thermal system (100, 200) with a sensor means for detecting ambient parameters, and operating parameters and conditions of the thermal system of the electric vehicle, wherein the thermal system (100, 200) includes cooling and / or heating components with and without electric driven components and electric driven auxiliary components; providing a reinforcement learning based control means (300) configured to generate an output (304) for operating the thermal system (100, 200) based on a control means input (302) including at least one user request and / or at least one parameter detected by the sensor means; providing the control means input (302) which is selected from a group of parameters defining target air conditions in the cabin, ambient conditions, thermal system conditions and vehicle states, providing the reinforcement learning based control means (300) so that it is able to capture a relation between the control means input (302), in a form of a state, and the output (304), in a form of at least one action, of the reinforcement learning based control means (300) as input to the thermal system (100, 200), which comprise operating parameters of the thermal system and / or temperature conditions, for implementing real-time operation by computing an optimization cost function for the reinforcement learning based control means (300) and to receive outputs (306) of the thermal system (100, 200) as the control means input for the reinforcement learning based control means (300) as reward; and providing the outputs (306) of the thermal system (100, 200) so that they comprise system performance parameters of the thermal system (100, 200) when exploring the thermal system (100, 200) for finding improved rewards, and wherein the thermal management method comprises a step of determining by the reinforcement learning based control means (300) control settings for real-time operation as the output (304) for operating the thermal system (100, 200) with improved rewards in view of the at least one user request while avoiding a risky operational area.
14. The thermal management method according to claim 13, wherein the system performance parameters include at least a coefficient of performance (COP) of the thermal system (100, 200) and a safety term for avoiding the risky operational area of the thermal system (100, 200), and optionally a driving comfort term, for instance to guarantee cooling requested.
15. The thermal management method according to claim 13 or 14, wherein, in a safety term for avoiding the risky operational area of the outputs (306) of the thermal system (100, 200), a control barrier function, which approaches an infinite value close to a boundary of a predefined safe set of states, is reflected and causes states of the thermal system (100, 200) to lie within the predefined safe set of states.
16. The thermal management method according to claim 13, 14 or 15, additionally comprising providing a digital twin (100’, 200’) of the thermal system (100, 200) in simulation for initially calibrating the reinforcement learning based control means (300) prior to the step of determining by the reinforcement learning based control means (300) control settings for the real-time operation.
17. The thermal management method according to claim 16, wherein the at least one action defines a safe policy to collect data for initial calibration while a target policy is updated towards an optimal policy considering a safety term.
18. The thermal management method according to one of claims 13 through 16, wherein the reinforcement learning based control means (300) is provided so that the reinforcement learning based control means is able to use a model-free learning approach.
19. The thermal management method according to claim 18, wherein the reinforcement learning based control means (300) uses tabular Q-learning or Deep-Q-learning, wherein a Q-value is an estimation of a quality of an interrelationship between the state as the control means input (302) and the action as the output (304).
Citation Information
Patent Citations
Electric vehicle energy management and distribution method
CN110962684A
Automated climate control system
US20180134118A1
Artificial intelligence in conditioning or thermal management of electrified powertrain
US20200376927A1
Heat management system for an electrified motor vehicle
WO2022233524A1