Managing a thermal environment in a vehicle
The method uses reinforcement learning to train an AI model for vehicle thermal management systems, optimizing user comfort and energy efficiency by learning optimal control strategies for HVAC, heated seats, and radiant panels.
Patent Information
- Application Number
- FR2024007410
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2026-01-09
AI Technical Summary
Existing thermal management systems in vehicles lack efficient automation for optimizing user thermal comfort and energy consumption, often leading to increased energy usage when manually adjusted by users.
A computer-implemented method using reinforcement learning to train an artificial intelligence model for controlling thermal management components, such as HVAC systems, heated seats, and radiant panels, to achieve a given thermal comfort level while minimizing energy consumption.
The method enables flexible compromise between user comfort and energy efficiency by learning optimal control strategies through trial and error, achieving rapid convergence to desired thermal comfort with reduced energy consumption.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Thermal environment management in a vehicle technical field
[0001] The invention relates to the management of a thermal environment in a vehicle. In particular, the invention relates to a method for training an artificial intelligence model for use in a vehicle to control at least one thermal management component of a vehicle in order to optimize the thermal comfort of a vehicle user, a method for controlling at least one thermal management component of a vehicle and the associated system. Previous art
[0002] A user's thermal comfort can be defined as the condition that reflects a user's satisfaction with their thermal environment and is subjectively assessed. Thermal comfort is one of the factors taken into account when designing the thermal environments of vehicles. Indeed, users spend a significant amount of time in their vehicles, and the perceived thermal comfort contributes to their well-being in the vehicle.
[0003] To achieve thermal comfort, the user can manually activate one or more thermal components via a human-machine interface. Among the thermal components that may be present in a vehicle are air conditioning systems, for example, a heating, ventilation, and air conditioning (HVAC) system. In such an air conditioning system, a heat transfer fluid passes through a fluid circuit comprising a compressor, a heat exchanger called a condenser, an expansion device, and a second heat exchanger called an evaporator.
[0004] There are also reversible air conditioning systems; these systems can absorb heat energy from the outside air at a heat exchanger, called an evaporator-condenser, and release it into the passenger compartment, notably by means of a dedicated heat exchanger. In this case, the air conditioning system functions as a heat pump.
[0005] Other thermal regulation devices may also be present in the vehicle, such as radiant panels, heated seats, a heated steering wheel and / or contact heating surfaces located for example on an armrest.
[0006] The vehicle's thermal system which manages the thermal regulation device(s) is designed to achieve various objectives such as thermal comfort management, energy consumption reduction and air quality.
[0007] Various solutions have been proposed for managing a thermal regulation device to achieve user thermal control. However, these solutions do not allow for the efficient automation of thermal regulation device management, particularly when taking into account the external and / or internal thermal environment of the vehicle, such as the vehicle's external temperature, the level of sunlight, the vehicle's speed, the humidity level, the wind intensity, the vehicle's internal temperature, the humidity level inside the vehicle, etc.
[0008] Furthermore, all these thermal regulation devices are very energy-intensive and significantly increase the vehicle's energy consumption, particularly in cold outside temperatures. This is especially true when the thermal regulation devices are activated and adjusted manually by the vehicle user. Indeed, very often, in an attempt to achieve thermal comfort as quickly as possible, the user activates all the devices simultaneously and sets them to a high power level. Summary of the invention
[0009] The invention aims in particular to offer an improvement in the control of one or more known thermal management components in order to optimize the thermal comfort of a vehicle user.
[0010] To this end, the invention relates to a computer-implemented method for training an artificial intelligence model for use in a vehicle to control at least one thermal management component of a vehicle in order to optimize the thermal comfort of a vehicle user, the method comprising a step of introducing training data into an artificial intelligence model to train said artificial intelligence model by means of a reinforcement learning agent, in order to control said at least one thermal management component to achieve a parameter relating to a given thermal comfort, the training data comprising training data sets, each training data set comprising at least one parameter related to the thermal environment of the user and / or the vehicle.
[0011] This method consists of training an artificial intelligence model using a reinforcement learning agent to control at least one thermal management component of a vehicle to achieve a parameter related to a given level of thermal comfort. Reinforcement learning relies on trial and error strategies to arrive at a flexible compromise that can be configured for user comfort.
[0012] According to a particular embodiment, the method is implemented to control at least two thermal management components of the vehicle and to further optimize the energy consumed by the thermal management components.
[0013] According to these characteristics, reinforcement learning allows for a configurable compromise between user comfort and reduced energy consumption. Furthermore, the training is designed to control several thermal management components, including simultaneously.
[0014] The reinforcement learning agent can be coupled to an environment module capable of capturing the state of the environment, the training consisting of training the artificial intelligence model by means of the reinforcement learning agent on the basis of a state of the environment observed in the environment module.
[0015] The training may include the following steps:
[0016] - a prediction step by the learning agent through reinforcement of an action control of the thermal management component(s), depending on the observed environmental conditions and the parameter relating to the given thermal comfort to be achieved, and
[0017] - an execution step of the control action controlling the component(s) of thermal management.
[0018] The training may further consist of training the artificial intelligence model using the reinforcement learning agent on the basis of a reward provided by the environment module, the reward being determined according to the state of the observed environment and the parameter relating to the given thermal comfort.
[0019] Reward-linked learning ensures accelerated convergence towards the given thermal comfort when the thermal management components are activated and can also lead to substantial energy savings during the operation of these components.
[0020] The prediction step may also be a function of the reward provided by the environment module.
[0021] The training may further include a reception stage by the learning agent through reinforcement of the observed environment state and the reward following the execution of the predicted control action.
[0022] The reward provided by the environment module may consist of:
[0023] (i) reward the learning agent by reinforcement when the state of the environment observed in the environment module approaches the parameter relating to the given thermal comfort to be achieved;
[0024] (ii) penalize the learning agent by reinforcement when the state of the environment observed in the environment module deviates from the parameter relating to the given thermal comfort to be achieved.
[0025] The reward provided by the environment module may further consist of:
[0026] (i) rewarding the learning agent by reinforcement further when the state of the environment observed in the environment module shows a decrease in energy consumption;
[0027] (ii) penalize the reinforcement learning agent when the state of the environment observed in the environment module shows an increase in energy consumption.
[0028] According to a particular embodiment, the observed state of the environment has reached the given thermal comfort parameter when the reward converges to a maximum value.
[0029] The training of the artificial intelligence model using the reinforcement learning agent may include one or more training episodes to learn a strategy that enables the achievement of the given thermal comfort parameter, the strategy comprising a set of control actions, an episode being defined by a set of training data.
[0030] The training steps can be repeated until the observed state of the environment has reached the given parameter relating to thermal comfort.
[0031] The given thermal comfort parameter can be defined by a range of values.
[0032] The observed state of the environment may include a state of thermal comfort for the user. The user's state of thermal comfort may be within a range of values, the central value of the range corresponding to the value of the given thermal comfort parameter.
[0033] The observed state of the environment may further include a state of energy consumption.
[0034] According to one embodiment, the training steps are further repeated until the energy consumption state of the observed environment has reached an optimal energy consumption state.
[0035] According to one embodiment, the observed state of the environment has reached a state of optimal energy consumption when the reward converges to a maximum value.
[0036] The training steps can be repeated until the user's thermal comfort state falls within a range of values for the given thermal comfort parameter, and then the training steps can be repeated. until the energy consumption state of the environment has reached the optimal energy consumption state.
[0037] Said at least one parameter related to the user's thermal environment may include at least one parameter related to a state of thermal comfort of the user.
[0038] Said at least one parameter related to the thermal environment of the vehicle may include at least one parameter related to the internal thermal environment of the vehicle and / or at least one parameter related to the external thermal environment of the vehicle.
[0039] Said at least one thermal management component may include at least one heating, ventilation and air conditioning device.
[0040] Said at least one thermal management component may further include at least one radiant panel and / or at least one heated seat.
[0041] According to one embodiment, at least one thermal management component is a simulated component.
[0042] The training of the artificial intelligence model using the reinforcement learning agent can be carried out using a soft actor-critical algorithm, a DQN (Deep Q-Network) algorithm, or a proximal policy optimization (PPO) algorithm.
[0043] The invention also relates to a computer-implemented method for controlling at least one thermal management component of a vehicle in order to optimize the thermal comfort of a vehicle user. The method comprises:
[0044] - a step of obtaining at least one piece of data related to the thermal environment of the user and / or the vehicle,
[0045] - an input step of said at least one data point related to the thermal environment of the user and / or the vehicle in an artificial intelligence model, the artificial intelligence model having been trained to control said at least one thermal management component to achieve a parameter relating to a given thermal comfort,
[0046] - to bring the trained artificial intelligence model to control said at least a thermal management component to achieve the given thermal comfort parameter.
[0047] Thus, the method for controlling at least one thermal management component of a vehicle makes it possible to achieve the given thermal comfort parameter in order to optimize the thermal comfort of a vehicle user, in particular in an optimal time.
[0048] The method can also be implemented to control at least two thermal management components of the vehicle and to optimize the energy consumed by the thermal management components.
[0049] The method may further include a step of bringing the trained artificial intelligence model to control said thermal management components to achieve a state of optimal energy consumption.
[0050] Said at least one data related to the thermal environment of the user and / or the vehicle may include at least one data related to a state of thermal comfort of the user and / or at least one data related to the environment of the vehicle.
[0051] Said at least one data related to the thermal environment of the vehicle may include data related to the internal thermal environment of the vehicle and / or data related to the external thermal environment of the vehicle.
[0052] Said at least one thermal management component may include at least one heating, ventilation and air conditioning device.
[0053] Said at least one thermal management component may further comprise at least one radiant panel and / or at least one heated seat.
[0054] The invention also relates to a computer-implemented method according to the method for controlling at least one thermal management component of a vehicle previously described, characterized in that the artificial intelligence model was trained according to the method to train an artificial intelligence model for use in a vehicle as previously described.
[0055] The invention also relates to a computer system comprising means for implementing the method for training an artificial intelligence model for use in a previously described vehicle or a method for controlling at least one thermal management component of a previously described vehicle.
[0056] The invention also relates to a computer program comprising instructions which, when executed, cause a device or computer system to execute a method for training an artificial intelligence model for use in a previously described vehicle or the method for controlling at least one thermal management component of a previously described vehicle.
[0057] The invention also relates to a computer-readable medium on which the previously described computer program is stored.
[0058] The invention also relates to an artificial intelligence model for controlling at least one thermal management component of a vehicle in order to optimize the thermal comfort of a vehicle user according to the method for training an artificial intelligence model for use in a vehicle previously described or the method for controlling at least one thermal management component of a vehicle previously described.
[0059] The invention also relates to a system for a vehicle comprising at least one thermal management component and a control device, the device of control implementing the method to control at least one thermal management component of a vehicle previously described.
[0060] Said at least one thermal management component may include at least one heating, ventilation and air conditioning device.
[0061] Said at least one thermal management component may further comprise at least one radiant panel and / or at least one heated seat. List of Figures
[0062] Other features and advantages of the invention will become more apparent upon reading the following description, given by way of illustrative and non-limiting example, and the accompanying figures, among which:
[0063] Fig. 1 is a schematic representation of a part of a vehicle, including the passenger compartment, comprising several thermal management components.
[0064] The [Fig.2] is a schematic diagram illustrating a Markov decision process.
[0065] Fig. 3 is a schematic diagram illustrating an example of a deep reinforcement learning process.
[0066] Fig. 4 is a flowchart showing an example of a method for training an artificial intelligence model according to certain embodiments.
[0067] Fig. 5 is a schematic diagram showing an example of a human body divided into sections having respective thermal comfort indices.
[0068] Fig. 6 is a schematic diagram illustrating an example of a reinforcement learning process according to the present invention.
[0069] Figures 7A, 7B, 7C and 7D show graphs illustrating a reward episode, the evolution of the thermal comfort index, the duration of discomfort and energy consumption as a function of time, respectively.
[0070] Figures 8A and 8B show the curves of the thermal comfort index and energy consumption for a range of outside temperatures.
[0071] Fig. 9 shows the curves of the time required to reach a target comfort level for different combinations of use of thermal management components.
[0072] Figure 10 shows the energy consumption curves for different combinations of use of thermal management components.
[0073] The [Fig. 11] is a flowchart illustrating an example of a method for controlling one or more thermal management components of a vehicle, according to certain embodiments. Detailed description of the invention
[0074] The following embodiments are examples. Although the description refers to one or more embodiments, this does not necessarily mean that each reference relates to the same embodiment, or that the features apply only to a single embodiment. Simple features from different embodiments can also be combined or interchanged to provide other embodiments.
[0075] Figure 1 represents a passenger compartment 10 of a motor vehicle. The vehicle may be an internal combustion engine vehicle, an electric vehicle, or a hybrid vehicle, that is, a vehicle that can be powered by an electric motor and an internal combustion engine. The passenger compartment 10 includes a driver's seat 15. It may include another front passenger seat and rear passenger seats. The vehicle includes a thermal management system 20. The thermal management system includes at least one thermal management component.
[0076] Such a thermal management component is, for example, a heating, ventilation, and air conditioning (HVAC) unit 25. Outlets 30 of heat-treated air from the HVAC unit, in particular a head air outlet 30a and a foot air outlet 30b, allow the heat-treated air from the HVAC unit 25 to be blown into the passenger compartment 10. The user can direct the flow of heat-treated air at the outlets 30 as needed. In a particular example of an HVAC unit 25, this unit may include an additional electric heating device, thus providing supplementary heating. This additional electric heating device includes, in particular, radiant elements.According to a particular embodiment, the additional electric heating device may further include resistive heating elements, for example, positive temperature coefficient resistors, enabling the conversion of an electric current into thermal energy.
[0077] A thermal management component may also be, for example, a heated seat device 30 installed in the driver's seat 15. Other heated seat devices may be installed, for example, in the front passenger seat and / or in the rear passenger seats.
[0078] A thermal management component may include, for example, at least one radiant panel 35. Radiant panels could be placed anywhere inside the passenger compartment 10. For example, radiant panels could be located above the occupant's windows 40 35a and / or under the occupant's feet 35b. Other radiant panels could be installed, for example, on the front passenger side and / or at the rear passenger level.
[0079] In order to optimize the thermal comfort of a vehicle user, the present invention relates to a method for training an artificial intelligence model for use in a vehicle to control at least one thermal management component 25, 30, 35 of a vehicle. The training consists of training an artificial intelligence model by means of a reinforcement learning agent to control said at least one thermal management component to achieve a parameter relating to a given thermal comfort level.
[0080] Reinforcement learning refers to a goal-oriented optimization process guided by an impact signal. Thus, learning is carried out efficiently using information collected in real time from interaction with the environment, namely, for example, the vehicle's interior 10. Reinforcement learning tasks can be formulated according to a Markov decision process.
[0081] A Markov decision process is shown in [Fig. 2]. Such a process involves an agent 210 (i.e., a reinforcement learning algorithm), in particular a rational agent, receiving observations on the state 230 of an environment 220 (which may be a real or simulated environment) and rewards 240, and performing actions 250 that will have an impact on the environment 220. The agent then seeks to learn a matching function between the state of the environment and an action that will maximize the expected reward. This matching function is called the strategy or policy of the reinforcement learning system.
[0082] According to the Markov decision process, the state st of the environment 220 at time t is provided to agent 210. Based on this state, the agent generates an action at from a set of available actions A to modify the state of the environment. The generated action at is then executed, and the environment 220 transitions to a new state st+i. Furthermore, a reward rt+i associated with the modification of the environment is determined. Agent 210 then receives the new state of the environment st+i and the reward rt+i associated with executing the action in the environment. The reward rt+i is then used by agent 210 to determine the next action. The goal of the reinforcement agent is to learn a policy that maximizes the expected reward. This process is repeated until convergence, that is, until the reward can no longer be improved.
[0083] Deep reinforcement learning is an approach to reinforcement learning that uses deep neural networks (DNNs) as an integral part of the agent. Figure 3 shows an example of a deep reinforcement learning system.
[0084] The present invention relates, in a first aspect, to training an artificial intelligence model for use in a vehicle to control at least one thermal management component. The training consists, in particular, of training an artificial intelligence model using a reinforcement learning agent to control at least one thermal management component of a vehicle, in order to achieve a parameter called a "parameter relating to a given thermal comfort level." The parameter relating to a given thermal comfort level may be a specific temperature in an area of a vehicle or a user's thermal comfort level within the vehicle. In one embodiment, the parameter relating to a given thermal comfort level is defined by a range of values. For example, it may include a temperature range or a range of values for the thermal comfort level.
[0085] The artificial intelligence model is trained using training data fed into the model. The training data comprises training datasets, and each training dataset includes at least one parameter related to the thermal environment of the user and / or the vehicle. In other words, each dataset represents a thermal comfort scenario from which the thermal management component(s) must be controlled to improve thermal comfort and achieve the parameter related to a given thermal comfort level.
[0086] Said at least one parameter related to the thermal environment of the user and / or the vehicle may include at least one parameter related to a state of thermal comfort of the user and / or at least one parameter related to the thermal environment of the vehicle.
[0087] A parameter related to a user's thermal comfort state may include at least one user comfort index, in particular a user physiological comfort index. The physiological comfort index is, for example, a data point representing the thermal comfort felt by the user. A user's comfort index is described in particular in application FR3129629A1.
[0088] A parameter related to the thermal environment of the vehicle may include at least one data point related to the internal thermal environment of the vehicle and / or at least one data point related to the external thermal environment of the vehicle.
[0089] The data related to the vehicle's external thermal environment may represent the vehicle's external temperature and / or the level of sunlight and / or the vehicle's speed and / or the humidity level and / or the wind intensity. In particular, the data related to the vehicle's external thermal environment is representative of the vehicle's external thermal environment and may represent at least one meteorological condition of the vehicle's environment.
[0090] The data related to the internal thermal environment of the vehicle may represent the internal temperature of the vehicle, namely of the passenger compartment and / or the humidity level present in the vehicle.
[0091] When the vehicle includes at least two thermal management components, as illustrated in [Fig.1], the method for training an artificial intelligence model can be implemented to control the plurality of thermal management components in order to optimize the thermal comfort of a vehicle user and the energy consumed by the thermal management components.
[0092] The artificial intelligence model is trained using a reinforcement learning agent that can be coupled to an environment module capable of capturing the state of the environment. An environment module 45 is illustrated in [Fig. 1]. It includes, for example, a device for determining the physiological comfort of the user 50 and / or a device for obtaining at least one piece of data related to the thermal environment of the vehicle 55. Thus, the environment module 45 determines at least one piece of data related to the thermal environment of the user and / or the vehicle.Said at least one data point relating to the thermal environment of the user and / or the vehicle includes at least one data point relating to a state of thermal comfort of the user determined for example by the device for determining a physiological comfort of the user 50 and / or at least one data point relating to the thermal environment of the vehicle determined for example by the device for obtaining at least one data point relating to the thermal environment of the vehicle 55. In particular, said at least one data point relating to the thermal environment of the vehicle includes at least one data point relating to the interior thermal environment of the vehicle and / or at least one data point relating to the exterior thermal environment of the vehicle.
[0093] The environment module can be linked to a real environment or a simulated environment, the environment being able to include the vehicle's passenger compartment or a part of the vehicle's passenger compartment.
[0094] The training then consists of training the artificial intelligence model using the reinforcement learning agent on the basis of a state of the environment observed in the environment module.
[0095] An example of a method for training an artificial intelligence model will now be described with reference to [Fig. 4]. This artificial intelligence model will be trained to control at least one thermal management component 25, 30, 35 of a vehicle as illustrated in [Fig. 1]. According to this method, the vehicle illustrated in [Fig. 1] is either a real vehicle or a simulated vehicle for training the artificial intelligence model. In other words, the thermal environment the vehicle will either be real and include physical thermal management components or simulated using simulated thermal management components.
[0096] The training of the artificial intelligence model can be carried out using a soft actor-critical algorithm, or a DQN-Deep Q-Network algorithm, or a proximal policy optimization (PPO) algorithm.
[0097] In particular, the soft actor-critic algorithm is a model-free, policy-free reinforcement learning algorithm that combines two main components: an actor and a critic. The actor is a neural network that produces a stochastic policy, that is, a probability distribution over the possible actions for each state. The critic is another neural network that estimates the value function, that is, the expected return for each state-action pair.
[0098] The DQN network algorithm is a type of reinforcement learning algorithm that uses a neural network to approximate the agent's value function. The value function estimates the expected return (cumulative reward) of acting in a given state. A DQN network algorithm learns to update the value function using a technique that compares the actual and expected rewards of each action and minimizes the difference.
[0099] The Proximal Policy Optimization (PPO) algorithm is a type of reinforcement learning algorithm that adopts a control function that prevents the policy update of an agent from being too large or too small for better learning stability.
[0100] In this example, the method is implemented by computer and is executed using a simulated thermal environment of a vehicle comprising one or more simulated thermal management components. The simulated thermal environment can be stored as a functional model unit. Based on input data for parameterizing / adjusting the simulated thermal management components, a simulated thermal environment allows the behavior of these components to be simulated, thus simulating a new state of the environment.
[0101] In other examples, the process can be carried out in a real environment including physical thermal management components.
[0102] The example will be illustrated with regard to the simulated vehicle shown in [Fig.1], which includes several thermal management components 25, 30, 35.
[0103] The illustrated method also applies when the vehicle includes a single thermal management component such as, for example, a heating, ventilation and air conditioning (HVAC) system.
[0104] The method begins by introducing training data into an S401 artificial intelligence model to train said artificial intelligence model by means of a reinforcement learning agent to control at least one thermal management component to achieve a parameter related to a given thermal comfort level. The training data comprises training datasets, each training dataset including at least one parameter related to the thermal environment of the user and / or the vehicle. In particular, a training dataset defines a specific scenario for training the artificial intelligence model using the reinforcement learning agent.
[0105] For example, a training data set may include the initial temperature outside the vehicle and / or the initial temperature inside the vehicle and / or an initial user thermal comfort level. The user thermal comfort level may be a user comfort index within a range of values. According to this example, a training data set may include the following parameters: an initial temperature parameter outside the vehicle with a value of -10°, an initial temperature parameter inside the vehicle with a value of 5°, and a parameter relating to an initial user thermal comfort index with a value of -2.
[0106] The parameter relating to a given thermal comfort level may, for example, include a thermal comfort index. This index may be a value, for example, the midpoint of a range of values representing the user's thermal comfort level. In one example, the range of values is between -4 and 4, corresponding to a perceived temperature for the user ranging from cold to warm, and the parameter relating to a given thermal comfort level corresponds to the value 0, this value corresponding to optimal thermal comfort.
[0107] Training can be carried out using the following steps. In particular, a prediction step by the learning agent through reinforcement of a control action of the thermal management component(s) S402, based on the state of the environment and the parameter relating to the given thermal comfort to be achieved. Initially, the state of the environment is represented by the set of training data.
[0108] In the example considered, the predicted control action may include an increase in the power supplied to the heating, ventilation, and air conditioning unit 25 and / or an increase in the heating level of the heated seat 30 and / or the activation of the radiant panel(s) 35. Indeed, in this example, since the user's initial thermal comfort index has a value of -2, the user's perception is one of cold. Therefore, the predicted control action will consist of an increase or activation of at least one thermal management component.
[0109] The training may further consist of training the artificial intelligence model using the reinforcement learning agent on the basis of a reward provided by the environment module 45, the reward is determined based on the state of the environment and the given thermal comfort parameter. In particular, the prediction step S402 can also be a function of the reward provided by the environment module.
[0110] The prediction step can be followed by an execution step of the control action controlling the thermal management component(s) S403.
[0111] The execution of the control action results in a change of the state of the environment to a new state, which may be closer to or further from the desired target state corresponding to the given thermal comfort parameter.
[0112] In the embodiment in which the thermal management components are simulated, the execution of the action is simulated and the new state of the environment is simulated.
[0113] The new state of the environment will then be observed by the environment module 45.
[0114] Determining the new state of the environment may include, in particular, determining the user's new state of thermal comfort, notably through the determination of the user's physiological comfort index. Physiological comfort expresses an overall thermal sensation, i.e., over the whole body, or a localized thermal sensation, i.e., over a specific area of the user's body, and depends in particular on the user's body heat, whether produced or absorbed in the area concerned, notably due to metabolic activity. It may also be based on the user's level of clothing.
[0115] The determination of the user's thermal comfort state can be carried out by the device for determining a physiological comfort of the user 50 of the environment module 45.
[0116] According to a particular embodiment, the user's physiological comfort index is determined from measurements of thermal or physiological quantities of different parts of the user's body, as illustrated in [Fig. 5]. Thus, the user's body is segmented in order to determine a local physiological comfort index for each body segment or for one or more body segments. The segments are, for example, the head 510, the chest 515, the legs 520, the feet 525, and the arms 530. Other, more or less precise, segmentations can be used. The user's physiological comfort index is then obtained from the local physiological comfort indices, for example, by averaging the values of the determined indices.
[0117] The new state of the environment may further include data related to the thermal environment inside the vehicle (e.g., the temperature inside the vehicle) and / or data related to the thermal environment outside the vehicle (e.g., the temperature outside the vehicle). For this purpose, the The new state of the environment can also be observed and captured by the device for obtaining at least one data related to the thermal environment of the vehicle 55 of the environment module 45 illustrated in [Fig.1].
[0118] In addition, the new state of the environment may include a state of energy consumption of the vehicle's thermal management components.
[0119] Following the execution of the predicted control action, the training may include a reception step S404 by the learning agent through reinforcement of the observed environment state and a reward.
[0120] The reward indicates to the reinforcement learning agent whether the control action performed has allowed the vehicle's thermal environment to approach, or even reach, the target defined by the given thermal comfort parameter, or to move away from this target. The reward can be determined, in particular by the environment module 45, based on the observed state of the environment and the parameter relating to a given thermal comfort.
[0121] Thus, the reward provided by the environment module may consist of:
[0122] (i) reward the learning agent by reinforcement when the state of the environment observed in the environment module approaches the parameter relating to the given thermal comfort to be achieved; and / or
[0123] (ii) penalize the reinforcement learning agent when the state of the environment observed in the environment module deviates from the parameter relating to the given thermal comfort to be achieved.
[0124] According to a particular embodiment, the awarding of a reward can be done according to the method described below when the objective is to achieve the parameter relating to a given thermal comfort level and to optimize the energy consumed by the thermal management components. In particular, the reward provided by the environment module also consists of:
[0125] (i) reward the learning agent by reinforcement further when the state of the environment observed in the environment module shows a decrease in energy consumption;
[0126] (ii) penalize the reinforcement learning agent when the state of the environment observed in the environment module shows an increase in energy consumption.
[0127] To this end, this method may include a weighting of the new observed state of the environment with regard to the two objectives, namely comfort and energy consumption. Thus, for example:
[0128] - if the user's thermal comfort level is outside an acceptable range [- D, D] represented by the parameter relating to a given thermal comfort, the penalty linked to energy consumption is minimized by using a low weighting, so that the reinforcement learning agent is encouraged to reach the desired comfort state (i.e., a value of 0 for the user's thermal comfort index) as soon as possible.
[0129] - if the user's thermal comfort level is within the acceptable range, The reinforcement learning agent is encouraged to stabilize in the comfort zone and focus on energy consumption, which is given a higher weighting.
[0130] The two states above can be expressed using the following expression: ™.....■■■■ < [ (l.....ssi . pO;i LmQ
[0131] In this expression, Pt is the energy consumption at time t, TCt is the value of the thermal comfort state at time t, D is a parameter delimiting the acceptable comfort zone, and alpha ("a") is a trade-off parameter that defines the priorities between energy consumption and comfort. A high value of a prioritizes reducing energy consumption over improving comfort, while a low value of a prioritizes improving comfort over reducing energy consumption. The parameter D can also be modified. For example, increasing the parameter D widens the range of thermal comfort values TC over which reducing energy consumption will take priority.
[0132] Steps S402-S404 can be repeated until the reinforcement learning agent determines that the observed environmental state has reached the parameter for a given thermal comfort. In particular, the reinforcement learning agent can determine that the parameter for a given thermal comfort has been reached when the reward converges to a maximum value.
[0133] If the reward does not converge to a maximum value, then the reinforcement learning agent will predict a new control action based on the new state of the environment and the reward at step S402, the predicted action will be executed S403 and the learning agent will receive the new state of the environment and the new reward S404.
[0134] In some embodiments, steps S402-S404 are repeated until the environmental state has reached the parameter relating to a given thermal comfort level and the energy consumption of the thermal management components has reached an optimal energy consumption state. The reinforcement learning agent can determine that the optimal energy consumption state has been reached when the reward converges to a maximum value.
[0135] According to a particular embodiment, the training steps are repeated until the user's thermal comfort state falls within a range of values for the given thermal comfort parameter, and then the training steps are repeated until the environment's energy consumption state reaches the optimal energy consumption state. Thus, the aim is first to achieve the user's thermal comfort and then, secondly, to optimize energy consumption.
[0136] The execution of the training process to completion can be called a "training episode." In some embodiments, several training episodes can be performed to train the artificial intelligence model. Each training episode can have a set of training data defining a particular scenario. For example, the artificial intelligence model can be trained to cope with a range of different outdoor temperatures.
[0137] The training episode(s) allow the artificial intelligence model to learn a strategy that enables it to achieve the given thermal comfort parameter, the strategy comprising a set of control actions, an episode being defined by a set of training data.
[0138] According to one embodiment, the deep learning agent is installed in the thermal management system 20 illustrated in [Fig. 1], in particular in a control device 60, which will control said at least one thermal management component of the vehicle. To this end, the control device 60 may include memory to store the process for training an artificial intelligence model. The control device 60 is connected to the thermal management components 25, 30, and 35.
[0139] The control device 60 may include at least one computer or processor and at least one memory in which a computer program is stored. The computer program is configured to implement the method for training an artificial intelligence model as described above. The computer program includes instructions that can be executed by the computer or processor that cause the control device 60 to execute the steps of the method for training an artificial intelligence model.
[0140] The computer program may also be stored in a computer-readable medium. The computer-readable medium may include memory for storing instructions. The memory may include any suitable memory for storing data and executable instructions, such as read-only memory, rewritable flash memory, and a hard disk drive.
[0141] Figure 6 illustrates a system capable of implementing the method for training an artificial intelligence model for use in a vehicle to control several thermal management components of a vehicle in order to optimize comfort thermal of a vehicle user and optimize the energy consumed by thermal management components, in particular by means of simulated thermal management components.
[0142] The system includes, in particular, an environment module 610, corresponding to environment module 45 of [Fig. 1]. The environment module 610 is capable of simulating the thermal environment of a vehicle or a part of a vehicle, such as the passenger compartment. For this purpose, the environment module includes, for example, a model of a heating, ventilation, and air conditioning (HVAC) system, a model of a radiant panel, and / or a model of a heated seat. According to the illustrated embodiment, the thermal management components are simulated. However, according to another embodiment, the thermal management components are real components. The system further includes a controller 620 capable of implementing a reinforcement learning agent.
[0143] This system is based on the Markov decision process. Thus, at each step, the reinforcement learning agent in the controller 620 collects an environmental state including data related to the thermal environment and data related to the energy consumption of the thermal management component(s) observed 630 in the environment module 610. The reinforcement learning agent then predicts a control action 640 for the control of the thermal management component(s).
[0144] The action consists of setting data controlling the thermal management components. For example, the setting of a heating, ventilation and air conditioning (HVAC) device is carried out by supplying a power of, for example, between -6000 and 6000 W, the setting of a heated seat is carried out by applying a heating level, the level being 0, 1, 2, 3 or 4 and the setting of one or more radiant panels consists of activating or deactivating the panel(s).
[0145] The control action 640 is then executed. The environment module 610 then transmits to the reinforcement learning agent of the controller 620 a reward 650 which reflects the evolution of the thermal environment determined by the environment module with regard to the parameter relating to a given thermal comfort, i.e. a penalty for energy consumption and thermal discomfort of the user or a positive reward for a reduction in energy consumption and improvement of the user's thermal comfort, as previously described.
[0146] Examples of results from training an artificial intelligence model using a reinforcement learning agent are now described. The parameters used in these experiments are as follows. First, the following algorithms were used for training the intelligence model artificial by means of the reinforcement learning agent: an actor-critical algorithm (Soft actor-critic), a DQN network algorithm (DQN - Deep Q-Network), and a proximal policy optimization (PPO) algorithm.
[0147] The training duration comprises 2 million time periods. Figures 7A, 7B, 7C, and 7D show the results of the experiments confirming that the reinforcement learning agent is capable of achieving episodic reward convergence and thermal comfort, minimizing the duration of discomfort, and reducing energy consumption. In particular, [Fig. 7A] illustrates reward convergence during learning, [Fig. 7B] illustrates the user's thermal comfort during learning, [Fig. 7C] illustrates the duration of discomfort for the user, and [Fig. 7D] illustrates energy consumption during learning.
[0148] Figures 8A and 8B show respectively the results relating to the thermal comfort index and energy consumption for an outside temperature range from -20° to 10°.
[0149] As shown in [Fig. 8A], the reinforcement learning agent is able to reach thermal comfort in a relatively short time and then stabilize within the comfort zone. The reinforcement learning agent is indeed able to reach thermal comfort smoothly, without overshooting. Furthermore, the energy consumption curve shown in [Fig. 8B] confirms that the reinforcement learning agent is able to accelerate the process of reaching the comfort zone and then reduce energy consumption within that zone.
[0150] Figure 9 shows the results relating to the time required to reach a parameter for a given thermal comfort level as a function of the outside temperature, using three different combinations of thermal management components. Figure 9 shows that the use of a heating, ventilation, and air conditioning (HVAC) system alone does not allow the parameter for a given thermal comfort level to be reached. However, the combined use of an HVAC system, heated seats, and radiant panels improves thermal comfort control by reducing the time required to reach the parameter for a given thermal comfort level.
[0151] Figure 10 shows the energy consumption curves over time for three different combinations of thermal management components. Figure 10 shows that controlling several thermal management components allows for more efficient energy use. The reinforcement learning agent optimally combines the heating, ventilation, and air conditioning system, the heated seat, and the radiant panels. It utilizes the heat sources additional measures to accelerate the convergence towards comfort and improve energy consumption efficiency.
[0152] Following the training process, the trained artificial intelligence model will be tested. In particular, the testing phase will consist of running the training process with test data in order to evaluate the artificial intelligence model. The test data is similar to training data. Indeed, the test data can be a subset of the training data that was not used during the learning phase.
[0153] Following the testing phase, a trained artificial intelligence model is suitable for use in a vehicle to control at least one thermal management component of the vehicle in order to optimize the thermal comfort of a vehicle user. Furthermore, the trained artificial intelligence model can also optimize the energy consumed by the thermal management components when it is capable of controlling a plurality of thermal management components.
[0154] Thus, the control of at least one thermal management component of a vehicle by the use of a trained artificial intelligence model, in particular trained according to the training method described above, will now be described.
[0155] As illustrated in [Fig. 1], a vehicle may include one or more thermal management components such as a heating, ventilation, and air conditioning system, heated seats, and radiant panels. In some examples, only one thermal management component may be controlled. According to some embodiments, only the user's thermal comfort is considered. According to other embodiments, several thermal management components are controlled. In this case, the thermal management components may be controlled in such a way as to optimize the energy consumption of the thermal management components while optimizing the user's thermal comfort.
[0156] The trained artificial intelligence model is installed in a vehicle to control at least one thermal management component of the vehicle in order to optimize the thermal comfort of a vehicle user. In particular, the artificial intelligence model is installed in a control device 60, illustrated in [Fig. 1], which will control said at least one thermal management component of the vehicle. To this end, the control device 60 may include memory to store the trained artificial intelligence model.
[0157] The control device 60 is connected to the thermal management components 25, 30 and 35 and uses the artificial intelligence model to control the thermal management components.
[0158] The control device 60 may include at least one computer or processor and at least one memory in which a computer program and The trained artificial intelligence model. The computer program is configured to implement a process for controlling at least one thermal management component of a vehicle. The computer program includes instructions that can be executed by the computer or processor, which lead the control device 60 to execute the steps of the process for controlling at least one thermal management component of a vehicle.
[0159] The computer program may also be stored in a computer-readable medium. The computer-readable medium may include memory for storing instructions. The memory may include any suitable memory for storing data and executable instructions, such as read-only memory, rewritable flash memory, and a hard disk drive.
[0160] The control device 60 is configured to control at least one thermal management component 25, 30, and 35 to achieve a parameter related to a given thermal comfort level. The control device 60 can also be configured to control at least two thermal management components 25, 30, and 35 to optimize the energy consumed by the thermal management components 25, 30, and 35. Controlling the thermal management components 25, 30, and 35 may involve modifying the power setpoints of the thermal management components. For example, the power supplied by the heating, ventilation, and air conditioning (HVAC) unit 25 may be a continuous value within a predetermined range (e.g., from 0 to 6000 W). The heat level of the seat heater may be a discrete value (e.g., 0, 1, 2, 3, or 4). The radiant panels may have a binary setpoint (e.g., YES or NO).
[0161] The parameter relating to a given thermal comfort can be a particular temperature in an area of the vehicle or a state of thermal comfort of a user in the vehicle defined for example by a thermal comfort index as described above.
[0162] In the present example, an environmental module 45 is also provided in the vehicle as illustrated in [Fig. 1]. The environmental module 45 includes a processor configured to determine the state of the environment, and in particular one or more data points related to the vehicle's thermal environment. The environmental module 45 is connected to the control device 60 and can provide it with the data points related to the vehicle's thermal environment. In other examples, the environmental module may be implemented in the control device.
[0163] The state of the environment may include a state of thermal comfort for the user, which may also be expressed using the thermal comfort index described above. Alternatively, or in addition, the state of the environment may include data related to the thermal environment inside the vehicle (e.g. the temperature inside the vehicle) and / or data related to the thermal environment outside the vehicle (e.g., the temperature outside the vehicle). The environmental status may include a representative energy consumption status of the thermal management components 25, 30, and 35.
[0164] For this purpose, the environment module 45 can receive data from the device for determining the physiological comfort of the user 50 and / or from the device for obtaining at least one piece of data related to the thermal environment of the vehicle 55.
[0165] The device for determining the physiological comfort of the user 50 may include measuring devices for determining the user's physiological comfort index. In addition, the device for obtaining at least one data point related to the thermal environment of the vehicle 55 may include a measuring device comprising one or more sensors such as a sunlight sensor, a temperature sensor, in particular a temperature sensor at an air outlet of the installation, a sensor for the temperature prevailing in the passenger compartment, or a humidity sensor.
[0166] An embodiment of a computer-implemented method for controlling at least one thermal management component 25, 30, 35 of a vehicle in order to optimize the thermal comfort of a vehicle user will now be described with reference to [Fig. 1 1]. The method is implemented, for example, in the control device 60.
[0167] The method includes obtaining one or more data points related to the thermal environment of the user and / or the vehicle SI 101. The data related to the thermal environment of the user and / or the vehicle are obtained, in particular, by the environment model 45 as previously described with regard to the drive method. For example, the control device 60 can obtain from the environment module a thermal comfort index for the user with a value of -2.
[0168] The process continues with a step of inputting said at least one data related to the thermal environment of the user and / or the vehicle into an artificial intelligence model installed in the vehicle S1102, in particular in the control device 60. The artificial intelligence model has been trained to control said at least one thermal management component to achieve a parameter relating to a given thermal comfort.
[0169] The process continues with a step consisting of bringing the trained artificial intelligence model to control said at least one thermal management component 25, 30, 35 to achieve the given thermal comfort parameter SI 103.
[0170] To do this, the trained artificial intelligence model can determine the actions to be executed for one or more of the thermal management components 25, 30, 35 on the basis of the parameter relating to a given thermal comfort and said at least data related to the thermal environment of the user and / or the vehicle in order to optimize the thermal comfort of a vehicle user.
[0171] For example, the trained artificial intelligence model can determine that the power of the heating, ventilation, and air conditioning unit 25 needs to be increased (e.g., by 1000 W). In some examples, the artificial intelligence model can also determine control actions to optimize the energy consumed by the thermal management components 25, 30, 35.
[0172] The control device 60 then commands the thermal management components 25, 30, 35, based on the actions determined by the trained artificial intelligence model. For example, the control device 60 can command the heating, ventilation and air conditioning device 25 to increase its power by 1000 W.
[0173] The trained artificial intelligence model can also determine a sequence of actions to be performed on the thermal management components, the actions being able to be sequenced in time.
[0174] The present invention also relates to a computer system comprising means for implementing the method for training an artificial intelligence model for use in a vehicle and / or the method for controlling at least one thermal management component of a vehicle.
[0175] The present invention also relates to a computer program comprising instructions which, when executed, cause a device or computer system to execute a method for training an artificial intelligence model for use in a vehicle and / or the method for controlling at least one thermal management component of a vehicle.
Claims
Demands
1. A computer-implemented method for training an artificial intelligence model for use in a vehicle to control at least one thermal management component (25, 30, 35) of a vehicle in order to optimize the thermal comfort of a vehicle user, the method comprising a step of introducing training data into an artificial intelligence model to train said artificial intelligence model (S401) by means of a reinforcement learning agent, in order to control said at least one thermal management component to achieve a parameter relating to a given thermal comfort, the training data comprising training data sets, each training data set comprising at least one parameter related to the thermal environment of the user and / or the vehicle.
2. A method according to claim 1, characterized in that the method is implemented to control at least two thermal management components of the vehicle and to further optimize the energy consumed by the thermal management components.
3. A method according to any one of the preceding claims, characterized in that the reinforcement learning agent is coupled to an environment module (45) capable of capturing the state of the environment, the training consisting of training the artificial intelligence model by means of the reinforcement learning agent on the basis of a state of the environment observed in the environment module.
4. Method according to the preceding claim, characterized in that the training comprises the following steps: - a prediction step by the learning agent by reinforcement of a control action of the thermal management component(s) (S402), according to the state of the observed environment and the parameter relating to the given thermal comfort to be achieved, and - a step of execution of the control action controlling the thermal management component(s) (S403).
5. A method according to claim 3 or claim 4, characterized in that the training further consists of training the artificial intelligence model by means of the learning agent by reinforcement based on a reward provided by the environment module, the reward being determined according to the observed state of the environment and the given parameter relating to thermal comfort.
6. A method according to claim 4 and claim 5, characterized in that the prediction step is further a function of the reward provided by the environment module.
7. A method according to claim 5 or claim 6, characterized in that the training further comprises a reception step by the learning agent through reinforcement of the observed environment state and the reward (S404) following the execution of the predicted control action.
8. A method according to any one of claims 5 to 7, characterized in that the reward provided by the environment module consists of: (i) rewarding the learning agent by reinforcement when the state of the environment observed in the environment module approaches the given thermal comfort parameter to be achieved; (ii) penalizing the learning agent by reinforcement when the state of the environment observed in the environment module deviates from the given thermal comfort parameter to be achieved.
9. A method according to claim 2 and the preceding claim, characterized in that the reward provided by the environment module further consists of: (i) rewarding the learning agent by reinforcement further when the state of the environment observed in the environment module shows a decrease in energy consumption; (ii) penalizing the learning agent by reinforcement when the state of the environment observed in the environment module shows an increase in energy consumption.
10. A method according to any one of the preceding claims 5 to 9, characterized in that the observed state of the environment has reached the given thermal comfort parameter when the reward converges to a maximum value.
11. A method according to claim 4, characterized in that the training of the artificial intelligence model by means of the agent Reinforcement learning includes one or more training episodes to learn a strategy that enables the achievement of the given thermal comfort parameter, the strategy comprising a set of control actions, an episode being defined by a set of training data.
12. A method according to any one of the preceding claims 3 to 11, characterized in that the training steps are repeated until the observed state of the environment has reached the given thermal comfort parameter.
13. A method according to any one of the preceding claims, characterized in that the parameter relating to the given thermal comfort is defined by a range of values.
14. A method according to any one of the preceding claims, characterized in that the observed state of the environment includes a state of thermal comfort for the user.
15. A method according to the preceding claim, characterized in that the user's thermal comfort state is within a range of values, the central value of the range corresponding to the value of the parameter relating to the given thermal comfort.
16. A method according to any one of the preceding claims, characterized in that the observed state of the environment further comprises a state of energy consumption.
17. A method according to the preceding claim, characterized in that the training steps are further repeated until the energy consumption state of the observed environment has reached an optimal energy consumption state.
18. A method according to the preceding claim and according to any one of claims 5 to 10, characterized in that the observed state of the environment has reached a state of optimal energy consumption when the reward converges to a maximum value.
19. A method according to claims 10, 13 and 17, characterized in that the training steps are repeated until the user's thermal comfort state is within a range of values of the given thermal comfort parameter and then the training steps are repeated until the energy consumption state of the environment has reached the optimal energy consumption state.
20. A method according to any one of the preceding claims, characterized in that said at least one parameter related to the user's thermal environment includes at least one parameter related to a state of thermal comfort of the user.
21. A method according to any one of the preceding claims, characterized in that said at least one parameter related to the thermal environment of the vehicle comprises at least one parameter related to the internal thermal environment of the vehicle and / or at least one parameter related to the external thermal environment of the vehicle.
22. A method according to any one of the preceding claims, characterized in that said at least one thermal management component comprises at least one heating, ventilation and air conditioning device (25).
23. A method according to any one of the preceding claims, characterized in that said at least one thermal management component further comprises at least one radiant panel (35) and / or at least one heated seat (30).
24. A method according to any one of claims 1 to 23, characterized in that at least one thermal management component is a simulated component.
25. A method according to any one of the preceding claims, characterized in that the training of the artificial intelligence model by means of the reinforcement learning agent is carried out using a soft actor-critic algorithm, or a DQN-Deep Q-Network algorithm, or a proximal policy optimization (PPO) algorithm.
26. A computer-implemented method for controlling at least one thermal management component (25, 30, 35) of a vehicle in order to optimize the thermal comfort of a vehicle user, the method comprising: - a step of obtaining at least one piece of data related to the thermal environment of the user and / or the vehicle (SI 101), - a step of inputting said at least one piece of data related to the thermal environment of the user and / or the vehicle into an artificial intelligence model (SI 102), the artificial intelligence model having been trained to control said at least one thermal management component to achieve a parameter relating to a given thermal comfort, - bring the trained artificial intelligence model to control said at least one thermal management component to achieve the parameter relating to the given thermal comfort (SI 103).
27. A method according to the preceding claim, characterized in that the method is further implemented to control at least two thermal management components of the vehicle and to optimize the energy consumed by the thermal management components.
28. A method according to the preceding claim, characterized in that the method further comprises a step of bringing the trained artificial intelligence model to control said thermal management components to achieve a state of optimal energy consumption.
29. A method according to any one of claims 26 to 28, characterized in that said at least one data related to the thermal environment of the user and / or the vehicle includes at least one data related to a state of thermal comfort of the user and / or at least one data related to the environment of the vehicle.
30. Method according to the preceding claim, characterized in that said at least one data related to the thermal environment of the vehicle comprises data related to the internal thermal environment of the vehicle and / or data related to the external thermal environment of the vehicle.
31. A method according to any one of claims 26 to 30, characterized in that said at least one thermal management component comprises at least one heating, ventilation and air conditioning device (25).
32. A method according to any one of claims 26 to 31, characterized in that said at least one thermal management component further comprises at least one radiant panel (35) and / or at least one heated seat (30).
33. A computer-implemented method according to any one of claims 26 to 32, characterized in that the artificial intelligence model was trained according to the method according to any one of claims 1 to 25.
34. Computer system comprising means for implementing the method according to any one of claims 1 to 25 or the method according to any one of claims 26 to 32.
35. A computer program comprising instructions which, when executed, cause a device or computer system to perform a process according to any one of claims 1 to 25 or any one of claims 26 to 32.
36. Artificial intelligence model for controlling at least one thermal management component of a vehicle in order to optimize the thermal comfort of a vehicle user according to the method according to any one of claims 1 to 25 or any one of claims 26 to 32.
37. System for a vehicle comprising at least one thermal management component and a control device (60), the control device implementing the method according to any one of claims 26 to 32.
38. System according to claim 37, characterized in that said at least one thermal management component comprises at least one heating, ventilation and air conditioning (HVAC) device.
39. System according to claim 37 or claim 38, characterized in that said at least one thermal management component further comprises at least one radiant panel and / or at least one heated seat.
Citation Information
Patent Citations
Thermal management system for vehicle passenger compartments and thermal management method implemented by such a system
FR3129629A1
Electric automobile air conditioner and passenger compartment heat management control method based on TD3 algorithm
CN116714411A
Pure electric vehicle passenger compartment air conditioner refrigeration control method based on DDPG
CN116729060A
Hybrid electric vehicle energy and heat management multi-agent cooperative control method
CN118124333A
CONTROL SYSTEM FOR AN ARTIFICIAL INTELLIGENCE-BASED INTEGRATED VEHICLE THERMAL MANAGEMENT SYSTEM AND METHOD FOR CONTROLLING THE SAME
DE112022000253T5