Managing a thermal environment in a vehicle
A reinforcement learning-based AI model optimizes thermal comfort and energy efficiency in vehicles by controlling HVAC systems, heated seats, and radiant panels, addressing inefficiencies in existing thermal management systems.
Patent Information
- Application Number
- PCT/EP2025/069246
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-08
- Filing Date
- 2025-07-07
- Publication Date
- 2026-01-15
AI Technical Summary
Existing thermal management systems in vehicles are inefficient and energy-intensive, lacking automation to optimize thermal comfort and energy consumption based on external and internal environmental factors.
A computer-implemented method using reinforcement learning to train an artificial intelligence model for controlling thermal management components, such as HVAC systems, heated seats, and radiant panels, to achieve a given thermal comfort level while minimizing energy consumption.
The method enables efficient automation of thermal management, optimizing user comfort and reducing energy consumption by dynamically adjusting thermal components based on environmental conditions.
Smart Images

Figure EP2025069246_15012026_PF_FP_ABST
Abstract
Description
[0001] Managing a thermal environment in a vehicle
[0002] technical field
[0003] The invention relates to the management of a thermal environment in a vehicle. In particular, the invention relates to a method for training an artificial intelligence model for use in a vehicle to control at least one thermal management component of a vehicle in order to optimize the thermal comfort of a vehicle user, a method for controlling at least one thermal management component of a vehicle and the associated system.
[0004] Previous art
[0005] A user's thermal comfort can be defined as the condition that reflects a user's satisfaction with their thermal environment and is subjectively assessed. Thermal comfort is one of the factors considered when designing the thermal environments of vehicles. Indeed, users spend a significant amount of time in their vehicles, and the perceived thermal comfort contributes to their well-being within the vehicle.
[0006] To achieve thermal comfort, the user can manually activate one or more thermal components via a human-machine interface. Among the thermal components that may be present in a vehicle are air conditioning systems, such as a heating, ventilation, and air conditioning (HVAC) system. In such an air conditioning system, a heat transfer fluid circulates through a fluid circuit comprising a compressor, a heat exchanger called a condenser, an expansion valve, and a second heat exchanger called an evaporator.
[0007] There are also reversible air conditioning systems; these systems can absorb heat energy from the outside air at a heat exchanger, called an evaporator-condenser, and transfer it into the passenger compartment, notably by means of a dedicated heat exchanger. In this case, the air conditioning system functions as a heat pump.
[0008] Other thermal regulation devices may also be present in the vehicle, such as radiant panels, heated seats, a heated steering wheel, and / or contact heating surfaces located, for example, on an armrest. The vehicle's thermal system, which manages the thermal regulation device(s), is designed to achieve various objectives, such as managing thermal comfort, reducing energy consumption, and improving air quality.
[0009] Various solutions have been proposed for managing a thermal regulation device to achieve user thermal control. However, these solutions do not allow for the efficient automation of thermal regulation device management, particularly when taking into account the external and / or internal thermal environment of the vehicle, such as the vehicle's exterior temperature, level of sunlight, vehicle speed, humidity level, wind intensity, internal vehicle temperature, humidity level inside the vehicle, etc.
[0010] Furthermore, all these thermal regulation systems are very energy-intensive and significantly increase the vehicle's energy consumption, especially in cold outside temperatures. This is particularly true when the thermal regulation systems are activated and adjusted manually by the vehicle's user. Indeed, very often, in an attempt to achieve thermal comfort as quickly as possible, the user activates all the systems simultaneously and sets them to a high power level.
[0011] Summary of the invention
[0012] The invention aims in particular to offer an improvement in the control of one or more known thermal management components in order to optimize the thermal comfort of a vehicle user.
[0013] To this end, the invention relates to a computer-implemented method for training an artificial intelligence model for use in a vehicle to control at least one thermal management component of a vehicle in order to optimize the thermal comfort of a vehicle user, the method comprising a step of introducing training data into an artificial intelligence model to train said artificial intelligence model by means of a reinforcement learning agent, in order to control said at least one thermal management component to achieve a parameter relating to a given thermal comfort, the training data comprising training data sets, each training data set comprising at least one parameter related to the thermal environment of the user and / or the vehicle.
[0014] This process involves training an artificial intelligence model using a reinforcement learning agent to control at least one thermal management component of a vehicle to achieve a parameter related to a given level of thermal comfort. Reinforcement learning relies on trial and error strategies to arrive at a flexible compromise that can be configured to optimize user comfort.
[0015] According to a particular embodiment, the process is implemented to control at least two thermal management components of the vehicle and to further optimize the energy consumed by the thermal management components.
[0016] Based on these characteristics, reinforcement learning allows for a configurable trade-off between user comfort and reduced energy consumption. Furthermore, the training is designed to control multiple thermal management components, including simultaneously.
[0017] The reinforcement learning agent can be coupled with an environment module capable of capturing the state of the environment, the training consisting of training the artificial intelligence model by means of the reinforcement learning agent based on an environment state observed in the environment module.
[0018] The training may include the following steps: a prediction step by the learning agent by reinforcement of a control action of the thermal management component(s), based on the observed state of the environment and the parameter relating to the given thermal comfort to be achieved, and an execution step of the control action controlling the thermal management component(s).
[0019] The training may also consist of training the artificial intelligence model using the reinforcement learning agent on the basis of a reward provided by the environment module, the reward being determined according to the state of the observed environment and the given parameter relating to thermal comfort.
[0020] Reward-linked learning ensures accelerated convergence to the given thermal comfort level when thermal management components are activated and can also lead to substantial energy savings during the operation of these components.
[0021] The prediction step can also be dependent on the reward provided by the environment module. The training may further include a reception step by the learning agent through reinforcement of the observed environment state and the reward following the execution of the predicted control action.
[0022] The reward provided by the environment module may consist of:
[0023] (i) reward the learning agent by reinforcement when the state of the environment observed in the environment module approaches the given thermal comfort parameter to be achieved;
[0024] (ii) penalize the reinforcement learning agent when the state of the environment observed in the environment module deviates from the parameter relating to the given thermal comfort to be achieved.
[0025] The reward provided by the environment module may also include:
[0026] (i) reward the learning agent by further reinforcement when the state of the environment observed in the environment module shows a decrease in energy consumption;
[0027] (ii) penalize the reinforcement learning agent when the state of the environment observed in the environment module shows an increase in energy consumption.
[0028] According to a particular embodiment, the observed state of the environment has reached the given thermal comfort parameter when the reward converges to a maximum value.
[0029] Training the artificial intelligence model using the reinforcement learning agent may include one or more training episodes to learn a strategy that enables the achievement of the given thermal comfort parameter, the strategy comprising a set of control actions, an episode being defined by a set of training data.
[0030] The training steps can be repeated until the observed environmental state has reached the given thermal comfort parameter.
[0031] The parameter relating to thermal comfort given can be defined by a range of values.
[0032] The observed state of the environment may include a state of thermal comfort for the user. The user's state of thermal comfort may fall within a range of values, with the central value of the range corresponding to the value of the given thermal comfort parameter.
[0033] The observed state of the environment may also include a state of energy consumption.
[0034] According to one embodiment, the training steps are further repeated until the energy consumption state of the observed environment has reached an optimal energy consumption state.
[0035] According to one embodiment, the observed state of the environment has reached a state of optimal energy consumption when the reward converges to a maximum value.
[0036] The training steps can be repeated until the user's thermal comfort state is within a range of values for the given thermal comfort parameter, and then the training steps can be repeated until the energy consumption state of the environment has reached the optimal energy consumption state.
[0037] Said at least one parameter related to the user's thermal environment may include at least one parameter related to a state of thermal comfort of the user.
[0038] Said at least one parameter related to the thermal environment of the vehicle may include at least one parameter related to the internal thermal environment of the vehicle and / or at least one parameter related to the external thermal environment of the vehicle.
[0039] Said at least one thermal management component may include at least one heating, ventilation and air conditioning device.
[0040] Said at least one thermal management component may further include at least one radiant panel and / or at least one heated seat.
[0041] According to one embodiment, at least one thermal management component is a simulated component.
[0042] Training the artificial intelligence model using the reinforcement learning agent can be achieved using a soft actor-critical algorithm, a deep Q-network (DQN) algorithm, or a proximal policy optimization (PPO) algorithm. The invention also relates to a computer-implemented method for controlling at least one thermal management component of a vehicle in order to optimize the thermal comfort of a vehicle user.The process includes: a step of obtaining at least one data related to the thermal environment of the user and / or the vehicle, a step of inputting said at least one data related to the thermal environment of the user and / or the vehicle into an artificial intelligence model, the artificial intelligence model having been trained to control said at least one thermal management component to achieve a parameter relating to a given thermal comfort, leading the trained artificial intelligence model to control said at least one thermal management component to achieve the parameter relating to the given thermal comfort.
[0043] Thus, the process for controlling at least one thermal management component of a vehicle makes it possible to reach the given thermal comfort parameter in order to optimize the thermal comfort of a vehicle user, in particular in an optimal time.
[0044] The process can also be implemented to control at least two thermal management components of the vehicle and to optimize the energy consumed by the thermal management components.
[0045] The process may further include a step of enabling the trained artificial intelligence model to control said thermal management components to achieve a state of optimal energy consumption.
[0046] Said at least one piece of data related to the thermal environment of the user and / or the vehicle may include at least one piece of data related to a state of thermal comfort of the user and / or at least one piece of data related to the environment of the vehicle.
[0047] Said at least one data point related to the thermal environment of the vehicle may include data related to the internal thermal environment of the vehicle and / or data related to the external thermal environment of the vehicle.
[0048] Said at least one thermal management component may include at least one heating, ventilation and air conditioning device.
[0049] Said at least one thermal management component may further comprise at least one radiant panel and / or at least one heated seat. The invention also relates to a computer-implemented method according to the method for controlling at least one thermal management component of a vehicle described above, characterized in that the artificial intelligence model was trained according to the method for training an artificial intelligence model for use in a vehicle as described above.
[0050] The invention also relates to a computer system comprising means for implementing the method for training an artificial intelligence model for use in a previously described vehicle or a method for controlling at least one thermal management component of a previously described vehicle.
[0051] The invention also relates to a computer program comprising instructions which, when executed, cause a device or computer system to perform a method for training an artificial intelligence model for use in a previously described vehicle or the method for controlling at least one thermal management component of a previously described vehicle.
[0052] The invention also relates to a computer-readable medium on which the previously described computer program is stored.
[0053] The invention also relates to an artificial intelligence model for controlling at least one thermal management component of a vehicle in order to optimize the thermal comfort of a vehicle user according to the method for training an artificial intelligence model for use in a vehicle previously described or the method for controlling at least one thermal management component of a vehicle previously described.
[0054] The invention also relates to a system for a vehicle comprising at least one thermal management component and a control device, the control device implementing the method to control at least one thermal management component of a vehicle previously described.
[0055] Said at least one thermal management component may include at least one heating, ventilation and air conditioning device.
[0056] Said at least one thermal management component may further include at least one radiant panel and / or at least one heated seat.
[0057] List of Figures Other features and advantages of the invention will become more apparent upon reading the following description, given by way of illustrative and non-limiting example, and the accompanying figures, among which:
[0058] Figure 1 is a schematic representation of a part of a vehicle, including the passenger compartment, comprising several thermal management components.
[0059] Figure 2 is a schematic diagram illustrating a Markov decision process.
[0060] Figure 3 is a schematic diagram illustrating an example of a deep reinforcement learning process.
[0061] Figure 4 is a flowchart showing an example of a process for training an artificial intelligence model according to certain embodiments.
[0062] Figure 5 is a schematic diagram showing an example of a human body divided into sections having respective thermal comfort indices.
[0063] Figure 6 is a schematic diagram illustrating an example of a reinforcement learning process according to the present invention.
[0064] Figures 7A, 7B, 7C and 7D show graphs illustrating a reward episode, the evolution of the thermal comfort index, the duration of discomfort and energy consumption as a function of time, respectively.
[0065] Figures 8A and 8B show the curves of the thermal comfort index and energy consumption for a range of outside temperatures.
[0066] Figure 9 shows the curves of the time required to reach a target comfort level for different combinations of use of thermal management components.
[0067] Figure 10 shows the energy consumption curves for different combinations of thermal management component usage.
[0068] Figure 11 is a flowchart illustrating an example of a method for controlling one or more thermal management components of a vehicle, according to certain embodiments.
[0069] Detailed description of the invention
[0070] The following are examples. Although the description refers to one or more embodiments, this does not necessarily mean that each reference relates to the same embodiment, or that the features apply only to a single embodiment. Simple features from different embodiments can also be combined or interchanged to provide other embodiments.
[0071] Figure 1 represents a passenger compartment 10 of a motor vehicle. The vehicle may be an internal combustion engine vehicle, an electric vehicle, or a hybrid vehicle, that is, a vehicle that can be powered by both an electric motor and an internal combustion engine. The passenger compartment 10 includes a driver's seat 15. It may include another front passenger seat and rear passenger seats. The vehicle includes a thermal management system 20. The thermal management system includes at least one thermal management component.
[0072] One such thermal management component is, for example, a heating, ventilation, and air conditioning (HVAC) unit 25. Outlets 30 of heat-treated air from the HVAC unit, specifically a head air outlet 30a and a foot air outlet 30b, allow the heat-treated air from the HVAC unit 25 to be supplied to the passenger compartment 10. The user can direct the flow of heat-treated air at the outlets 30 as needed. In a particular example of an HVAC unit 25, this unit may include an additional electric heating element, thus providing supplementary heating. This additional electric heating element includes, in particular, radiant heating elements.According to a particular embodiment, the additional electric heating device may further include heating resistive elements, for example positive temperature coefficient resistors, enabling the conversion of an electric current into thermal energy.
[0073] A thermal management component can also be, for example, a heated seat device 30 installed in the driver's seat 15. Other heated seat devices can be installed, for example, in the front passenger seat and / or in the rear passenger seats.
[0074] A thermal management component may include, for example, at least one radiant panel 35. Radiant panels could be placed anywhere inside the passenger compartment 10. For example, radiant panels could be located above the occupant's windows 40 35a and / or under the occupant's feet 35b. Other radiant panels could be installed, for example, on the front passenger side and / or at the rear passenger level.
[0075] To optimize the thermal comfort of a vehicle user, the present invention relates to a method for training an artificial intelligence model for use in a vehicle to control at least one thermal management component 25, 30, 35. The training consists of training an artificial intelligence model using a reinforcement learning agent to control said at least one thermal management component to achieve a parameter related to a given thermal comfort level.
[0076] Reinforcement learning refers to a goal-oriented optimization process guided by an impact signal. Learning is thus achieved efficiently through real-time information gathered from interaction with the environment, such as the vehicle's interior. Reinforcement learning tasks can be formulated according to a Markov decision process.
[0077] A Markov decision process is shown in Figure 2. Such a process involves an agent 210 (i.e., a reinforcement learning algorithm), in particular a rational agent, receiving observations about the state 230 of an environment 220 (which can be a real or simulated environment) along with rewards 240, and performing actions 250 that will impact the environment 220. The agent then seeks to learn a matching function between the state of the environment and an action that will maximize the expected reward. This matching function is called the strategy or policy of the reinforcement learning system.
[0078] According to Markov's decision process, the state St of environment 220 at time t is provided to agent 210. Based on this state, the agent generates an action at from a set of available actions A to modify the state of the environment. The generated action at is then executed, and environment 220 transitions to a new state st+i. Furthermore, a reward rt+i associated with the environmental modification is determined. Agent 210 then receives the new environmental state st+i and the reward rt+i associated with executing the action in the environment. The reward rt+i is then used by agent 210 to determine the next action. The reinforcement agent's objective is to learn a policy that maximizes the expected reward. This process is repeated until convergence, that is, until the reward can no longer be improved.
[0079] Deep reinforcement learning is an approach to reinforcement learning that uses deep neural networks (DNNs) as an integral part of the agent. Figure 3 shows an example of a deep reinforcement learning system. Deep reinforcement learning is an advanced machine learning method, at the intersection of deep learning and reinforcement learning. In this paradigm, an intelligent agent learns to make a sequence of optimal decisions by autonomously interacting with a complex and often large-scale environment. It uses deep neural networks to process raw, high-dimensional input data (such as images, sounds, or sensor data) to estimate the state of the environment and determine the best course of action.Based on a feedback mechanism, where actions are followed by rewards or penalties, the agent iteratively adjusts the weights of its neural network to develop a decision policy that maximizes the sum of expected rewards in the long term, thus enabling it to master complex tasks without explicit rule programming.
[0080] The present invention relates, in a first aspect, to training an artificial intelligence model for use in a vehicle to control at least one thermal management component. The training consists, in particular, of training an artificial intelligence model using a deep reinforcement learning agent to control at least one thermal management component of a vehicle, in order to achieve a parameter called a "parameter relating to a given thermal comfort level." The parameter relating to a given thermal comfort level can be a specific temperature in a zone of a vehicle or a user's thermal comfort level within the vehicle. In one embodiment, the parameter relating to a given thermal comfort level is defined by a range of values. For example, it can include a temperature range or a range of values for the thermal comfort level.
[0081] The artificial intelligence model is trained using training data fed into the model. This training data comprises training datasets, and each dataset includes at least one parameter related to the thermal environment of the user and / or the vehicle. In other words, each dataset represents a thermal comfort scenario from which the thermal management component(s) must be controlled to improve thermal comfort and achieve the parameter corresponding to that specific thermal comfort level.
[0082] Said at least one parameter related to the thermal environment of the user and / or the vehicle may include at least one parameter related to the user's thermal comfort level and / or at least one parameter related to the vehicle's thermal environment. A parameter related to the user's thermal comfort level may include at least one user comfort index, in particular a user physiological comfort index. The physiological comfort index is, for example, a data point representing the thermal comfort experienced by the user. A user's comfort index is described in particular in application FR3129629A1.
[0083] A parameter related to the thermal environment of the vehicle may include at least one piece of data related to the internal thermal environment of the vehicle and / or at least one piece of data related to the external thermal environment of the vehicle.
[0084] The data related to the vehicle's external thermal environment may represent the vehicle's outside temperature and / or the level of sunlight and / or the vehicle's speed and / or the humidity level and / or the wind intensity. In particular, the data related to the vehicle's external thermal environment is representative of the vehicle's external thermal environment and may represent at least one meteorological condition of the vehicle's environment.
[0085] The data related to the vehicle's internal thermal environment can represent the vehicle's internal temperature, namely that of the passenger compartment, and / or the humidity level present in the vehicle.
[0086] When the vehicle includes at least two thermal management components, as illustrated in Figure 1, the method for training an artificial intelligence model can be implemented to control the plurality of thermal management components in order to optimize the thermal comfort of a vehicle user and the energy consumed by the thermal management components.
[0087] The artificial intelligence model is trained using a reinforcement learning agent that can be coupled with an environment module capable of capturing the state of the environment. An environment module 45 is illustrated in Figure 1. It includes, for example, a device for determining the physiological comfort of the user 50 and / or a device for obtaining at least one piece of data related to the thermal environment of the vehicle 55. Thus, the environment module 45 determines at least one piece of data related to the thermal environment of the user and / or the vehicle.Said at least one data related to the thermal environment of the user and / or the vehicle includes at least one data related to a state of thermal comfort of the user determined for example by the device for determining a physiological comfort of the user 50 and / or at least one data related to the thermal environment of the vehicle determined for example by the device for obtaining at least one data related to the thermal environment of the vehicle 55. In particular, said at least one data related to the thermal environment of the vehicle includes at least one data related to the interior thermal environment of the vehicle and / or at least one data related to the exterior thermal environment of the vehicle.
[0088] The environment module can be linked to a real environment or a simulated environment, the environment being able to include the vehicle's passenger compartment or a part of the vehicle's passenger compartment.
[0089] The training then consists of training the artificial intelligence model using the reinforcement learning agent based on a state of the environment observed in the environment module.
[0090] An example of a method for training an artificial intelligence model will now be described with reference to Figure 4. This artificial intelligence model will be trained to control at least one thermal management component 25, 30, 35 of a vehicle as illustrated in Figure 1. According to this method, the vehicle illustrated in Figure 1 is either a real vehicle or a simulated vehicle for training the artificial intelligence model. In other words, the vehicle's thermal environment will be either real and include physical thermal management components or simulated using simulated thermal management components.
[0091] Artificial intelligence model training can be achieved using a soft actor-critical algorithm, a DQN (Deep Q-Network) algorithm, or a proximal policy optimization (PPO) algorithm.
[0092] In particular, the soft actor-critic algorithm is a policy-free, model-free reinforcement learning algorithm that combines two main components: an actor and a critic. The actor is a neural network that produces a stochastic policy, that is, a probability distribution of possible actions for each state. The critic is another neural network that estimates the value function, that is, the expected return for each state-action pair.
[0093] The DQN network algorithm is a type of reinforcement learning algorithm that uses a neural network to approximate the agent's value function. The value function estimates the expected return (cumulative reward) of acting in a given state. A DQN network algorithm learns to update the value function using a technique that compares the actual and predicted rewards of each action and minimizes the difference. The Proximal Policy Optimization (PPO) algorithm is a type of reinforcement learning algorithm that adopts a control function that prevents an agent's policy update from being too large or too small, thus improving the stability of the learning process.
[0094] In this example, the process is implemented by computer and is executed using a simulated thermal environment of a vehicle containing one or more simulated thermal management components. The simulated thermal environment can be stored as a functional model unit. Based on input data for configuring / adjusting the simulated thermal management components, a simulated thermal environment allows the behavior of these components to be simulated, thus simulating a new state of the environment.
[0095] In other examples, the process can be executed in a real-world environment including physical thermal management components.
[0096] The example will be illustrated with regard to the simulated vehicle shown in Figure 1, which includes several thermal management components 25, 30, 35.
[0097] The illustrated process also applies when the vehicle includes a single thermal management component such as, for example, a heating, ventilation and air conditioning (HVAC) system.
[0098] The process begins by feeding training data into an S401 artificial intelligence model to train the model, using a reinforcement learning agent, to control at least one thermal management component to achieve a parameter related to a given level of thermal comfort. The training data comprises training datasets, each containing at least one parameter related to the thermal environment of the user and / or the vehicle. Specifically, each training dataset defines a particular scenario for training the artificial intelligence model using the reinforcement learning agent.
[0099] For example, a training data set might include the initial temperature outside the vehicle and / or the initial temperature inside the vehicle and / or the user's initial thermal comfort level. The user's thermal comfort level could be a user comfort index within a range of values. In this example, a training data set might include the following parameters: an initial temperature outside the vehicle with a value of -10°C, an initial temperature inside the vehicle with a value of 5°C, and a parameter relating to the user's initial thermal comfort index with a value of -2.
[0100] The parameter relating to a given level of thermal comfort can, for example, include a thermal comfort index. This index can be a value, such as the midpoint of a range representing the user's thermal comfort level. For example, the range of values is between -4 and 4, corresponding to a perceived temperature for the user ranging from cold to warm, and the parameter relating to a given level of thermal comfort is 0, this value representing optimal thermal comfort.
[0101] Training can be carried out using the following steps. In particular, a prediction step by the learning agent through reinforcement of a control action of the S402 thermal management component(s), based on the environmental state and the parameter relating to the given thermal comfort to be achieved. Initially, the environmental state is represented by the set of training data.
[0102] Depending on the example considered, the predicted control action may include increasing the power supplied to the heating, ventilation, and air conditioning (HVAC) system 25 and / or increasing the heating level of the heated seat 30 and / or activating the radiant panel(s) 35. Indeed, in this example, since the user's initial thermal comfort index is -2, the user perceives a sensation of cold. Therefore, the predicted control action will consist of increasing or activating at least one thermal management component.
[0103] The training may also consist of training the artificial intelligence model using the reinforcement learning agent based on a reward provided by the environment module 45, the reward being determined according to the state of the environment and the given parameter relating to thermal comfort. In particular, the prediction step S402 may also be a function of the reward provided by the environment module.
[0104] The prediction step can be followed by an execution step of the control action controlling the S403 thermal management component(s).
[0105] The execution of the control action results in a change of the state of the environment to a new state, which may be closer to or further from the desired target state corresponding to the given thermal comfort parameter.
[0106] In the embodiment in which the thermal management components are simulated, the execution of the action is simulated and the new state of the environment is simulated. The new state of the environment will then be observed by the environment module 45.
[0107] Determining the new state of the environment may include determining the user's new state of thermal comfort, notably through the determination of the user's physiological comfort index. Physiological comfort expresses an overall thermal sensation, that is, throughout the entire body, or a localized thermal sensation, that is, in a specific area of the user's body, and depends in particular on the user's body heat, whether produced or absorbed in the area concerned, notably due to metabolic activity. It may also be based on the user's level of clothing.
[0108] The determination of the user's thermal comfort state can be carried out by the device for determining a physiological comfort of the user 50 of the environment module 45.
[0109] In one particular embodiment, the user's physiological comfort index is determined from measurements of thermal or physiological quantities of different parts of the user's body, as illustrated in Figure 5. The user's body is segmented to determine a local physiological comfort index for each body segment or for one or more body segments. Examples of such segments include the head (510), chest (515), legs (520), feet (525), and arms (530). Other, more or less precise, segmentations may be used. The user's physiological comfort index is then obtained from these local physiological comfort indices, for example, by averaging the determined index values.
[0110] The new environmental state may also include data related to the thermal environment inside the vehicle (e.g., the temperature inside the vehicle) and / or data related to the thermal environment outside the vehicle (e.g., the temperature outside the vehicle). To this end, the new environmental state may also be observed and captured by the device for obtaining at least one data point related to the vehicle's thermal environment 55 of the environmental module 45 illustrated in Figure 1.
[0111] In addition, the new state of the environment may include a state of energy consumption of the vehicle's thermal management components.
[0112] Following the execution of the predicted control action, the training may include a reception step S404 by the learning agent through reinforcement of the observed environment state and a reward.
[0113] The reward indicates to the reinforcement learning agent whether the executed control action has allowed the vehicle's thermal environment to approach, or even reach, the target defined by the given thermal comfort parameter, or to move away from this target. The reward can be determined, notably by the environment module 45, based on the observed state of the environment and the parameter relating to a given thermal comfort.
[0114] Thus, the reward provided by the environment module may consist of:
[0115] (i) reward the learning agent by reinforcement when the state of the environment observed in the environment module approaches the given parameter relating to thermal comfort to be achieved; and / or
[0116] (ii) penalize the reinforcement learning agent when the state of the environment observed in the environment module deviates from the parameter relating to the given thermal comfort to be achieved.
[0117] In one particular embodiment, the awarding of a reward can be done according to the method described below when the objective is to achieve the parameter relating to a given thermal comfort level and to optimize the energy consumed by the thermal management components. Specifically, the reward provided by the environment module also consists of:
[0118] (i) reward the learning agent by further reinforcement when the state of the environment observed in the environment module shows a decrease in energy consumption;
[0119] (ii) penalize the reinforcement learning agent when the state of the environment observed in the environment module shows an increase in energy consumption.
[0120] To achieve this, the method may involve weighting the new observed environmental state with respect to two objectives: comfort and energy consumption. For example:
[0121] - If the user's thermal comfort state is outside an acceptable range [-D, D] represented by the parameter relating to a given thermal comfort level, the penalty related to energy consumption is minimized by using a low weighting, so that the reinforcement learning agent is encouraged to reach the desired comfort state (i.e., a value of 0 for the user's thermal comfort index) as soon as possible. - If the user's thermal comfort state is within the acceptable range, the reinforcement learning agent is encouraged to stabilize within the comfort zone and focus on energy consumption, which is assigned a higher weighting.
[0122] The two states above can be expressed using the following expression:
[0123] In this expression, Pt is the energy consumption at time t, TCt is the value of the thermal comfort state at time t, D is a parameter delimiting the acceptable comfort zone, and alpha ("a") is a trade-off parameter that defines the priorities between energy consumption and comfort. A high value of a prioritizes reducing energy consumption over improving comfort, while a low value of a prioritizes improving comfort over reducing energy consumption. The parameter D can also be modified. For example, increasing the parameter D widens the range of thermal comfort values TC over which reducing energy consumption will take priority.
[0124] Steps S402-S404 can be repeated until the reinforcement learning agent determines that the observed environmental state has reached the parameter for a given thermal comfort level. Specifically, the reinforcement learning agent can determine that the parameter for a given thermal comfort level has been reached when the reward converges to a maximum value.
[0125] If the reward does not converge to a maximum value, then the reinforcement learning agent will predict a new control action based on the new state of the environment and the reward at step S402, the predicted action will be executed S403 and the learning agent will receive the new state of the environment and the new reward S404.
[0126] In some embodiments, steps S402-S404 are repeated until the environmental state reaches the parameter for a given thermal comfort level and the energy consumption of the thermal management components reaches an optimal energy consumption state. The reinforcement learning agent can determine that the optimal energy consumption state has been reached when the reward converges to a maximum value.
[0127] 18
[0128] REPLACEMENT SHEET (RULE 26) According to a particular embodiment, the training steps are repeated until the user's thermal comfort state falls within a range of values for the given thermal comfort parameter, and then the training steps are repeated until the environment's energy consumption state reaches the optimal energy consumption state. Thus, the initial objective is to achieve user thermal comfort, followed by an optimization of energy consumption.
[0129] The execution of the training process to completion can be called a "training episode." In some embodiments, several training episodes can be performed to train the artificial intelligence model. Each training episode can have a set of training data defining a particular scenario. For example, the artificial intelligence model can be trained to cope with a range of different outdoor temperatures.
[0130] The training episode(s) allow the artificial intelligence model to learn a strategy that enables it to achieve the given thermal comfort parameter, the strategy comprising a set of control actions, an episode being defined by a set of training data.
[0131] According to one embodiment, the deep learning agent is installed in the thermal management system 20 illustrated in Figure 1, specifically in a control device 60, which will control said at least one thermal management component of the vehicle. To this end, the control device 60 may include memory to store the process for training an artificial intelligence model. The control device 60 is connected to the thermal management components 25, 30, and 35.
[0132] The control device 60 may include at least one computer or processor and at least one memory in which a computer program is stored. The computer program is configured to implement the method for training an artificial intelligence model as described above. The computer program includes instructions that can be executed by the computer or processor, which cause the control device 60 to perform the steps of the method for training an artificial intelligence model.
[0133] The computer program can also be stored on a computer-readable medium. This computer-readable medium may include memory for storing instructions. The memory can comprise any suitable memory for storing data and executable instructions, such as read-only memory, rewritable flash memory, and a hard drive. Figure 6 illustrates a system capable of implementing the method to train an artificial intelligence model for use in a vehicle to control several thermal management components of a vehicle in order to optimize the thermal comfort of a vehicle user and optimize the energy consumed by the thermal management components, notably through the use of simulated thermal management components.
[0134] The system includes, in particular, an environment module 610, corresponding to environment module 45 in Figure 1. Environment module 610 is capable of simulating the thermal environment of a vehicle or a vehicle part such as the passenger compartment. For this purpose, the environment module includes, for example, a model of a heating, ventilation, and air conditioning (HVAC) system, a model of a radiant panel, and / or a model of a heated seat. In the illustrated embodiment, the thermal management components are simulated. However, in another embodiment, the thermal management components are real components. The system further includes a controller 620 capable of implementing a reinforcement learning agent.
[0135] This system is based on the Markov decision process. Thus, at each step, the reinforcement learning agent in the controller 620 collects an environmental state including data related to the thermal environment and data related to the energy consumption of the observed thermal management component(s) 630 in the environment module 610. The reinforcement learning agent then predicts a control action 640 for the control of the thermal management component(s).
[0136] The action consists of setting data controlling the thermal management components. For example, the setting of a heating, ventilation and air conditioning (HVAC) system is achieved by supplying a power output, for example, between -6000 and 6000 W; the setting of a heated seat is achieved by applying a heating level, the level being 0, 1, 2, 3 or 4; and the setting of one or more radiant panels consists of activating or deactivating the panel(s).
[0137] Control action 640 is then executed. The environment module 610 then transmits a reward 650 to the controller's reinforcement learning agent 620. This reward reflects the evolution of the thermal environment determined by the environment module with respect to the parameter relating to a given thermal comfort level; that is, a penalty for energy consumption and user thermal discomfort, or a positive reward for reduced energy consumption and improved user thermal comfort, as previously described. Examples of results from learning an artificial intelligence model using a reinforcement learning agent are now described. The parameters used in these experiments are as follows.First, the following algorithms were used to train the artificial intelligence model using the reinforcement learning agent: a soft actor-critical algorithm, a DQN (Deep Q-Network) algorithm, and a proximal policy optimization (PPO) algorithm.
[0138] The training duration comprises 2 million time periods. Figures 7A, 7B, 7C, and 7D show the experimental results confirming that the reinforcement learning agent is capable of achieving episodic reward convergence and thermal comfort, minimizing discomfort duration, and reducing energy consumption. Specifically, Figure 7A illustrates reward convergence during learning, Figure 7B illustrates user thermal comfort during learning, Figure 7C illustrates user discomfort duration, and Figure 7D illustrates energy consumption during learning.
[0139] Figures 8A and 8B show respectively the results relating to the thermal comfort index and energy consumption for an outside temperature range of -20° to 10°.
[0140] As shown in Figure 8A, the reinforcement learning agent is able to reach thermal comfort in a relatively short time and then stabilize within the comfort zone. Indeed, the reinforcement learning agent is able to reach thermal comfort smoothly, without overshooting. Furthermore, the energy consumption curve presented in Figure 8B confirms that the reinforcement learning agent is able to accelerate the process of reaching the comfort zone and subsequently reduce energy consumption within that zone.
[0141] Figure 9 shows the results for the time required to reach a parameter related to a given thermal comfort level as a function of the outside temperature, using three different combinations of thermal management components. Figure 9 shows that using a heating, ventilation, and air conditioning (HVAC) system alone does not achieve the parameter related to a given thermal comfort level. However, the combined use of an HVAC system, heated seats, and radiant panels improves thermal comfort control by reducing the time required to reach the parameter related to a given thermal comfort level.
[0142] Figure 10 shows the energy consumption curves over time for three different combinations of thermal management components. Figure 10 demonstrates that controlling multiple thermal management components allows for more efficient energy use. The reinforcement learning agent optimally combines the heating, ventilation, and air conditioning system, the heated seat, and the radiant panels. It leverages the additional heat sources to accelerate convergence to comfort and improve energy efficiency.
[0143] Following the training process, the trained artificial intelligence model will be tested. Specifically, the testing phase will consist of running the training process with test data to evaluate the AI model. The test data is similar to the training data; in fact, the test data can be a subset of the training data that was not used during the learning phase.
[0144] Following the testing phase, a trained artificial intelligence model is suitable for use in a vehicle to control at least one thermal management component in order to optimize the thermal comfort of a vehicle user. Furthermore, the trained artificial intelligence model can also optimize the energy consumed by thermal management components when it is capable of controlling multiple thermal management components.
[0145] Thus, the control of at least one thermal management component of a vehicle by the use of a trained artificial intelligence model, in particular trained according to the training process described previously, will now be described.
[0146] As illustrated in Figure 1, a vehicle may include one or more thermal management components such as a heating, ventilation, and air conditioning (HVAC) system, heated seats, and radiant panels. In some examples, only one thermal management component may be controlled. In some embodiments, only the user's thermal comfort is considered. In other embodiments, several thermal management components are controlled. In this case, the thermal management components may be controlled in such a way as to optimize their energy consumption while simultaneously optimizing user thermal comfort.
[0147] The trained artificial intelligence model is installed in a vehicle to control at least one thermal management component of the vehicle in order to optimize the thermal comfort of a vehicle user. Specifically, the artificial intelligence model is installed in a control device 60, illustrated in Figure 1, which will control said at least one thermal management component of the vehicle. To this end, the control device 60 may include memory to store the trained artificial intelligence model.
[0148] The control device 60 is connected to the thermal management components 25, 30 and 35 and uses the artificial intelligence model to control the thermal management components.
[0149] The control device 60 may include at least one computer or processor and at least one memory in which a computer program and the trained artificial intelligence model are stored. The computer program is configured to implement a method for controlling at least one thermal management component of a vehicle. The computer program includes instructions that can be executed by the computer or processor, which cause the control device 60 to perform the steps of the method for controlling at least one thermal management component of a vehicle.
[0150] The computer program can also be stored on a computer-readable medium. This computer-readable medium may include memory for storing instructions. The memory can include any type of memory suitable for storing data and executable instructions, such as read-only memory, rewritable flash memory, and a hard drive.
[0151] The control device 60 is configured to control at least one thermal management component 25, 30, and 35 to achieve a parameter related to a given thermal comfort level. The control device 60 can also be configured to control at least two thermal management components 25, 30, and 35 to optimize the energy consumed by these components. Controlling the thermal management components 25, 30, and 35 may involve modifying the power setpoints of the thermal management components. For example, the power output of the HVAC unit 25 may be a continuous value within a predetermined range (e.g., 0 to 6000 W). The heat level of the seat heater may be a discrete value (e.g., 0, 1, 2, 3, or 4). Radiant panels may have a binary setpoint (e.g., YES or NO).
[0152] The parameter relating to a given thermal comfort can be a particular temperature in an area of the vehicle or a state of thermal comfort of a user in the vehicle defined for example by a thermal comfort index as described above.
[0153] In this example, an environmental module 45 is also provided in the vehicle as illustrated in Figure 1. The environmental module 45 includes a processor configured to determine the state of the environment, and in particular one or more data points related to the vehicle's thermal environment. The environmental module 45 is connected to the control device 60 and can provide it with the data point(s) related to the vehicle's thermal environment. In other examples, the environmental module may be implemented within the control device.
[0154] The environmental state may include a user thermal comfort state, which may also be expressed using the thermal comfort index described above. Alternatively, or in addition, the environmental state may include data related to the thermal environment inside the vehicle (e.g., the temperature inside the vehicle) and / or data related to the thermal environment outside the vehicle (e.g., the temperature outside the vehicle). The environmental state may also include an energy consumption state representative of the energy consumption of thermal management components 25, 30, and 35.
[0155] For this purpose, the environment module 45 can receive data from the device for determining the physiological comfort of the user 50 and / or from the device for obtaining at least one piece of data related to the thermal environment of the vehicle 55.
[0156] The device for determining the user's physiological comfort 50 may include measuring devices for determining the user's physiological comfort index. Furthermore, the device for obtaining at least one data point related to the vehicle's thermal environment 55 may include a measuring device comprising one or more sensors such as a sunlight sensor, a temperature sensor, including a temperature sensor at an air outlet of the system, a sensor for the temperature inside the passenger compartment, or a humidity sensor.
[0157] An embodiment of a computer-implemented process for controlling at least one thermal management component 25, 30, 35 of a vehicle in order to optimize the thermal comfort of a vehicle user will now be described with reference to Figure 11. The process is implemented for example in the control device 60.
[0158] The method comprises obtaining one or more data points related to the thermal environment of the user and / or the vehicle S1101. These data points are obtained, in particular, by the environment model 45 as previously described in relation to the training method. For example, the control device 60 can obtain a thermal comfort index for the user of -2 from the environment module. The method continues with a step of inputting said at least one data point related to the thermal environment of the user and / or the vehicle into an artificial intelligence model installed in the vehicle S1102, specifically in the control device 60. The artificial intelligence model has been trained to control said at least one thermal management component to achieve a parameter related to a given thermal comfort level.
[0159] The process continues with a step consisting of bringing the trained artificial intelligence model to control said at least one thermal management component 25, 30, 35 to achieve the given thermal comfort parameter S1103.
[0160] To do this, the trained artificial intelligence model can determine the actions to be executed for one or more of the thermal management components 25, 30, 35 based on the parameter relating to a given thermal comfort and said at least one data related to the thermal environment of the user and / or the vehicle in order to optimize the thermal comfort of a vehicle user.
[0161] For example, the trained artificial intelligence model can determine that the power of the heating, ventilation, and air conditioning unit 25 needs to be increased (e.g., by 1000 W). In some examples, the artificial intelligence model can also determine control actions to optimize the energy consumed by the thermal management components 25, 30, 35.
[0162] The control device 60 then commands the thermal management components 25, 30, 35, based on actions determined by the trained artificial intelligence model. For example, the control device 60 can command the heating, ventilation, and air conditioning unit 25 to increase its power by 1000 W.
[0163] The trained artificial intelligence model can also determine a sequence of actions to be performed on the thermal management components, with the actions being sequenced over time.
[0164] The present invention also relates to a computer system comprising means for implementing the method for training an artificial intelligence model for use in a vehicle and / or the method for controlling at least one thermal management component of a vehicle.
[0165] The present invention also relates to a computer program comprising instructions which, when executed, cause a device or computer system to perform a method for training an artificial intelligence model for use in a vehicle and / or the method for controlling at least one thermal management component of a vehicle.
Claims
Claims 1. A computer-implemented method for training an artificial intelligence model for use in a vehicle to control at least one thermal management component (25, 30, 35) of a vehicle in order to optimize the thermal comfort of a vehicle user, the method comprising a step of introducing training data into an artificial intelligence model to train said artificial intelligence model (S401) by means of a reinforcement learning agent, in order to control said at least one thermal management component to achieve a parameter relating to a given thermal comfort, the training data comprising training data sets, each training data set comprising at least one parameter related to the thermal environment of the user and / or the vehicle.
2. A method according to claim 1, characterized in that the method is implemented to control at least two thermal management components of the vehicle and to further optimize the energy consumed by the thermal management components.
3. A method according to any one of the preceding claims, characterized in that the reinforcement learning agent is coupled to an environment module (45) capable of capturing the state of the environment, the training consisting of training the artificial intelligence model by means of the reinforcement learning agent on the basis of a state of the environment observed in the environment module.
4. Method according to the preceding claim, characterized in that the training comprises the following steps: a prediction step by the learning agent by reinforcement of a control action of the thermal management component(s) (S402), according to the state of the observed environment and the parameter relating to the given thermal comfort to be achieved, and a step of execution of the control action controlling the thermal management component(s) (S403).
5. Method according to claim 3 or claim 4, characterized in that the training further consists of training the artificial intelligence model by means of the reinforcement learning agent on the basis of a reward provided by the environment module, the reward being determined according to the state of the observed environment and the parameter relating to the given thermal comfort.
6. A method according to claim 4 and claim 5, characterized in that the prediction step is further a function of the reward provided by the environment module.
7. Method according to claim 5 or claim 6, characterized in that the training further comprises a reception step by the learning agent by reinforcement of the state of the observed environment and of the reward (S404) following the execution of the predicted control action.
8. A method according to any one of claims 5 to 7, characterized in that the reward provided by the environment module consists of: (i) reward the learning agent by reinforcement when the state of the environment observed in the environment module approaches the given thermal comfort parameter to be achieved; (ii) penalize the reinforcement learning agent when the state of the environment observed in the environment module deviates from the parameter relating to the given thermal comfort to be achieved.
9. A method according to claim 2 and the preceding claim, characterized in that the reward provided by the environment module further consists of: (i) reward the learning agent by further reinforcement when the state of the environment observed in the environment module shows a decrease in energy consumption; (ii) penalize the reinforcement learning agent when the state of the environment observed in the environment module shows an increase in energy consumption.
10. A method according to any one of the preceding claims 5 to 9, characterized in that the observed state of the environment has reached the given thermal comfort parameter when the reward converges to a maximum value.
11. Method according to claim 4, characterized in that the training of the artificial intelligence model by means of the reinforcement learning agent comprises one or more training episodes to learn a strategy that enables the achievement of the given thermal comfort parameter, the strategy comprising a set of control actions, an episode being defined by a set of training data.
12. A method according to any one of the preceding claims 3 to 11, characterized in that the training steps are repeated until the observed state of the environment has reached the given thermal comfort parameter.
13. A method according to any one of the preceding claims, characterized in that the parameter relating to the given thermal comfort is defined by a range of values.
14. A method according to any one of the preceding claims, characterized in that the observed state of the environment includes a state of thermal comfort for the user.
15. Method according to the preceding claim, characterized in that the user's thermal comfort state is within a range of values, the central value of the range corresponding to the value of the parameter relating to the given thermal comfort.
16. A method according to any one of the preceding claims, characterized in that the observed state of the environment further comprises a state of energy consumption.
17. A method according to the preceding claim, characterized in that the training steps are further repeated until the energy consumption state of the observed environment has reached an optimal energy consumption state.
18. A method according to the preceding claim and according to any one of claims 5 to 10, characterized in that the observed environmental state has reached a state of optimal energy consumption when the reward converges to a maximum value.
19. Method according to claims 10, 13 and 17, characterized in that the training steps are repeated until the user's thermal comfort state is within a range of values of the given thermal comfort parameter and then the training steps are repeated until the energy consumption state of the environment has reached the optimal energy consumption state.
20. A method according to any one of the preceding claims, characterized in that said at least one parameter related to the user's thermal environment includes at least one parameter related to a state of thermal comfort of the user.
21. A method according to any one of the preceding claims, characterized in that said at least one parameter related to the thermal environment of the vehicle comprises at least one parameter related to the internal thermal environment of the vehicle and / or at least one parameter related to the external thermal environment of the vehicle.
22. A method according to any one of the preceding claims, characterized in that said at least one thermal management component comprises at least one heating, ventilation and air conditioning device (25).
23. Method according to any one of the preceding claims, characterized in that said at least one thermal management component further comprises at least one radiant panel (35) and / or at least one heated seat (30).
24. A method according to any one of claims 1 to 23, characterized in that at least one thermal management component is a simulated component.
25. A method according to any one of the preceding claims, characterized in that the training of the artificial intelligence model by means of the reinforcement learning agent is carried out using a soft actor-critic algorithm, or a DQN-Deep Q-Network algorithm, or a proximal policy optimization (PPO) algorithm.
26. A computer-implemented method for controlling at least one thermal management component (25, 30, 35) of a vehicle in order to optimize the thermal comfort of a vehicle user, the method comprising: a step of obtaining at least one data related to the thermal environment of the user and / or the vehicle (S1101), a step of inputting said at least one data related to the thermal environment of the user and / or the vehicle into an artificial intelligence model (S1102), the artificial intelligence model having been trained to control said at least one thermal management component to achieve a parameter relating to a given thermal comfort, making the trained artificial intelligence model control said at least one thermal management component to achieve the parameter relating to the given thermal comfort (S1103).
27. Method according to the preceding claim, characterized in that the method is further implemented to control at least two thermal management components of the vehicle and to optimize the energy consumed by the thermal management components.
28. A method according to the preceding claim, characterized in that the method further comprises a step of bringing the trained artificial intelligence model to control said thermal management components to achieve a state of optimal energy consumption.
29. A method according to any one of claims 26 to 28, characterized in that said at least one data related to the thermal environment of the user and / or the vehicle includes at least one data related to a state of thermal comfort of the user and / or at least one data related to the environment of the vehicle.
30. Method according to the preceding claim, characterized in that said at least one data related to the thermal environment of the vehicle includes data related to the internal thermal environment of the vehicle and / or data related to the external thermal environment of the vehicle.
31. A method according to any one of claims 26 to 30, characterized in that said at least one thermal management component comprises at least one heating, ventilation and air conditioning device (25).
32. A method according to any one of claims 26 to 31, characterized in that said at least one thermal management component further comprises at least one radiant panel (35) and / or at least one heated seat (30).
33. A computer-implemented method according to any one of claims 26 to 32, characterized in that the artificial intelligence model was trained according to the method according to any one of claims 1 to 25.
34. Computer system comprising means for implementing the method according to any one of claims 1 to 25 or the method according to any one of claims 26 to 32.
35. A computer program containing instructions which, when executed, cause a device or computer system to perform a process according to any one of claims 1 to 25 or any one of claims 26 to 32.
36. Artificial intelligence model for controlling at least one thermal management component of a vehicle in order to optimize the thermal comfort of a vehicle user according to the method according to any one of claims 1 to 25 or any one of claims 26 to 32.
37. System for a vehicle comprising at least one thermal management component and a control device (60), the control device implementing the method according to any one of claims 26 to 32.
38. System according to claim 37, characterized in that said at least one thermal management component comprises at least one heating, ventilation and air conditioning (HVAC) device.
39. System according to claim 37 or claim 38, characterized in that said at least one thermal management component further comprises at least one radiant panel and / or at least one heated seat.