METHOD FOR SIZING VEHICLE COMPONENTS TO OPTIMIZE VEHICLE ENERGY CONSUMPTION
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- VITESCO TECHNOLOGIES GMBH
- Filing Date
- 2023-02-20
- Publication Date
- 2026-04-24
AI Technical Summary
Existing optimization methods for vehicle energy consumption, such as dynamic programming and reinforcement learning, require significant simplifications that limit their effectiveness, leading to suboptimal or unfeasible calculations due to high computational demands, and fail to identify the most optimal control strategies for vehicle elements.
A method utilizing a reinforcement learning agent initialized by transfer learning and behavior cloning, combined with a dynamic programming approach, to efficiently determine a control strategy close to the theoretical optimal control for vehicle elements, allowing rapid and precise evaluation of energy consumption across various dimensions.
The method significantly reduces calculation time and improves the accuracy of identifying optimal vehicle element dimensions and control strategies, ensuring thorough evaluation of each dimension's potential and understanding parameter impacts on energy consumption.
Abstract
Description
Description Title of the invention: DIMENSIONING METHOD OF ELEMENTS OF A VEHICLE TO OPTIMIZE THE VEHICLE ENERGY CONSUMPTION
[0001] — The present invention relates to a computer-implemented method of dimen- sioning of at least one vehicle element in order to optimize consumption in energy of this vehicle. The invention aims in particular to select a dimen- optimal operation, as well as an associated vehicle element control, enabling the vehicle's energy consumption to be minimized. By "consumption in vehicle energy" means within the meaning of the present invention a consumption in fuel, in hydrogen, in electrical energy, in electric current delivered by the branches of an electronic charger of an electric battery of the vehicle, in number of restarts of a generator set in a series hybrid vehicle (also called "range extender" in English), or even a combined consumption of several of these elements. The process is particularly suitable for electric vehicles or hybrids. More specifically, and in a non-limiting manner, the method is adapted in particular to hybrid truck type vehicles equipped with a fuel cell and of an electric machine connected to an electric battery. The invention further relates to a method of controlling at least one such vehicle element, the method comprising a sizing sub-process as described above; as well as a computer program product for implementing the dimensioning process tion.
[0002] In a motor vehicle, it is known to optimize energy consumption of the powertrain on a given or planned route. Such optimization can be carried out on fuel, on electrical energy, on hydrogen consumption (in the case of a vehicle equipped with a fuel cell) or on two or three of these criteria at a time.
[0003] It may also be useful to have to size certain elements of the vehicle in order to further optimize this energy consumption of the vehicle, such elements being able for example to consist of a fuel tank or hydrogen, an electric battery and / or super-capacitors capable of providing electric power, a heat engine powered by the fuel tank or a fuel cell powered by the hydrogen tank, or even a machine electric powered by electrical energy supplied by the battery and / or super- capacities. By "sizing vehicle elements" we mean in the following description the act of selecting a combination of parameter values from a set of possible parameter value combinations, the parameters being relative to the vehicle elements, each parameter being capable of having several distinct values, and each combination of parameter values being associated with a theoretical optimal control of the vehicle elements. Each dimensioning thus corresponds to a distinct parameter value combination. The theoretical optimal control is the control which makes it possible, among all possible controls, to optimize the energy consumption of the vehicle. The parameters relating to the vehicle elements are, for example, the number of cells in the electric battery, the size of the electric battery, the number of stack layers in the fuel cell, the number of cells for each stack layer in the fuel cell, the threshold dead zone of the fuel cell, etc.The associated optimal control then consists, for example, of determining, on a given or planned route, the distribution of electrical power between the fuel cell and the electric battery at any time. Usually, for a specific dimensioning of the vehicle elements and in order to determine the theoretical optimal control (mentioned above) associated with this dimensioning, it is known to use optimization methods of the dynamic programming algorithm type, or the so-called Pontryagin Maximum Principle (PMP) method. However, the main disadvantage of these optimization methods is that, to determine the best dimensioning and its associated optimal control, the definition of the problem (initially of algorithmic complexity) must be simplified in order to use these methods appropriately, because without this the problem becomes insoluble (in other words impossible to solve exhaustively unless a considerable computation time is required, resulting in extremely high computing power and memory size, and in any case incompatible with the needs of the sector).With such a simplification constraint, it is only possible to calculate optimal controls, only for a few specific dimensions of the vehicle elements. Under these conditions, some dimensions are favored to the detriment of other dimensions, and only these first dimensions are then evaluated among all possible dimensions for the vehicle elements, in order to find the best possible dimensioning. However, the most optimal control strategy can change depending on the dimensioning (combination of parameter values) selected for the vehicle elements. This therefore prevents the most optimal control strategy (associated with a dimensioning) to be identified with certainty in a reasonable time in order to optimize the vehicle's energy consumption.It is also possible to use a reinforcement learning algorithm, however such an algorithm does not guarantee the optimality of the identified solutions. Other known methods involve using a basic control strategy to evaluate a large number of designs for vehicle elements. However, such a control strategy is far (in terms of solution optimality) from the optimal control strategies for most designs. In order to optimize the vehicle's energy consumption, it is then possible to determine the best design for the selected basic control strategies, but it is likely that many designs will underperform due to an inefficient and unsuitable control strategy. Most designs are thus not evaluated to their true potential, and the main disadvantage of this strategy is that it misses out on better designs and more optimal solutions to the problem. There is therefore a need to have a method for sizing at least one vehicle element to optimize the vehicle's energy consumption, which allows for the efficient and rapid calculation of an optimal (or near-optimal) control for any new sizing (even if the latter involves the modification of a single element). To this end and according to a first aspect, the invention relates to a computer-implemented method for dimensioning at least one vehicle element to optimize the energy consumption of said vehicle, the dimensioning consisting of a selection of a combination of parameter values from a set of possible combinations of parameter values, the parameters relating to said at least one vehicle element, each parameter being capable of having several distinct values, each combination of parameter values being associated with a theoretical optimal control of said at least one vehicle element making it possible to optimize the energy consumption of the vehicle, the vehicle comprising elements consisting of: a fuel or hydrogen tank, an electric battery and / or supercapacitors capable of supplying electrical energy,a heat engine powered by the fuel tank or a fuel cell powered by the hydrogen tank, and at least one electric machine powered by electrical energy supplied by the battery and / or the supercapacitors, the method comprising an initial step of randomly selecting a first combination of parameter values or a first subset of combinations of parameter values, then a step of determining the theoretical optimal control associated with the first combination of parameter values or with each combination of the first subset of combinations of parameter values, the energy consumption of the vehicle being calculated for each theoretical optimal control thus determined, associated with said first combination of parameter values or with a combination of said first subset of combinations of parameter values, , the method further comprising the steps of: - training a reinforcement learning agent stored in memory means of said computer, on the basis of said first combination of parameter values or said first subset of combinations of parameter values associated with their respective theoretical optimal command(s), said reinforcement learning agent thus trained providing a set of knowledge; - selection of a second combination of parameter values or a second subset of combinations of parameter values; - determination, for said second combination of parameter values or for each combination of said second subset of combinations of parameter values, of a command close to the theoretical optimal command associated with said combination, via a transfer, by a transfer learning algorithm stored in the memory means of said computer, of said set of knowledge on said second combination of parameter values or on each combination of said second subset of combinations of parameter values, allowing re-training of the reinforcement learning agent; - calculation of the vehicle's energy consumption for each command thus determined during the previous step; - looping the previous steps of selection, determination and calculation, until a predetermined criterion is reached; and - selection of a combination of parameter values for which said pre-determined criterion is met. By "control close to the theoretical optimal control" we mean a combination that approximates the performance obtained by the theoretical optimal control. This control is good enough (in view of the criterion to be optimized) to be able to compare the dimensions impartially. By using a reinforcement learning agent, initialized by a transfer learning algorithm, the sizing method according to the invention is faster in terms of the calculation of the command as well as in terms of the cost of the calculation. The calculation time required by the processing means of the computer to identify a sizing solution close to optimality is thus considerably reduced. Indeed, the reinforcement learning agent, after having been trained with one or more theoretical optimal commands, provides a set of knowledge. For each new sizing explored, the transfer learning algorithm then exploits the set of knowledge provided by the reinforcement learning agent (typically by copying the weights of the neural network when the reinforcement learning agent is a neural network). neurons) to quickly determine a command close to the theoretical optimal command. This allows considerable time savings and better precision for the determination of each command associated with a given dimension. Furthermore, the sizing method according to the invention makes it possible to calculate a control close to optimality for each dimension explored, in a reasonable time, and thus makes it possible to actually evaluate (i.e. at its true potential) the performance of each dimension explored, in order to determine the best sizing (associated with its control close to the theoretical optimal control) in terms of energy consumption of the vehicle. Each dimension is thus correctly evaluated with a control strategy adapted and specific to this dimension, and not with a sub-optimal control strategy as is the case in certain methods of the prior art.Finally, the dimensioning method according to the invention makes it possible (thanks to the downstream use of artificial intelligence tools) to better understand which parameters influence the energy consumption of the vehicle (and what their precise degree of influence is), in order to better understand and better grasp the energy system constituted by the vehicle. Preferably, said reinforcement learning agent is a neural network. According to one embodiment of the invention, the step of training the reinforcement learning agent is initialized via the implementation of a behavior cloning algorithm stored in the memory means of said computer. Such a behavior cloning algorithm allows the reinforcement learning agent to best imitate the theoretical optimal control determined during the previous step (for example by copying the states of this theoretical optimal control). The use of such a behavior cloning algorithm thus makes it possible to initialize the reinforcement learning agent on an optimal solution. Advantageously, the step of selecting a second combination of parameter values or a second subset of combinations of parameter values is carried out via the implementation of an optimizer stored in the memory means of said computer. Such an optimizer makes it possible to implement an intelligent selection of the dimensions during the step of selecting a second combination of parameter values or a second subset of combinations of parameter values. Indeed, each dimension explored being associated with a cost of a cost function (the latter being relative to the energy consumption of the vehicle), the optimizer makes it possible to explore “zones” of dimensions (in the multi-dimensional space of the parameter values) for which the value of the cost function is low. The use of such an optimizer is all the more efficient as the number of parameters is high (and therefore the size of the multi-dimensional space explored for the sizing is large). This intelligent exploration carried out by the optimizer constitutes an additional lever making it possible to further accelerate the determination of a sizing associated with its order (close to optimality), with a view to optimizing the energy consumption of the vehicle. Preferably, the optimizer is selected from the group consisting of: a genetic algorithm, a tree-structured Parzen estimator algorithm, an evolutionary algorithm with covariance matrix adaptation, and a reinforcement learning agent. According to a particular technical characteristic of the invention, the step of determining the theoretical optimal control associated with the first combination of parameter values or with each combination of the first subset of combinations of parameter values is carried out via the implementation of a dynamic programming method or a method known as the Pontryagin Maximum Principle. Advantageously, the method further comprises a step of implementing a dynamic programming method or a method known as the Pontryagin Maximum Principle to determine the theoretical optimal control associated with the current selection of a combination of parameter values, said step being implemented according to a predetermined step of combinations of selected parameter values. This makes it possible to periodically ensure (according to said predetermined step) that the dimensioning method according to the invention is always close to the theoretical optimal control. This ensures control over any possible deviation of the solutions (combinations of parameter values) identified, with respect to the optimality of the control. According to a particular technical characteristic of the invention, said predetermined criterion consists of reaching a threshold value by the calculated energy consumption of the vehicle, or of reaching a predefined number of combinations of selected parameter values. The invention also relates to a method, implemented in a computer embedded in a vehicle, for controlling at least one vehicle element by sending an instruction to said element, said at least one vehicle element being chosen from the group consisting of: a fuel or hydrogen tank, an electric battery and / or super-capacitors capable of supplying electrical energy, a heat engine powered by the fuel tank or a fuel cell powered by the hydrogen tank, and at least one electric machine powered by electrical energy supplied by the battery and / or the super-capacitors, the method comprising a sub-method for dimensioning said at least one a vehicle element as described above, said setpoint corresponding to the command associated with the selected combination of parameter values, for which said predetermined criterion is reached. The invention also relates to a computer program product downloadable from a communication network and / or recorded on a computer-readable medium and / or executable by a processor, remarkable in that it comprises a set of program code instructions which implement the method as described above when executed on a processing unit of a computing device. Embodiments of the present invention will be described below, by way of non-limiting examples, with reference to the appended figures in which: - [Fig.1] schematically illustrates a vehicle according to the invention, the vehicle being equipped with an on-board computer; and - [Fig.2] is a flowchart representing a computer-implemented method of dimensioning elements of the vehicle of [Fig.1], according to the present invention. Referring to [Fig. 2], the present invention relates to a computer-implemented method for dimensioning at least one element of a vehicle 2 to optimize the energy consumption of this vehicle 2. The computer (which is not shown in the figures for reasons of clarity) conventionally comprises processing means and memory means (the latter being able to store a software application or computer program capable of implementing the dimensioning method when executed by the processing means of the computer). The vehicle 2 (visible in [Fig. 1]) is for example (but not limited to) a hybrid vehicle and comprises an on-board computer 4.In addition to the computer 4, such a hybrid vehicle 2 also conventionally comprises a thermal engine M (in the case of a hybrid vehicle 2 of the “thermal-electric” type), often called an “ICE” engine for “Internal Combustion Engine” in English, or a fuel cell P operating for example on hydrogen (in the case of a hybrid vehicle 2 of the “hydrogen-electric” type). Such a hybrid vehicle 2 further comprises at least one electrical machine Mz, often called an “EMA machine” for “Electrical Machine” in English, a fuel tank or a hydrogen tank 10 and an electrical power supply battery 30 (or super-capacitors in a variant not shown). The heat engine M is in particular capable of being powered from the fuel supplied by the fuel tank 10, and the fuel cell P is capable of being powered by the hydrogen tank 10. The heat engine M also comprises a system for exhausting the exhaust gases emitted during the combustion of the air and fuel mixture in the heat engine M. The electric machine Mz is capable of being powered by the electrical energy supplied by the battery 30 and / or by the fuel cell P directly. The term "system" refers to the set of elements mounted in the vehicle 2, capable of storing, consuming or producing electrical energy, fuel or hydrogen. For example, the system comprises all of the devices described previously: the heat engine M or the fuel cell P, the electric machine My, the hydrogen fuel tank 10 and the battery 30. The system is characterized by a set of parameters relating to these elements, each parameter being capable of having several distinct values. The parameters relating to the elements M, P, M£, 10, 30 of the vehicle 2 are for example the number of cells of the electric battery 30, the size of the electric battery 30, the number of stack layers of the fuel cell P, the number of cells for each stack layer of the fuel cell P, the threshold dead zone of the fuel cell P, etc. The sizing of the elements M, P, Mz, 10, 30 of the vehicle 2 then consists of selecting, from these parameters, a combination of parameter values making it possible to optimize the energy consumption of the vehicle 2. Each combination of parameter values (or sizing) is associated with a theoretical optimal control of the elements of the vehicle 2 making it possible to optimize this energy consumption of the vehicle 2.By "theoretical optimal combination" we mean the control which allows, among all the possible controls for optimizing the energy consumption of the vehicle 2, to optimize this energy consumption as best as possible. The computer memory means used to implement the method store a reinforcement learning agent and a transfer learning algorithm (not shown in the figures). The reinforcement learning agent is for example a neural network trained during a step 44 of the method which will be detailed later. Preferably, the computer memory means used to implement the method also store a behavior cloning algorithm and an optimizer (not shown in the figures). The optimizer is preferably chosen from the group consisting of: a genetic algorithm, a tree-structured Parzen estimator algorithm, an evolution algorithm with covariance matrix adaptation, and a reinforcement learning agent. With reference to [Fig. 2], an embodiment of the method for dimensioning the elements M, P, Mr, 10, 30 of the vehicle 2 according to the invention will now be described, implemented by computer. Initially, the software application or computer program stored in the memory means of the computer is executed by the latter's processing means. A list of parameters relating to the elements M, P, M£, 10, 30 of the vehicle 2, as well as all the values that can be taken by these parameters, is provided as input to the computer. The dimensioning process is for example implemented (in other words simulated) on a predetermined driving cycle of the vehicle 2. The method comprises an initial step 40 of randomly selecting a first combination of parameter values or a first subset of combinations of parameter values. The selection 40 may for example be a random, pseudo-random selection, or a selection imposed by an operator on purpose upstream of the method. The method comprises a following step 42 of determining the theoretical optimal control associated with the first combination of parameter values or with each combination of the first subset of combinations of parameter values. At the end of this determination step 42, the energy consumption of the vehicle 2 is calculated for each theoretical optimal control thus determined, which is associated with the first combination of parameter values or with a combination of the first subset of combinations of parameter values. Preferably, this determination step 42 is carried out via the implementation of a dynamic programming method or a method known as the Pontryagin Maximum Principle. The method comprises a subsequent step 44 of training the reinforcement learning agent on the basis of the first combination of parameter values or the first subset of combinations of parameter values associated with their respective theoretical optimal control(s). Preferably, this step 44 of training the reinforcement learning agent is initialized via the implementation of the behavior cloning algorithm. Such a behavior cloning algorithm allows the reinforcement learning agent to best imitate the theoretical optimal control determined during the previous step 42 (for example by copying the states of this theoretical optimal control). The use of such a behavior cloning algorithm thus makes it possible to initialize the reinforcement learning agent on an optimal solution.The training step 44 of the reinforcement learning agent is then carried out by exploiting the knowledge already acquired via the behavior cloning algorithm and by training the reinforcement learning agent via one or more reinforcement learning algorithm(s). The training step 44 of the reinforcement learning agent consists for example of supervised learning training. At the end of this training step 44, the reinforcement learning agent provides a set of knowledge. The method comprises a following step 46 of selecting a second combination of parameter values or a second subset of combinations of parameter values. Preferably, this selection step 46 is carried out via the implementation of the optimizer. To do this, the optimizer explores, for example, dimensioning “zones” (in the multi-dimensional space of parameter values) for which the value of the cost function is low (the latter being relative to the energy consumption of the vehicle 2). The stability of the criterion to be optimized (in this case the energy consumption of the vehicle 2) can be added to the cost function and evaluated by the optimizer during step 46.More precisely, this stability can be evaluated locally by the optimizer, and make it possible to define a sampling horizon for the second combination of parameter values or the second subset of combinations of parameter values selected, thus making it possible to select combinations of parameter values very close to each other in terms of stability of the criterion to be optimized. The method comprises a following step 48 of determining, for the second combination of parameter values or for each combination of the second subset of combinations of parameter values, a control close to the theoretical optimal control associated with this combination.This determination 48 is carried out via a transfer, by the transfer learning algorithm, of the set of knowledge provided by the reinforcement learning agent at the end of the training step 44 (typically by copying the weights of the neural network), on the second combination of parameter values or on each combination of the second subset of combinations of parameter values. It should be noted that, to do this, the transfer learning algorithm assumes that a solution close to optimality for the second combination of parameter values or for each combination of the second subset of combinations of parameter values is close to the theoretical optimal control found for the first combination of parameter values or for each combination of the first subset of combinations of parameter values.Transferring the knowledge set on the second combination of parameter values or on each combination of the second subset of parameter value combinations enables re-training of the reinforcement learning agent (e.g. via one or more reinforcement learning algorithm(s)). The method comprises a following step 49 of calculating the energy consumption of the vehicle 2 for each command thus determined during the previous step 48. The method comprises a following step 50 of evaluating the achievement of a predetermined criterion. Such a criterion consists, for example, of the achievement of a threshold value by the energy consumption of the vehicle 2 calculated during step 49, in the execution time of the process or in reaching a predefined number of combinations of parameter values selected during step 46. Alternatively, the predetermined criterion may consist in the fact that the value of the energy consumption of the vehicle 2, calculated during step 49, stagnates in its optimization. If, during the evaluation step 50, the predetermined criterion is not met, the selection steps 46, determination steps 48 and calculation steps 49 are looped again. The selection steps 46, determination steps 48 and calculation steps 49 are then looped again in the same manner until the predetermined criterion is met. Preferably, as illustrated in [Fig. 2], the method may comprise an intermediate step 51 of implementing a dynamic programming method or a method known as the Pontryagin Maximum Principle to determine the theoretical optimal control associated with the current selection 48 of a combination of parameter values. This step 51 is for example implemented according to a predetermined step of combinations of parameter values selected during step 48, in other words all the NI combinations of parameter values selected, N1 being a strictly positive integer. N1 is for example (but in a non-limiting manner) equal to 10, 20 or 50.Alternatively or in addition, this step 51 can be implemented if the optimizer has selected, during step 46, a current parameter value combination whose command calculated during step 48 is too far from the command calculated for the previously selected parameter value combination. Alternatively or in combination, this step 51 can be implemented if the energy consumption of the vehicle 2 calculated during step 49 for the current parameter value combination is considerably different from that calculated for the previous parameter value combination. This may be the case, for example, if the sizing of the elements M, P, M£, 10, 30 of the vehicle 2 intrinsically has a major impact on the energy consumption of the vehicle 2, or if the learning of the reinforcement learning agent is not yet sufficient. If, during the evaluation step 50, the predetermined criterion is reached, the method comprises a final step 52 of selecting a combination of parameter values for which the criterion is reached. This combination of parameter values selected during the final step 52 then corresponds to the dimensioning chosen for the elements M, P, Mz, 10, 30 of the vehicle 2. The computer 4 is configured to send an instruction to at least one of the elements M, P, Mz, 10, 30 of the vehicle 2, in order to control this element according to a control law. For this purpose, the computer 4 comprises a processor capable of implementing a set of instructions making it possible to carry out this function. The present The invention also relates to a method for controlling the elements M, P, Mç, 10, 30 of the vehicle 2, implemented by the computer 4, and comprising the sub-method for dimensioning the elements M, P, M£, 10, 30 as described previously with reference to [Fig. 2]. The instruction sent by the computer 4 to at least one of the elements M, P, Mç, 10, 30 of the vehicle 2 then corresponds to the command associated with the combination of parameter values selected during the final step 52, for which the predetermined criterion is reached. The sizing method according to the invention is faster in calculating the control associated with the sizing, as well as in the cost of the calculation, while offering a resolution with low optimality error in terms of sizing and associated control. Finally, the sizing method according to the invention makes it possible (thanks to the downstream use of artificial intelligence tools) to better understand which parameters influence the energy consumption of the vehicle (and what is their precise degree of influence), in order to better understand and better grasp the energy system constituted by the vehicle 2. The method for controlling the elements M, P, Mz, 10, 30 of the vehicle 2 according to the invention, comprising the dimensioning sub-method as described previously, allows optimal control of the elements M, P, Mr, 10, 30 of the vehicle 2 in order to minimize the energy consumption of the vehicle 2.
Claims
Claims
1. A computer-implemented method of dimensioning at least an element (M, P, Mz, 10, 30) of vehicle (2) to optimize the energy consumption of said vehicle (2), the sizing consisting of a selection of a combination of parameter values among a set of possible parameter value combinations, the parameters being relative to said at least one element (M, P, M;, 10, 30) of vehicle (2), each parameter being capable of presenting several distinct values, each combination of parameter values being associated with a theoretical optimal control of said at least one element vehicle allowing to optimize the energy consumption of the vehicle (2), the vehicle (2) comprising elements consisting of: a fuel or hydrogen tank (10), an electric battery and / or supercapacitors (30) capable of supplying electrical energy, a thermal engine (M) powered by the fuel tank (10) or a fuel cell (P) powered by the hydrogen tank (10), and less one electric machine (M£) powered by electric energy provided by the battery and / or supercapacitors (30), the method comprising an initial step (40) of randomly selecting a first combination of parameter values or a first subset of combinations of parameter values, then a step (42) of determining- mination of the theoretical optimal order associated with the first combination of parameter values or to each combination of the first subset of parameter value combinations, the vehicle energy consumption (2) being calculated for each theoretical optimal control thus determined, associated with said first combination of parameter values or to a combination said first subset of combinations of parameter values, characterized in that the method further comprises the steps of: - training (44) a stored reinforcement learning agent in memory means of said computer, on the basis of said first combination of parameter values or said first sub- set of combinations of parameter values associated with their respective theoretical optimal command(s), said agent of reinforcement learning thus trained providing a set of knowledge; - selection (46) of a second combination of parameter values or a second subset of combinations of pa- values meters; - determination (48), for said second combination of values of parameters or for each combination of said second subset of combinations of parameter values, of a command close to the theoretical optimal order associated with said combination, via a transfer, by a transfer learning algorithm stored in the memory means of said computer, of said set of knowledge on said second combination of parameter values or on each combination of said second subset of combinations of values of parameters, allowing re-training of the learning agent by reinforcement; - calculation (49) of the vehicle's energy consumption for each order thus determined during the previous step; - looping of the previous selection steps (46), determination (48) and calculation (49), until a predetermined criterion is reached; And - selection (52) of a combination of parameter values for which said predetermined criterion is reached.
2. The method of claim |, wherein said learning agent by reinforcement is a neural network. |Claim 3] A method according to claim 1 or 2, wherein step (44) training of the reinforcement learning agent is ini- tialized via the implementation of a cloning algorithm of com- behavior stored in the memory means of said computer.
4. Method according to any one of claims 1 to 3, wherein step (46) of selecting a second combination of pa- values meters or a second subset of value combinations of parameters is performed via the implementation of an optimizer stored in the memory means of said computer.
5. The method of claim 4, wherein the optimizer is chosen from the group consisting of: a genetic algorithm, an algorithm Tree-structured Parzen estimator, an evolutionary algorithm with covariance matrix adaptation, and a learning agent by reinforcement.
6. A method according to any one of claims 1 to 5, wherein step (42) of determining the theoretical optimal control associated with the first combination of parameter values or each combination of the first subset of combinations of parameter values is done through the implementation of a method of dynamic programming or a method called the Principle of Pontryagin Maximum.
7. A method according to any one of claims 1 to 6, wherein the method further comprises a step (51) of implementing a dynamic programming method or a so-called method of Pontryagin's Maximum Principle for Determining Order theoretical optimum associated with the current selection of a combination of parameter values, said step being implemented according to a step predetermined number of combinations of selected parameter values.
8. A method according to any one of claims 1 to 7, wherein said predetermined criterion consists of the achievement of a threshold value by the vehicle energy consumption (2) calculated, or upon reaching a predefined number of combinations of selected parameter values tioned.
9. Method, implemented in a computer (4) embedded within a vehicle (2), controlling at least one element (M, P, My, 10, 30) of vehicle (2) by issuing an instruction to said element {M, P, M, 10, 30), said at least one element (M, P, My, 10, 30) of vehicle (2) being chosen from the group consisting of: a tank of fuel or hydrogen (10), an electric battery and / or super- capacities (30) capable of supplying electrical energy, a motor thermal (M) powered by the fuel tank (10) or a fuel cell fuel (P) supplied by the hydrogen tank (10), and least one electric machine (Mz) powered by electrical energy provided by the battery and / or the super-capacitors (30), characterized in that that the method comprises a sub-method for dimensioning said at least one at least one element (M, P, M,, 10, 30) of vehicle (2) according to one any of claims 1 to 8, said instruction corresponding to the command associated with the combination of parameter values se- selected, for which said predetermined criterion is met.
10. | Computer program product downloadable from a network of communication and / or recorded on a computer-readable medium and / or executable by a processor, characterized in that it comprises a set of program code instructions that implement the method according to any one of claims 1 to 8 when they are executed on a processing unit of a computing device.