Method, system and computer program product for autonomous calibration of an electric powertrain

DE102022104313B4Active Publication Date: 2025-08-14DR ING H C F PORSCHE AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102022104313
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-23
Publication Date
2025-08-14
Estimated Expiration
2042-02-23

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for the autonomous calibration of an individual electric drive train (10), wherein sensors and / or measuring devices for determining parameters (p i ) of properties (e i ) of the individual electric drive train (10), comprising: - Creating (S10) a training model (TM) for an electric drive train (10) by a learning reinforcement agent (320) by means of simulated observations (b1, b2, ..., b n ), wherein the learning reinforcement agent (320) uses a reinforcement learning algorithm, wherein the simulated observations (b1, b2, .... b n ) in particular the current, the voltage, the torque and / or the speed of an electric motor and / or the state of charge of a battery of the electric drive train (10); - modifying (S20) the training model (TM) of the learning reinforcement agent (320) by means of real observations (b r1 , br2 , ..., b rn ) of a real ideal-typical drive train (10) for creating a simulated model (M) for the real ideal-typical electric drive train (10), wherein the real observations (b r1 , b r2 , .... b rn ) measured values ​​of parameters (p i ) of a property (e i ) of the real ideal-typical drive train (10), which are determined by the sensors or stored in a database (250), and wherein the simulated model (M) represents target states (sm t1 , sm t2 , ..., sm tn ) contains; - Determining (S30) at least one state (s i ) of an individual real electric drive train (10) by a state module (350), wherein a state (s i ) by parameter (p i ) such as data and / or measurements of at least one property (e i ) of the electric drive train (10), - Transmitting (S40) the status (s i ) to the learning reinforcement agent (320); - Determining (S50) calibration results (450) for the individual real electric drive train (10) from the learning reinforcement agent (320) by comparing the state (s i ) with at least one target state (sm ti ) of the simulated model (M).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method, a system and a computer program product for autonomously calibrating an electric drive train.

[0002] The calibration of control devices for electric drive trains using artificial intelligence methods, particularly reinforcement learning, is well known. An electric drive train has one or more electric motors powered by an electrical power supply, such as a battery or a fuel cell. Using power electronics such as an inverter, the output voltage of the electrical power supply is converted into alternating current to supply the electric motor with the required current and voltage according to the respective drive situation.Electric drives are used for a wide variety of functions and tasks, such as traction drives in motor vehicles, industrial trucks, railways, or in industry for assembly machines, or as lifting drives, or in the field of robotics, as well as for a wide variety of household appliances. Compared to other drive types such as hydraulic or pneumatic, an electric drive has the advantage of efficient controllability of the electric motor as an electromechanical energy converter in terms of torque and speed using controllable power electronics. By controlling the power electronics, the motor's performance is adapted to the respective task, for example, for a desired acceleration when driving a motor vehicle. The control of the power electronics, in turn, depends on the motor configuration and other parameters, such as the weight of a motor vehicle.

[0003] However, in the known reinforcement learning methods for calibrating an electric powertrain, a model of a real-life electric powertrain is provided to a learning reinforcement agent, which is not modified by the learning reinforcement agent. The model can be described using physical-mathematical equations, for example, or created on a data-driven basis, for example, using neural networks. Another approach is to create a model based on Markov decision processes. Regardless of the type of underlying model for an electric powertrain, the model is always provided to the learning reinforcement agent. This means that the learning reinforcement agent uses the given model to plan its actions. The learning reinforcement agent therefore does not act completely independently, as the selection of its actions depends on the model used.

[0004] The model is typically created by experts such as engineers and represents an environment that the learning reinforcement agent draws on. However, creating the model, which in the case of a powertrain reflects its dynamic behavior, for example, with regard to the voltage and current used depending on a traffic situation, is complex and difficult, so that the model sometimes does not represent the actual behavior of an electric powertrain and is therefore inaccurate. Furthermore, creating a model for an environment within a learning reinforcement process is associated with a considerable amount of time and therefore expense. This means that the learning results of the learning reinforcement agent also depend on the underlying model and therefore have only limited validity.

[0005] CN 112632860 A discloses a method for identifying model parameters of a power transmission system based on reinforcement learning. The reinforcement learning method for identifying model parameters of the power transmission system avoids local optimizations and exhibits a high convergence rate because it is based on a step-by-step identification process.

[0006] US 2019 / 0378036 A1 discloses a control method for motor vehicles based on reinforcement learning techniques. A reinforcement learning control unit is trained both on the basis of a simulated ground vehicle environment during a simulation mode and on the basis of a motor vehicle environment during an operating mode of a motor vehicle.

[0007] DE 10 2019 215 530 A1 discloses an operating strategy for a hybrid vehicle with an electric motor and an internal combustion engine, which is based on reinforcement learning methods.

[0008] DE 10 2019 208 262 A1 discloses a method for determining a control strategy for a technical system. The control strategy is created and executed based on model parameters of a control model, using reinforcement learning to find the control strategy.

[0009] EP 3 825 263 A1 discloses a method for the computer-implemented configuration of a controlled drive configuration of a logistics system, wherein a control function is determined by means of reinforcement learning.

[0010] DE 10 2020 118 805 A1 discloses a method for autonomously constructing and / or designing at least one component of a device. A state of the component is determined by a state module, wherein a state is defined by parameters such as data and / or measured values ​​of at least one property of the component. The state is transmitted to a reinforcement learning agent that uses a reinforcement learning algorithm. By selecting a calculation function and / or an action based on a policy, at least one parameter is modified, and a modeled value for the property is calculated using the modified parameter. The policy of the reinforcement learning agent is changed based on a reward until the target state is reached.

[0011] The object underlying the invention is to create a method, a system and a computer program product for the autonomous calibration of an electric drive train, which is characterized by high reliability, safety and accuracy and is easy to implement.

[0012] According to the present invention, a method, a system and a computer program product are proposed which enable autonomous calibration of an electric powertrain to thereby provide the basis for reliable and accurate control of the electric powertrain.

[0013] This object is achieved according to the invention with respect to a method by the features of patent claim 1, with respect to a system by the features of patent claim 8, and with respect to a computer program product by the features of patent claim 11. The further claims relate to preferred embodiments of the invention.

[0014] According to a first aspect, the invention provides a method for autonomously calibrating an individual electric drive train, wherein sensors and / or measuring devices are provided for determining parameters of properties of the individual electric drive train. The method comprises the following method steps: - Creating a training model for an electric drive train by a learning reinforcement agent using simulated observations, wherein the learning reinforcement agent uses a reinforcement learning algorithm, and wherein the simulated observations include in particular the current, voltage, torque and / or speed of an electric motor and / or the state of charge of a battery of the electric drive train; - modifying the training model of the learning reinforcement agent using real observations of a real ideal-typical powertrain to create a simulated model for the real ideal-typical electric powertrain, wherein the real observations represent measured values ​​of parameters of a property of the real ideal-typical powertrain determined by sensors or stored in a database, and wherein the simulated model contains target states; - Determining at least one state of an individual real electric drive train by a state module, wherein a state is defined by parameters such as data and / or measured values ​​of at least one property of the electric drive train, - Communicating the state to the learning reinforcement agent; - Determining calibration results for the individual real electric powertrain from the learning reinforcement agent by comparing the state with at least one target state of the simulated model.

[0015] In an advantageous embodiment, it is provided that for the creation of a training model for an electric drive train by a learning reinforcement agent by means of simulated observations, an environment module is provided which comprises at least a state sub-module, a reward sub-module and a strategy sub-module.

[0016] In a further development, it is provided that the state sub-module generates states based on the simulated observations.

[0017] In a further embodiment, determining calibration results comprises the following method steps: - selecting a calculation function and / or an action based on a policy for a state for modifying at least one parameter of the learning reinforcement agent; - Calculating a modeled value for the property using the modified parameter; - Calculating a new state from an environment module based on the modeled value for the property; - Comparing the new state with the target state and assigning a deviation for the comparison result in the state module; - Determining a reward from a reward module for the comparison result; - Adjust the policy of the learning reinforcement agent based on the reward, where upon convergence of the policy, the optimal action for the calculated state is returned, and upon non-convergence of the policy, a further calculation function and / or a further action for a state with a modification of at least one parameter is selected by the learning reinforcement agent until the target state is reached.

[0018] Advantageously, a positive action A+, which increases the value for a parameter, a neutral action A0, in which the value of the parameter remains the same, and a negative action A-, in which the value of the parameter decreases, are provided.

[0019] In one embodiment, the reward module comprises a database or matrix for evaluating the actions.

[0020] In particular, at least one algorithm of the learning reinforcement agent is designed as a Markov decision process, temporal difference learning (TD-learning), Q-learning, SARSA, Monte Carlo simulation or actor-critic.

[0021] According to a second aspect, the invention provides a system for autonomously calibrating an individual electric powertrain. The system comprises an input module, a learning reinforcement module, an output module, and sensors and / or measuring devices for determining parameters of characteristics of the individual electric powertrain. The learning reinforcement module comprises a learning reinforcement agent using a reinforcement learning algorithm, an action module, an environment module, a state module, and a reward module.The learning reinforcement agent is designed to create a training model for an electric drive train using simulated observations, and to modify the training model using real observations of a real ideal-typical drive train to create a simulated model for the real ideal-typical electric drive train, wherein the simulated observations include in particular the current, the voltage, the torque and / or the speed of an electric motor and / or the state of charge of a battery of the electric drive train and the real observations represent measured values ​​of parameters of a property that are determined by sensors or that are stored in a database, and wherein the simulated model contains target states.The state module is configured to determine at least one state of an individual real electric powertrain, wherein a state is defined by parameters such as data and / or measured values ​​of at least one property of the electric powertrain, and to transmit the state to the learning reinforcement agent. The learning reinforcement agent is configured to determine calibration results for the individual real electric powertrain by comparing the state with at least one target state of the simulated model.

[0022] In a further development, it is provided that the environment module comprises at least a state sub-module, a reward sub-module and a strategy sub-module.

[0023] In a further embodiment, it is provided that the state sub-module is designed to generate states based on the simulated observations.

[0024] According to a third aspect, the invention provides a computer program product comprising executable program code configured to carry out the method according to the first aspect when executed.

[0025] The invention is explained in more detail below with reference to embodiments shown in the drawing.

[0026] It shows: Fig. 1 is a block diagram illustrating an embodiment of a system according to the invention; Fig. 2 a flow chart explaining the individual method steps of a method according to the invention; Fig. 3 is a block diagram of a computer program product according to an embodiment of the third aspect of the invention.

[0027] Additional features, aspects and advantages of the invention or embodiments thereof will become apparent from the detailed description taken in conjunction with the claims.

[0028] Fig. 1 shows a system 100 according to the invention for autonomously calibrating an electric drive train 10. An electric drive train 10 has one or more electric motors that are supplied with energy by an electrical power supply, such as a battery or a fuel cell. By means of power electronics such as an inverter, the output voltage of the electrical power supply is converted into alternating voltage in order to supply the electric motor with the required current and voltage according to the respective drive situation. Electric drives are used for a variety of functions and tasks, such as traction drives in motor vehicles, industrial trucks, railways, or in industry for assembly machines, or as lifting drives, or in the field of robotics, as well as for a variety of household appliances.Compared to other drive types such as hydraulic or pneumatic, an electric drive offers the advantage of efficient control of the electric motor as an electromechanical energy converter in terms of torque and speed using controllable power electronics. By controlling the power electronics to the respective task, the motor's performance is adapted, for example, to achieve a desired acceleration when driving a motor vehicle. The control of the power electronics, in turn, depends on the motor configuration and other parameters, such as the weight of the vehicle.

[0029] The system 100 according to the invention is based on reinforcement learning methods and comprises an input module 200, a learning reinforcement module 300, and an output module 400. The learning reinforcement module 300 comprises a learning reinforcement agent (LV agent) 320, an action module 330, an environment module 340, a state module 350, and a reward module 370.

[0030] The input module 200, the learning reinforcement module 300 and the output module 400 may each be provided with a processor and / or a memory unit.

[0031] In the context of the invention, a "processor" can be understood, for example, as a machine or an electronic circuit. A processor can be, in particular, a central processing unit (CPU), a microprocessor, or a microcontroller, for example, an application-specific integrated circuit or a digital signal processor, possibly in combination with a memory unit for storing program instructions, etc. A processor can also be understood as a virtualized processor, a virtual machine, or a soft CPU.It may, for example, also be a programmable processor which is equipped with configuration steps for carrying out the said method according to the invention or is configured with configuration steps such that the programmable processor implements the inventive features of the method, the component, the modules, or other aspects and / or partial aspects of the invention.

[0032] In the context of the invention, a "storage unit" or "storage module" and the like can be understood to mean, for example, a volatile memory in the form of random-access memory (RAM), a permanent memory such as a hard drive or a data storage device, or, for example, a removable storage module. However, the storage module can also be a cloud-based storage solution.

[0033] In the context of the invention, a "module" can be understood, for example, as a processor and / or a memory unit for storing program instructions. For example, the processor is specifically configured to execute the program instructions in such a way that the processor and / or the control unit performs functions to implement or realize the method according to the invention or a step of the method according to the invention.

[0034] In the context of the invention, “data” means both raw data and already processed data, for example from measurement results from sensors or from simulation results.

[0035] Reinforcement learning is based on the LV agent 320 being able to i ∈ S from a set of available states at least one action a i ∈ A from a set of available actions. The choice of the selected action a iis based on a strategy or policy. For the selected action a i the LV agent 320 receives a reward r i ∈ R from the reward module 370. The states s i ∈ S is received by the agent 320 from the state module 350, which can be accessed by the LV agent 320. The strategy is determined based on the received rewards r i adjusted by the LV agent 320. The strategy specifies which action a i ∈ A from the set of available actions for a given state s i ∈ S is to be selected from the set of available states. This creates a new state s i+1 generated, for which the LV agent 320 receives a reward r i+1 A strategy thus defines the assignment between a state s i and an action a i so that the strategy determines the choice of action to be performed i for a state s iThe goal of the LV agent 320 is to collect the achieved rewards r i ,r i+1 , ..., r i+n to maximize.

[0036] In the action module 330, the actions selected by the LV agent 320 are i carried out. Through an action a i For example, an adjustment of a value of a parameter p i ∈ P from the set of parameters for at least one property e i a technical component of the electric drive train. Preferably, the action is a i is one of the actions A(+), A(0) and A(-). A positive action A(+) is an action that changes the value for a parameter p i increased, a neutral action A(0) is an action in which the value of the parameter p i remains the same, while with a negative action A(-) the value of the parameter p i reduced.

[0037] The environment module 340 calculates based on the selected action a i and taking into account previously defined constraints the states s i ∈ S. The boundary conditions can also include economic aspects such as the cost structure, energy costs, environmental impact, availability or the delivery situation.

[0038] A state of i ∈ S is thus determined by choosing certain values ​​for parameter p i of properties e i of the electric drive train 10. The properties e i For example, it can be a voltage response, an electrical resistance, or a characteristic curve for the torque / speed behavior of an electric motor in the electric drive train. A parameter value p i gives the specific voltage or torque for this property e i again.

[0039] In the state module 350, a deviation Δ between a target state s t and the calculated state s i The final state is reached when the calculated states s i equal to or greater than the target states s t are.

[0040] In the reward module 370, the degree of deviation Δ between the calculated value for the state s i and the target value of the state s t a reward r i Since the degree of deviation Δ depends on the selection of the respective action A(+), A(0), A(-), the reward r is preferably calculated in a matrix or a database of the respective selected action A(+), A(0), A(-). i assigned. A reward r i preferably has the values ​​+1 and -1, whereby a small or positive deviation Δ between the calculated state s i and the target state s tis rewarded with +1 and thus reinforced, while a significant negative deviation Δ is rewarded with -1 and thus negatively evaluated. However, it is also conceivable that values ​​> 1 and values ​​< 1 are used.

[0041] Preferably, a Markov decision process is used as the algorithm for the LV agent 320. However, it may also be provided to use a Temporal Difference Learning (TD-Learning) algorithm. An LV agent 320 with a TD-Learning algorithm does not adjust the actions A(+), A(0), A(-) only when it receives the reward, but after each action a i based on an estimated expected reward. Furthermore, algorithms such as Q-learning and SARSA, as well as Actor-Critic or Monte Carlo simulations, are also conceivable. These algorithms enable dynamic programming and strategy adaptation through iterative procedures.

[0042] In addition, the LV agent 320 and / or the action module 330 and / or the environment module 340 and / or the state module 350 and / or the reward module 370 contain calculation methods and algorithms for i for mathematical regression methods or physical model calculations that require a correlation between selected parameters p i ∈ P from a set of parameters and the target states s t describe. For the mathematical functions f tThese can be statistical methods such as mean values, minimum and maximum values, lookup tables, models of expected values, linear regression methods or Gaussian processes, fast Fourier transformations, integral and differential calculus, Markov methods, probability methods such as Monte Carlo methods, temporal difference learning, but also extended Kalman filters, radial basis functions, data fields, or even convergent neural networks, deep neural networks, feedback / recurrent neural networks or convolutional neural networks. Based on the actions a i and the rewards r i the LV agent 320 and / or the action module 330 and / or the environment module 340 and / or the state module 350 and / or the reward module 370 selects / selects for a state s i one or more of these calculation functions f i out of.

[0043] A neural network consists of neurons arranged in multiple layers and connected to one another in various ways. A neuron is capable of receiving information at its input from outside or from another neuron, evaluating the information in a specific way, and forwarding it in a modified form to another neuron at the neuron output or outputting it as the final result. Hidden neurons are arranged between the input neurons and output neurons. Depending on the network type, there may be multiple layers of hidden neurons. They ensure the forwarding and processing of information. Output neurons ultimately deliver a result and transmit it to the outside world. The arrangement and connection of the neurons gives rise to different types of neural networks, such as feedforward networks, recurrent networks, or convolutional neural networks.A convolutional neural network (CNN) has multiple convolutional layers and is well-suited for machine learning and artificial intelligence (AI) applications in the field of pattern recognition. These networks can be trained using unsupervised or supervised learning.

[0044] While in a classic environment module 340 a model of an electric drive train 10 is specified, which represents the target states s t1 , s t2 , .... , s tn According to the present invention, the learning reinforcement agent 320 independently and autonomously develops the model of the electric drive train 10. The model of the electric drive train 10 is developed through a plurality of actions. i ∈ A is learned by the learning reinforcement agent 320 and then forms the basis for the calibration of a real electric powertrain 10 by the learning reinforcement module 300.

[0045] The inventive concept thus consists in calibrating a real electric drive train 10 using model-based reinforcement learning, in which the model of the electric drive train 10 does not have to be present, but is modeled by the LV agent 320 itself. The model of the electric drive train 10 created by the LV agent 320 does not simulate the physics or dynamics of the electric drive train 10 in detail; rather, the model is developed using a multitude of interactions between actions, states, and rewards executed by the LV agent 320. The question posed by the LV agent 320 is therefore always which states exist and what happens when it performs an action for a specific state, and what the reward looks like when it performs an action for this specific state.

[0046] To create a model of an electric drive train 10, the invention provides that the environment module 340 has at least three submodules. The first submodule is designed as a state submodule 342, the second submodule as a reward submodule 343, and the third submodule as a strategy submodule 344.

[0047] The state submodule 342 represents states su1, su2 ..., su n that the LV agent 320 can select, with the selected state being j then the state in which the LV agent 320 is currently located. A state su j is simulated and is based on simulated observations b1, b2, .... b n , which are supplied to the state submodule 342 in the form of input data 220 from the input module 200. The LV agent 320 learns the states su1, su2 ..., su n of the state submodule 342 by collecting the observations b1, b2, .... b n. For the collected observations b1,b2, .... b n he designs a model that contains the states su1,su2 ...,su n in which it can be located, and which is a function of the collected observations b1, b2, .... b n He uses neural networks in particular to develop the model. For the observations b1, b2, .... b n For example, this can be the current, voltage, torque, and speed of an electric motor or the charge state of a battery of the electric drive train 10. Possible states su1, su2 ..., su n of the state sub-module 342 are thus derived from these simulated observations b1,b2, .... b n , such as a torque or speed of an electric motor.

[0048] The reward submodule 343 assigns the determined states su1, su2 ..., su n Rewards ru1, ru2, ...., ru n to.

[0049] The strategy sub-module 344 develops a strategy for determining new states su 1+1 , su 2+1 ..., su n+1 by suggesting which actions a j from the a1, a2, ..., a n Actions from the action submodule 330 to the old states su1, su2 ..., su n are to be applied. By applying the actions a1, a2, ..., a n new conditions will be su 1+1 , su 2+1 ..., su n+1 generated, which are then fed back to the state sub-module 342. In the reward sub-module 343, the newly determined states are 1+1 , su 2+1 ..., su n+1 in turn rewards ru 1+1 , ru 2+1 , ...., ru n+1 assigned.

[0050] The environment module 340 performs the calculations until a stable state level has been reached. This state level can be a target state.tj or a variety of target states t1 , su t2 ..., su tn for the LV agent 320. The result of the environment module 340 thus consists of the calculated target states su t1 , su t2 ..., su tn , which represent a training model TM of the electric drive train 10.

[0051] For the training phase, any or selected simulated observations b1, b2, .... b n as input data 220. From this input data 220, the LV agent 320 autonomously develops a first training model TM of the electric drive train 10. This model is defined by the target states su t1 , su t2 ..., su tn and the strategy used is described.

[0052] The training phase is followed by the modeling phase, in which the training model TM is transformed into a model M of a real electric drive train 10. The real electric drive train 10 is an ideal-typical embodiment in which a desired dynamic is given, for example, with regard to the ratio of torque and speed. In the modeling phase, real observations b are sent to the state sub-module 342 from the input module 200. r1 , b r2 , .... b rn as data 230, from which the real states are r1 , su r2 , ..., su rn generated. The real observations b r1 , b r2 , .... b rn measured parameter values ​​p i of a property e ithat have been determined by sensors not described in detail here. Preferably, the parameter values ​​are stored in a database 250 that is connected to the input module 200.

[0053] In the reward module 343, a deviation Δ between the real states is now r1 , su r2 , ..., su rn and the target states generated during the training phase t1 , su t2 , ..., su tn In addition, in the reward module 343, the degree of deviation Δ between the real state su ri and the target value of the target state su ti a reward r i+1 assigned.

[0054] The strategy sub-module 344 develops due to the new rewards r 1+1 , r 2+1 , ..., r n+1 a changed strategy for identifying new conditions 1+1 , su 2+i ..., su n+1by suggesting which actions a j from the a1, a2, ..., a n from the action submodule 330 to the old target states su t1 , su t2 ..., su tn should be applied. The final state is reached when the generated states are t1+1 , su t2+1 , ..., su tn+1 equal to or greater than the real states su r1 , su r2 , ..., su rn are, since the training model TM was then transformed into a model M, which represents a real ideal-typical electric drive train.

[0055] This model M of a real electric drive train 10 now represents the target states sm t1 , sm t2 , ...., sm tn available with which a calibration of an individual real electric drive train 10 can be carried out by the LV agent 320.

[0056] For this purpose, 350 values ​​of parameters pi of properties e i of an individual electric drive train 10 from the input module 200 in the form of real data 240. The parameter values ​​p i can be measured by sensors not described in detail here. These sensors include, in particular, pressure sensors, torque sensors, speed sensors, acceleration sensors, speed sensors, capacitive sensors, inductive sensors, and temperature sensors.

[0057] A state of i ∈ S of an individual electric drive train 10 is thus determined by selecting values ​​of parameters p i of properties e i defined. The properties e i For example, it can be a voltage response, an electrical resistance, or a characteristic curve for the torque / speed behavior of an electric motor in the electric drive train. A parameter value p igives the specific voltage or torque for this property e i again.

[0058] The LV agent selects for these states s1, s2, ..., s n as described above, perform actions (A+), (A0) and (A-) to adapt to the target states sm t1 , sm t2 , ...., s mtn of the generated model M. The environment module 340 calculates a i and taking into account previously defined constraints, new states s i+1 ∈ S. The boundary conditions can also include economic aspects such as the cost structure, energy costs, environmental impact, availability or the delivery situation.

[0059] In the state module 350, a deviation Δ between a target state s t and the calculated state s i+1In the reward module 370, the degree of deviation Δ between the calculated value for the state s i+1 and the target value of the state sm t a reward r i assigned.

[0060] Then a second cycle begins, in which the LV agent 320 performs another action i+1 and / or another calculation function f i+1 and / or another parameter p i+1 selected according to the defined strategy or policy. The result is again fed to the state module 350, and the result of the comparison is evaluated in the reward module 370. The LV agent 320 repeats the calibration process for all planned actions a i , a i+1 , ..., a i+n , calculation functions f i , f i+1 , ..., f i+n and parameter p i , p i+1 , ..., p i+n until the greatest possible agreement between a calculated state si+1 , s i+2 , ..., s i+n and a target state sm ti is reached. Preferably, the final state of the calibration is reached when the deviation Δ is in the range of + / -5%. The LV agent 320 thus optimizes its behavior and thus the strategy according to which an action a i is selected until the calculated states s i+1 , s i+2 , ..., s i+n converge. The final state is reached when the calculated states s i+1 , s i+2 , ..., s i+n equal to or greater than the target states sm1, sm2, ..., sm n The calibration result can be output in the form of output data 450 on the output module 400. The input module 200 and the output module 400 can be integrated into a hardware device such as a computer, a tablet, a smartphone, etc.

[0061] In particular, it can be provided that the calculation results in the form of states, actions, rewards, and strategies are stored in a cloud computing infrastructure and are each available via the Internet. The LV agent 320, the action module 330, the environment module 340, the state module 350, and the reward module 370 have the necessary technical interfaces and protocols for accessing the cloud computing infrastructure. This can increase computing efficiency because the access options and speeds to previously calculated states, actions, rewards, and strategies are simplified.

[0062] In Fig. 2 shows the process steps for autonomously calibrating an individual electric drive train 10.

[0063] In a step S10, a training model TM for an electric drive train 10 is created by a learning reinforcement agent 320 using simulated observations b1, b2, ..., b n created, wherein the learning reinforcement agent 320 uses a reinforcement learning algorithm.

[0064] In a step S20, the training model TM is modified by the learning reinforcement agent 320 using real observations b r1 , b r2 , ... , b rn a real ideal-typical drive train 10 for creating a simulated model M for the real ideal-typical electric drive train 10, wherein the simulated model M target states sm1, sm2, ..., sm n contains.

[0065] In a step S30, at least one state s i of an individual real electric drive train 10 is determined by a state module 350, wherein a state s i by parameter p isuch as data and / or measurements of at least one property e i of the electric drive train 10 is defined.

[0066] In a step S40, the state s i transmitted to the learning reinforcement agent 320.

[0067] In a step S50, calibration results 450 for the individual real electric drive train 10 are obtained by the learning reinforcement agent 320 by comparing the state s i with at least one target state sm ti of the simulated model M.

[0068] Fig. 3 schematically illustrates a computer program product 900 comprising executable program code 950 configured to perform the method according to the first aspect of the present invention when executed.

[0069] With the method and system 100 according to the present invention, an electric drivetrain 10 can thus be reliably calibrated using reinforcement learning methods without requiring a detailed model of a real electric drivetrain 10 in the environment module 340 of the learning reinforcement module 300. Rather, the modeling of a real electric drivetrain is performed independently and autonomously by the LV agent 320. As a result, the target states to be achieved during calibration are specified by the model created by the LV agent. The target states are more precise and therefore enable improved calibration. With the present invention, reliable calibration of electric drivetrains can thus be performed in a short time and at reduced cost. Reference symbol 10 electric drivetrain 100 systems 200 input module 220 simulated data 230 real data 240 real data 250 database 300 Learning Reinforcement Module 320 Learning Reinforcement Agent 330 Action Module 340 Environment module 342 State submodule 343 Reward submodule 344 Strategy submodule 350 State module 370 Reward Module 400 output module 450 calibration results 900 computer program product 950 program code

Claims

[1] Method for the autonomous calibration of an individual electric drive train (10), wherein sensors and / or measuring devices for determining parameters (p i ) of properties (e i ) of the individual electric drive train (10), comprising: - Creating (S10) a training model (TM) for an electric drive train (10) by a learning reinforcement agent (320) by means of simulated observations (b1, b2, ..., b n ), wherein the learning reinforcement agent (320) uses a reinforcement learning algorithm, wherein the simulated observations (b1, b2, .... b n ) in particular the current, the voltage, the torque and / or the speed of an electric motor and / or the state of charge of a battery of the electric drive train (10); - modifying (S20) the training model (TM) of the learning reinforcement agent (320) by means of real observations (b r1, b r2 , ..., b rn ) of a real ideal-typical drive train (10) for creating a simulated model (M) for the real ideal-typical electric drive train (10), wherein the real observations (b r1 , b r2 , .... b rn ) measured values ​​of parameters (p i ) of a property (e i ) of the real ideal-typical drive train (10), which are determined by the sensors or stored in a database (250), and wherein the simulated model (M) represents target states (sm t1 , sm t2 , ..., sm tn ) contains; - Determining (S30) at least one state (s i ) of an individual real electric drive train (10) by a state module (350), wherein a state (s i ) by parameter (p i ) such as data and / or measurements of at least one property (e i ) of the electric drive train (10), - Transmitting (S40) the status (s i ) to the learning reinforcement agent (320); - Determining (S50) calibration results (450) for the individual real electric drive train (10) from the learning reinforcement agent (320) by comparing the state (s i ) with at least one target state (sm ti ) of the simulated model (M). [2] Method according to claim 1, wherein for the creation of a training model (TM) for an electric drive train (10) by a learning reinforcement agent (320) by means of simulated observations (b1, b2, ..., b n ) an environment module (340) is provided, which comprises at least a state sub-module (342), a reward sub-module (343) and a strategy sub-module (344). [3] Method according to claim 2, wherein the state sub-module (342) generates states (su1, su2 ..., su n ) are generated based on the simulated observations (b1, b2, .... bn ) are based on. [4] The method according to claim 1, wherein determining calibration results comprises the following method steps: - Selecting a calculation function (f i ) and / or an action (a i ) based on a policy for a state (s i ) for the modification of at least one parameter (p i ) from the learning reinforcement agent (320); - Calculate a modeled value for the property (e i ) using the modified parameter (p i ); - Calculate a new state (s i+1 ) by an environment module (340) based on the modeled value for the property (e i ); - Compare the new state (s i+1 ) with the target state (sm t ) and assigning a deviation (Δ) for the comparison result in the state module (350); - Determining a reward (r i) by a reward module (370) for the comparison result; - Adjusting the policy of the learning reinforcement agent (320) based on the reward (r i ), where, if the policy converges, the optimal action (a j ) for the calculated state (s j ) and, in case of non-convergence of the policy, another calculation function (f j ) and / or another action (a j+1 ) for a state (s j+1 ) with a modification of at least one parameter (p j ) is selected by the learning reinforcement agent (320) until the target state (sm t ) is reached. [5] Method according to one of claims 1 to 4, wherein a positive action (A+) which determines the value for a parameter (p i ) increases, a neutral action (A0) in which the value of the parameter (p i ) remains the same, and a negative action (A-), in which the value of the parameter (pi ) are provided for. [6] Method according to one of the preceding claims 1 to 5, wherein the reward module (370) comprises a database or matrix for evaluating the actions (a i ) includes. [7] Method according to one of claims 1 to 6, wherein the at least one algorithm of the learning reinforcement agent (320) is designed as a Markov decision process, temporal difference learning (TD-learning), Q-learning, SARSA, Monte Carlo simulation or actor-critic. [8] A system (100) for autonomously calibrating an individual electric drive train (10), comprising an input module (200), a learning reinforcement module (300), an output module (400) and sensors and / or measuring devices for determining parameters (p i ) of properties (e i) of the individual electric drive train (10), wherein the learning reinforcement module (300) comprises a learning reinforcement agent (320) using a reinforcement learning algorithm, an action module (330), an environment module (340), a state module (350), and a reward module (370); wherein the learning reinforcement agent (320) is configured to create a training model (TM) for an electric drive train (10) by means of simulated observations (b1, b2, ..., b n ) and to test the training model (TM) using real observations (b r1 , b r2 , ..., b rn ) of a real ideal-typical drive train (10) to create a simulated model (M) for the real ideal-typical electric drive train (10), wherein the simulated observations (b1, b2, .... b n) in particular the current, the voltage, the torque and / or the speed of an electric motor and / or the state of charge of a battery of the electric drive train (10) and the real observations (b r1 , b r2 , .... b rn ) measured values ​​of parameters (p i ) of a property (e i ) which are determined by the sensors or which are stored in a database (250), and wherein the simulated model (M) represents target states (sm t1 , sm t2 , ..., sm tn ); wherein the state module (350) is designed to contain at least one state (s i ) of an individual real electric drive train (10), wherein a state (s i ) by parameter (p i ) such as data and / or measurements of at least one property (e i ) of the electric drive train (10) is defined, and the state (s i) to the learning reinforcement agent (320); and wherein the learning reinforcement agent (320) is designed to obtain calibration results (450) for the individual real electric drive train (10) by comparing the state (s i ) with at least one target state (sm ti ) of the simulated model (M). [9] The system (100) of claim 8, wherein the environment module (340) comprises at least a state sub-module (342), a reward sub-module (343), and a strategy sub-module (344). [10] System (100) according to claim 9, wherein the state sub-module (342) is configured to store states (su1, su2 ..., su n ) based on the simulated observations (b1, b2, .... b n ) are based on. [11] A computer program product (900) comprising an executable program code (950) configured to carry out the method according to any one of claims 1 to 7 when executed.

Citation Information

Patent Citations

  • System and method for autonomously constructing and / or designing at least one component for a part

    DE102020118805A1