Method for calibrating a target metric of an ECU function
A reinforcement learning agent trains a machine learning algorithm to predict calibration parameter changes, addressing the inefficiencies of manual and existing automated methods, achieving robust and efficient control unit function calibration.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-07
- Publication Date
- 2026-04-09
AI Technical Summary
Calibrating control unit functions with significant dead time and numerous operating parameters is time-consuming and costly, often requiring manual trial-and-error methods, and existing automated optimization agents require restarts if no convergence occurs.
A method using a reinforcement learning agent to train a machine learning algorithm to predict changes in calibration parameters based on input signals, optimizing the control unit function by defining the parameterization problem as a sequential decision task, particularly a Markov decision problem, and employing a reward function to guide the training process.
This approach provides robust and efficient calibration of control unit functions with reduced time and cost, ensuring stable and optimal performance by continuously adapting to input signals and adjusting calibration parameters.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a method for calibrating a target metric of a control unit function.
[0002] Control unit functions require calibration to perform their task. This task might involve controlling a system so that a target metric remains within a predefined value range. For example, an internal combustion engine can be regulated so that a lambda sensor reading falls within a certain range. Calibrating systems with significant dead time and / or a large number of operating parameters is time-consuming, as it is performed manually by an experienced engineer using a trial-and-error approach.
[0003] It is also known to use automated optimization agents for calibration, which search for a global optimum for calibration or determine the optimum using a gradient descent method. For example, a method for operating an internal combustion engine is known from DE 10 2020 116 488 B3. In this method, an air-fuel ratio is controlled using a neural network. The neural network can be trained with a strengthening learning method to predict a lambda value.
[0004] However, such optimization agents must be restarted if no convergence occurs.
[0005] The object of the invention is to provide a method for calibrating a control unit function that provides robust calibration and thus reduces time and costs.
[0006] The problem is solved by the features of the independent claims. Advantageous further developments are the subject of the dependent claims and the following description.
[0007] According to a first aspect, a method for calibrating a target metric of an ECU function of a system with a multitude of operating parameters, in particular an ECU function for controlling an air-fuel ratio of an internal combustion engine, comprises at least the following steps: using at least one calibration parameter in the at least one ECU function, wherein the at least one calibration parameter was determined by means of at least one trained learning algorithm from at least one input signal provided by the ECU function; wherein the algorithm is trained beforehand by means of at least the following step: solving the at least one sequential decision task with an agent for reinforcement learning to train the learning algorithm.
[0008] This method provides calibration of the ECU function using a machine learning algorithm trained on a sequential decision task for parameterizing the ECU function. In contrast to the prior art, the parameterization problem for the system with its multitude of operating parameters is defined as a sequential decision task. The reinforcement learning agent solves this sequential decision task by training the machine learning algorithm. The agent with the machine learning algorithm is therefore trained to predict, at least based on an input signal provided by the ECU function, a change in the calibration parameter that can be used for the ECU function in a subsequent time interval.This approach does not directly train the machine learning algorithm based on the system's operating parameters to provide a calibration parameter. Instead, the machine learning algorithm can be trained to predict changes in the calibration parameter, at least from input signals provided by the control unit function. The input signal can, for example, contain one or more control signals for the system. In this way, a robust machine learning algorithm can be provided for calibrating the control unit function, and it can be trained with a manageable time investment and correspondingly manageable costs.
[0009] The control unit function can have a closed control loop.
[0010] The adaptive algorithm can be implemented as a neural network. The input signal can be provided at the input level of the neural network. An output level of the neural network can provide an output signal that can indicate at least one predicted change in a calibration parameter.
[0011] According to some embodiments, it is conceivable that the sequential decision task describing the calibration problem can be designed as a Markov decision problem, preferably an approximation.
[0012] This allows all, or nearly all, possible subsequent states and their relationships to the current state to be known for every state the system is in during training. The closer the sequential decision task is to a Markov decision problem, the more robust the training of the learning algorithm can be.
[0013] According to some embodiments, solving the sequential decision task could include at least the following training steps: Determining at least one training input signal for the learning algorithm by means of at least one simulation of a control unit function of the target metric, which uses at least one training calibration parameter; Predicting at least one change in the training calibration parameter using the learning algorithm based on the training input signal; Determining at least one reward value based on at least one changed control signal, which is determined from a simulation based on the changed training calibration parameter; Changing the agent and / or the learning algorithm based on the reward value;Repeat the training steps with at least one different training input signal if the reward value does not converge within a target interval and / or if a predefined maximum number of repetitions of the training steps has not been reached.
[0014] First, a simulation of an ECU function can be performed, using training calibration parameters. The simulation can then, for example, provide at least one training control signal from which the training input signal for the machine learning algorithm can be generated. Based on the training input signal, the machine learning algorithm then predicts a change in the training calibration parameters. Based on this change, the simulation can determine at least one modified control signal. At least this modified control signal can then form the basis for determining a reward value. Afterward, the agent and / or the machine learning algorithm can be modified. If the termination criterion is not met, the steps described above are performed with a modified training input signal.These procedural steps do not optimize the machine learning algorithm to determine a specific training calibration parameter from certain training input signals. Instead, the machine learning algorithm is presented with a new training input signal in each iteration and uses this to determine changes in the training calibration parameter. The reward value depends on the result of the simulation with the modified training calibration parameter. A reward function can be used by the agent to train the machine learning algorithm. The agent receives a high reward if the algorithm outputs stable and optimal values or improvements in those values. Otherwise, the agent receives only a small reward. Furthermore, the reward value can be determined over the entire training process. This allows the resulting machine learning algorithm to be trained more robustly and efficiently.
[0015] According to some embodiments, it is conceivable that the simulation in the step of determining at least one training input signal uses at least one training operating parameter of the system that can be changed by a random disturbance, and / or that the simulation can be changed by changing a setpoint for the target metric.
[0016] This allows a different training input signal to be used for each repetition, one that differs from previously used training input signals. This differentiation can be achieved by randomly perturbing real and / or artificially generated measurement data. Alternatively or additionally, the differentiation can be achieved by shifting the target value for the objective metric, which, for example, causes the real and / or artificially generated measurement data to be processed differently by the simulation.
[0017] According to some embodiments, it is conceivable that in the step of predicting at least one change in the training calibration parameter, a weighting value can be determined with which a change in the training calibration parameter predicted by the learning algorithm can be weighted in order to obtain the changed training calibration parameter.
[0018] For example, when using the weighting value to change training calibration parameters, a difference from values in a table containing existing values for the training operating parameters can be taken into account. This allows for a further statistical analysis of the training operating parameters, which can then be incorporated into the changes to the training calibration parameters.
[0019] According to some embodiments, it is conceivable that in the step of determining at least one reward value, the reward value is based on at least one training target metric, which can be based on at least one training operating parameter changed by the changed control signal.
[0020] At least with the training operating parameter, which can result from the control signal from the simulation obtained through the modified training calibration parameters, a training target metric can be determined. In the example of an internal combustion engine, the air-fuel ratio can serve as the target metric. Depending on where the air-fuel ratio, or the training target metric, lies within or outside a target metric interval, the reward value can be set. In this way, the training can be kept efficient, as a reward value is continuously determined, motivating the agent to improve its decision-making performance throughout the entire process.
[0021] According to some embodiments, it is conceivable that the reward value is determined using a reward function, which may have a logarithmic function.
[0022] Using a logarithmic function to determine the reward value allows for a larger change in the reward value compared to linear or quadratic relationships in reward functions when the training target metric changes due to at least one modified training calibration parameter. Furthermore, even within the target interval for the reward value, a sensitive change in the reward value can occur depending on a small change in the training target metric, for example, in the per mille range, which is mathematically more difficult to achieve with comparable approaches.
[0023] According to some embodiments, it is conceivable that at least one predefined initial training calibration parameter can be provided for the simulation for a first execution of the training steps.
[0024] Using a predefined initial training calibration parameter for the simulation, the probability of stable execution of the training steps in the first run can be determined. The initial training calibration parameters may still be suboptimal, allowing for further improvement potential through training.
[0025] According to some embodiments, it is conceivable that the training input signal can be based on real measurement data as training operating parameters.
[0026] The training of the learning algorithm can then be based on simulations that reflect real-world conditions.
[0027] According to some embodiments, it is conceivable that the training input signal could be based on artificially generated measurement data to which noise has been added as training operating parameters.
[0028] If insufficient or no real measurement data is available for certain operating states of the system, training input signals can be artificially generated. This can involve using artificially generated measurement data, which can then be subjected to noise. For example, a Gaussian function can be used for this purpose.
[0029] According to another aspect, a system is described comprising at least one internal combustion engine, at least one sensor for detecting a target metric, in particular an air-fuel ratio, and at least one control unit for the control unit function of the internal combustion engine based on sensor signals from the sensor, wherein the control unit is configured to operate the trained learning algorithm for calibrating the control unit function of the internal combustion engine, which was trained according to the procedure explained above.
[0030] The advantages, effects, and further development opportunities of the system arise from the advantages, effects, and further development opportunities of the procedure described above. To avoid repetition, reference is therefore made to the preceding description in this regard.
[0031] The invention is described below with reference to an exemplary embodiment and the accompanying drawing. The drawing shows: Fig. 1. A flowchart of the process; Fig. 2. A flowchart of the training steps of the procedure; Fig. 3 a schematic diagram showing a system trajectory of the operating parameters; Fig. 4. A schematic diagram of the reward function; and Fig. 5 a schematic representation of the system.
[0032] The procedure for calibrating a control unit function of a target metric of a system with a multitude of operating parameters, referred to below as the entirety of the system by the reference symbol 100.
[0033] The system can, for example, be an internal combustion engine, where the control unit function can be designed to regulate an air-fuel ratio, the so-called lambda value.
[0034] According to step 102, the control unit function uses at least one calibration parameter determined by a trained, self-learning algorithm. This algorithm determines at least one calibration parameter from at least one input signal that can be provided by the control unit function and could, for example, be a control signal. Furthermore, the algorithm can also use other parameters, such as operating parameters of the system, to determine the calibration parameters.
[0035] According to a further step 104, the algorithm may have been trained before step 102 using an agent for reinforcement learning. The agent can solve a sequential decision task for training the learning algorithm, which describes the problem of calibrating the control unit function.
[0036] Furthermore, the agent can, for example, be a software agent that runs on a computer. The learning algorithm can also be executed as software on a computer.
[0037] The sequential decision task can be further developed as at least an approximation of a Markov decision problem. The training of the machine learning algorithm is thus embedded in a sequential decision task that is solved by the agent. In particular, if the sequential decision task is at least an approximation of a Markov decision problem, the agent can be influenced by a reward function in such a way that it solves the decision task as efficiently as possible by adapting the machine learning algorithm. The learning progress of the machine learning algorithm is therefore not directly assessed, for example, by means of a cost function, but rather controlled by the agent's indirect assessment. Furthermore, a sum of the reward values can be calculated that can take into account all repetitions and changes made since the start of the training.Both bad and good reward values have long-term consequences for training, thus improving the efficiency of the training.
[0038] Step 104 can optionally be the one described in the Fig. The training steps 106-116 shown in the illustrations are shown.
[0039] According to step 106, a simulation of an ECU function of the target metric can first be performed. The simulation can use at least one training calibration parameter and at least one training operating parameter as input, where the training operating parameter can represent a parameter of the system. The output of the simulation can be at least one control signal for the system, which can influence the target metric. At least the control signals from the simulation can be used as a training input signal that can be provided to the machine learning algorithm.
[0040] The system's training parameters can be generated from real measurement data. Alternatively or additionally, the training parameters can be generated from artificially generated measurement data. This artificially generated data can then be further processed with noise.
[0041] Before simulating the ECU function, the training operating data can be modified by a random disturbance. The simulation then uses this randomly disturbed training operating data to generate the corresponding control signals. Alternatively or additionally, a target value for the simulated ECU function can be changed for the target metric before the simulation is performed.
[0042] In this way, training input signals can be generated that can be changed randomly, so that the sequential decision task can be performed with new training input signals at each repetition.
[0043] The machine learning algorithm can be implemented as a neural network. It is advantageous if the machine learning algorithm is implemented as a deep neural network that can be scaled.
[0044] According to step 108, the learning algorithm can predict at least one change in the training calibration parameter upon provision of the training input signal.
[0045] Furthermore, in step 108, a weighting value can be determined, which can be used to weight a change in the training calibration parameter predicted by the machine learning algorithm. This allows a modified training calibration parameter to be generated.
[0046] The weighting parameter can be determined, for example, by first statistically processing the mean and / or median of the observables, such as the control signal of the simulated and / or real system and / or operating parameters of the system. An observable is understood to be information about the task observed from the agent's perspective. A system trajectory that can represent the observable can be used for this purpose.
[0047] Trajectory 52 is exemplified in Fig. Figure 3 is shown in a diagram 30, which can display a grid. The grid points can represent table values from a lookup table. For a center of the observable, marked by coordinates 32 and 34 in trajectory 52, a distance to the nearest table values can be calculated. The nearest table values can be represented, for example, by grid points 44, 48, 50, and 52. The distances can be represented accordingly by the values 36, 38, 40, and 42. The other grid points in diagram 30 can represent further table values. The weighting value can be calculated from the distances.
[0048] Multiple observables can be used with correspondingly multiple lookup tables to determine at least one weighting value.
[0049] Using the weighted output of the machine learning algorithm, which can represent changes in the training calibration parameter, the simulation can detect at least one changed control signal. From this changed control signal, further changed training operating parameters can be determined.
[0050] According to the subsequent step 110, at least one reward value can be determined based on the at least one modified control signal and optionally also based on the training operating parameters. For example, a training target metric can be determined using the modified training operating parameters resulting from the modified control signal. The training target metric could be, for example, a modified lambda value, such as a modified air-fuel ratio if the system has an internal combustion engine.
[0051] The reward value can be calculated using a reward function 56, which is exemplified in Fig. 4 is shown. Diagram 54 in Fig. Figure 4 shows, on the right axis of the target metric, an example of the Lambda value and on the vertical axis the reward value r, which depends on Lambda, where Lambda in this embodiment can be constant and can have the value 1.
[0052] The target value is 1 in this example. The reward values can advantageously range from -2 to +1 with a lower limit r_low and an upper limit r_high, as exemplified in Fig. Figure 4 illustrates this. The training target metric, specifically the lambda value, should lie between the limits 58 and 60, which represent the target interval. If the target metric lies within these two limits, the reward value is positive. If the target metric lies outside these two limits, the reward value is negative.
[0053] The reward function 56 can have logarithmic components. This means that, compared to linear or quadratic relationships, the reward value changes significantly even with only slight changes in the target metric. For example, a change in the target metric in the per mille range can result in a measurable percentage change in the reward value. Furthermore, a sensitive adjustment of the reward value within the target interval can also be configured.
[0054] This means that the quality of the control unit function can react very sensitively to small changes or fluctuations in the target metric.
[0055] Furthermore, the reward value can be designed to be the sum of all previous reward values. With many repetitions of the training steps, this provides continuous rewards, which promotes an improvement in the accuracy of the learning algorithm across all steps.
[0056] According to a further step 112, the agent and / or the machine learning algorithm can be modified based on the reward value. Alternatively or additionally, the determined control signals, the training operating parameters, and the determined changes in the training calibration parameters can be used as the basis for a change in the machine learning algorithm and / or the agent.
[0057] According to step 114, training steps 106-112 are repeated if the reward value does not converge within the target interval. Convergence of the reward value can be defined as occurring over at least two repetitions of training steps 106-112. However, any number of repetitions, except for a single repetition, is conceivable for defining convergence.
[0058] Alternatively, the repetition can be stopped in step 114 if a maximum number of repetitions of training steps 106-112 has been reached.
[0059] Before training steps 106-112, an additional optional step 116 can be performed. In this step 116, at least one initial training calibration parameter for the simulation can be provided to initialize the training. The initial training calibration parameter can include calibration parameters that ensure stable simulation execution. Furthermore, the initial training calibration parameter may still be suboptimal for use in the ECU function, and therefore can be optimized.
[0060] The above-described procedure 100 can be used in a system that is in Fig. Figure 5 is shown schematically. The system as a whole is designated by reference numeral 10.
[0061] System 10 can include an internal combustion engine 12. Furthermore, system 10 can include at least one sensor 14 for acquiring a target metric and a control unit 16 for the internal combustion engine 12. The sensor 14 can be configured as at least one wideband lambda sensor. The control unit 16 can be configured to read the sensor 14. The target metric can, for example, be the air-fuel ratio in an exhaust pipe 18 upstream of a catalyst 20 of the internal combustion engine 12.
[0062] The control unit 16 can be trained to use the trained learning algorithm described above to determine calibration parameters.
[0063] The control unit 16 and the system 10 can also be configured separately. The control unit 16 can be an integral part of a control system and can have calibration parameters.
[0064] In some embodiments, the agent can be located outside the control unit 16 on other computing units.
[0065] Furthermore, both online and offline operation of System 10 or Procedure 100 is conceivable.
[0066] The example described above does not in any way limit the invention. Rather, the invention can be modified in numerous ways. All features of the invention described above can be essential to the invention, either alone or in combination. Reference symbol list 10 System 12 Internal combustion engine 14 Sensor 16 Control unit 18 Exhaust pipe 20 catalyst 30 Diagram 32 Coordinate 34 Coordinate 36 distance 38 distance 40 distance 42 distance 44 grid point 46 grid point 48 grid points 50 grid points 52 Trajectory 54 Diagram 56 Reward function 58 Limit value 60 Limit value QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] DE 10 2020 116 488 B3
[0003]
Citation Information
Patent Citations
Method for operating an internal combustion engine, control unit and internal combustion engine
DE102020116488B3