Control of an actuator based on reinforcement learning

The method addresses the challenge of parameterizing control loops for complex systems by using reinforcement learning to select discrete control parameters, improving actuator control and training success by clearly distinguishing output signals.

DE102023209638B4Active Publication Date: 2025-06-12ZF FRIEDRICHSHAFEN AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102023209638
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-09-29
Publication Date
2025-06-12
Estimated Expiration
2043-09-29

AI Technical Summary

Technical Problem

The parameterization of control loops for complex systems like electrohydraulic actuators is challenging due to non-linearities and measurement inaccuracies, leading to inefficient manual adaptation of control parameters.

Method used

A method using reinforcement learning to generate control parameters, where a software agent selects discrete control parameters to maximize a reward, improving control and reinforcement learning by clearly distinguishing output signals despite measurement inaccuracies.

Benefits of technology

The method enhances the control and training success of actuators by allowing the software agent to reliably differentiate the influence of parameter changes, leading to improved actuator control and broader applicability of reinforcement learning methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for controlling an actuator based on control parameters generated by reinforcement learning, as well as an associated device and an associated training method, is disclosed. The method for controlling an actuator based on control parameters generated by reinforcement learning comprises the steps of determining (S1) a control parameter set with a plurality of discrete control parameters for the respective actuator, receiving (S2) a control parameter generated by a software agent, mapping (S3) the control parameter to one of the control parameters of the control parameter set, and controlling (S4) the actuator based on the control parameter. The invention also relates to a method for improved training of a software agent for controlling an actuator and a device for carrying out this method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field

[0001] The present invention relates to a method for controlling an actuator based on control parameters generated by reinforcement learning and to an associated device. State of the art

[0002] In practice, the parameterization of control loops is a complex process. Due to the complexity of the system, the parameterization of the control loops usually requires considerable effort and effort. This is usually done using so-called sample parameters, which then have to be manually adjusted.

[0003] For example, choosing the right control parameters for controlling electrohydraulic actuators in transmissions is challenging and time-consuming due to non-linearities and influences from component variations.

[0004] DE 10 2007 001 024 A1 relates to a method for computer-aided control and / or regulation of a technical system, in particular a gas turbine, as well as a corresponding computer program product. With the method according to the invention, a simulation model can be created based on just a few measured data, which can then be used to determine which learning or optimization method is particularly suitable for controlling or regulating the system. Description of the invention

[0005] In a first aspect, the present invention relates to controlling an actuator based on control parameters generated by means of reinforcement learning. An actuator can be any actuator and in particular an electrohydraulic actuator. Reinforcement learning is understood to mean so-called reinforcement learning, in which a software agent selects suitable control parameters in the form of actions through interaction with the actuator, such as the electrohydraulic actuator, in order to maximize a reward. In this context, the reward describes the quality of the selected control parameters. Various methods for reinforcement learning are generally known to those skilled in the art.

[0006] Such reinforcement learning methods typically operate within a continuous range of values ​​and sometimes select the output control parameters in very fine detail. If these values ​​are then input into the actuator and subsequently used for reinforcement learning, the influence of the change in the control parameters can no longer be correctly evaluated due to measurement inaccuracies in the actuator output, and thus the influence of the individual control parameters cannot be clearly distinguished. Thus, if the uncertainty is too great, the training success can be reduced or, in the worst case, completely prevented. Accordingly, a discretization will be performed below.

[0007] In a first step, a control parameter set with a large number of discrete control parameters is determined for the respective actuator. A control parameter set thus consists of a large number of control parameters, each of which corresponds to values ​​with which the actuator can be controlled. These control parameters generally have a certain distance from one another, also known as resolution. For example, a control parameter set can be defined as follows: a diskret= {20, 40, 60, 80, 100, 120, 140, 160}. In the above example, the resolution is therefore 20. The resolution, or the distance between the control parameters, must be selected such that controlling the actuator with adjacent control parameters from the control parameter set results in distinguishable outputs from the actuator. This means that the control parameters from the control parameter set must be selected such that the actuator output can be clearly distinguished, even taking measurement errors into account.

[0008] Then, in a further step, a control parameter is received from the aforementioned software agent.

[0009] The received control parameter is then mapped to one of the control parameters of the control parameter set. A mapping here refers to the assignment of the received control parameter to a control parameter using a predefined mapping function.

[0010] Based on this control parameter obtained in this way, the actuator can then be controlled.

[0011] For this purpose, the control parameter can be applied to the actuator or entered into it in another way.

[0012] By appropriately discretizing the continuous output signal of the software agent, it is possible to improve both the control and any subsequent reinforcement learning that may follow later, since such discretization makes the output signals clearly distinguishable, thus resulting in both better control and an improvement in reinforcement learning based on the actuator's outputs. The software agent can thus reliably distinguish the influence of parameter changes and draw correct conclusions for further increasing the reward and thus improving the actuator's control. At the same time, the present invention enables a broader set of methods for such control by enabling reinforcement learning methods that have continuous output signals.

[0013] In one aspect, in the determining step, a distance between the control parameters is determined such that an input of two adjacent control parameters into the actuator results in outputs with spaced value ranges at the actuator. As already explained above, such a selection of the control parameters results in a spacing of the outputs and thus in outputs in which the value ranges of the outputs for different input signals do not overlap and are thus clearly distinguishable from one another.

[0014] In a further aspect, the following steps are carried out in the mapping step. First, the control parameter is normalized based on a previously known action space of the software agent. Each software agent has a respective previously known action space, i.e., a range within which the output values ​​of the software agent can fall. This action space is then normalized by mapping it to a predetermined value range. This can, for example, be a value range between 0 and 1. Then, in a further step, the normalized control parameter is multiplied by a value based on the number of control parameters in the control parameter set. This can either be the number of control parameters in the control parameter set itself, but a correction factor can also be subtracted from the number of control parameters in the control parameter set. This correction factor is preferably in the range 0 < Δ << 0.01.Accordingly, the multiplication output is a number whose integer portion corresponds to the number of the control parameter in the control parameter set to be selected. Consequently, the decimal places are then removed, also known as truncation.

[0015] In the next step, the control parameter located at this position in the control parameter set is then selected.

[0016] Accordingly, any software agent can be used with the aforementioned implementation, since the software agent can be used independently of its action space and can be mapped to a rule parameter of the rule parameter set using the same procedure.

[0017] According to a further aspect, the method further comprises the step of inputting the output of the actuator as input to the software agent for performing reinforcement learning. Accordingly, reinforcement learning can be improved because the outputs, and thus the influences of the input signals, are clearly distinguishable from one another. This method can thus also be referred to as a method for improving reinforcement learning used for controlling an actuator.

[0018] Accordingly, a method for improved training of a software agent for controlling an actuator based on control parameters generated by reinforcement learning may, according to a further aspect, comprise the steps discussed below. First, the method comprises the steps of determining a control parameter set with a plurality of discrete control parameters for the respective actuator, receiving a control parameter generated by a software agent, mapping the control parameter to one of the control parameters of the control parameter set, and controlling the actuator based on the control parameter. These steps have already been discussed above.

[0019] The method then further comprises the step of determining parameters as the actuator's response to the input. Such parameters can be measured values ​​that represent the actuator's behavior. The parameters thus define a state of the actuator that results from the parameter input. A corresponding discretization of the parameters is also possible here.

[0020] The parameters are then entered into the reinforcement learning system's software agent, which is programmed so that a desired behavior of the actuator's parameters leads to an increased reward. For example, if the selected control parameter leads to a desired behavior of the actuator, a positive or increased reward is awarded. Furthermore, the software agent is programmed to output a control parameter for the actuator.

[0021] The steps of receiving, mapping, controlling, determining parameters, and inputting are then repeated until a termination condition is reached. This leads to optimization of the control parameter. At the same time, the software agent learns which change in the control parameter in which state—that is, for which parameters—leads to an increased reward and thus to better control of the actuator or its component.

[0022] Regarding the aforementioned aspects, it should be noted that these are not all limited to one control parameter, but can also be applied to several control parameters and thus the software agent can also output several control parameters.

[0023] Furthermore, one aspect comprises an apparatus for carrying out the method according to one of the preceding claims. Short description of the characters Fig. 1 shows the influence of directly applying the output of software agents with continuous outputs as a control signal for an actuator; Fig. 2 shows a flowchart of a method according to a first aspect; Fig. 3 shows an example of a mapping of a control parameter to the control parameter set according to one aspect; Fig. 4 shows the change in the output of an actuator resulting from the method. Detailed description of embodiments

[0024] Fig. Figure 1 shows the effects of directly inputting a continuous output signal from a software agent as a control parameter to an actuator. An input control parameter is shown on the x-axis, and a corresponding actuator output is shown on the y-axis. The outputs in areas a and c can be clearly distinguished. However, due to measurement uncertainties, the effect of the parameter changes in area b cannot be distinguished.

[0025] This will now be done with the Fig. 2 procedures shown as a flow chart.

[0026] Here, in step S1, a control parameter set with a plurality of discrete control parameters is first determined for an actuator. Such a determination is carried out in such a way that a distance between the control parameters is determined such that inputting two adjacent control parameters to the actuator results in outputs with spaced value ranges at the actuator.

[0027] This is exemplified in Fig. 4. In the left figure of the Fig. Figure 4 again shows the result of an actuator's output when a continuously changing input signal is input, and thus when a software agent's output is input directly into the actuator. Due to measurement uncertainties and similar factors at the actuator, the effects of a change in the input parameter on the actuator's output cannot be differentiated and thus cannot be determined.

[0028] If the control parameters of the control parameter are selected in such a way that the input of two adjacent control parameters into the actuator leads to outputs with spaced value ranges, the result is the Fig. 4, where a clear effect of entering two adjacent control parameters on the actuator output can be seen. This improves the actuator control, but at the same time, reinforcement learning is also improved, as the software agent does not learn false correlations.

[0029] To utilize the continuous output of the software agent, a control parameter generated by a software agent is received in step S2. This is then mapped to one of the control parameters of the control parameter set in step S3.

[0030] A possible method of mapping should be described with reference to Fig. 3. In section a of the Fig. 3 shows the action space of the software agent. This can assume values ​​between x and y. The action space of a software agent depends on the algorithm of the software agent and is known in advance. This action space is then normalized in the step shown in section b. In this case, a mapping to the value range between 0 and 1 is performed. Then, by multiplying by a value resulting from the number of control parameters in the control parameter set, a mapping to an index in the control parameter set is performed. If necessary, a corresponding correction factor, which is preferably in the range 0 < Δ << 0.01, can be deducted. Mathematically, a truncation is also performed in order to obtain only the integer part of the index after multiplying by the number of control parameters in the control parameter set.Then, based on the corresponding index in the control parameter set, the corresponding control parameter in section d in the . Fig. 3 selected.

[0031] The control parameter thus obtained can be used in step S4 to control the actuator based on the control parameter.

[0032] Accordingly, this method allows the use of any reinforcement learning method, including those with continuous output, for controlling actuators.

[0033] Furthermore, the method presented here also includes inputting the actuator output as input to the software agent for performing reinforcement learning, as shown in step S5. For this purpose, in step S5, parameters are determined as the actuator's response to the input. Such parameters represent the actuator's behavior. These parameters are then input back into the software agent, which learns the actuator's behavior by repeatedly outputting one or more control parameters and evaluating the response based on the received parameters.

[0034] Discretization improves learning because no false correlations, which can result from measurement inaccuracies, for example, are learned. Reference symbol S1 Determining a control parameter set S2 Receiving a control parameter; S3 Mapping the control parameter to one of the control parameters S4 Control of the actuator S5 Enter the output of the actuator as input to the software agent

Claims

[1] Control of an actuator based on control parameters generated by reinforcement learning, comprising the steps: - Determining (S1) a control parameter set with a plurality of discrete control parameters for the respective actuator; - receiving (S2) a control parameter generated by a software agent; - mapping (S3) the control parameter to one of the control parameters of the control parameter set; and - Control (S4) of the actuator based on the control parameter, where - in the determining step (S1), a distance between the control parameters is determined such that an input of two adjacent control parameters into the actuator leads to outputs with spaced value ranges at the actuator. [2] The method according to claim 1, wherein - the mapping step (S3) comprises the following steps: - Normalizing the control parameter based on a previously known action space of the software agent; - Multiplication of the standardized control parameter by a value based on the number of control parameters in the control parameter set; - removing decimal places; and - Select the control parameter located at this position in the control parameter set. [3] Method according to one of the preceding claims, further comprising: - Inputting (S5) the output of the actuator as input to the software agent to perform reinforcement learning. [4] Method for improved training of a software agent for controlling an actuator based on control parameters generated by reinforcement learning, comprising the steps: - Determining (S1) a control parameter set with a plurality of discrete control parameters for the respective actuator; - receiving (S2) a control parameter generated by a software agent; - mapping (S3) the control parameter to one of the control parameters of the control parameter set; - Controlling (S4) the actuator based on the control parameter; - Determining parameters as the actuator's response to the input; - inputting (S5) the parameters into the software agent of the reinforcement learning system, wherein the software agent is programmed such that a desired behavior of the parameters of the actor leads to an increased reward; and - Repeatedly performing the step of receiving (S2), mapping (S3), controlling (S4), determining parameters, and inputting (S5) until a termination condition is reached, wherein in the step of determining (S1) a distance between the control parameters is determined such that an input of two adjacent control parameters into the actuator leads to outputs with spaced value ranges at the actuator. [5] Apparatus for carrying out the method according to any one of the preceding claims.

Citation Information

Patent Citations

  • Technical system i.e. gas turbine, controlling and / or regulating method, involves selecting learning and / or optimization procedure suitable for technical system, and regulating and / or controlling system by selected procedure

    DE102007001024A1