Method of controlling an actuator by a control system and control system of an actuator

The method improves actuator control reliability by using fuzzy sets and reinforcement learning to generate consistent control rules, addressing the inconsistency issues in hybrid machine learning algorithms.

FR3144327B1Active Publication Date: 2025-12-05THALES SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2022014407
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-12-23
Publication Date
2025-12-05
Estimated Expiration
2042-12-23

AI Technical Summary

Technical Problem

Existing control systems for actuators face reliability issues due to inconsistent rules generated by hybrid machine learning algorithms, making it difficult to identify and address malfunctions effectively.

Method used

A method involving fuzzy sets and reinforcement learning is employed to generate a control law for actuators, where fuzzy sets are modified based on truth values and temporal difference errors, ensuring consistent rule generation and robust control.

Benefits of technology

This approach enhances the reliability and robustness of actuator control by ensuring consistent rule definitions, reducing errors, and optimizing memory usage through redundant rule elimination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000022_0000
    Figure 00000022_0000
  • Figure 00000023_0000
    Figure 00000023_0000
  • Figure 00000024_0000
    Figure 00000024_0000
Patent Text Reader

Abstract

Method for controlling an actuator by a control system and control system for an actuator. The present invention relates to a method for controlling an actuator comprising: - a phase for generating a control law for the actuator with a performance objective in mind, the generation phase comprising the steps of: - collecting signals from measurement sensors, - creating fuzzy sets corresponding to the collected signals, - generating the associated rule or creating new fuzzy sets, - implementing reinforcement learning on the fuzzy sets including an operation to modify the position of the center and the width of the fuzzy sets, - converting the result of the reinforcement learning into a set of control rules forming the control law, and - a control phase during which the control system applies the control law. Figure for the abstract: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method for controlling an actuator by a control system and control system for an actuator

[0001] The present invention relates to a method of controlling an actuator by a control system. It also relates to a control system for an actuator.

[0002] In the field of control systems, it is common to seek to control a process, often based on physical quantities, with the objective of regulating the behavior of a system.

[0003] The development of transparent and interpretable artificial intelligence models has become essential in recent years, as AI is becoming an indispensable tool in a growing number of fields. These models allow for superior performance compared to models requiring complete knowledge of a system.

[0004] Trustworthy artificial intelligence models are increasingly critical in critical or sensitive areas. To inspire confidence in the system operator or user, two properties are required: the user's ability to anticipate the system's decisions and the ease of understanding a result or decision made by the system that does not conform to what was expected. This implies that the artificial intelligence model created must be interpretable.

[0005] To this end, it is known to perform machine learning of an interpretable model for control systems. Several machine learning algorithms are applicable to these systems, including evolutionary algorithms and reinforcement learning algorithms.

[0006] In particular, hybrid algorithms are used between rule-based models and evolutionary or reinforcement learning algorithms.

[0007] However, in practice, it is observed that similar rules obtained by these algorithms are not consistent with each other.

[0008] Thus, for the following two rules:

[0009] - Rule 1: IF low temperature AND high humidity THEN element speed slow,

[0010] - Rule 2: IF high temperature AND high humidity THEN element speed fast,

[0011] It is observed that the definition of high humidity is not the same between the two rules.

[0012] Furthermore, due to the overlapping of the rules, it is not possible to determine easily identify the causes of the control system malfunction.

[0013] There is therefore a need for a method of controlling an actuator by a control system with better reliability.

[0014] To this end, the description describes a method for controlling an actuator by a control system, the control method comprising:

[0015] - a phase of generating a control law for the actuator, the law of command controlling the actuator to achieve a performance target based on values ​​taken by measuring sensors, the generation phase including the steps of:

[0016] - collection of signals from measurement sensors,

[0017] - creation of fuzzy sets corresponding to the collected signals,

[0018] - when a truth value associated with a fuzzy set exceeds a threshold value, generate the associated rule and otherwise, create new fuzzy sets,

[0019] - implementation of reinforcement learning on fuzzy sets including multiple iterations of applying an action and updating action-state values, to determine the actions to be taken to achieve the performance objective based on fuzzy sets,

[0020] the implementation comprising at each iteration an operation of modifying at least one of the center position and the width of the fuzzy sets to obtain modified fuzzy sets on which reinforcement learning is applied at the next iteration,

[0021] - conversion of the result of reinforcement learning into a set of control rules forming the control law, and

[0022] - a control phase during which the control system applies the law of control on the actuator.

[0023] A truth value associated with a fuzzy set represents the probability that the fuzzy set is true. In binary logic, the truth value would be 0 or 1, whereas in fuzzy logic, as in the case of the present invention, the truth value is a real value between 0 and 1.

[0024] According to particular embodiments, the control method has one or more of the following characteristics, taken individually or in all technically possible combinations:

[0025] - during the modification operation, both the position of the center and the width of Each fuzzy set is modified.

[0026] - during the modification operation, the errors of time differences are calculated the porelles of each of the different state-action associations and the modified position of the center depend on the calculated temporal difference errors.

[0027] - the modified position of the center is the current position of the center ad- summation of the arithmetic mean of the temporal difference errors.

[0028] - during the modification operation, the expected value of the time differences is calculated pores of each of the different state-action associations and the modified width is a function of the expected temporal differences.

[0029] - the modified width is the sum of the current width plus the average arithmetic of expected values ​​of time differences.

[0030] - the implementation further includes an operation to test the usefulness of the sup pressure of each fuzzy set, the fuzzy set being removed if the test is successful.

[0031] The description also relates to a control system adapted to implement:

[0032] - a phase of generating a control law for the actuator, the law of command controlling the actuator to achieve a performance target based on values ​​taken by measuring sensors, the generation phase including the steps of:

[0033] - collection of signals from measurement sensors,

[0034] - creation of fuzzy sets corresponding to the collected signals,

[0035] - when a truth value associated with a fuzzy set exceeds a threshold value, generate the associated rule and otherwise, create new fuzzy sets,

[0036] - implementation of reinforcement learning on fuzzy sets including multiple iterations of applying an action and updating action-state values, to determine the actions to be taken to achieve the performance objective based on fuzzy sets,

[0037] the implementation comprising at each iteration an operation of modifying at least one of the center position and the width of the fuzzy sets to obtain modified fuzzy sets on which reinforcement learning is applied at the next iteration,

[0038] - conversion of the result of reinforcement learning into a set of control rules forming the control law, and

[0039] - a control phase during which the control system (12) applies the law of control on the actuator.

[0040] In this description, the expression "specific to" means interchangeably "suited for", "adapted to" or "configured for".

[0041] Some features and advantages of the invention will become apparent from the following description, given solely by way of non-limiting example, and made with reference to the accompanying drawings, in which:

[0042] - [Fig.1] [Fig.1] is a schematic representation of an automated system,

[0043] - [Fig.2] [Fig.2] is a schematic representation of the different positions of a pendulum,

[0044] - [Fig.3] [Fig.3] graphically illustrates the creation of a new fuzzy set in the context of the pendulum in [Fig.2],

[0045] - [Fig.4] [Fig.4] is a schematic representation of a modification operation Determining the position of the center and the width of fuzzy sets in the implementation of a DFQL algorithm,

[0046] - [Fig.5] [Fig.5] is a schematic representation of a modification operation fication of the position of the center and the width of fuzzy sets in the implementation of a part of a control method for a part of the system of [Fig. 1], and

[0047] - [Fig.6], [Fig.6] is a schematic representation of a system and a product computer program.

[0048] A control method is now described.

[0049] The control method is a method for controlling an actuator 10 by a system control 12.

[0050] The actuator 10 is, in the context of automated systems, an element specifically designed to perform an action.

[0051] For example, an actuator 10 is a cylinder, a rotor or an effector.

[0052] The control system 12 is a system specifically designed to control the actuator 10, in particular a controller. Control is achieved by sending a control signal.

[0053] The actuator 10 and the control system 12 are part of an automated system 14 as schematically represented in [Fig.1].

[0054] The automated system 14 also includes measuring sensors 16 (a variant with a single sensor is possible) and a comparison unit 18.

[0055] The measuring sensors 16 are suitable for measuring elements relating to the environment of the actuator 10 or to its operation.

[0056] The output of the measuring sensors 16 is compared to a setpoint C which is sent to the comparison unit 18.

[0057] In the example proposed, the comparison unit 18 is a subtractor suitable for obtaining the difference between the setpoint C and the outputs of the measuring sensors 16.

[0058] It is assumed here that the setpoint C and the outputs of the measuring sensors 16 have been made comparable by a processing ensuring that these elements are expressed in the same space.

[0059] For example, if the setpoint C is a voltage setpoint and the measuring sensor 16 is a temperature, either the setpoint C is converted into temperature or the temperature signal is converted into voltage.

[0060] Some specific examples of control of an automated system 14 are given below:

[0061] - in the case of anti-lock brakes, brake control in situations manageable depending on the speed and acceleration of the car and the speed and wheel acceleration,

[0062] - in the case of automatic transmission, fuel injection control and Ignition timing depends on throttle position, coolant temperature, or engine speed.

[0063] - in the case of automatic vehicles, speed control based on the engine load, driving style and road conditions,

[0064] - in the case of a photocopier, the control of the drum tension according to image density, humidity, and temperature,

[0065] - in the case of cruise control, the accelerator control to adjust the speed and acceleration of the car, or

[0066] - in the case of a dishwasher, the control of the cleaning cycle, the strategies of rinsing and washing depending on the number of dishes and the amount of food on the dishes.

[0067] Examples could also be derived from other very different contexts such as microwave ovens, computers, engraving, trains or autonomous robots, to name just a few.

[0068] Before describing the control method further, it is appropriate to introduce some concepts on which the present method is based.

[0069] It is possible to describe an automated system according to its state on the one hand and the actions that the system can perform.

[0070] In such a formalism, each action-state pair is associated with a real number which is called the q-value (or sometimes the q-value). The q-value measures the quality of an action performed in a given state of the system.

[0071] It can be noted that for cases with a limited number of pairs, it is possible to establish tables of q values.

[0072] This is no longer possible when the number of states becomes too large. This is particularly the case when the actions and / or states are continuous and not discrete. Therefore, an approximation function is used to limit the memory space occupied and the processing time. This approximation function then gives a value q even for action-state pairs that are not observed in practice.

[0073] A fuzzy inference system allows operation in the case of continuous spaces for actions and / or states. A fuzzy inference system is often referred to by the abbreviation FIS, in reference to the corresponding English term "Fuzzy Inference System." Therefore, in the following, such a system will be referred to as an FIS system.

[0074] More precisely, fuzzy logic allows the conversion of numerical data into semantic data. This means that fuzzy logic allows the conversion of a numerical value into a degree of truth of belonging to a semantic value.

[0075] Thus, a numerical data point U is translated into a truth value zr( of the variable semantics 77 77'J designates the j-th semantic variable of the i-th rule.

[0076] The indices i and j are here two non-null natural numbers. More precisely, with regard to the index i, this index varies from 1 to n, n being the number of rules.

[0077] A FIS system is a set of fuzzy rules, the fuzzy rules being conjunctions of semantic variables.

[0078] In general, an indexed fuzzy rule i is written in the following form:

[0079] IF premise THEN conclusion

[0080] Where SI / ALORT denotes the IF / THEN operation in English terminology.

[0081] The premises are the conjunction of fuzzy sets characterizing a property of each coordinate of the input vector.

[0082] The premise of a rule i can generally be written as:

[0083] Siest77 AND ... AND snest

[0084] In this expression “AND” denotes the AND operation in English terminology.

[0085] Furthermore, s corresponds to the state vector of the environment. Each coordinate of the vector s is denoted Si, so that Si is the first coordinate of the state vector s and sn the nth coordinate of the state vector s.

[0086] The state vector s is received by collecting data from measurement sensors 16.

[0087] The strength of a premise is defined as the degree of membership of s in the fuzzy set resulting from the conjunction of the fuzzy sets of the premise. This is written mathematically as

[0088] 9;=

[0089] The conclusions vary depending on the actions chosen for each of these premises. The conclusion is generally written as follows:

[0090] a is Ki

[0091] Ki here denotes the value of the action. For example, if the action is to increase the temperature, Ki takes a value corresponding to the temperature increase to be applied to the environment.

[0092] In summary, the notation used here reflects the fact that a is the value corresponding to the conclusion of rule i.

[0093] It is conceivable that the FIS system uses a reinforcement algorithm, and more specifically a Q-learning algorithm better known by the corresponding English name of “Q-Learning”.

[0094] Q-Learning allows learning a strategy that indicates which action to take in each state of the system. Q-Learning works by learning an action-state value function, denoted Q, which determines the potential gain, that is, the long-term reward obtained by choosing a certain action in a certain state while following an optimal policy. When this value function

[0095]

[0096]

[0097]

[0098]

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105] If the action-state is known or learned by the agent, the optimal strategy can be constructed by selecting the maximum value action for each state, that is, by selecting the action a that maximizes the value Q(s,a) when the agent is in state s. Because this algorithm is used for fuzzy sets, it is common to say that the FIS system uses an FQL algorithm. The abbreviation FQL stands for "Fuzzy Q-Learning," which could be literally translated as Q-learning for fuzzy sets. The FQL algorithm is thus an approach to learning a set of fuzzy rules by reinforcement. In this approach, the agents represent different rules that are evaluated by different q values. Furthermore, FQL uses zero-order Takagi-Sugeno FIS. The order refers to the degree of the conclusions, which are polynomials in the coefficients of the input vector. Since the conclusions are constants, they are directly associated with a sample of actions A = {Ab, ..., AZ] taken from the set of possible actions in the environment. This sample is rule-independent, ensuring that the q-values ​​in the lookup table are associated with the premises in the rows and with each action in the sample in the columns. Finally, the activation function takes the form: 5x5^0. 1] (¾ .v) y, Vi is the strength of the premise mentioned above. It is expressed in terms of the operator dedicated to conjunction, for example x. In this case, _ || where The sk are the coordinates of the input vector. These mathematical elements correspond to the fact that the actions (the conclusions of the rules) are constants. The DFQL algorithm is an improvement on this algorithm, using a self-tuning FIS system based on reinforcement signals. The abbreviation DFQL stands for "Dynamic Fuzzy Q-Learning," which could be literally translated as dynamic Q-learning for fuzzy sets. The DFQL algorithm is thus an extension of the FQL algorithm allowing the creation of the set of rules of the FIS system during training. Within the DFQL algorithm, actions are associated with premises based on q-values ​​(or q). These values ​​are calculated throughout the learning phase in a lookup table (more often referred to by its English name) which contains the actions in column and row premises. This table constitutes Q. The activation function calculates the degree of truth of the premise.

[0106] Thus, unlike the FQL algorithm, the DFQL algorithm does not involve the intervention of a business expert either at the start or during training and only creates the rules that seem useful.

[0107] For this, instead of relying on a set S fixed at the start, this is a variable set St which is continuously adapted during additional steps compared to Q-Leaming.

[0108] This DFQL algorithm comprises six successive steps.

[0109] These six successive steps are as follows:

[0110] - first step: observation of the data,

[0111] - second step: generation of rules where applicable,

[0112] - third step: selection and application of an action,

[0113] - fourth step: updating the q values,

[0114] - fifth step: refinement of the fuzzy sets used by the premises, and

[0115] - sixth step: removal of fuzzy sets deemed unnecessary if any exist.

[0116] To better understand this sequence of six steps, it will now be detailed through a particular example.

[0117] This example is that of controlling a pendulum. The goal of this example is to bring the pendulum to equilibrium (shown as a dashed line in [Fig.2]) from a random position and to maintain it in equilibrium by applying a force to the pendulum.

[0118] Such an example corresponds to an environment in which both inputs and actions are continuous.

[0119] The first step aims to observe data to understand the state of the environment.

[0120] In this example, there are three input data.

[0121] The first input is the cosine of the angle between the pendulum and the equilibrium position. This first input is thus written as cos( 6) ■

[0122] The second input is the sine of the angle between the pendulum and the equilibrium position. This second input is thus written as sin(8)-

[0123] The third input data is the angular velocity of the pendulum. This third input data is thus written as 0.

[0124] These are transmitted to the learning agent.

[0125] The learning agent represents the FIS. The learning agent takes the inputs and decides on an action and, depending on the reward, it updates the function Q.

[0126] The second step corresponds to a generation of rules where applicable.

[0127] This second step is actually implemented only if the data The input data does not trigger any rules. More specifically, the condition for implementing the second step is that the values ​​of the input data do not allow a certain truth value threshold to be exceeded for a dimension of the state.

[0128] In such a scenario, fuzzy sets corresponding to the input data are created. The newly created fuzzy set is centered on the input value corresponding to the dimension.

[0129] An example is shown in Figure 3 where the two fuzzy sets present for the dimension sin(0) do not allow reaching the minimum truth value threshold (left graph in Figure 3). This leads to the creation of a new fuzzy set centered on the value sin(0) received as input (right graph in [Fig. 3]).

[0130] During the third step, reinforcement learning via the q value allows us to obtain which action is to be implemented while fuzzy rules indicate the associated weighting.

[0131] In this case, the F value of -2 is weighted by 0.5, the T value of 0 is weighted by 0.1 and the T value of 2 is weighted by 0.5.

[0132] With the previous notations, the values ​​of -2, 0 or 2 correspond to a K( while the values ​​of 0,1 or 0,5 correspond to a .

[0133] The fourth step consists of an update.

[0134] More specifically, the q values ​​are updated to adapt the choice of action made in the third step according to the reward received.

[0135] The reward is, for example, defined by experts in the field according to the objective sought.

[0136] The fifth refinement step consists of modifying the centers and widths of the fuzzy sets in order to better match the environment to be controlled.

[0137] With reference to [Fig.4] (top part), the centers of the updates of the fuzzy sets are averaged for each rule in which the fuzzy set intervenes.

[0138] We can denote Pu the position of the center corresponding to the set at time t and Pjt the position of the center corresponding to the set 77 at the same time t.

[0139] In the present case, for illustrative purposes, it is assumed that Pjj ~ Pjj.

[0140] The fifth step allows us to obtain respectively the position of the center Pij+\ corresponding to the set at time t+1 and that of the center Pp+v corresponding to the set at the same time t+1.

[0141] As can be seen in Figure 4 (upper part), a correction specific to each center is applied, which is mathematically written for the position of the center corresponding to the set: l°'«l

[0143] In this equation, (Ap),f corresponds to the value of the position update of the center of the Gaussian corresponding to the set at time t.

[0144] Using a similar notation, it also comes down to the position of the center corresponding to set 77:

[0145]

[0146] Because the corrections (Ap) 7 and (Ap) are not the same, at time t+1, the position of the center corresponding to the set is different from the position of the center ^: / +1 corresponding to the set ^1.

[0147] With reference to [Fig.4] (lower part), in the same way as the centers, the widths are updated but with an additional constraint on the update to control distinguishability.

[0148] We can note the width corresponding to set 1 at time t and the width corresponding to set 77 at the same time t.

[0149] In the present case, for illustrative purposes, it is assumed that — .

[0150] The fifth step allows us to obtain respectively the width corresponding to the set l at time t+1 and the width corresponding to the set at the same time t+1.

[0151] As can be seen in Figure 4 (lower part), a correction specific to each width is applied, which is mathematically written for the width corresponding to set 77:

[0152] <7,,+i= a^+ (Acr)^

[0153] In this equation, ( Aor) corresponds to the value of the update of the width of the Gaussian corresponding to the set 77 at time t.

[0154] With a similar notation, it also comes down to the width corresponding to set 77:

[0155] cr / f+i = (Arr)

[0156] Because the corrections (Aoj^et (A <j)^.f ne sont pas les mêmes, à l’instant t+1, la largeur 77correspondant à l’ensemble est différente de la largeur correspondant à l’ensemble

[0157] The sixth step consists of removing fuzzy sets deemed unnecessary after the fifth refinement step.

[0158] The deletion uses, for example, a criterion based on half the size (i) of the set of possible values ​​of the variable j when it is an interval [aj, . This half the size CJ is written:

[0159] cj = (bj-aj) / 2

[0160] The suppression criterion is that Çc,) > 3d where 3d is a suppression threshold.

[0161] This corresponds to the fact that a fuzzy set is modeled by a Gaussian. If This Gaussian is too wide, that is to say it is above a threshold, the fuzzy set is removed because this set will disrupt the system.

[0162] This criterion amounts to increasing the width of the fuzzy set. If the fuzzy set is too wide, it disturbs neighboring fuzzy sets and must therefore be removed.

[0163] The process which will now be described uses the previous steps with some refinements.

[0164] The method includes a phase of generating a control law for the actuator 10 and a control phase.

[0165] The generation phase aims to obtain a control law controlling the actuator 10 with a view to a performance objective from values ​​taken by measuring sensors 16.

[0166] The generation phase includes a collection step, a creation step, a generation step, an implementation step and a conversion step.

[0167] During the collection step, signals from the sensors are collected.

[0168] During the first creation step, fuzzy sets corresponding to the collected signals are created.

[0169] Thus, during the first creation step, the center is the value received from the collected signal (that from the measurement sensors 16) and the width is fixed to a predetermined value.

[0170] This predetermined value is, for example, chosen according to the number of linguistic variables desired for each rule.

[0171] Similar to what has been described previously for the algorithm In DFQL, when a truth value associated with a fuzzy set exceeds a threshold value, a step is implemented to generate the associated rule, and otherwise, new fuzzy sets are created.

[0172] Still with reference to the DFQL algorithm, during the implementation step, reinforcement learning with fuzzy sets is then implemented, comprising several iterations of an application of an action and updates of state values ​​to determine the actions to be performed to obtain the objective as a function of the fuzzy sets.

[0173] The reader is invited to refer to the initial concepts for more information on the operations of this step.

[0174] However, in the example described, the preceding fifth step, that is, the operation of modifying at least one of the center position and the width of the fuzzy sets, is implemented differently.

[0175] By way of illustration, it is assumed in a non-limiting manner in what follows, that both the position of the center and the width are modified.

[0176] Unlike the case of the DFQL algorithm where the modification operation is While the modification operation differs depending on the rules, it is common to all rules using the fuzzy set.

[0177] More precisely, the centers p and the widths o of the fuzzy sets are modified according to the reward received by applying the action a in the state

[0178] The first consists of harmonizing the updates of all the modifications made to the Pij and averaging them, as shown in [Fig.5].

[0179] More specifically, for the case of the center (upper part of [Fig.5]), the temporal difference errors of each of the different action-state associations are calculated and the modified position of the center depends on the calculated temporal difference errors.

[0180] The modified position of the center is then the sum of the current position of the center plus the arithmetic mean of the time difference errors.

[0181] Using the notation corresponding to that used for the description of Figure 4, the position of the center corresponding to the set at time t+1 is written:

[0182] (Ap) f+(AP) Pi.M~ Pü* 2

[0183] In this equation, (Ap) corresponds to the value of the update of the position of the center of the Gaussian corresponding to the set at time t while (Ap)^.? corresponds to the value of the update of the position of the center of the Gaussian corresponding to the set at time t.

[0184] Using a similar notation, it also applies to the position of the center corresponding to the set:

[0185] (Ap).+(Ap) / t Pj.t+ 2-----

[0186] Even if the corrections (Ap) and (Ap) are not the same, at time t+1, the position of the center Pj^+i corresponding to the set is identical to the position of the center Pij+i corresponding to the set

[0187] Similarly, for the case of width (bottom part of [Fig.5]), the width is also modified as a function of the expected temporal differences.

[0188] More precisely, here, the modified width is the sum of the current width plus the arithmetic mean of the time difference expectancies.

[0189] Using the notation corresponding to that used for the description of Figure 4, the width corresponding to the assembly at time t+1 is written:

[0190] (A <r).+^).,

[0191] In this equation, (Acr) corresponds to the value of the update of the width of the Gaussian corresponding to the ensemble at time t while (Acr) corresponds to the value of the update of the width of the Gaussian corresponding to the set at time t.

[0192]

[0193]

[0194]

[0195]

[0196]

[0197]

[0198]

[0199]

[0200]

[0201]

[0202]

[0203]

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210] Using a similar notation, it also applies to the width corresponding to the whole: (^),-.,+(^) / , Even if the corrections (Acr) and (A <t) ne sont pas les mêmes, à l’instant t+1, la largeur correspondant l’ensemble est identique ^ +1 According to an alternative embodiment, it is also proposed to take into account distinguishability during the update by modulating the effects on centers and widths by a coefficient. For example, the following formula can be used: / Pu / >+1 / / . \* / \ / pi; \ In ¢7,-. / ) / JT \ 0 (AlnH) / \ ln p., / With, : (a; y) six <y 1 otherwise And 7T* = minijrr^p^ In this case, it is required that the distinguishability of fuzzy sets is at least 9a. This yields the set of actions to be performed to achieve the objective based on the fuzzy sets. During the conversion stage, this result of reinforcement learning is converted into a set of control rules forming the control law. This is simply a rewrite. The generation phase just described is relatively simple to implement since it does not require modeling of the physical system. It also does not involve using prior knowledge, and certainly not the expertise of a specialist. This generation phase makes it possible to obtain a control law with better reliability. To fully understand this, it is helpful to return to the case of the DFQL algorithm. In this case, the updates are different according to the rules and contribute to the differentiation of definitions of fuzzy sets having the same semantic meaning. The refinement of the DFQL algorithm therefore does not take into account the sharing of Fuzzy sets with multiple rules. From the first update, an initially shared fuzzy set yields two very similar sets, which significantly increases the number of rules, slows down calculations, and, most importantly, impoverishes the performance of the algorithm, which no longer associates a state coordinate with a given characteristic. Put another way, this means that if and are shared between premises i and j (that is, their membership function is equal) and their errors are different, THEN and are no longer shared in the next iteration.

[0211] In the described process, the average of the updates, according to the rules, allows us to have only one update common to all rules using the fuzzy set.

[0212] The method ensures that for a given semantic variable, only one fuzzy set is used to define it. This allows the rules to be verifiable and contributes to increasing the robustness of the control law.

[0213] Controlling the evolution of the width of widths also makes it possible to ensure the distinguishability of fuzzy sets and therefore of the associated semantic variables.

[0214] Distinguishing within the local partition is defined from the normalized membership function of the fuzzy set is distinguishable <^3 Sj, =

[0216] 77 is the normalized membership function that represents the proportion to which The conclusion associated with this fuzzy set, ultimately linked to the premise, is considered in relation to the others. If it is not distinguishable, then we cannot find the exact cause of an error; therefore, our contribution must verify distinguishability.

[0217] Here again, this allows the rules to be verifiable and contributes to increasing the robustness of the control law.

[0218] In addition, there is a gain in memory due to the elimination of redundant rules.

[0219] During the control phase, the control system 12 applies the control law thus obtained to benefit from its best robustness in the operation of the automated system.

[0220] To improve this effect, the implementation further includes a test operation of the usefulness of deleting each fuzzy set, the fuzzy set being deleted if the test is validated.

[0221] The process which has just been described can be implemented by a device 100 as shown in [Fig.6].

[0222] The interaction between the device 110 and the computer program product 112 enables the implementation of the generation phase of the control process, which is thus a phase implemented by computer.

[0223] Device 110 is a desktop computer. Alternatively, device 110 is a gold rack-mounted dining device, laptop, tablet, personal digital assistant (PDA) or smartphone.

[0224] In specific embodiments, the computer is adapted to operate in real time and / or is in an embedded system, in particular in a vehicle such as an aircraft.

[0225] In the case of [Fig.6], the device 110 comprises a computing unit 114, a user interface 116 and a communication device 118.

[0226] The computing unit 114 is an electronic circuit designed to manipulate and / or transform data represented by electronic or physical quantities in registers of the device 110 and / or memories into other similar data corresponding to physical data in register memories or other types of display devices, transmission devices or storage devices.

[0227] As specific examples, the computing unit 114 includes a single-core or multi-core processor (such as a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller and a digital signal processor (DSP)), a programmable logic circuit (such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device (PLD) and programmable logic arrays (PLAs)), a state machine, a logic gate and discrete hardware components.

[0228] The computing unit 114 includes a data processing unit 120 adapted for processing data, in particular by performing calculations, memories 122 adapted for storing data and a reader 124 adapted for reading computer-readable media.

[0229] The user interface 116 includes an input device 126 and an output device 128.

[0230] The input device 126 is a device enabling the system user to enter information or commands on the device 110.

[0231] In [Fig.6], the input device 126 is a keyboard. Alternatively, the input device 126 is a pointing device (such as a mouse, a touchpad and a graphics tablet), a speech recognition device, an eye tracker or a haptic (motion analysis) device.

[0232] The output device 128 is a graphical user interface, that is to say a display unit designed to provide information to the user of the device 110.

[0233] In [Fig. 1], the output device 128 is a display screen enabling a visual presentation of the output. In other embodiments, the output device is a printer, an augmented and / or virtual display unit, a loudspeaker or other sound-generating device for presenting the output in audible form, a unit producing vibrations and / or odors, or a unit adapted to produce a signal electric.

[0234] In a specific embodiment, the input device 126 and the output device 128 are the same component forming human-machine interfaces, such as an interactive screen.

[0235] The communication device 118 enables unidirectional or bidirectional communication between the components of the device 110. For example, the communication device 118 is a bus communication system or an input / output interface.

[0236] The presence of the communication device 118 allows that, in certain embodiments, the components of the device 110 are distant from each other.

[0237] The computer program product 112 includes a computer-readable medium 132.

[0238] The computer-readable medium 132 is a tangible device readable by the reader 124 of the computing unit 114.

[0239] In particular, the computer-readable medium 132 is not a transient signal in itself, such as radio waves or other freely propagating electromagnetic waves, such as light pulses or electronic signals.

[0240] Such a computer-readable storage medium 132 is, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any combination thereof.

[0241] As a non-exhaustive list of more specific examples, computer-readable storage medium 132 is a mechanically coded device, such as punched cards or embossed structures in a groove, a floppy disk, a hard disk, read-only memory (ROM), random-access memory (RAM), read-only erasable memory (EROM), electrically erasable and readable memory (EEPROM), a magneto-optical disk, static random-access memory (SRAM), a compact disc (CD-ROM), a digital multipurpose disc (DVD), a USB key, a floppy disk, a flash memory, a solid-state drive (SSD), or a PC card such as a PCMCIA memory card.

[0242] A computer program is stored on the computer-readable storage medium 132. The computer program comprises one or more sequences of stored program instructions.

[0243] Such program instructions, when executed by the data processing unit 120, cause the execution of process steps.

[0244] For example, the form of program instructions is a form of source code, a computer-executable form, or any intermediate form between source code and a computer-executable form, such as the form resulting from the conversion of Source code via an interpreter, assembler, compiler, linker, or locator. Alternatively, program instructions are microcode, firmware instructions, state definition data, integrated circuit configuration data (e.g., VHDL), or object code.

[0245] Program instructions are written in any combination of one or more languages, for example an object-oriented programming language (FORTRAN, C++, JAVA, HTML), a procedural programming language (C language for example).

[0246] Alternatively, the program instructions are downloaded from an external source via a network, as is particularly the case for applications. In this case, the computer program product includes a computer-readable data carrier on which the program instructions are stored or a data carrier signal on which the program instructions are encoded.

[0247] In each case, the computer program product 112 includes instructions that can be loaded into the data processing unit 120 and adapted to cause the process to be executed when executed by the data processing unit 120. Depending on the embodiment, the execution is carried out wholly or partly on the device 110, i.e. a single computer, or in a system distributed among several computers (in particular via the use of cloud computing).< / y>

Claims

Demands

1. A method for controlling an actuator (10) by a control system (12), the control method comprising: - a phase of generating a control law for the actuator (10), the control law controlling the actuator (10) with a view to a performance objective based on values ​​taken by measurement sensors (16), the generation phase comprising the steps of: - collecting signals from the measurement sensors (16), - creating fuzzy sets corresponding to the collected signals, - when a truth value associated with a fuzzy set exceeds a threshold value, generating the associated rule and otherwise, creating new fuzzy sets, - implementing reinforcement learning on fuzzy sets comprising several iterations of applying an action and updating the action-state values, to determine the actions to be performed to obtain the performance objective based on the fuzzy sets,the implementation comprising at each iteration a modification operation of at least one of the center position and the width of the fuzzy sets to obtain modified fuzzy sets on which reinforcement learning is applied at the next iteration, - conversion of the result of the reinforcement learning into a set of control rules forming the control law, and - a control phase during which the control system (12) applies the control law on the actuator (10), in which, during the modification operation, the expected time differences of each of the different state-action associations are calculated and the modified width is a function of the expected time differences.

2. A control method according to claim 1, wherein, during the modification operation, both the position of the center and the width of each fuzzy set is modified.

3. A control method according to claim 1 or 2, wherein, during the modification operation, the time difference errors of each of the different state-action associations are calculated and the modified position of the center depends on the calculated time difference errors.

4. A control method according to claim 3, wherein the changed position of the center is the sum of the current position of the center plus the arithmetic mean of the time difference errors.

5. A control method according to any one of claims 1 to 4, wherein the modified width is the sum of the current width plus the arithmetic mean of the expected time differences.

6. A control method according to any one of claims 1 to 5, wherein the implementation further comprises a test operation for the usefulness of removing each fuzzy set, the fuzzy set being removed if the test is validated.

7. Control system (12) adapted to implement: - a phase of generating a control law for the actuator (10), the control law controlling the actuator (10) with a view to a performance objective based on values ​​taken by measurement sensors (16), the generation phase comprising the steps of: - collecting signals from the measurement sensors (16), - creating fuzzy sets corresponding to the collected signals, - when a truth value associated with a fuzzy set exceeds a threshold value, generating the associated rule and otherwise, creating new fuzzy sets, - implementing reinforcement learning on fuzzy sets comprising several iterations of applying an action and updating the action-state values, to determine the actions to be performed to obtain the performance objective based on the fuzzy sets,the implementation comprising at each iteration a modification operation of at least one of the center position and the width of the fuzzy sets to obtain modified fuzzy sets on which reinforcement learning is applied at the next iteration, - conversion of the result of the reinforcement learning into a set of control rules forming the control law, and - a control phase during which the control system (12) applies the control law on the actuator (10), in which, during the modification operation, the expected values ​​of the time differences of each of the different state-action associations are calculated and the modified width is a function of the expected values ​​of, temporal differences.