System, method, and computer-readable storage medium for calibrating a feedback controller

By employing a Kalman filter to adaptively calibrate control parameters in real-time, the challenges of inefficient manual calibration and safety concerns in dynamic machines are addressed, resulting in improved control quality and stability.

JP7693110B2Active Publication Date: 2025-06-16MITSUBISHI ELECTRIC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024524508
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-16
Filing Date
2022-05-19
Publication Date
2025-06-16
Estimated Expiration
2042-05-19

AI Technical Summary

Technical Problem

Current methods for calibrating feedback controllers in dynamic machines are inefficient and often require manual intervention or iterative learning, which are not suitable for all applications, especially those requiring continuous control or prioritizing safety.

Method used

The use of a Kalman filter to adaptively calibrate control parameters in real-time, allowing for online updates and adjustments based on performance goals, while ensuring stability and safety through safety checks.

Benefits of technology

This approach enables efficient and automated calibration of feedback controllers, improving control quality and ensuring safety and stability in dynamic machine operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693110000029
    Figure 0007693110000029
  • Figure 0007693110000030
    Figure 0007693110000030
  • Figure 0007693110000031
    Figure 0007693110000031
Patent Text Reader

Abstract

A system for controlling the operation of a machine to perform a task is disclosed. The system submits a sequence of control inputs to the machine and receives a feedback signal. The system further determines, at each control step, a current control input for controlling the machine based on the feedback signal including a current measurement of a current state of the system by applying a control policy, and converts the current measurement to a current control input based on current values ​​of control parameters in a set of control parameters of the feedback controller by applying the control policy. Furthermore, the system can iteratively update a state of the feedback controller defined by the control parameters using a prediction model that predicts values ​​of the control parameters and a measurement model that updates the prediction values ​​and generates current values ​​of the control parameters that describe the sequence of measurements according to a performance goal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to control systems, and more specifically to a system for calibrating a feedback controller , Method, and Computer-Readable Storage Medium .

Background Art

[0002] Currently, there are various dynamic machines that can operate in unstructured and uncertain environments. In fact, these dynamic machines are inherently more complex in order to operate in unstructured and uncertain environments. Since dynamic machines are inherently complex and operate in increasingly unstructured or uncertain environments, the need to automate the design and calibration process of dynamic machines becomes more important. In particular, the control of many dynamic machines, such as autonomous vehicles or robots, generally involves various conflicting specifications and therefore requires a significant amount of manual calibration effort. Furthermore, calibration is usually performed during the production stage, and since the operating conditions of dynamic machines change throughout their lifespan, it is often difficult to adjust the controllers associated with dynamic machines at a later stage.

[0003]

[0004] There are several currently available methods aimed at automating controller calibration and enabling the controller to adapt to the operation and operating conditions of dynamic machines. However, these available methods focus on learning from human experts or iterative learning tasks through trial-and-error search. Therefore, these available methods may only be suitable for applications that are suitable for iterative learning. As an example, these available methods can be used in robots for manipulating objects. However, these available methods do not provide controller calibration in inherently more continuous control applications such as autonomous driving. Furthermore, trial-and-error search is often not suitable for machines that prioritize safety. Also, the requirement of having a human demonstrator limits the amount of automation.Therefore, there is a need for a system that can automatically calibrate a controller in an efficient and realizable manner.

Summary of the Invention

[0005] The object of some embodiments is to repeatedly calibrate a controller in real time and use the calibrated controller to control the operation of a machine. Examples of machines can include vehicles (e.g., autonomous vehicles), robot assemblies, motors, elevator doors, HVAC (Heating, Ventilating, and Air-Conditioning) systems, etc. Examples of the operation of a machine can include operating a vehicle according to a specific trajectory, operating an HVAC system according to specific parameters, operating a robotic arm according to a specific task, and opening and closing an elevator door, but are not limited thereto. Examples of controllers can include PID (Proportional Integral Derivative) controllers, optimal controllers, neural network controllers, etc. Hereinafter, "controller" and "feedback controller" may be used interchangeably without distinction as having the same meaning.

[0006] To calibrate a feedback controller, some embodiments use a Kalman filter. However, a Kalman filter is generally used when estimating state variables that define the state of a machine, which may be physical quantities such as position, velocity, etc. For this reason, the purpose of some embodiments is to transform or adapt a Kalman filter to estimate control parameters of a feedback controller for controlling a machine, as opposed to state variables of the machine. State variables define the state of the machine being controlled, while control parameters are used to calculate control commands. Examples of control parameters are gains of a feedback controller, such as gains in a PID controller, and / or parameters of the physical structure of the machine, such as the mass of a robotic arm, or the friction between the tires of a vehicle and the road. In particular, control parameters should not be confused with control variables that define the input and output to a control law or control policy executed by a feedback controller, such as the value of a voltage for controlling an actuator. In other words, input control variables are mapped to output control variables based on a control law defined by control parameters. The mapping may be analytical or may be based on a solution to an optimization problem.

[0007] In many control applications, the control parameters are known in advance and fixed, i.e., remain constant during control. For example, the mass of a robotic arm can be measured or known from the specifications of the robot, the friction of the tires can be limited or selected, and the gains of the controller can be adjusted in the laboratory. However, fixing the control parameters in advance may not be optimal in some applications, and in some other applications where it may be necessary to control a machine with control parameters having uncertainties, it may even be impractical.

[0008] Some embodiments are based on the recognition that the principle of tracking state variables provided by a Kalman filter can be extended or adapted to track control parameters. In fact, while control is a process rather than a machine, it is recognized that control can be treated as if it were a virtual machine having a virtual state defined by control parameters. According to this intuitive knowledge, a Kalman filter can iteratively track control parameters if the prediction model used by the Kalman filter during the prediction stage can predict control parameters for which the measured values of the state of the machine can be explained according to the measurement model.

[0009] In particular, since the prediction model and the measurement model are provided by the designer of the Kalman filter, this flexibility enables the Kalman filter to be adapted to different types of control objectives. For example, in some embodiments, the prediction model is a constant or identity model that predicts that the control parameter will not change within the variance of the process noise. In fact, such predictions are common to many control applications that use fixed control parameters. In addition or alternatively, some embodiments define a prediction model that can predict at least some parameters based on a predefined relationship with other parameters. For example, some embodiments can predict changes in tire friction based on the current speed of the vehicle. In this configuration of the Kalman filter, the process noise controls how fast the control parameter changes over time.

[0010] In any case, such a prediction model directs the main effort for tracking control parameters to the measurement model and adds flexibility to vary the update of the measurement model based on the control objective. In particular, such flexibility enables the measurement model to be varied to control different machines, but also enables the measurement model to be varied at different times or in different states during the control of the same machine.

[0011] Therefore, in various embodiments, the measurement model uses performance goals that evaluate online the performance of controlling the operation of a closed-loop machine, and then uses this to adapt control parameters to improve the operation of the closed-loop machine measured with respect to the performance goals. In particular, the performance goals have a very flexible structure and can be different from the objectives of an optimal controller. This is beneficial because the optimal control cost function has a restricted structure due to its real-time use. For example, the cost function often needs to be differentiable and convex to be suitable for numerical optimization. Additionally, the performance goals can change at different times of control according to the same optimal control objective. Furthermore, the optimal control objective or other control parameters can change as a function of the state of the machine at different times or according to the same performance goals.

[0012] Thus, the advantages of the Kalman filter are extended to the recursive estimation of control parameters. These advantages include that the Kalman filter can (i) adapt parameters online during machine operation, (ii) be robust to noise due to filter-based design, (iii) maintain the guarantee of the safety of closed-loop operation, (iv) be computationally efficient, (v) require reduced data storage due to recursive implementation, and (vi) be easy to implement, and therefore make it attractive for industrial applications.

[0013] Some embodiments are based on the recognition that in many applications, some control parameters need to be adjusted together while being dependent on each other. For example, the gains of a PID controller need to be adjusted together to obtain the desired performance and ensure safe operation, and the weights of the cost function for optimal control need to be adjusted together because they define the trade-off between multiple, sometimes conflicting objectives. The filter coefficients used in an H ∞ controller or a dynamic output feedback controller need to be adjusted together to ensure performance and stability requirements.

[0014] Generally, calibrating interdependent parameters is a more difficult problem because this interdependence adds another variable to be considered. Therefore, having multiple interdependent parameters to calibrate can increase the complexity of calibration. However, some embodiments are based on the recognition that such interdependence of calibrated control parameters can be adjusted statistically in a natural way by adjusting the Kalman gain that applies different weights to the updates of different parameters.

[0015] Some embodiments are based on the recognition that the control parameters used in a feedback controller depend on the state of the machine. Some embodiments address this state dependence by using a linear combination of basis functions that are a function of the state of the machine. In fact, a Kalman filter can be implemented to adjust the coefficients of the basis functions, which are then used to generate the control parameters. In addition or alternatively, some embodiments use state-dependent regions in combination with the basis functions. In each region, the control parameters are calculated as a linear combination of the basis functions. The Kalman filter can adjust both the coefficients of the basis functions in each region and the region that determines which set of basis functions is used to calculate the control parameters.

[0016] In different embodiments, the machine being controlled has different uncertainties in linear or non-linear dynamics and control parameters with different ranges. Some embodiments address these differences by selecting different types of implementations of the Kalman filter and / or different variances for the process and / or measurement noise.

[0017] For example, one embodiment uses an extended Kalman filter (EKF) to calculate the Kalman gain. The EKF numerically calculates the gradient of the performance objective with respect to the control parameters. The EKF is useful for problems where the performance objective is distinguishable with respect to the state of the machine because it calculates the gradient using two gradients: (i) the gradient of the performance objective with respect to the state of the machine and (ii) the gradient of the dynamic machine state with respect to the control parameters. The gradient of the performance objective with respect to the state of the machine is calculated by the designer. The gradient of the machine state with respect to the control parameters is calculated using a model that defines the structure of the feedback controller and the dynamics of the machine.

[0018] In addition or alternatively, one embodiment uses an unscented Kalman filter (UKF) to calculate the Kalman gain. The UKF estimates the gradient of the performance objective with respect to the control parameters using function evaluations of the performance objective. In this case, the UKF can calculate the sigma points that are the realizations of the control parameters. Then, the gradient is estimated using the evaluations of the performance objective for all the sigma points in combination with the joint probability distribution of the control parameters. The UKF is useful for differentiable and non-differentiable performance objectives because it estimates the gradient using function evaluations.

[0019] Some embodiments are based on the understanding that while online iterative updating of the control parameters of a feedback controller can improve the quality of control, it is accompanied by additional challenges. For example, online updating of the control parameters during the operation of a machine can introduce discontinuities in the control. However, some embodiments are based on the recognition that such discontinuities can be addressed by implementing control commands so as to satisfy the constraints on the operation of the machine. These constraints can be established by verifying the control parameters so as to satisfy established control-theoretic properties.

[0020] In addition to or instead of this, some embodiments are based on the recognition that online updating of control parameters may destabilize the operation of the machine. For example, if the control law or control policy is represented by a differential equation (ODE) with respect to the control parameters, changing the control parameters may compromise the stability of the equilibrium of the ODE. To address this new problem that may be introduced by the Kalman filter in different embodiments, some embodiments perform a safety check, e.g., a stability check on the control policy using the value of the control parameters generated by the Kalman filter. Further, the control parameters in the control policy may be updated only if the stability check is satisfied.

[0021] For example, the stability check is satisfied if there exists a Lyapunov function for the control policy with the updated control parameters. The existence of the Lyapunov function can be confirmed in a number of ways. For example, some embodiments solve an optimization problem aimed at finding and / or proving the existence of the Lyapunov function. In addition to or instead of this, one embodiment checks whether the updated control parameters result in a decrease in the cost of the state with respect to the performance goal over the entire history of the state and input. In addition to or instead of this, another embodiment checks whether the updated control parameters preserve the proximity of the machine to its origin. The proximity to the origin is preserved based on the recognition that the cost associated with the end point of the prediction horizon of the control policy with the updated parameters is bounded, e.g., by the ratio of the maximum eigenvalue to the minimum eigenvalue of a positive definite matrix defining the terminal cost.

[0022] In addition, some embodiments are based on the recognition that when the control parameters generated by the Kalman filter do not meet the safety check, the control parameters of the feedback controller should not be updated with the output of the Kalman filter, but the Kalman filter itself should not be restarted, and even if the control parameters of the Kalman filter are different from the control parameters of the feedback controller, the iteration should continue using the newly generated control parameters. If the control parameters of the Kalman filter meet the safety check during some of the subsequent iterations, the safe control parameters of the Kalman filter update the old control parameters of the feedback controller. In this way, the embodiments guarantee the stability of control in the presence of online updates of control parameters.

[0023] Accordingly, one embodiment discloses a system for controlling the operation of a machine to perform a task. The system includes a transceiver configured to submit a sequence of control inputs to the machine and receive a feedback signal including a corresponding sequence of measurements, where each measurement indicates a state of the machine caused by the corresponding control input. The system further includes a feedback controller configured to determine, at each control step, a current control input for controlling the machine based on a feedback signal including a current measurement of the current state of the machine by applying a control policy, where the feedback controller converts the current measurement into the current control input based on the current value of a control parameter within a set of control parameters of the feedback controller by applying the control policy. Further, the system includes a Kalman filter configured to iteratively update a state of the feedback controller defined by the control parameters by using a prediction model for predicting a value of a control parameter affected by process noise and a measurement model for updating a predicted value of the control parameter based on a sequence of measurements affected by measurement noise, to generate a current value of the control parameter that explains the sequence of measurements according to a performance goal.

[0024] Accordingly, another embodiment discloses a method for controlling the operation of a machine to perform a task. The method includes submitting a sequence of control inputs to the machine and receiving a feedback signal including a sequence of corresponding measurement values, where each measurement value indicates a state of the machine caused by a corresponding control input. The method further includes, in each control step, determining a current control input for controlling the machine based on the feedback signal including the current measurement value of the current state of the machine by applying a control policy, where applying the control policy converts the current measurement value into the current control input based on the current value of a control parameter within a set of control parameters of a feedback controller. The method further includes generating a current value of a control parameter that describes the sequence of measurement values according to a performance goal by repeatedly updating the state of the feedback controller defined by the control parameter using a prediction model that predicts the value of the control parameter affected by process noise and a measurement model that updates the predicted value of the control parameter based on a sequence of measurement values affected by measurement noise.

[0025] Accordingly, yet another embodiment discloses a non-transitory computer-readable storage medium having a program executable by a processor for implementing a method of controlling the operation of a machine to perform a task. The method includes submitting to the machine a sequence of control inputs and receiving a feedback signal including a sequence of corresponding measurement values, each measurement value indicating a state of the machine caused by a corresponding control input. The method further includes, in each control step, determining a current control input for controlling the machine based on the feedback signal including the current measurement value of the current state of the machine by applying a control policy, where applying the control policy converts the current measurement value into the current control input based on the current value of a control parameter within a set of control parameters of a feedback controller. The method further includes generating a current value of a control parameter that describes the sequence of measurement values according to a performance goal by repeatedly updating the state of the feedback controller defined by the control parameter using a prediction model that predicts a value of a control parameter affected by process noise and a measurement model that updates the predicted value of the control parameter based on a sequence of measurement values affected by measurement noise.

Brief Description of the Drawings

[0026]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 2E

Figure 2F

Figure 3

Figure 4A

Figure 4B

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 7

Figure 8A

Figure 8B

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0027] In the following description, for the sake of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent to one of ordinary skill in the art, however, that the present disclosure may be practiced without these specific details. In other instances, well-known devices and methods are shown in block diagram form only to avoid obscuring the present disclosure.

[0028] As used in this specification and the claims, the terms “for example,” “for instance,” and “such as,” as well as the verbs “comprising,” “having,” “including,” and other forms of these verbs, when used in conjunction with a listing of one or more components or other items, are to be construed as open-ended, meaning that the listing is not to be considered as excluding further components or items. The term “based on” means at least in part based on. Further, it should be understood that the style and terminology used in this specification are for the purpose of description and should not be considered limiting unless specifically defined. Any headings used in this specification are for convenience only and have no legal or limiting effect.

[0029] FIG. 1 shows an overview of the principle of a Kalman filter according to some embodiments of the present disclosure. The Kalman filter 100 is a process (or method) that uses a series of measurements observed over a period of time, including statistical noise and other inaccuracies, to generate an estimated value of an unknown variable. In fact, these generated estimated values of the unknown variable can be more accurate than the estimated values of the unknown variable generated using a single measurement. The Kalman filter 100 generates an estimated value of the unknown variable by estimating the joint probability distribution over the unknown variable.

[0030] In a scenario as a specific example, the series of measurement values used by the Kalman filter 100 can be the measurement values 102 associated with the state variables of the dynamic machine. Thus, in this scenario as a specific example, the Kalman filter 100 may be used in generating the state estimation 104 of the dynamic machine. As used herein, the state variable may be a variable that mathematically describes the "state" of the dynamic machine. The state of the dynamic machine describes the dynamic machine sufficiently to determine its future behavior (e.g., motion) in the absence of any external force affecting the dynamic machine. By way of example, the state estimation value 104 can be an estimated value of speed, position, and / or other such physical quantities. In fact, these state estimation values 104 are required in applications such as navigation guidance and control of vehicles, particularly aircraft, spacecraft, and dynamically positioned ships.

[0031] The Kalman filter 100 is a two-step process including a prediction step and an update step. In the prediction step, the Kalman filter 100 uses a prediction model to predict the current state along with those uncertainties governed by the process noise. By way of example, the prediction model may be artificially designed such that the prediction model takes into account the effect of the process noise (e.g., assumption 108) for reducing the uncertainty in the state while predicting the current state. In fact, the predicted current state may be represented by a joint probability distribution over the current state. In some example embodiments, the prediction model may use the model 106 of the dynamic machine to predict the current state. As used herein, the model 106 of the dynamic machine may be a mathematical formula that associates the state of the dynamic machine with (i) the previous state of the dynamic machine and (ii) the control input to the dynamic machine. Examples of the model 106 are as follows.

Number

[0032] In an update step, when the result of the next measurement (inevitably impaired with some error including random noise) is observed, the predicted state is updated according to a measurement model that is affected by measurement noise. Measurement noise can control the error in the measurement. Also, measurement noise may be included in assumption 108. The measurement model may be designed such that the measurement model aims to align the prediction with the measured value. For example, the measurement model may update the joint probability distribution over the current state using a weighted average, where more weight is given to more certain estimates.

[0033] The output of the Kalman filter 100 may be a state estimate value 104 that maximizes the likelihood of the received measurement value 102 of the state considering assumptions 108 about noise (e.g., process noise and measurement noise) and the model 106 of the dynamic machine. As an example, assumptions 108 about noise may include a mathematical noise model aimed at reducing the inaccuracy of the state and the measured value. The Kalman filter 100 is a recursive process that can be executed in real time using only the current measurement value and the previously calculated state and its uncertainty matrix, and no additional past information is required.

[0034] Some embodiments are based on the recognition that the principles provided by the Kalman filter 100 for estimating the state of a dynamic machine can be extended or adapted to estimate the virtual state of a virtual machine. In other words, the Kalman filter 100 for estimating the state of a dynamic machine can be extended to a Kalman filter 110 for estimating the virtual state of a virtual machine. In particular, since the prediction model and the measurement model are provided by the designer of the Kalman filter 100, this flexibility allows the Kalman filter 100 to be adapted or extended to the Kalman filter 110.

[0035] In many control applications, the control parameters that define the state of a controller may be known in advance and fixed, i.e., remain constant during the control of a dynamic machine. Examples of control parameters include the gain of a controller, such as the gain in a PID controller, and / or parameters of the physical structure of a dynamic machine, such as the mass of a robotic arm, or the friction between the tires of a vehicle and the road. For example, the mass of a robotic arm can be measured or known from the specifications of the robot, the tire friction can be limited or selected, and the gain of the controller can be adjusted in the laboratory. However, fixing the control parameters in advance may not be optimal in some applications, and rather, in some other applications where it may be necessary to control a dynamic machine with control parameters having uncertainty, it may even be impractical.

[0036] Therefore, an object of some embodiments is to extend or adapt the Kalman 100 to a Kalman filter 110 that estimates the control parameter 112 that defines the state of a controller. In these embodiments, the virtual state is the state defined by the control parameter, and the virtual machine is the controller. To extend the Kalman filter 100 to the Kalman filter 110, in the prediction step, the prediction model affected by the process noise may be adapted to predict the control parameter using the transition model 116 of the control parameter 112. The process noise in the Kalman filter 110 can control not how inaccurate the state is, but how fast the control parameter changes over time. Thus, the assumption 118 can be designed. Further, the transition model 116 may be artificially designed.

[0037] In an update step, a measurement model affected by measurement noise may be adapted to evaluate the performance of predicted control parameters in the control of a dynamic machine based on performance target 114. Further, the measurement model may be adapted to update predicted control parameters based on the evaluation. In particular, performance target 114 has a very flexible structure and may be different from the target of the controller.

[0038] Thereby, Kalman filter 110 can estimate control parameter 112 based on assumption 118 regarding how fast the control parameter changes in the presence of an error from performance target 114. In fact, the output of Kalman filter 110 is control parameter estimate 112 that maximizes the likelihood of received performance target 114 considering (i) assumption 118 and (ii) transition model 116. As an example, a control system using the principle of Kalman filter 110 is as described in the detailed description of FIG. 2A.

[0039] FIG. 2A shows a block diagram of a control system 200 for controlling the operation of a dynamic machine 202 according to some embodiments of the present disclosure. Some embodiments are based on the recognition that the purpose of the control system 200 is to control a dynamic machine 202 in an engineering process. For this reason, the control system 200 can be operably coupled to the dynamic machine 202. Hereinafter, "control system" and "system" may be used interchangeably without distinction as having the same meaning. Hereinafter, "dynamic machine" and "machine" may be used interchangeably without distinction as having the same meaning. Examples of the machine 202 may include vehicles (e.g., autonomous vehicles), robotic assemblies, motors, elevator doors, HVAC (heating, ventilation, and air conditioning) systems, and the like. For example, the vehicle may be an autonomous vehicle, an aircraft, a spacecraft, a dynamically positioned ship, or the like. Examples of the operation of the machine 202 may include, but are not limited to, operating a vehicle according to a specific trajectory, operating an HVAC system according to specific parameters, operating a robotic arm according to a specific task, and opening and closing an elevator door.

[0040] System 200 may include at least one processor 204, a transceiver 206, and a bus 208. Additionally, system 200 may include memory. The memory may be implemented as a storage medium such as RAM (Random Access Memory), ROM (Read Only Memory), a hard disk, or any combination thereof. By way of example, the memory can store instructions executable by at least one processor 204. The at least one processor 204 may be implemented as a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. The at least one processor 204 may be operably connected to the memory and / or the transceiver 206 via the bus 208. According to certain embodiments, the at least one processor 204 may be configured as a feedback controller 210 and / or a Kalman filter 212. Thus, the feedback controller 210 and the Kalman filter 212 may be implemented within a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. Alternatively, the feedback controller 210 may be implemented external to system 200 and may communicate with system 200. In this configuration, system 200 may be operably coupled to the feedback controller 210, and the feedback controller may be coupled to machine 202. For example, the feedback control unit 210 may be a PID (Proportional-Integral-Derivative) controller, an optimal controller, a neural network controller, or the like.

[0041] According to one embodiment, the feedback controller 210 may be configured to determine a sequence of control inputs for controlling the machine 202. For example, the control inputs may in some cases be associated with physical quantities such as voltage, pressure, force, torque, etc. In an example of an embodiment, the feedback controller 210 may determine a sequence of control inputs such that the sequence of control inputs changes the state of the machine 202 to perform a particular task, such as tracking a reference. Once the sequence of control inputs is determined, the transceiver 206 may be configured to submit the sequence of control inputs as the input signal 214. As a result, the state of the machine 202 may be changed according to the input signal 214 to perform a particular task. By way of example, the transceiver 206 may be an RF (radio frequency) transceiver or the like.

[0042] Furthermore, the state of the machine 202 can be measured using one or more sensors installed on the machine 202. The one or more sensors can send a feedback signal 216 to the transceiver 206. The transceiver 206 can receive the feedback signal 216. In an example of an embodiment, the feedback signal 216 may include a sequence of measured values corresponding to each of the sequence of control inputs. By way of example, the sequence of measured values may be measured values of the state output by the machine 202 according to the sequence of control inputs. Thus, each measured value in the sequence of measured values may indicate the state of the machine 202 caused by the corresponding control input. Each measured value in the sequence of measurements may in some cases be associated with physical quantities such as current, flow rate, speed, position, and / or others. In this way, the system 200 can iteratively submit a sequence of control inputs and receive a feedback signal. In an example of an embodiment, to determine the sequence of control inputs in the current iteration, the system 200 uses the feedback signal 216 that includes a sequence of measured values indicating the current state of the machine 202.

[0043] To determine the sequence of control inputs in the present iteration, the feedback controller 210 may be configured to determine a current control input for controlling the machine 202 based on a feedback signal 216 that includes a current measurement of the current state of the machine at each control step. According to one embodiment, to determine the current control input, the feedback controller 210 may be configured to apply a control policy. As used herein, a control policy may be a set of mathematical formulas that map all or a subset of the states of the state of the machine 202 to control inputs. This mapping may be analytical or may be based on a solution to an optimization problem. In response to the application of the control policy, the current measurement of the current state may be converted into a current control input based on the current values of the set of control parameters within the set of control parameters of the feedback controller 210. As used herein, the control parameters may be (i) the gains of the feedback controller 210 and / or (ii) the parameters of the physical structure of the machine 202. For example, if the feedback controller 210 corresponds to a PID controller, the set of control parameters includes the proportional gain, integral gain, and derivative gain of the PID controller. For example, the parameters of the physical structure of the machine 202 may include the mass of the robotic arm or the friction between the tires of the vehicle and the road. In particular, the control parameters should not be confused with the control input that is the output of the control policy. According to one embodiment, the current values of the control parameters may be generated by the Kalman filter 212. By way of example, the Kalman filter 212 that generates the control parameters is as described in the detailed description of FIG. 2B.

[0044] FIG. 2B shows a Kalman filter 212 for generating control parameters according to some embodiments of the present disclosure. FIG. 2B is described in relation to FIG. 2A. According to an embodiment, the Kalman filter 212 may be configured to iteratively update the state of the feedback controller 210. According to an embodiment, the state of the feedback controller 210 is determined by control parameters. Thus, the purpose of the Kalman filter 212 is to iteratively generate control parameters. In an example of an embodiment, the Kalman filter 212 may iteratively generate control parameters using a prediction model 218 and a measurement model 220. As an example, the prediction model 218 and the measurement model 220 may be artificially designed.

[0045] To generate the control parameters in the current iteration (e.g., at time step k), the prediction model 218 may be configured to predict the value of the control parameter using prior knowledge 218a of the control parameter. As an example, the prior knowledge 218a of the control parameter may be generated in the previous iteration (e.g., at time step k-1). The prior knowledge 218a of the control parameter may be a joint probability distribution (or Gaussian distribution) over the control parameter in the previous iteration. The joint probability distribution over the control parameter in the previous iteration is the mean θ k-1|k-1 , and variance (or covariance) P k-1|k-1 which can be determined by. As an example, the joint probability distribution in the previous iteration may be generated based on the joint probability distribution generated in a past iteration (e.g., at time step k-2) and / or a model of the feedback controller 210 (e.g., transition model 116).

[0046] According to one embodiment, the value of the control parameter predicted in the current iteration may be the joint probability distribution 218b (or Gaussian distribution 218b). As an example, the output of the prediction model 218 may be the joint probability distribution 218b when the prediction model 218 is configured to predict a plurality of control parameters. Alternatively, the output of the prediction model 218 may be the Gaussian distribution 218b when the prediction model 218 is configured to predict a single control parameter. As an example, the joint probability distribution 218a is the mean θ k|k-1 , and the variance (or covariance) P k|k-1 may be determined by. For example, while predicting a single control parameter, the Gaussian distribution output by the prediction model 218 is as shown in FIG. 2C.

[0047] FIG. 2C shows a Gaussian distribution 224 representing one specific control parameter according to some embodiments of the present disclosure. FIG. 2C is described in relation to FIG. 2B. The Gaussian distribution 224 may be predicted by the prediction model 218. As an example, the Gaussian distribution 224 may correspond to the Gaussian distribution 218b. The Gaussian distribution 224 is the mean 228 (for example, the mean θ k|k-1 ) and the variance 226 (for example, the variance P k|k-1 ) may be determined by, and the mean 228 determines the center position of the Gaussian distribution 224, and the variance 226 determines a measure of the spread (or width) of the Gaussian distribution 224.

[0048] Referring again to FIG. 2B, according to one embodiment, the prediction model 218 may be affected by process noise. As used herein, process noise may be an assumption (e.g., assumption 118) that defines how fast the control parameters change over time. The process noise can control how fast the control parameters change over time within the variance defined by the process noise. The process noise may be artificially designed. For example, if the prediction model 218 is affected by process noise, the prediction model 218 may output multiple Gaussian distributions for one particular control parameter, and the multiple Gaussian distributions may have different variances defined within the variance of the process noise. As an example, the multiple Gaussian distributions output by the prediction model 218 for one particular control parameter are as shown in FIG. 2D.

[0049] FIG. 2D shows Gaussian distributions 230, 232, and 234 having different variances, according to some embodiments of the present disclosure. FIG. 2D is described in relation to FIG. 2B. The Gaussian distributions 230, 232, and 234 may be predicted by the prediction model 218. Each of these Gaussian distributions 230, 232, and 234 may have a different variance from each other, but the means 236 of the Gaussian distributions 230, 232, and 234 may be constant. The Gaussian distribution with (i) a small variance and (ii) a mean 236 having the highest probability among the other Gaussian distributions may be a correct prediction of the control parameter. As an example, the Gaussian distribution 230 may represent a correct prediction of the control parameter.

[0050] Referring again to FIG. 2B, in this manner, the prediction model 218 subject to process noise may be configured to predict the value of the control parameter, which is output as a joint probability distribution 218b (or a Gaussian distribution 218b). Once the joint probability distribution 218b is output by the prediction model 218 in the current iteration, the measurement model 220 may be configured to update the predicted value of the control parameter for generating the current value of the control parameter based on the sequence of measurements 220a. In an example embodiment, the sequence of measurements 220a may be a sequence of measurements received by the transceiver 206. As an example, the sequence of measurements 220a used by the measurement model 220 is as shown in FIG. 2E.

[0051] FIG. 2E illustrates the evolution 238 of the state of the machine 202 over time, according to some embodiments of the present disclosure. FIG. 2E is described in conjunction with FIG. 2A and FIG. 2B. By way of example, the evolution 238 of the state of the machine 202 may be obtained from one or more sensors installed on the machine 202. By way of example, if the current time is t0, the measurement model 220 may use N state measurements 240 to update the predicted values ​​of the control parameters. The N state measurements 240 may correspond to the sequence of measurements 220a. The N state measurements 240 may be obtained from a time t0 in the past. -N Measurement value x associated with t-N Starting from t0, the measurement x associated with the current time t0 t0 2E, consider a measurement model 220 with N state measurements 240 for only one state. However, if the machine 202 is associated with more than one state, the measurement model 220 may use the N measurements for all states within the same time frame.

[0052] Referring back to FIG. 2B, some embodiments are based on the recognition that a sequence 220a of measurements obtained from one or more sensors may not be accurate due to sensor defects, other noise (e.g., random noise), etc. For this reason, the measurement model 220 may be made to be affected by measurement noise. As used herein, measurement noise is a noise model that can be used to reduce the inaccuracy of the measurements 220a caused by sensor defects, other noise, etc. By way of example, measurement noise can be artificially designed.

[0053] In an example of an embodiment, the measurement model 220 affected by measurement noise may be configured to update a predicted value of a control parameter based on the sequence 220a of measurements. To update the predicted value, the measurement model 220 may be configured to calculate a model mismatch between the sequence 220a of measurements and a model of the machine 202 (e.g., model 106). Further, the measurement model 220 may be configured to simulate the evolution (e.g., state measurement) of the machine 202 using the predicted control value, the model of the machine 202, and the calculated model mismatch. For example, the simulated evolution (i.e., state measurement) may be similar to the sequence 220a of measurements. Further, the measurement model 220 may be configured to evaluate the simulated evolution of the machine 202 according to the performance goal 220b and generate a current value of the control parameter. Since the current value of the control parameter is generated based on the evaluation of the simulated evolution, which may be similar to the sequence 220a of measurements, the current value of the control parameter can explain the sequence 220a of measurements. For example, the measurement model 220 that updates the predicted value of the control parameter is graphically shown in FIG. 2F.

[0054] FIG. 2F shows a schematic diagram 242 for updating a predicted value of a control parameter according to some embodiments of the present disclosure. FIG. 2F is described in relation to FIG. 2B. The schematic diagram 242 includes a predicted Gaussian distribution 244, a control parameter 246 (or a value of the control parameter), and an updated Gaussian distribution 248. As an example, the predicted Gaussian distribution 244 may be a Gaussian distribution 218b defined by a mean θ k|k-1 and a variance P k|k-1 . As an example, the control parameter 246 may be a control parameter that can be used to control the machine 202 to achieve a specific trajectory with respect to the performance target 220b. Further, the control parameter 246 may be derived from the predicted Gaussian distribution 244, in which case the measured value is close to zero probability with respect to the predicted Gaussian distribution 244. For this reason, the measurement model 220 may update the predicted Gaussian distribution 244 so that the predicted Gaussian distribution 244 approaches the updated Gaussian distribution 248. In other words, the measurement model 220 updates the mean and variance associated with the predicted Gaussian distribution 244 to the mean (e.g., mean θ k|k ) and variance (e.g., variance P k|k) , the variance corresponding to the updated Gaussian) corresponding to the updated Gaussian distribution 248.

[0055] Referring back to FIG. 2B, in this way, the measurement model 220 may update the predicted value of the control parameter based on the sequence 220a of measured values to generate the current value of the control parameter according to the performance target 220b. In an example of an embodiment, the performance target 220b may be different from the control policy of the feedback controller 210 used to determine the control input. This is beneficial because the control policy has a structure that is limited due to its real-time application. For example, the cost function often needs to be differentiable and convex so that the cost function can be suitable for numerical optimization. However, the performance target 220b may change at different times of control according to the same control policy.

[0056] According to an embodiment, the measurement model 220 may output the current values of the generated control parameters as a joint probability distribution 220d (or Gaussian distribution 220d), which define the quantity 220c, for example, with a mean θ k|k and a variance P k|k- The Kalman filter 212 may repeat the procedure for generating the control parameters in the next iteration 222 (e.g., at time step k + 1).

[0057] In this way, the Kalman filter 212 may iteratively generate control parameters that can be used to iteratively update the state of the feedback controller 210. The updated state of the feedback controller 210 may then be used to determine a control input for controlling the operation of the machine 202. Since the Kalman filter 212 uses the joint probability distribution of the control parameters (e.g., prior knowledge 218a) to iteratively generate the control parameters instead of recomputing the control parameters using the entire data history, the Kalman filter 212 can efficiently generate the control parameters for controlling the operation of the machine 202. Further, since the system 200 only needs to store the prior knowledge of the control parameters instead of the entire data history, the data to be stored in the memory of the system 200 can also be reduced. Thus, the memory requirements of the system 200 can be reduced.

[0058] Some embodiments are based on the recognition that when one or more of the control parameters depend on other (different) control parameters among the same control parameters, the Kalman filter 212 must calibrate the control parameters together. For example, in a PID controller, the gains of the PID controller are interdependent, so the gains must be calibrated together.

[0059] Generally, calibrating these interdependent control parameters can be difficult because the interdependence may add additional variables during calibration. In such situations, the Kalman filter 212 may be configured as described in the detailed description of FIG. 3.

[0060] FIG. 3 shows a block diagram of a Kalman filter 212 for calibrating a plurality of interdependent control parameters, according to some embodiments of the present disclosure. FIG. 3 is described in relation to FIG. 2B. According to an embodiment, when a control parameter corresponds to a plurality of interdependent control parameters, the Kalman filter 212 may be configured to adjust a Kalman gain 300 to calibrate the control parameter. As an example, if one or more of the control parameters included in the control parameter depend on other control parameters among the same control parameters, the control parameter can be referred to as a multiple interdependent control parameter. As used herein, "adjusting the Kalman gain 300" may indicate applying different weights to the control parameters. To calibrate the multiple interdependent control parameters, the Kalman filter 212 may adjust the Kalman gain 300 to apply more weight to one or more control parameters that depend on other control parameters than to other control parameters. Further, the Kalman filter 212 may be configured to simultaneously update the control parameters using a measurement model 220 to output calibrated interdependent control parameters 302. As an example, the Kalman filter 212 can calculate the Kalman gain 300 as described in the detailed description of FIGS. 4A and / or 4B.

[0061]

Number

[0062]

Number

[0063]

Number

[0064]

Number

[0065]

Number

[0066]

Number

[0067]

Number

[0068]

Number

[0069] Furthermore, the Kalman filter 212 can output an updated joint probability distribution determined by the mean θ k and the variance P k|k as a control parameter for controlling the machine.

[0070]

Number

[0071]

Number

[0072]

Number

[0073]

Number

[0074]

Number

[0075]

Number

[0076]

Number

[0077] Furthermore, the Kalman filter 212 can output an updated joint probability distribution defined by the mean θ k and the variance P k|k as control parameters for controlling the machine.

[0078] FIG. 5 shows a method 500 for calibrating state-dependent control parameters according to some embodiments of the present disclosure. FIG. 5 is described in relation to FIGS. 2A and 2B. Some embodiments are based on the recognition that a set of control parameters of the feedback controller 212 may include at least some control parameters that depend on the state of the machine 202. For example, the friction of a vehicle's tires may depend on the vehicle's speed. Hereinafter, "at least some control parameters that depend on the state of the machine" and "state-dependent control parameters" may be used interchangeably without distinction as having the same meaning. When the set of control parameters includes state-dependent control parameters, calibration of the control parameters can be difficult because these state-dependent control parameters can vary continuously with respect to the state of the machine. In these embodiments, the Kalman filter 212 can execute a method 500 for calibrating state-dependent control parameters.

[0079]

Number

[0080] In step 504, the Kalman filter 212 can predict a state-dependent control parameter within the variance determined by process noise based on the algebraic relationship with the state of the machine 202. As an example, the prediction model 218 of the Kalman filter 212 may be designed (or described) such that the prediction model 218 predicts a state-dependent control parameter within the variance determined by process noise based on the algebraic relationship with the state of the machine 202. For example, when the algebraic relationship of the state-dependent control parameter corresponds to a linear combination of the state-dependent control parameter and a basis function, the prediction model 218 may be configured to check whether the basis function is determined by two or more state-dependent regions. If the basis function is not determined by two or more state-dependent regions, the prediction model 218 may be configured to predict the coefficient of the basis function.

[0081]

Number

[0082]

Number

[0083] FIG. 6A shows a block diagram of a system 200 for controlling the operation of a machine 202, according to some other embodiments of the present disclosure. FIG. 6A is described in relation to FIGS. 2A and 2B. Some embodiments are based on the recognition that an online update of control parameters may destabilize the operation of the machine 202. For example, when a control law or control policy is represented by a differential equation of the control parameters (e.g., an ordinary differential equation (ODE)), a change (update) of the control parameters may compromise the stability of the equilibrium of the differential equation. For this reason, the system 200 may further include a safety verification module 600. As an example, the safety verification module 600 may be implemented within at least one processor 204. Alternatively, the safety verification module 600 may be a software module stored in a memory that can be executed by at least one processor 204. According to an embodiment, the safety verification module 600 may be configured to execute a safety verification method using the value of the control parameters generated by the Kalman filter 212 to ensure the safe operation of the machine 202. As an example, the safety verification method executed by the safety verification module 600 is as described in the detailed description of FIG. 6B.

[0084] FIG. 6B shows a safety verification method executed by a safety verification module 600, according to some embodiments of the present disclosure. FIG. 6B is described in relation to FIG. 6A. In step 602, the safety verification module 600 can obtain the value of the control parameters (e.g., the current value) generated by the Kalman filter 212.

[0085] In step 604, the safety confirmation module 600 can confirm whether the value of the control parameter generated by the Kalman filter 212 satisfies the safety confirmation according to the control policy. In other words, the safety confirmation module 600 can confirm whether the value of the control parameter by the Kalman filter 212 provides stable control of the machine 202 when the machine 202 is controlled by the feedback controller 210 according to the control policy updated with the control parameter generated by the Kalman filter 212. To confirm whether the control parameter generated by the Kalman filter 212 satisfies the safety confirmation, the safety confirmation module 600 can use the previous state, the sequence of measurements (e.g., measurement sequence 220a), and / or the model of the machine 202 (e.g., model 106).

[0086] For example, if there exists a Lyapunov function for the control policy updated with the control parameters generated by the Kalman filter 212, the safety check is satisfied. In some embodiments, the existence of the Lyapunov function can be proven by solving an optimization problem aimed at finding the Lyapunov function. In one embodiment, while controlling the machine 202 using the feedback controller 210 updated with the control parameters generated by the Kalman filter 212, if a decreasing cost of the state of the machine 202 is realized with respect to the performance goals for the entire history of the state and the sequence of measurements, the safety check is satisfied. In another embodiment, while controlling the machine 202 using the feedback controller 210 updated with the control parameters generated by the Kalman filter 212, if the proximity (or boundedness) of the state of the machine 202 to the origin is realized, the safety check is satisfied. In yet another embodiment, if a combination of the decreasing cost of the state of the machine 202 and the proximity of the state of the machine 202 to the origin is realized, the safety check is satisfied. Thus, the safety check may include one or a combination of the decreasing cost of the state of the machine 202 and the proximity of the state of the machine 202. As an example, the safety check module 600 that performs the safety check is as described in the detailed description of FIG. 6C.

[0087]

Number

[0088] Referring again to FIG. 6B, if the value of the control parameter generated by the Kalman filter 212 does not meet the safety check, the safety check module 600 may hold the control parameter of the feedback controller 210 in step 608. In other words, if the safety check fails, the safety check module 600 may not update the control parameter of the feedback controller 210 with the control parameter generated by the Kalman filter 212. Further, the Kalman filter 212 may be configured to repeatedly generate new values of the control parameter until the safety check is satisfied. In particular, even if the control parameter generated by the Kalman filter 212 does not meet the safety check, the Kalman filter 212 should not be restarted.

[0089] If the value of the control parameter generated by the Kalman filter 212 meets the safety check, the safety check module 600 can update the control parameter of the feedback controller 210 with the control parameter generated by the Kalman filter 212 in step 606. In this way, the safety check module 600 can ensure the stability of the control during the online update of the control parameter of the feedback controller 210.

[0090] Referring again to FIG. 6A, when the control parameter of the feedback controller 210 is updated, the feedback controller 210 may be configured to determine a control input for controlling the operation of the machine 202 by applying the control policy updated with the control parameter generated by the Kalman filter 212.

[0091] Some embodiments are based on the understanding that while online updating of the control parameters of the feedback controller 210 can improve the quality of control, it is accompanied by additional challenges. For example, online updating of the control parameters during operation of the machine 202 may introduce discontinuities in control. Some embodiments are based on the recognition that the discontinuities in control can be processed to implement control commands that satisfy the constraints on the operation of the machine 202. To this end, the feedback controller 210 may be configured to determine a control input (e.g., the current control input) using a control command that satisfies the constraints on the operation of the machine 202. In other words, the feedback controller 210 may be configured to process the discontinuities in control by determining a control input that is affected by the constraints on the operation of the machine 202. As an example, the control command satisfies the constraints on the operation of the machine 202 when the control parameters (e.g., the current control parameters) satisfy the control theory characteristics. For example, the control theory characteristics may be specified by a designer.

[0092]

Number

[0093] In addition or instead, the performance goal 200b may include a cost function for a particular control parameter (rather than a state or control input), e.g., y k = θ nom , h(θ k ) = θ k and θ nom defines the nominal value of any or all of the control parameters.

[0094] In some embodiments, measurement model 220 may be configured to select one from a list of performance goals based on one or a combination of the state of the machine and the state of the environment surrounding the machine. For example, depending on the state of the machine and / or the control parameters used to control the machine, measurement model 220 can identify, from the list of performance goals, a cost function that significantly degrades control performance compared to other cost functions in the list of performance goals. Further, measurement model 220 can select the identified cost function as performance goal 220a. As an example, if cost function 700 significantly degrades control performance compared to other cost functions in the list of performance goals, measurement model 220 can select cost function 700 as performance goal 200a. Further, measurement model 220 can update the control parameters by optimizing (e.g., minimizing) cost function 700 while generating the control parameters.

[0095] Some embodiments are based on the recognition that if the control parameters are independent of the state of the machine, bounds on the uncertainty of the control parameters can be pre-determined. Hereinafter, "the control parameters are independent of the state of the machine" and "state-independent control parameters" may be used interchangeably without distinction as having the same meaning. In these embodiments, Kalman filter 212 can select one or a combination of a performance goal, measurement noise, or process noise based on the bounds. Further, some possible state-dependent control parameters and state-independent control parameters are as shown in FIG. 8A.

[0096]

Number

[0097]

Number

[0098] [Number]

[0099] [Number]

[0100] [Number]

[0101] [Number]

[0102] H ∞ In the case of the H∞ controller 812, the Kalman filter 212 can estimate the filter coefficients of the pre-compensator and the post-compensator, which are used to determine the trade-off between performance and robustness. For example, the H∞ controller 812 is shown in FIG. 8B.

[0103] [Number]

[0104] Here, H ∞ The H∞ controller 812 may be calculated by minimizing the maximum magnitude of the frequency response of the machine 202.

[0105] FIG. 9 shows a schematic diagram of a system 200 for controlling a motor 900 according to some embodiments of the present disclosure. In this example, the feedback controller of the system 200 may be a PID controller for controlling the motor 900. The system 200 can receive a position or velocity signal 904 (e.g., the feedback signal 216) from a sensor 902 installed on the motor 900. Further, the system 200 can receive a desired position or velocity signal 906 (e.g., a control command) and calculate an error signal 908. Further, the system 200 has a proportional gain kP , the integral gain k I , and the derivative gain k D can be calibrated. Further, the system 200 can determine the control input 910 by using the error signal 908 and applying a control policy. As an example, the control policy is the sum of three components, for example, the proportional component 912a obtained by multiplying the proportional gain k P calibrated to the error, the integral component 912b obtained by integrating the error and multiplying the integral gain k I calibrated to the integrated error, and the derivative component 912c obtained by differentiating the error with respect to time and multiplying the derivative gain k D calibrated to this derivative. Further, the system 200 can submit the determined control input 910 to the motor to control the motor 900.

[0106] FIG. 10 shows a schematic diagram of a system 200 for controlling a vehicle 1000 according to some embodiments of the present disclosure. In this example, the system 200 can control the vehicle 1000 to stay in the center of the lane 1002. The system 200 can receive position and / or speed signals from the sensor 1004. The system 200 can further calibrate one or more control parameters associated with the control policy. For example, the control parameter may be the friction between the tires of the vehicle 1000 and the road. Further, the system 200 can determine the control input by applying a control policy updated with one or more calibrated control parameters. As an example, the system 200 can determine the control input so that the control input causes the vehicle 1000 to stay in the center of the lane 1002. For example, the determined control input may be the steering angle 1006 that causes the vehicle 1000 to travel on the vehicle trajectory 1008 converging to the center of the lane 1002.

[0107] The above description provides only embodiments as specific examples and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the above description of embodiments as specific examples provides those skilled in the art with an explanation that enables one or more embodiments as specific examples to be realized. What is intended are various changes that can be made to the functions and configurations of the elements without departing from the spirit and scope of the disclosed subject matter as recited in the appended claims.

[0108] Specific details are provided in the above description to obtain a thorough understanding of the embodiments. However, those skilled in the art will understand that the embodiments can be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in block diagram form to avoid obscuring the embodiments with unnecessary details. In other instances, well-known processes, structures, and techniques may be shown without unnecessary details to avoid obscuring the embodiments. Further, like reference numerals and names in the various drawings indicate like elements.

[0109] Also, individual embodiments may be described as a process shown as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. A flowchart may describe operations as a sequential process, but many of the operations can be performed in parallel or simultaneously. Further, the order of the operations can be interchanged. A process may end when its operations are completed, but may have other steps not discussed or included in the figure. Further, not all operations in any specifically described process may occur in all embodiments. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, the end of the function may correspond to returning the function to the calling function or the main function.

[0110] Furthermore, embodiments of the disclosed subject matter may be implemented, at least in part, either manually or automatically. Manual or automatic implementation may be performed or at least assisted through the use of a machine, hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the required tasks may be stored on a machine-readable medium. A processor may perform the required tasks.

[0111] The various methods or processes outlined herein may be encoded as software executable on one or more processors employing any one of a variety of operating systems or platforms. Additionally, such software may be written using any of a number of suitable programming languages and / or programming or scripting tools, and may be compiled as executable machine language code or intermediate code that is executed on a framework or virtual machine. Typically, the functionality of program modules may be combined or distributed as desired in various embodiments.

[0112] Embodiments of the present disclosure may be implemented as a method, and an example thereof is provided. The order of operations performed as part of this method may be determined in any suitable manner. Accordingly, embodiments may be configured so that operations are performed in an order different from that illustrated, which may include performing some operations simultaneously, even though in the illustrated embodiments some operations are shown as a series of operations. Although the present disclosure has been described with reference to some preferred embodiments, it should be understood that various other adaptations and modifications can be made within the spirit and scope of the present disclosure. Accordingly, it is the aspect of the appended claims to cover all such variations and modifications that fall within the true spirit and scope of the present disclosure.

Claims

1. A system for controlling the operation of a machine to execute a task, the system comprising: a transceiver configured to submit a sequence of control inputs to the machine and receive a feedback signal including a corresponding sequence of measurement values, each measurement value indicating a state of the machine caused by the corresponding control input, the system further comprising: in each control step, a feedback controller configured to determine a current control input for controlling the machine based on the feedback signal including the current measurement value of the current state of the machine by applying a control policy, the feedback controller converting the current measurement value into the current control input based on a current value of a control parameter within a set of control parameters of the feedback controller by applying the control policy, the system further comprising: a Kalman filter, the Kalman filter configured to iteratively update the state of the feedback controller determined by the control parameter by using a prediction model for predicting a value of the control parameter affected by process noise and a measurement model for updating a predicted value of the control parameter based on the sequence of measurement values affected by measurement noise, so as to generate the current value of the control parameter associated with the sequence of measurement values according to a performance target; The system, wherein the prediction model is an identity matrix configured such that a prediction variance of the control parameter is constant and is determined by the process noise.

2. The system according to claim 1, wherein the Kalman filter is further configured to adjust a Kalman gain to calibrate a plurality of interdependent control parameters.

3. The system according to claim 1, wherein the prediction model is configured to predict at least some control parameters within a variance determined by the process noise based on an algebraic relationship with the state of the machine.

4. The system according to claim 1, wherein the performance target for updating the control parameter is different from the control policy of the feedback controller.

5. The system according to claim 1, wherein the performance target includes a cost function that defines a deviation of the state of the machine from a reference state of the machine, and the measurement model is configured to update the control parameter by optimizing the cost function.

6. The system according to claim 1, wherein the measurement model is further configured to select one from among different performance targets based on one or a combination of the state of the machine and the state of the environment surrounding the machine.

7. The system according to claim 1, wherein the performance target includes one or a combination of (i) a cost function that defines a deviation of the state from a reference state, (ii) a cost function for the state that exceeds an optimal operating region, (iii) a cost function when the reference state is overshot by a specific value, (iv) a cost function for oscillations of the state, and (v) a cost function when the state changes between time steps.

8. The control parameter according to claim 1 includes one or a combination of (i) one or more gains of the feedback controller, (ii) one or more structural parameters of the machine, (iii) one or more coefficients of one or more filters used by the feedback controller, or (iv) one or more weights of a neural network controller.

9. The system according to claim 1, wherein, in order to generate the control parameter, the Kalman filter is configured to update coefficients of basis functions in one or more state-dependent regions.

10. The system according to claim 1, wherein, in order to generate the control parameter, the Kalman filter is configured to update coefficients of basis functions in a plurality of state-dependent regions together with boundaries separating the plurality of state-dependent regions.

11. The system according to claim 1, wherein the Kalman filter is an extended Kalman filter (EKF) configured to calculate a Kalman gain by calculating a gradient of the performance target.

12. The system according to claim 1, wherein the Kalman filter is an unscented Kalman filter (UKF) configured to calculate a Kalman gain by evaluating the control parameter with respect to the performance target.

13. An uncertainty boundary of at least one of the control parameters is predetermined, and the Kalman filter is further configured to select one or a combination of the process noise, the measurement noise, or the performance target based on the uncertainty boundary. The system according to claim 1.

14. The system according to claim 1, wherein the feedback controller is configured to determine the current control input affected by the constraints on the operation of the machine and thereby handle the discontinuity of the control.

15. Further comprising a safety confirmation module, the safety confirmation module executes a confirmation associated with whether the value of the control parameter generated by the Kalman filter satisfies the safety confirmation according to the control policy, The system according to claim 1, wherein when the safety confirmation is satisfied, the control parameter of the feedback controller is updated using the control parameter generated by the Kalman filter.

16. The system according to claim 15, wherein when the value of the control parameter generated by the Kalman filter does not satisfy the safety confirmation, the Kalman filter is further configured to repeatedly generate a new value of the control parameter until the safety confirmation is satisfied.

17. The system according to claim 15, wherein the safety confirmation includes one or a combination of boundedness of the state to the origin and decreasing cost of the state.

18. A method for controlling the operation of a machine to perform a task, the method comprising: submitting a sequence of control inputs to the machine; receiving a feedback signal including a sequence of corresponding measurement values, each measurement value indicating a state of the machine caused by the corresponding control input, and the method further comprising: in each control step, determining a current control input for controlling the machine based on the feedback signal including the current measurement value of the current state of the machine by applying a control policy, and by applying the control policy, converting the current measurement value into the current control input based on the current value of the control parameter within a set of control parameters of the feedback controller, and the method further comprising: Including the step of generating the current value of the control parameter that explains the sequence of the measured values according to a performance goal by repeatedly updating the state of the feedback controller determined by the control parameter, using a prediction model that predicts the value of the control parameter affected by process noise and a measurement model that updates the predicted value of the control parameter based on the sequence of the measured values affected by measurement noise. The method, wherein the prediction model is an identity matrix configured such that the predicted variance of the control parameter is constant with the variance determined by the process noise. **Claim 19** A non-transitory computer-readable storage medium having a program executable by a processor implemented to execute a method for controlling the operation of a machine to perform a task, the method comprising: Submitting a sequence of control inputs to the machine; Receiving a feedback signal including a sequence of corresponding measured values, each measured value indicating the state of the machine caused by the corresponding control input, and the method further comprising: In each control step, determining a current control input for controlling the machine based on the feedback signal including the current measured value of the current state of the machine by applying a control policy, and by applying the control policy, converting the current measured value into the current control input based on the current value of the control parameter within a set of control parameters of the feedback controller, and the method further comprising: Including the step of generating the current value of the control parameter that explains the sequence of the measured values according to a performance goal by repeatedly updating the state of the feedback controller determined by the control parameter, using a prediction model that predicts the value of the control parameter affected by process noise and a measurement model that updates the predicted value of the control parameter based on the sequence of the measured values affected by measurement noise. The prediction model is a non-transitory computer-readable storage medium that is an identity matrix configured assuming that the predicted variance of the control parameter is constant with the variance determined by the process noise.

Citation Information

Patent Citations

  • Predictive control unit

    JP2000056805A

  • Vehicular state estimating device, vehicular state estimating method, vehicular suspension control device, and automobile

    JP2010195323A