Controller with neural network and improved stability

By introducing a combination of projection functions and robust control into deep reinforcement learning, the output of the neural network is mapped to a stable control space, solving the stability problem of deep reinforcement learning in safety-critical fields and realizing stable control in complex dynamic systems.

CN113448244BActive Publication Date: 2026-03-17ROBERT BOSCH GMBH +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-23
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Deep reinforcement learning lacks stability guarantees in safety-critical areas, especially when faced with sensor readings or disturbances, which may lead to system instability and make it difficult to provide effective control strategies in robust control and complex dynamic systems.

Method used

By employing a projection function to map the control strategy of the neural network to a stable control strategy space, and by using stability equations such as the Lyapunov function to constrain the controller's output, combined with deep reinforcement learning and robust control, the controller is ensured to remain stable in the face of uncertainties and disturbances.

Benefits of technology

It provides stable control strategies for complex dynamic systems, ensuring system stability under worst-case conditions while maintaining efficient control performance, and is suitable for a variety of machines such as drones, robotic arms and autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113448244B_ABST
    Figure CN113448244B_ABST
Patent Text Reader

Abstract

A controller with neural networks and improved stability. Some embodiments are for controllers used to generate control signals for computer-controlled machines. A neural network can be applied to current sensor signals, configured to map the sensor signals to raw control signals. A projection function can be applied to the raw control signals to obtain stable control signals for controlling computer-controlled machines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The currently disclosed subject matter relates to controllers for generating control signals for computer-controlled machines, control methods for generating control signals for computer-controlled machines, training methods for training neural networks for use in controlling computer-controlled machines, and training systems for training neural networks for use in controlling computer-controlled machines. Background Technology

[0002] Despite recent remarkable advances and state-of-the-art performance for many tasks, deep reinforcement learning still has limited applications in the "safety-critical" domain, where incorrect actions, both during training and at runtime, can have a significant impact on the controlled system. In contrast, the field of robust control, dating back decades, has been able to provide strict bounds on when a controller will succeed or fail in controlling a system of interest. Specifically, robust control techniques offer provable guarantees that the resulting system will be stable if the controlled system can be properly bounded in some way. However, the simple, often linear, nature of the policies generated by provably robust control techniques frequently limits performance in typical scenarios.

[0003] Robust control studies the design of feedback controllers for dynamic systems with uncertainties and / or unknown disturbances. Specifically, robust control strategies aim to design controllers that guarantee performance in the worst-case scenario for uncertainties or disturbances in the system. Particularly... In control, the goal is to stabilize the system while attenuating the impact of external disturbances on some performance outputs (such as LQR costs). This effect is mapped to the output by the disturbance. Gain characterization, the perturbation mapped to the output Gain is defined as the output The ratio between the norm and the disturbance. Despite providing performance guarantees, robust control strategies are often overly conservative due to the need to consider worst-case scenarios or being limited to linear class controllers.

[0004] Many classes of robust control problems—even many initially formulated in the frequency domain (e.g., alternative ways of characterizing dynamical systems)—can ultimately be formulated using linear matrix inequalities (see, for example, "Linear Matrix Inequalities in System and Control Theory" by Stephen Boyd, Laurent El Ghaoui, Eric Feron, and Venkataramanan Balakrishnan). The resulting control problems can thus often be formulated as semidefinite procedures that can be solved by using readily available software to generate controllers for reasonably sized domains.

[0005] WO 93 / 00618 discloses an adaptive control system that uses a neural network to provide adaptive control when the plant is operating within the normal operating range, but switches to other types of control when the plant operating conditions move outside the normal range.

[0006] Reinforcement learning, particularly deep reinforcement learning, where optimal control policies are approximated by neural networks, has shown impressive results in learning a variety of complex control tasks. However, due to its lack of safety guarantees, deep reinforcement learning is primarily applied to simulated environments or highly controlled real-world problems where potential system failures are either inexpensive or impossible. For deep reinforcement learning methods to be employed in safety-critical settings, it is necessary to couple them with some form of safety guarantee, as addressed in this work. Indeed, it has been found that control policies approximated by neural networks can fail entirely in practice, for example, when faced with sensor readings or disturbances that are slightly different from those encountered during training. In particular, it has been found that the stability of the control policy can completely vanish under adversarial perturbations, even if the adversarial perturbations are bounded in magnitude.

[0007] There is an expectation for safe reinforcement learning in which control policies can be implemented while maintaining system stability. Furthermore, it is desirable to use global uncertainty features, for example, in robust control, rather than employing assumptions about local smoothness regarding dynamics or costs. Summary of the Invention

[0008] Having an improved controller for generating control signals for computer-controlled machines would be advantageous. For example, a machine may have multiple components that interact in complex ways. These components can be considered, for example, a dynamic system or approximated therefrom.

[0009] Controllers can apply neural networks to sensor signals. For example, using direct search or reinforcement learning to train a neural network, it can control a machine in a certain way, such as bringing it to a specific state or making its state follow a trajectory. Control can also be limited to a portion of the state, such as the position of a vehicle like a car or drone; for example, controlling a car to stay in a lane, or controlling a drone to stay in a 2D plane, and so on. However, instead of using the output of the neural network, for example, the original control signal directly controls the machine, a projection function is applied to the original control signal to obtain a stable control signal. The projection function maps the original control policy space to a stable control policy space of predefined stable control policies. The effect is that neural network-based control is restricted to predefined stable control policies. For example, the control policy can be predefined as stable by requiring it to satisfy a stability equation, such as a decreasing Lyapunov function. Stable control signals can be used to control computer-controlled machines. Neural networks can be trained with projection functions to select good control policies, but only under the restriction of using only stable control policies.

[0010] In this embodiment, the projection function is a continuous, piecewise differentiable function. This has the advantage that reinforcement learning can be performed, for example. Note that even if the projection function is implemented as an optimization layer, it can still be a continuous, piecewise differentiable function. Reinforcement learning can tune the neural network to reduce the loss function. For example, the loss function can measure how close the system comprising the neural network is to reaching a goal, such as reaching a state, following a trajectory, maintaining constraints on all or some state variables, etc.

[0011] For example, a neural network can be tuned to reduce the loss function over multiple learning steps. These learning steps can be repeated, for example, until the desired accuracy is achieved, until a predefined number of iterations, until the loss function no longer decreases, and so on. Many learning algorithms, including backpropagation, can be used to tune the neural network.

[0012] An interesting advantage of using neural networks at the front end is the ability to use a wide variety of sensor inputs. Processed information, such as the position and velocity of one or more parts of the machine, can be fed to the neural network, but instead, the network can learn to derive that information itself. For example, a sensor system can include one or more of the following: image sensors, radar, lidar, pressure sensors, etc. Therefore, sensor signals can include 2D data, such as image data. For example, for a robotic arm controller, sensor data could include one or more images showing the robotic arm.

[0013] This controller can be used in a wide variety of control tasks. For example, the machine can be one of the following: a robot, a vehicle such as a car, a home appliance, a power tool, a manufacturing machine, a drone such as a quadcopter, etc. In this embodiment, the controller provides guarantees of robust control, but also offers the capability of deep reinforcement learning. The stability specifications generated by traditional robust control methods under different system uncertainty models can be respected, while simultaneously achieving improved control.

[0014] For example, such as Robust control methods generate both a stable controller and a Lyapunov function, proving the stability of the system under certain worst-case perturbations. These specifications can be used to construct a new class of reinforcement learning policies that project a nominal, for example, original control policy into a stable controller space specified by the robust Lyapunov function. The original control policy can be nonlinear and neural network-based. The result is a nonlinear control policy that can be trained using deep reinforcement learning, yet remains stable under the same conditions as simple, conventional robust control policies.

[0015] The illustrated embodiment improves upon the conventional LQR method, while remaining stable even under the worst-case perturbations of the underlying dynamics system, unlike non-robust LQR and neural network methods.

[0016] The controller and training system can be electronic devices and may include a computer. The control methods described herein can be applied to a wide range of practical applications. Such practical applications include the control of computer-controlled machines, including fully or partially autonomous vehicles, robotic arms, and drones.

[0017] Another aspect is a training system configured to train a neural network used in a controller. Another aspect is the control method and the training method. Embodiments of this method can be implemented on a computer as a computer-implemented method, or in dedicated hardware, or a combination of both. Executable code for embodiments of the method can be stored on a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Preferably, the computer program product includes non-transitory program code stored on a computer-readable medium for executing embodiments of the method when the program product is executed on a computer.

[0018] In an embodiment, the computer program includes computer program code that, when run on a computer, is adapted to perform all or part of the steps of an embodiment of the method. Preferably, the computer program is embodied on a computer-readable medium. Attached Figure Description

[0019] Additional details, aspects, and embodiments will be described by way of example only with reference to the accompanying drawings. Elements in the figures are illustrated for simplicity and clarity and are not necessarily drawn to scale. In the figures, elements corresponding to those already described may have the same reference numerals. In the drawings,

[0020] Figure 1 An example embodiment of a controller for generating control signals for a computer-controlled machine is illustrated schematically.

[0021] Figure 2a An example of an embodiment of the physical system is illustrated schematically.

[0022] Figure 2b An example of an embodiment of a robotic arm is illustrated schematically.

[0023] Figure 3 An example embodiment of a controller for generating control signals for a computer-controlled machine is illustrated schematically.

[0024] Figure 4 An example embodiment of a training system for training a neural network for use in controlling a computer-controlled machine is illustrated schematically.

[0025] Figure 5 An example embodiment of a control method for generating control signals for a computer-controlled machine is illustrated schematically.

[0026] Figure 6 An example embodiment of a training method for training a neural network for use in controlling a computer-controlled machine is illustrated schematically.

[0027] Figure 7a A computer-readable medium having a writable portion including a computer program, according to an embodiment, is illustrated schematically.

[0028] Figure 7b The illustration schematically shows a representation of a processor system according to an embodiment.

[0029] Figure 1-4 , Figure 7a , Figure 7b List of reference numbers in the text:

[0030] The following list of reference numerals and abbreviations is provided to facilitate the interpretation of the drawings and should not be construed as limiting the claims.

[0031] 100 Controller

[0032] 110 Computer-controlled machines

[0033] 111 Machine-learnable systems

[0034] 120 Sensor Input Interface

[0035] 122 Sensor System

[0036] 130 Neural Network Storage Devices and Neural Networks

[0037] 150 processor system

[0038] 152 neural network units

[0039] 154 projection units

[0040] 156 training units

[0041] 160 control interface

[0042] 170 training sets

[0043] 200 physical systems

[0044] 210 Movable parts

[0045] 212 Actuator

[0046] 214 Sensors

[0047] 220 Controller

[0048] 230 camera

[0049] 240 robotic arms

[0050] 241 links

[0051] 242 tools

[0052] 243 connector

[0053] 300 controller

[0054] 330 processor

[0055] 340 memory

[0056] 350 communication interface

[0057] 1000 computer-readable media

[0058] 1010 writable portion

[0059] 1020 Computer Program

[0060] 1110 (one or more) integrated circuits

[0061] 1120 Processing Unit

[0062] 1122 Memory

[0063] 1124 Application-Specific Integrated Circuit

[0064] 1126 Communication Components

[0065] 1130 Interconnection

[0066] 1140 processor system. Detailed Implementation

[0067] While the subject matter currently disclosed allows for many different embodiments, there are one or more specific embodiments shown in the accompanying drawings and to be described in detail herein. It should be understood that this disclosure is to be considered as an example of the principles of the subject matter currently disclosed and is not intended to limit it to the specific embodiments shown and described. Hereinafter, for understanding purposes, elements of the embodiments are described in operation. However, it will be clear that the corresponding elements are arranged to perform the functions described as being performed by them.

[0068] Furthermore, the subject matter disclosed herein is not limited to the embodiments, but also includes every other combination of features described herein or referenced in mutually different dependent claims.

[0069] Figure 1 An example embodiment of a machine-learnable system 111 is illustrated schematically. System 111 includes a controller 100 and a computer-controlled machine 110. For example, controller 100 may be configured to generate control signals for machine 110. The machine may include multiple interacting components that may interact in a complex, e.g., nonlinear, manner, for example, in a so-called dynamic system. The dynamic system may be explicitly modeled, although this is not necessary. For example, the dynamic system may be nonlinear or have time-varying parameters. Such dynamic systems are difficult to control, for example, to reliably control to a desired state or to follow a desired trajectory. Although neural networks can be trained to successfully control such systems, neural networks have the disadvantage that their solutions do not provide any guarantees. Although neural networks are often very efficient, they may have unexpected failures, for example, compared to conventional controllers. For example, the state faced by the neural network or the sensor input it faces—which may even be slightly different from what was known from training—may suddenly become unstable. Clearly, such a situation is undesirable for real-world applications. In the embodiment, a controller is provided that is efficient and provides guarantees regarding its stability.

[0070] For example, in one embodiment, the machine includes a so-called trolley lever. In this case, components may include a trolley and a lever. Control signals may include position, speed, and / or direction, etc., used to steer the trolley. The trolley and lever have complex interactions and can be considered as a dynamic system.

[0071] Computer-controlled machine 110 can be any machine in which control is a problem. For example, the machine can include one or more of the following: robots, vehicles, household appliances, power tools, manufacturing machines, drones, etc. For example, a drone can be controlled to remain stable at a specific point in the air, or to remain in a 2D plane, or to follow a specific trajectory. For example, in a vehicle, the controller can be configured for cruise control, such as keeping the vehicle at a predefined speed; or for trajectory assistance, such as lane assist. For example, control signals can control speed and steering to maintain the vehicle's position within a lane.

[0072] Machine 100 can be computer-controlled by controller 100. Machine 100 may include an additional computer for controlling machine 110, for example, by means of stable control signals dependent on controller 100.

[0073] For example, controller 100 may include control interface 160 configured to transmit stable control signals to computer-controlled machine 110 to control it. Control can be direct, for example, the output of controller 100 can directly control, for example, the actuator of machine 110, but it can also be indirect. For example, the control signal can be to increase the force by a certain amount at a point, such as on the rotor. In such cases, machine 110 may include a separate controller to translate the control signals from controller 100 into a form suitable for machine 110 (e.g., its actuators, etc.).

[0074] System 111 may include a sensor system 122 configured to sense the computer-controlled machine 110. Controller 100 may include a sensor input interface 120 configured to receive sensor signals from sensor system 122. For example, sensor system 122 may include one or more of the following: image sensor, radar, lidar, pressure sensor. Sensor signals may include 1D sensor signals such as pressure sensors, configured to measure pressure at a specific point. Sensor signals may also include 2D sensor signals, such as image sensors. For example, one or more cameras may record the position of a robotic arm, or the position of a vehicle in its lane, etc. Interestingly, controller 100 may directly use such high-dimensional sensor signals.

[0075] In one embodiment, controller 110 receives sensor signals and thereby calculates the next action, such as an updated value for actuator inputs. For control purposes, controller 100 can be repeatedly applied to the current sensor signals obtained from sensor system 122 to repeatedly obtain stable control signals for controlling computer-controlled machine 110. For example, the actions to be performed can be discretized. The frequency of updating actions can vary for different embodiments. For example, this could be multiple times per second, such as 10 times, for example, in cases where the dynamic system changes rapidly, or where external disturbances can change rapidly, or even more. Actions can be updated more slowly, for example, once per second, or once every 2 seconds, and so on.

[0076] In one embodiment, controller 100 may be configured to compute control signals to cause machine 110 to be in a desired state. For example, this state may be a predefined state, such as a zero state or a static state. In another embodiment, the controller is configured to acquire, for example, receive a target state, in which case controller 100 is configured to cause the machine to reach the target state. The target state may be generated by controller 110, for example, to cause the machine state to follow a trajectory; the controller may include a target state interface for receiving the target state, for example, from another computer or operator.

[0077] In an embodiment, controller 100 may be configured to calculate control signals to cause a portion of the machine 110 to be in a desired state. For example, this state could be the position and speed of a vehicle, such as a car, and the control signals might be configured to maintain speed without affecting steering, for example, in the case of cruise control. For a drone, for example, the control signals might keep the drone stable in a 2D plane without further restricting its movement within that plane. However, in an embodiment, the complete state of the machine may be controlled by the controller.

[0078] In one embodiment, the controller 100 is configured to change the target state along a trajectory, causing the machine to follow some trajectory. For example, this could be used to turn a vehicle along a trajectory, or to move a robotic arm along a trajectory, and so on. For example, at a first frequency... It can update actions based on the current sensor signal, and at the second frequency The target state is updated to the current target state. Depending on the application, these two frequencies may be equal or different. In an embodiment, .

[0079] Figure 2aAn example embodiment of a computer-controlled machine 200, such as a physical system, is schematically illustrated. A movable component 210 is shown, which can be a known movable component, such as a joint, linkage, block, shaft, wheel, gear, etc. In particular, the movable component may include multiple joints. The machine can be configured with actions that can be performed on the machine. For example, the actions may include rotational and / or translational joints. The joints may be connected to other moving components, such as linkages. For example, the joints may be: prismatic, such as a slider, translating only; cylindrical, such as rotating and sliding; or rotary, such as rotating only.

[0080] Figure 2a An actuator 212 and a controller 220, such as an embodiment of controller 100, are also shown. Actuator 212 is configured to receive control signals from controller 220 to cause changes in the position parameters of a movable part. For example, some parts may be displaced, rotated, etc. Sensor system 214, such as sensor 214, provides feedback on the machine's state. For example, sensor 214 may include one or more image sensors, force sensors, etc.

[0081] For example, the machine can be a robot, a vehicle, a home appliance, a self-driving car, a power tool, a manufacturing machine, etc. Machine 200 is at least partially under computer control. For example, machine 200 can be configured to enter a specific target state. Machine 200 can have modes in which a human operator, for example, controls machine 200 wholly or partially. Controller 220 can be used to control the system. The planning task can be continuous, discrete, or discretized. The observation can be continuous and high-dimensional, for example, image data.

[0082] For example, the machine could be a vehicle, such as a car, configured for autonomous driving. Typically, the car makes continuous observations, such as distance to an exit, distance to a vehicle ahead, and its exact location within the lane. The vehicle's controller can be configured to determine how to act and where to drive, for example, changing lanes. Sensors can be used, for example, to learn the car's state in its environment, and the driving actions the car might take affect its state.

[0083] For example, the machine can be a manufacturing machine. A manufacturing machine can be arranged with a continuous or discrete set of actions, possibly with continuous parameters. For example, actions can include the tool to be used and / or parameters for the tool, such as the action of drilling at a specific location or angle. For example, a manufacturing machine can be a computer numerical control (CNC) machine. A CNC machine can be configured for the automated control of machining tools by means of a computer, such as drill bits, drilling tools, lathes, 3D printers, etc. A CNC machine can process a piece of material such as metal, plastic, wood, ceramic, or composite material to meet specifications.

[0084] The observable state of manufacturing machines is often continuous and noisy, such as the noisy location of a drill bit or the noisy location of a part being manufactured. Embodiments can be configured to plan how the part is created, for example, first drilling a hole at that location, then performing some second action, and so on.

[0085] In an embodiment, the controller is configured to turn the machine 200 to a stationary state, such as a state suitable for a power outage, for example, where it is safe for an operator to approach the machine, for example, for maintenance.

[0086] Figure 2b An example of an embodiment of a robotic arm 240 is schematically illustrated. In this embodiment, the machine includes a robotic arm. An example of a robotic arm is shown in... Figure 2b As shown in the diagram. For example, a robotic arm may include one or more links 241, a tool 242, and a connector 243. The robotic arm may be controlled, for example, by a controller 220 to actuate the connector. The robotic arm 240 may be associated with one or more sensors, some of which may be integrated with the arm. The sensor may be a camera, such as camera 230. Optionally, the sensor output including the image may be an observable state of the robotic arm. In an embodiment, the controller is configured to turn the robotic arm to a stationary state, etc.

[0087] Even if the observations are continuous and high-dimensional, such as the state of all joints, sensor measurements such as camera images, etc., the controller can calculate safe actions to achieve the desired state and / or follow the desired trajectory.

[0088] Back Figure 1 The controller 100 may include a processor system 150, which is configured to apply neural networks and projection layers to the sensor signals.

[0089] For illustrative purposes, the neural network and the projection layer are shown and discussed separately; however, it should be understood that in embodiments, the neural network may have multiple layers, one of which is a projection layer. Typically, the projection layer will be the last layer in such a neural network, but this is not strictly necessary. For example, additional processing can be performed on the output of the neural network, such as transforming from one domain to another, e.g., scaling the signal, transforming the coordinate system, etc. Preferably, the additional processing preserves the stability improvements provided by the projection layer.

[0090] For example, system 111 may include a neural network storage device configured to store a trained neural network 130. Interestingly, the neural network 130 defines the original control strategy for the computer-controlled machine 110. For example, the control strategy is a function that maps sensor signals to actions. If the neural network 130 is trained without a projection layer, the output of the neural network can be used as a control signal. However, such a controller may malfunction if the machine or sensor values ​​venture into unexpected ranges.

[0091] The neural network storage device may or may not be part of the controller 100. For example, neural network parameters may be stored externally, such as in cloud storage, or locally, such as in a local storage device, such as a memory, hard disk drive, etc.

[0092] For a dynamical system, a large set of stable potential policies can be defined. For example, a specific desired target state, such as the origin, is guaranteed to be reached as time increases. Such stability is typically defined under certain assumptions about perturbations that can be expected in the machine. Clearly, if external perturbations are allowed to be unrestricted, then stable policies cannot exist under any circumstances. However, there are cases where, for example, a norm limit for stable policies can be defined for external perturbations. Although stable policies have the advantage of providing guaranteed stability, such as reaching a specific target state at a certain point in time, or arbitrarily approaching it, there is usually no indication of which stable policy will do this efficiently. System 150 is advantageously provided with a projection function that maps the input to a space of stable policies defined for the machine, for example, for the dynamical system. The projection function maps the original control policy space to the stable control policy space of the predefined stable control policies. Thus, by applying the projection, the original control signal is mapped to a stable control signal.

[0093] Therefore, the neural network is free to learn any raw control signal it deems suitable; however, the output of the neural network is passed through a projection function, such as one contained in a projection layer. This results in a neural network with a large parameter space and efficient learning capabilities, but this neural network guarantees to provide a control policy that satisfies certain conditions, such as stability criteria. Essentially, the resulting neural network is free to learn any control policy, as long as it is stable.

[0094] For example, processor 150 may be configured with neural network unit 152 and projection unit 154. For example, neural network unit 152 may be configured to apply neural network 130 to sensor signals to obtain raw control signals. Note that if neural network 130 has been trained together with the projection layer, the raw control signals cannot be directly applied to machine 110. Instead, the raw control signals can be considered as indices or pointers to the stable policy space.

[0095] Neural network unit 152 and projection unit 154 can be configured to fix a target state, such as a static state, but instead, they can be configured to accept the target state as input and to control the computer-controlled machine 110 to reach the target state. For example, the target state can be a parameter of a projection function, but used as neural network input for neural network 130.

[0096] The sensor signal input to the neural network 130 and the raw control signal output of the neural network 130 can be represented as, for example, a vector comprising real-valued numbers. For example, the values ​​can be scaled between 0 and 1. For example, the neural network can include known neural network layers such as convolutional layers, ReLU layers, pooling layers, etc. The dimension of the raw control signal can be the same as the dimension-stable control space. Typically, the dimension of the raw control signal can be the same as the output dimension of the projection function. For example, in an embodiment, the input to the neural network includes image data, and the neural network includes convolutional layers.

[0097] The output of the projection function represents a stable control signal, such as the input to a computer-controlled machine 110. For example, the stable control signal can be a vector. Elements of the vector can be relative actions, such as increasing distance, voltage, angle, etc. Elements of the vector can also be absolute actions, such as setting speed or distance to a specific value.

[0098] A projection function can be configured to map control policies, such as arbitrary control policies, to a space of stable control policies. Such a space can be defined in various ways. For example, stability equations can typically be imposed for stable control policies. For instance, for a stable control policy, it might be required that the Lyapunov function is decreasing. As a mapping from one space to another, a projection can be considered a continuous, piecewise differentiable function; this means that learning algorithms can propagate through the projection function.

[0099] Depending on the machine's dynamics, various projection layers are possible. An example of a projection layer would be to limit the control signal to a safe value. For example, the original control signal... It can be allowed, as long as it is for vectors and threshold , However, if the original control signal Exceeding these safety limits reduces its value, for example, compared to the value... Proportional. For example, for sensor input x, stable control signal... Original control signals Vector Sum The following projection functions can be defined. :

[0100]

[0101] Neural network 130 can implement functions Note that the projection layer described above can be implemented within a neural network layer.

[0102] In another example, the projection function can be used to solve an optimization problem. The solution to the optimization problem is piecewise differentiable with respect to the original control signal. In general, the projection function can be chosen relative to the dynamical system. The dynamical system can be defined, for example, as a set of stable control policies in control theory. A neural network can be taught to select an appropriate stable control policy to reduce the loss function.

[0103] Figure 3 An example embodiment of the controller 300 is illustrated schematically. For example, Figure 3The controller 300 can be used to control the machine 110. The controller 300 may include a processor system 330, a memory 340, and a communication interface 350. The controller 300 can communicate with the machine 110, external storage devices, input devices, output devices, and / or one or more sensors via a computer network. The computer network may be the Internet, an intranet, a LAN, a WLAN, etc. The system includes a connection interface arranged to communicate as needed within or outside the system. For example, the connection interface may include connectors—such as wired connectors, Ethernet connectors, optical connectors, etc.—or wireless connectors—such as antennas, such as Wi-Fi, 4G, or 5G antennas.

[0104] For example, the execution of controllers such as controller 100, controller 300, etc., can be implemented in a processor system, such as one or more processor circuits, for example, a microprocessor, examples of which are shown herein. Figure 1 and Figure 3 The diagrams illustrate functional units that can be functional units of a processor system. For example, the diagrams can serve as blueprints for possible functional organization of a processor system. In these diagrams, processor circuitry(s) is not shown as separate from the units. For example, the illustrated functional units can be implemented wholly or partially as computer instructions stored at the controller, for example, in electronic memory, and can be executed by the controller's microprocessor. In hybrid embodiments, the functional units are partially implemented in hardware, for example as coprocessors, such as neural network coprocessors, and partially stored in software and executed on the controller. Network parameters and / or training data can be stored locally or can be stored in cloud storage.

[0105] Figure 4 An example embodiment of a training system 101 for training a neural network for use in controlling a computer-controlled machine is illustrated schematically. For example, training system 101 can be used to train neural network 130 used in controller 100. If desired, controller 300 can also be adapted to run training system 101 or, instead, controller 100, by providing access to training data and / or software configured to perform training functions.

[0106] For example, training system 101 may include a training data interface for accessing training data. For example, training data may include sensor signals representing a sensor system that senses a computer-controlled machine. Sensor data may be obtained, in whole or in part, from the actual machine 110, for example, through measurements using the sensor system; sensor data may also be obtained, in whole or in part, through simulation, such as a simulation of machine 110.

[0107] System 101 may include a neural network storage device configured to store a neural network that defines an initial control strategy for a computer-controlled machine, such as a storage device in controller 100. Training system 101 is configured for the neural network and projection function, for example, as controller 100. For example, system 101 may include a neural network unit 152 and a projection unit 154, as in controller 100, except that the neural network of neural network unit 152 has not yet been fully trained.

[0108] Training system 101, such as its processor system, can be configured to apply a neural network to sensor signals, such as those obtained from training data 170, to obtain the neural network output, such as the original control signal, and apply a projection function to it to obtain a stable control signal. Because of the projection layer, the resulting control signal will be stable, but this does not mean it has any advantage. Therefore, a loss function can be applied to the stable control signal to obtain a value representing the goodness of the control policy defined by the combination of the neural network and the projection layer. For example, a machine can be simulated over a time period and its progress toward a target state can be determined. During this time period, actions may be updated or not. The neural network parameters are then modified to improve the loss function. For example, backpropagation can be applied to the loss function, the projection layer, and the neural network.

[0109] Embodiments of system 100 or 101 may include a communication interface, which may be selected from various alternatives. For example, the interface may be a network interface to a local area network or a wide area network such as the Internet, a storage interface to an internal or external data storage device, a keyboard, an application programming interface (API), etc. Systems 100 and 101 may have a user interface, which may include known components such as one or more buttons, a keyboard, a display, a touchscreen, etc. The user interface may be configured to accommodate user interactions for configuring the system, training a network on a training set, or applying the system to new sensor data, setting target states, setting trajectories, etc.

[0110] Storage devices can be implemented as electronic storage devices such as flash memory, or magnetic storage devices such as hard disks. A storage device can include multiple discrete memories that together constitute a storage device. A storage device can include temporary storage, such as RAM. A storage device can be cloud storage. A storage device can be used to store neural networks, software, training data, and so on.

[0111] Systems 100 and / or 101, 110 can be implemented in a single device or in a potentially distributed system. Typically, systems 100 and 101 include a microprocessor that executes appropriate software stored at the system; for example, this software may have been downloaded and / or stored in a corresponding memory, such as volatile memory like RAM or non-volatile memory like flash memory. Alternatively, the system can be implemented wholly or partially in programmable logic, such as a field-programmable gate array (FPGA). The system can be implemented wholly or partially as a so-called application-specific integrated circuit (ASIC), such as an integrated circuit (IC) customized for its specific purpose. For example, the circuit can be implemented in CMOS using hardware description languages ​​such as Verilog, VHDL, etc. In particular, systems 100 and 101 may include circuitry for evaluating neural networks.

[0112] Processor circuitry can be implemented in a distributed manner, for example, as multiple sub-processor circuits. Storage devices can be distributed across multiple distributed sub-storage devices. Some or all of the memory can be electronic memory, magnetic memory, etc. For example, storage devices can have volatile and non-volatile components. Some storage devices may be read-only.

[0113] The following illustrations show several additional optional refinements, details, and embodiments.

[0114] Consider controlling machines characterized as nonlinear, continuous-time dynamical systems, such as (1) of which Marked in time t state, It is a control input. It is an external (possibly random) disturbance term, and in which Marked in time t status x The time derivatives of these dynamics can be written in alternative, potentially time-varying linearized forms, which may yield robust control specifications that guarantee system stability. Given such robust control specifications, one can machine nonlinear policies, such as those based on deep neural networks, which can be proven to satisfy these specifications while optimizing some objective of interest.

[0115] In an embodiment, the learning of a provably robust nonlinear controller with reinforcement learning can be performed as follows: considering nonlinearity, such as that based on neural networks... Parameterized strategy class It can learn parameters, such as those of a neural network. This ensures that the projection of the final policy onto a provably stable set of controllers optimizes a control objective (within an infinite time range). Formally, people search for... To optimize

[0116]

[0117] in, It is a performance target. Characterizes the set of (potentially nonlinear) policies that are stabilized under a given robust control specification, and The projection onto this set. For example, It is the original control signal, such as the output of a neural network, and It is a stable control signal obtained by projecting and mapping the original control signal.

[0118] The projection operator can be implemented in a differentiable manner, for example, by using differentiable convex optimization layers, see, for example, the paper "Differentiable Convex Optimization Layers" by A. Agrawal, B. Amos, S. Barratt, S. Boyd, S. Diamond, and J. Zico Kolter. Therefore, to optimize this problem, one can train the neural network by projecting its results onto a set of stability criteria and then propagating the gradients of the corresponding loss through both the projection and the neural network, using any of these methods (e.g., direct policy search or virtually any reinforcement learning algorithm). The above problem can be viewed as having an infinite time range and being continuous in time, but in practice, it can be optimized in discrete time over a finite time range. By transforming the output of the neural network, one can adopt an expressive policy class to optimize the target of interest, while ensuring that the final policy will stabilize the system during both training and testing.

[0119] Besides being designed for stability, it is often desirable to design controllers that optimize certain performance objectives. Different types of objectives can be employed in different settings. For example, in In this paradigm, the objective is to minimize the output of a given system relative to any (potentially unbounded) perturbation. The ratio between norms, i.e., the perturbation-to-output mapping. Gain. For time-dependent signals. , of Norm can be defined as For example, considering the traditional "linear quadratic regulator" (LQR) objective, for some... and It is given by the following formula

[0120]

[0121] Minimizing the LQR cost while obeying the relevant dynamic equations can be transformed into a convex optimization problem.

[0122]

[0123] in This is the set of constraints characterizing a linear controller that ensures the stability of the dynamics within the relevant domain, as described below. In this example, Q yes s × s Positive definite matrix, and R yes a × a Positive definite matrix.

[0124] Equation 4 defines the solution to the LQR problem. Given some specific dynamic description, such as a norm-bounded LDI or a polyhedron, the optimization can be solved. Thus, one can obtain the optimal linear policy and the Lyapunov function. With the Lyapunov function, this defines a stable policy set that includes the optimal linear policy, but also includes nonlinear policies. The original control output of the neural network can be projected onto this policy set that stabilizes the system, as has been demonstrated, for example, by solving equation (4) to obtain the Lyapunov function.

[0125] Consider a nonlinear, continuous-time dynamical system as presented in the governing equation (1). It is often possible and convenient to write the dynamics in the following alternative (potentially time-varying) linearized form.

[0126]

[0127] in , and , and among them It is a term that can capture both deviations from linearity and any external disturbances. Furthermore, It can depend on itself and Although this dependency is omitted from the notation for brevity, this type of model is called Linear Differential Inclusion (LDI). Several sub-cases of this model encompass many existing traditional robust control methods. Two such cases are discussed in detail below.

[0128] LDI with bounded norm.

[0129] In this setting, the dynamics are assumed to be of the following form.

[0130]

[0131] in A , B andG It is time-invariant, and the disturbance term It is arbitrary, but it is known to obey certain norm boundedness conditions, that is, for some (again, time-invariant) matrices... ,

[0132]

[0133] For the purposes of the controller, additional regional distinctions are included. Special cases will be useful.

[0134] Assuming a time-invariant linear control strategy One can formulate a specification for a policy set via a set of linear matrix inequalities, which will stabilize the system under any worst-case perturbation of that class. At a high level, these methods will yield controller gains. K and the second-order Lyapunov function This ensures that the Lyapunov function guarantees the exponential stability of the resulting system under any perturbation that follows the norm bound (7), where for certain design parameters... Exponential stability is defined by the following conditions.

[0135] .

[0136] It can be mathematically derived that if one can find matrices that satisfy the following linear matrix inequalities... , and

[0137]

[0138] but and These are the gain of the stabilized linear controller and the quadratic Lyapunov function, respectively, which guarantee the exponential stability of NLDI(6)–(7). The above equations provide the convex set of the stabilized linear controller. Then the convex set can be used in the optimization problem (4). To design a controller that further optimizes the LQR objective. For this domain, one can additionally... Optimized above.

[0139] LDI polyhedron.

[0140] Another setting of interest is the polyhedral LDI (PLDI) setting. Here, the dynamics take the following form.

[0141]

[0142] in and They are matrices that can change arbitrarily over time and must only be subject to the constraint that they lie within the convex hull of some point set.

[0143]

[0144] Among them, for , , ,and Indicates the convex hull.

[0145] Similar to the above, one can solve for variables and The above set of linear matrix inequalities is used to design a stable linear controller. and the second-order Lyapunov function , specifically

[0146] .

[0147] As mentioned earlier, the resulting controller and quadratic Lyapunov function are derived from... and Parameterization. The above equations again provide the convex set of the stabilized linear controller. It can be used, for example, in equation (4).

[0148] Nonlinear control policies, potentially parameterized by deep neural networks, can be obtained that are guaranteed to obey the same stability conditions enforced by the robustness specification illustrated above. Although it is difficult to write out conditions that generally characterize the stability of nonlinear controllers globally, sufficient conditions for stability can be created, for example, by ensuring that stability criteria hold, such as by reducing the Lyapunov function of a given robustness specification by a given policy.

[0149] Output of the nominal nonlinear controller For example, the raw control signal generated by a neural network in response to sensor signals can be projected into a control action space that is provably stable under a given Lyapunov function and robustness specification. Since the details of this projection will vary depending on the type of robustness specification, several example projections are given for different settings. To simplify the notation, x and u of t Dependencies are suppressed, but note that these are still continuous time quantities as before.

[0150] Stability in NLDI

[0151] To ensure the exponential stability of NLDI, one can ensure that the control policy satisfies a sufficiently reduced Lyapunov function for any perturbation located within the allowable set (7). Given such a Lyapunov function—which can be obtained as described above—one can derive the convex set of the nonlinear control strategy that satisfies this criterion. For example, one can mathematically derive the stability parameters for NLDI systems (6)–(7). and having satisfying (9) P Lyapunov function People can Defined as for a given state Control action set that satisfies the following formula

[0152] .

[0153] therefore, It satisfies the sufficient reduction criterion The non-empty set of the controller. Furthermore, since the above inequality represents a second-order cone constraint, therefore... yes A convex set in [the context of a mathematical expression]. In the definition... In this situation, people are concerned about Given a value, the projection function can be obtained by solving problem (4) under constraint (9). P Therefore, people have a certain... x It can be achieved through nominal strategy The following stabilization strategy is obtained by projecting the image.

[0154] .

[0155] This projection can then be used in the loop of neural network optimization (2), i.e., where The projection does not necessarily have a closed form, but it can be achieved by optimization using, for example, differentiable convex optimization layers from the convex optimization library CVXPY, see, for example, “Differentiable Convex Optimization Layers” by A. Agrawal, B. Amos, S. Barratt, S. Boyd, S. Diamond and J. Zico Kolter.

[0156] exist Stability in NLDI under certain conditions

[0157] The special cases of the above problems are among them. In the case of disturbances, that is,w The norm bound is independent of the control policy. This form of NLDI appears in many common settings, such as... w This describes settings where the linearization error is characterized, but the original dynamics depend only linearly on the controller. In this case, the projection of the nominal policy onto the set of stable controllers can be derived in a closed-form. Under this set of forms, the stable controller is first proposed. In this situation, people can Defined as for a given state Control action set that satisfies the following formula ,

[0158] .

[0159] The above inequality represents a linear constraint, such that... yes The convex set in.

[0160] Interestingly, people's views on Given a value, the optimization problem (4) can be solved under constraint (9) and then... Projecting onto the corresponding constraint set to obtain the projection P In this specific case, the projection has a closed form and can be implemented via a single ReLU operation. Specifically, in the definition and In this situation, people obtain

[0161] .

[0162] As previously stated, this projection can then be used in neural network optimization (2), i.e., where .

[0163] Stability in PLDI

[0164] To ensure the exponential stability of PLDI, one can similarly ensure that the control policy is effective for any (11) in the allowable set. To sufficiently reduce the Lyapunov function A convex set of nonlinear control strategies that satisfy this criterion can be derived. For example, consider the PLDI system (10)–(11), the stability parameters and having satisfying (12) P Lyapunov function .Will Defined as for a given state Control action set that satisfies the following formula

[0165] .

[0166] It satisfies the sufficient reduction criterion The controller's non-empty set. Furthermore, since the above inequality represents a linear constraint, therefore... yes The convex set in.

[0167] Two detailed embodiments are described below, demonstrating improved stability even against disturbances: a pushrod task and a quadcopter domain.

[0168] Pushcart lever.

[0169] In the cart-pushing task, the objective is to balance an inverted pendulum that is stationary on top of a cart. The state of the system can be defined as follows: ,in It's the stroller location, and It is the angular displacement of the pendulum relative to its perpendicular position. People sought to address this by applying a horizontal force to the cart. To enable the system to Stabilization is achieved. For length... and quality The pendulum and its mass The dynamics of the system for the trolley are given by the following equation:

[0170]

[0171] in It is due to the acceleration caused by gravity. This system can be defined... The system is then linearized about its equilibrium point and written as NLDI.

[0172]

[0173] in It is the linearization error. Assume... x and u Within the neighborhood of the origin, this linearization error can be obtained numerically as a matrix. C and D The system is bounded. In this case, the dynamical system is defined as NLDI because the NLDI formula produces a much smaller problem description. However, it is also possible to use polyhedral uncertainties. For the LQR objective, the matrix... Q and R It is randomly generated.

[0174] Planar quadcopter.

[0175] In a planar quadcopter setup, the goal is to stabilize the quadcopter in a two-dimensional plane. The state of the system can be defined as... ,in( ) is the position of the quadcopter in the vertical plane, and It is its roll (i.e., the angle relative to the horizontal position). One seeks to control the force provided by the left and right thrusters of the quadcopter. The amount to make the system in Stabilization. Assuming action. u The baseline force provided by the thruster by default The addition of [something] is to prevent the quadcopter from crashing. For vehicles with mass... m Thruster torque arm and moment of inertia about the roll axis J The dynamics of the quadcopter system are given by the following equation:

[0176]

[0177] in The system can be linearized in a similar manner to that used for the pusher arm, for example, as in equation (19). The dynamics are related to... u The dependence is linear, which makes people have a certain degree of confidence in the final NLDI. As in the previous setup, one could randomly generate the LQR target matrix. Q and R For example, by sampling each entry of each matrix from a standard normal distribution iid.

[0178] In the examples of both the pushrod and the quadcopter, the nominal nonlinear control strategy can be constructed as follows: ,in K Obtained through robust LQR optimization, and among which It is a neural network. To construct the relevant projection, one can use... K When obtained P Value. Within it. In cases such as (e.g., for the cart handle example), the projection (13) can be implemented using a differentiable optimization layer. For example, for a quadcopter, projection (14) can be implemented via ReLU. Robust policies are trained via direct policy search. Each epoch comprises a 1-second time range discretized from 0.005 seconds, and the network is trained until its performance against the hold set has not improved over 100 epochs. Although direct policy search is used in this embodiment, the general approach is agnostic to the specific training method and can be deployed alongside other deep reinforcement learning paradigms.

[0179] The robust neural network-based approach is compared with a robust (linear) LQR controller, a non-robust neural network trained via direct policy search, and a standard non-robust (linear) LQR controller. This comparison is made not only in the original settings (e.g., under the original dynamics) but also against generated test-time perturbations. Performance was evaluated by minimizing the decrease in the Lyapunov function (see Appendix 7). All methods were evaluated over a time range of 1 second with a discretization of 0.005 seconds.

[0180] Table 1 shows the performance of these methods for the domain of interest. The results are reported as the integral of the quadratic loss over the test state set within a specified time frame, or marked "X" to indicate cases where the relevant method becomes unstable (and therefore the loss becomes inf, NaN, or many orders of magnitude larger than other methods). These results illustrate the fundamental advantages of robust NN methods. In all cases, robust NNs outperform robust LQR methods (e.g., linear controllers that also provide stability guarantees) for the original dynamics (which are the targets being optimized). Meanwhile, as expected, traditional (non-robust) LQR methods and non-robust NNs often perform better within the original nominal dynamics (because they are optimized for expected rather than worst-case performance). Original nominal dynamics are indicated by O; adversarial dynamics are indicated by A.

[0181]

[0182] However, it is worth noting that when adversarial perturbations are employed, these perturbations remain within the allowable norm bounds and are therefore effective perturbations, while non-robust LQR and neural network methods may diverge or perform very poorly. In contrast, both robust neural networks and robust LQR methods remain stable even under these perturbations.

[0183] Both robust and non-robust neural network methods converge to their final performance levels fairly quickly. Non-robust neural networks frequently become unstable under adversarial dynamics very early in the process. Overall, these results demonstrate that the embodiments can learn more expressive policies than traditional robust LQRs, while guaranteeing that these policies will be stable.

[0184] Figure 5An example embodiment of a control method (500) for generating control signals for a computer-controlled machine is illustrated schematically. Method 500 may be implemented by a computer. The computer-controlled machine comprises multiple parts interacting in a dynamic system. The control method may include...

[0185] - Receive (510) sensor signals from the sensor system of the computer-controlled machine, which indicate the current state of the computer-controlled machine.

[0186] - Apply a neural network (520) to the current sensor signal, the neural network defining the original control strategy for the computer-controlled machine, the neural network being configured to map the sensor signal to the original control signal.

[0187] - Apply the projection function (530) to the original control signal to obtain a stable control signal. This projection function maps the original control strategy space to the stable control strategy space of a predefined stable control strategy.

[0188] - Causes (540) stable control signals to control computer-controlled machines.

[0189] Figure 6 An example embodiment of a training method (600) for training a neural network for use in controlling a computer-controlled machine is illustrated schematically. The computer-controlled machine comprises multiple parts interacting in a dynamic system. Method 600 may be computer-implemented. The training method may include...

[0190] - Receive (610) sensor signals representing a sensor system sensing a computer-controlled machine, the sensor signals indicating the current state of the computer-controlled machine,

[0191] - Apply a neural network (620) to the current sensor signal, the neural network defining the original control strategy for the computer-controlled machine, the neural network being configured to map the sensor signal to the original control signal.

[0192] - Apply the projection function (630) to the original control signal to obtain a stable control signal. This projection function maps the original control strategy space to the stable control strategy space of a predefined stable control strategy.

[0193] - Calculate (640) the training parameters of the neural network for stabilizing the control signal and reducing the loss.

[0194] For example, the control and training methods can be computer-implemented. For example, accessing training data and / or receiving input data can be accomplished using a communication interface, such as an electronic interface, network interface, or memory interface. For example, storing or retrieving parameters can be accomplished using an electronic storage device, such as a memory or hard disk drive, storing network parameters. For example, applying the neural network to the training data and / or adjusting the stored parameters to train the network can be accomplished using an electronic computing device, such as a computer.

[0195] During training and / or application, a neural network can have multiple layers, which may include, for example, convolutional layers. For instance, a neural network can have at least 2, 5, 10, 15, 20, or 40 hidden layers or more, and so on. The number of neurons in a neural network can be, for example, at least 10, 100, 1000, 10000, 100000, 100000 or more, and so on.

[0196] Many different ways of performing this method are possible, as will be apparent to those skilled in the art. For example, the steps may be performed in the order shown, but the order of the steps may vary or some steps may be performed in parallel. Furthermore, other method steps may be inserted between steps. The inserted steps may represent a refinement of the method described herein, or they may be unrelated to the method. For example, some steps may be performed at least partially in parallel. Moreover, a given step may not be fully completed before the next step begins.

[0197] Embodiments of this method can be executed using software that includes instructions for causing a processor system to perform methods 500 and / or 600. The software may only include those steps taken by a specific sub-entity of the system. The software can be stored in a suitable storage medium, such as a hard disk, floppy disk, memory, optical disk, etc. The software can be transmitted as a signal along a wired or wireless data network, or using, for example, the Internet. The software can be made available for download and / or remote use on a server. Embodiments of this method can be executed using a bitstream arranged to be configured for programmable logic, such as a field-programmable gate array (FPGA), to perform the method.

[0198] It will be appreciated that the currently disclosed subject matter also extends to computer programs, particularly computer programs on or within a carrier, suitable for putting the currently disclosed subject matter into practice. The program may be in the form of source code, object code, intermediate source code, and object code, such as in a partially compiled form, or in any other form suitable for use in an implementation of an embodiment of the method. Embodiments related to the computer program product include computer-executable instructions corresponding to each processing step of at least one of the elaborated methods. These instructions may be subdivided into subroutines and / or stored in one or more files that may be statically or dynamically linked. Another embodiment related to the computer program product includes computer-executable instructions corresponding to each device, unit, and / or portion of at least one of the elaborated systems and / or products.

[0199] Figure 7a A computer-readable medium 1000 is illustrated having a writable portion 1010 including a computer program 1020, which includes instructions for inducing a processor system to perform control and / or training methods according to embodiments. The computer program 1020 may be embodied on the computer-readable medium 1000 as a physical mark or by magnetization of the computer-readable medium 1000. However, any other suitable embodiments are conceivable. Furthermore, it will be appreciated that although the computer-readable medium 1000 is shown herein as an optical disc, the computer-readable medium 1000 may be any suitable computer-readable medium, such as a hard disk, solid-state storage, flash memory, etc., and may be non-recordable or recordable. The computer program 1020 includes instructions for inducing a processor system to perform the control and / or training methods.

[0200] Figure 7b A schematic representation of a processor system 1140 according to an embodiment of a controller and / or training system or device is shown. The processor system includes one or more integrated circuits 1110. Figure 7b The diagram schematically illustrates the architecture of one or more integrated circuits 1110. Circuit 1110 includes a processing unit 1120, such as a CPU, for running computer program components to perform methods according to embodiments and / or implement modules or units thereof. Circuit 1110 includes a memory 1122 for storing programming code, data, etc. A portion of the memory 1122 may be read-only. Circuit 1110 may include a communication element 1126, such as an antenna, a connector, or both, etc. Circuit 1110 may include an application-specific integrated circuit 1124 for performing some or all of the processing defined in the method. Processor 1120, memory 1122, application-specific IC 1124, and communication element 1126 may be interconnected to each other via interconnect 1130 (e.g., a bus). Processor system 1110 may be arranged for contact and / or contactless communication using antennas and / or connectors, respectively.

[0201] For example, in one embodiment, the processor system 1140, such as a controller and / or training system / device, may include processor circuitry and memory circuitry, with the processor arranged to execute software stored in the memory circuitry. For example, the processor circuitry may be an Intel Core i7 processor, an ARM Cortex-R8, etc. In another embodiment, the processor circuitry may be an ARM Cortex M0. The memory circuitry may be ROM circuitry or non-volatile memory such as flash memory. Alternatively, the memory circuitry may be volatile memory, such as SRAM memory. In the latter case, the device may include a non-volatile software interface arranged to provide the software, such as a hard disk drive, a network interface, etc.

[0202] A processor system may include multiple microprocessors configured to independently execute the methods described herein, or to execute steps or subroutines of the methods described herein, such that the multiple processors cooperate to implement the functions described herein. Furthermore, in the case of implementing devices and / or systems in a cloud computing system, various hardware components may belong to separate machines. For example, a processor may include a first processor in a first server and a second processor in a second server.

[0203] It should be noted that the embodiments mentioned above are illustrative and not limiting of the subject matter currently disclosed, and those skilled in the art will be able to devise many alternative embodiments.

[0204] In the claims, any reference marks placed between parentheses should not be construed as limiting the claims. The use of the verb "comprising" and its variations does not exclude the presence of elements or steps other than those recited in the claims. The article "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. Expressions such as "at least one of..." when preceding a list of elements indicate the selection of all elements or any subset of elements from that list. For example, the expression "at least one of A, B, and C" should be understood to include only A, only B, only C, both A and B, both A and C, both B and C, or all A, B, and C. The subject matter currently disclosed can be implemented by hardware comprising several different elements, and by a suitably programmed computer. In a device claim enumerating several components, several of these components can be embodied by the same item of hardware. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used advantageously.

[0205] In the claims, reference numerals enclosed in parentheses refer to reference numerals in the accompanying drawings of exemplary embodiments or formulas of embodiments, thus increasing the comprehensibility of the claims. These reference numerals should not be construed as limiting the claims.

Claims

1. A controller (100) for generating control signals for a computer-controlled machine (110), the computer-controlled machine (110) comprising a plurality of interacting components (210), the controller (100) comprising - a sensor input interface (120) configured to receive sensor signals from a sensor system (122) sensing the computer-controlled machine, the sensor signals being indicative of a current state of the computer-controlled machine - a neural network storage (130) configured for storing a trained neural network, the neural network defining an original control policy for the computer-controlled machine, - a processor system (150) configured to - apply the neural network to the current sensor signals, the neural network being configured to map the sensor signals to original control signals, - apply a projection function to the original control signals to obtain stable control signals, the projection function mapping an original control policy space to a stable control policy space of predefined stable control policies, - cause the stable control signals to control the computer-controllable machine, characterized in that For a sensor input x, a stable control signal π(x), a raw control signal a vector p and a threshold value h, the projection function is defined as in the setup, the dynamics are of the form where A, B and G are time-invariant and the disturbance term w(t) is arbitrary but known to satisfy certain norm boundedness conditions as follows: ||w(t)||2≤||Cx(t) + Du(t)||2, for time-invariant matrices and D = 0, where the vector p is defined as p T = 2x T PB and the threshold is defined as η≡ -ax T Px- 2||G T Px||2||Cx||2- 2x T PAx, where a > 0 is a given value and where P = S -1 where and satisfy the following linear matrix inequality 2. The controller of claim 1, wherein, the neural network is repeatedly applied to current sensor signals obtained from the sensor system (122) to repeatedly obtain stable control signals for controlling the computer-controllable machine.

3. The controller of any of the preceding claims 1-2, wherein, the sensor system (122) comprises one or more of an image sensor, radar, lidar, pressure sensor.

4. The controller according to any of the preceding claims 1-2, wherein the neural network is further configured to receive a target state as input and configured to control the computer-controllable machine to reach the target state.

5. The controller according to any of the preceding claims 1-2, wherein the plurality of components interact according to a dynamics system.

6. The controller according to any of the preceding claims 1-2, wherein the computer-controllable machine comprises one or more of a robot (240), a vehicle, a household appliance, a power tool, a manufacturing machine, a drone.

7. The controller according to any of the preceding claims 1-2, wherein the stable control policy space is defined by a stability equation.

8. The controller according to any of the preceding claims 1-2, wherein the original control signals and the stable control signals are represented as vectors, elements of the vectors representing stable control signals representing inputs to the computer-controlled machine.

9. The controller according to any of the preceding claims 1-2, wherein the original signal is a sum of a linear controller (Kx) and a neural network .

10. A control method (500) for generating control signals for a computer-controlled machine, the computer-controlled machine comprising a plurality of components interacting in a dynamics system, the control method comprising - receiving (510) sensor signals from a sensor system (122) sensing the computer-controlled machine, the sensor signals being indicative of a current state of the computer-controlled machine - applying (520) a neural network to the current sensor signals, the neural network defining an original control policy for the computer-controlled machine, the neural network being configured to map the sensor signals to original control signals, - applying (530) a projection function to the original control signal to obtain a stable control signal, the projection function mapping the original control policy space to a stable control policy space of predefined stable control policies, - causing (540) the stable control signal to control a computer controllable machine, characterized in that For a sensor input x, a stable control signal π(x), a raw control signal a vector p and a threshold value h, the projection function is defined as in the setup, the dynamics are of the form where A, B and G are time-invariant and the disturbance term w(t) is arbitrary but known to satisfy certain norm boundedness conditions as follows: ||w(t)||2≤||Cx(t) + Du(t)||2for time-invariant matrices and D = 0, where the vector p is defined as p T ≡ 2x T PB and the threshold is defined as η≡ -αx T Px- 2||G T Px||2||Cx||2- 2x T PAx, where a > 0 is a given value and where P = S -1 where and satisfy the following linear matrix inequality 11. A training method (600) for training a neural network for use in controlling a computer controlled machine, the computer controlled machine comprising a plurality of components interacting in a dynamic system, the training method comprising - receiving (610) sensor signals representative of a sensor system (122) sensing the computer controlled machine, the sensor signals being indicative of a current state of the computer controlled machine, - applying (620) a neural network to the current sensor signals, the neural network defining an original control policy for the computer controlled machine, the neural network being configured to map the sensor signals to original control signals, - applying (630) a projection function to the original control signal to obtain a stable control signal, the projection function mapping the original control policy space to a stable control policy space of predefined stable control policies, - computing (640) a loss for the stable control signal and training parameters of the neural network reducing the loss, characterized in that For a sensor input x, a stable control signal π(x), a raw control signal a vector p and a threshold value h, the projection function is defined as in the setup, the dynamics are of the form where A, B and G are time-invariant and the disturbance term w(t) is arbitrary but known to satisfy certain norm boundedness conditions as follows: ||w(t)||2≤||Cx(t) + Du(t)||2for time-invariant matrices and D = 0, where the vector p is defined as p T = 2x T PB and the threshold is defined as η≡ -αx T Px- 2||G T Px||2||Cx||2- 2x T PAx, where a > 0 is a given value and where P = S -1 where and satisfy the following linear matrix inequality 12. A training system for training a neural network for use in controlling a computer controlled machine, the computer controlled machine comprising a plurality of components interacting in a dynamic system, the training system comprising - a training data interface for receiving sensor signals representative of a sensor system (122) sensing the computer controlled machine, the sensor signals being indicative of a current state of the computer controlled machine, - a neural network storage (130) configured for storing a neural network, the neural network defining an original control policy for the computer controlled machine, - a processor system (150) configured for - applying a neural network to the current sensor signals, the neural network defining an original control policy for the computer controlled machine, the neural network being configured to map the sensor signals to original control signals, - applying a projection function to the original control signal to obtain a stable control signal, the projection function mapping the original control policy space to a stable control policy space of predefined stable control policies, - computing a loss for the stable control signal and training parameters of the neural network reducing the loss, characterized in that For a sensor input x, a stable control signal π(x), a raw control signal a vector p and a threshold value h, the projection function is defined as in the setup, the dynamics are of the form where A, B and G are time-invariant and the disturbance term w(t) is arbitrary but known to satisfy certain norm boundedness conditions as follows: ||w(t)||2≤||Cx(t) + Du(t)||2 for time-invariant matrices and D = 0, where the vector p is defined as p T ≡ 2x T PB and the threshold is defined as η≡ -αx T Px-2||G T Px||2||Cx||2-2x T PAx, where a > 0 is a given value and where P = S -1 where and satisfy the following linear matrix inequality 13. A transitory or non-transitory computer readable medium (1000) comprising data (1020) representing instructions, which instructions, when executed by a processor system (150), cause the processor system (150) to perform a method according to claim 10 and / or 11.

Citation Information

Patent Citations

  • Stable adaptive neural network controller

    WO1993000618A1