Neural network controller
A neural network controller trained via neuro-evolution optimization addresses the limitations of traditional control methods by adapting to complex industrial systems, ensuring stable and efficient operation across varying conditions.
Patent Information
- Application Number
- PCT/GB2025/050045
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2025-01-10
- Publication Date
- 2025-07-17
AI Technical Summary
Existing control methods for complex industrial systems, such as PID controllers and Model Predictive Control (MPC), are limited in their ability to adapt to non-linear, time-variant, and uncertain processes, leading to inefficiencies and instability, while AI/ML controllers face challenges in data availability and computational intensity, making them impractical for APC applications.
A method involving a neural network controller generated through neuro-evolution optimization, trained on a system model with a training curriculum, verified against control challenges, and deployed in hardware tests to ensure robustness and adaptability.
The neural network controller effectively adapts to complex industrial systems, providing stable control across varying conditions and unanticipated events, outperforming traditional methods in efficiency and adaptability.
Smart Images

Figure GB2025050045_17072025_PF_FP_ABST
Abstract
Description
CONTROLLER TECHNICAL FIELD
[0001] The present disclosure relates to a controller for controlling a dynamic system, particularlysuited to industrial systems. Aspects of the invention relate to a method of generating a controller, a neural network for a controller, and a controller. BACKGROUND
[0002] In almost every industrial process, process controllers are required to stabilise and optimiseprocess performance.
[0003] At present, in a large number of industrial systems, proportional-integral-derivative (PID)controllers are used for process control. PID controllers control process variables such as pressure, temperature, and flow, among many others, using a feedback control loop mechanism, as shown by the block diagram in Figure 1. The PID controller compares the output process variable to a desired set point for that variable, and based on a set of linear mathematical functions, calculates a required input for the system to bring the process variable closer to the desired set point. PID controllers have proven to be very effective in controlling relatively simple processes involving controlled returns to set-points in single input / single output systems which have mostly linear dynamics and have low uncertainty, for example certain classes of valve control or motor control systems.
[0004] However, since PID controllers are based on relatively simple control strategies which aretypically tuned or optimised for a limited scope of operation, they are not suited to controlling morecomplex systems efficiently. If a process is non-linear, time-variant, has discontinuous statechanges, has unanticipated perturbations, or has a high number of inputs and outputs, PID controllers may not be able to accurately control the process and the control may start to degrade as the system moves beyond the narrow scope of optimisation. To enable PIDs to work for more complex systems, and dynamically adjust set-points in industrial processes, PIDs may be layered on top of each other (as cascaded controllers), or combined with more advanced methodologies such as fuzzy logic and Model Predictive Control (MPC)(discussed later). However, when using strategies where there are cascaded PIDs or PIDs combined with another control method, there is an inherent trade-off between set-point tracking and system stability. As such, more complex systems require advanced process control (APC), for which other control techniques must be used.
[0005] In some complex systems, a lookup table of PID controllers may be used for APC, whichallows different set-points to be included for changing states of the system. This is also known as gain-scheduling. In gain-scheduling, a lookup table is created that contains a database of PID controllers, each with different constant, or ‘set-point’ settings. Therefore, as the system changes, the PID controller used for controlling the system can switch, leading to a state switch within the system. However, state-switching requires that the system detects which PID to use based on current conditions, which may in itself be a difficult inference problem. Additionally, fast switching between one PID controller to another can lead to system instability and oscillations. Furthermore, gain scheduling is limited to a known range of variability in the system, and does not account for wide ranging uncertainties and nonlinearities. Therefore, the efficacy of PID lookup tables in controlling complex systems is limited.
[0006] A second option for APC, and currently the most widely used method for APC, is ModelPredictive Control (MPC). In MPC, a mathematical model of the process is generated, and this can be based on parametric physical models, on differential equations, or can be a data-driven model. During operation, the MPC iteratively calculates the best sequence of outputs based on measurements of the state of the system. However, even in applications where MPC works well(typical applications are when very high frequency response rates are not required), there aresignificant disadvantages.
[0007] MPC has low robustness: it is not practical to include any unexpected changes in the systembeing controlled (for example changes in equipment or equipment wear) mathematically in the model. Such changes, when they occur, require extensive model redevelopment or they lead todecay in the performance of the system.
[0008] A further important disadvantage of MPC in the context of a controller is that it iscomputationally intensive (the entire model must be re-run at each process step, and run continually in the control loop). It is therefore apparent that there are several limitations when using MPC for APC.
[0009] Recently, Artificial Intelligence and Machine Learning (AI / ML) techniques have beenconsidered as potential control methods for APC. However, systems requiring APC, for example power generation and chemical engineering systems, are frequently responsible for applications of substantial economic value, and often there is a human safety risk in such systems. Thus, a high degree of certainty in the system, and consequently in the controller, is essential. At present, AI / ML networks have not been widely adopted for APC for a number of reasons.
[0010] Firstly, AI / ML control methods for APC require large datasets to train the neural networks toensure that the networks perform satisfactorily on edge cases. For complex systems requiring APC, a huge amount of data would be required to optimally train the traditional AI / ML controller, since data sets for all possible eventualities are required. However, in industrial processes, it is too damaging, dangerous, and expensive to carry out system failures and obtain data for such eventualities, and so that data is simply not available. Indeed, if the AI / ML controller needs to be able to deal with a catastrophic failure, as is the case in industrial processes, there would need to be several thousand catastrophic failures to provide the data for training the network, which is simply not feasible.
[0011] Furthermore, training with large datasets results in large, many layered models with highcomputational intensity. Therefore, similar to MPC discussed above, networks are expensive to run,and the costs of running an AI / ML network may prohibit the use of the network.
[0012] Finally, AI / ML solutions often do not adapt well to real world environments when used in APC,which relates to there not being data available across the full range of eventualities for training the AI / ML controller, as discussed above. The AI / ML network may be trained, but then whenimplemented in a real-world process, fails because it does not extrapolate to this real-worldenvironment. Therefore, AI / ML controllers have difficulty adapting to process changes, unanticipatedevents, and modes of failure. Typically, this means that present AI / ML controllers are restricted toautopiloting stable conditions and providing suggestions to operators. For these reasons, AI / ML controllers that are currently available are not practical for APC, and thus have not been widely adopted for this application.
[0013] As outlined above, while APC is essential to ensure smooth running of various industrialprocesses, there is not yet an automatic method that is robust and cost effective, while beingcomputationally efficient. Additionally, many APC methods currently used in industry are notoptimised or tuned correctly. Optimisation is extremely important for delivering significant benefits, including increased throughput, increased yield, and reduction in energy consumption. An APC method that can be optimised for a particular industrial process is therefore highly desirable.
[0014] An objective of the current invention is therefore to provide a controller that is not subject tothe limitations described above, and that can transfer effectively into real-world applications to be used for APC of industrial systems. SUMMARY OF THE INVENTION
[0015] According to a first aspect of the present invention, there is provided a method of generatinga dynamic system controller for controlling a dynamic system, the dynamic system controller comprising a neural network, the method comprising: modelling the dynamic system to generate a system model; defining a training set of control challenges for evaluating candidate neural networks, the control challenges being derived from the system model; generating an optimised neural network using a neuro-evolution optimiser from a population of candidate neural networks, wherein the population of candidate neural networks are evaluated using the training set of control challenges; verifying the optimised network against a verification set of control challenges derived from the system model; deploying the optimised neural network into a hardware test controller for the system; evaluating performance of the hardware test controller in a hardware test; and, in the event the hardware test controller fails the hardware test, generating a further optimised neural network and repeating the verification, deployment and evaluation steps; or, in the event the hardware test controller passes the hardware test, deploying the optimised network into the controller for controlling the system.
[0016] The present invention provides a method a generating a dynamic system controller for adynamic system in which the system is initially modelled to define a system model. A set of training control challenges are then derived from the system model and used to evaluate a population of candidate neural networks.
[0017] The system model provides an environment in which the neural network can be trained, anda series of test / training scenarios can be created that would otherwise be impractical or too expensive to provide using a real dynamic system. A gradient-free optimiser in the form of a neuro- evolution optimiser is used to evaluate the candidate neural networks to generate an optimised neural network. The optimised neural network is then verified against a verification set of control challenges in order to detect overfitting and to test the performance of the network against more extreme challenges including challenges which exceed the expected difficulty in controlling the real dynamic system.
[0018] The optimised neural network, once verified, is then deployed into a hardware test controllerand a hardware test is undertaken. If the optimised neural network fails the hardware test a further optimised neural network may be generated and, for example, the system model may be adjusted based on the reasons for the hardware test failure. If the optimised neural network passes the hardware test then it may be deployed into a dynamic system controller for controlling the dynamic system.
[0019] It is noted that the generation of the optimised neural network may also comprise anassessment against a validation set of control challenges in order to detect overfitting or under- generalisation of the network. Following validation the method may move to the verification process described above which may test extrapolation behaviour of the optimised network and its performance against edge cases.
[0020] The dynamic system may comprise an actuator for actuating a joint within a robotic arm, afeedback control loop mechanism for controlling pressure, temperature of flow variables, a motor within an unmanned aerial vehicle, any system where a PID controller or MPC controller may conventionally be used.
[0021] Conveniently, generating a further optimised neural network may comprise updating thesystem model and / or defining a further set of control challenges. In this manner the reasons for a previous optimised neural network failing the hardware test can be incorporated into the system model or the curriculum control challenges for the next optimised neural network. The hardware test may comprise testing the dynamic system controller in a range of operational scenarios and the further set of control challenges may be defined in dependence on the operational scenario that resulted in the failed hardware test.
[0022] The neural network may comprise neurons and synapses configured to form a directed graphstructure and wherein at least some synapses within the neural network are configured to have strengths that can change during network operation. The graph structure of the neural network allows loops and arbitrary directed graphs to be encoded with resultant less neural structure than in traditional neural networks. This enables the neural network to be able to learn and handle more complex relationships with fewer computational units.
[0023] The training set of control challenges may comprise a series of control challenges thatsuccessively increase in complexity. A training curriculum may be defined to guide the evolution of the neural network.
[0024] The verification set of control challenges may comprise control challenges that exceed theexpected difficulty of controlling the dynamic system. The neural network may be challenged on more extreme scenarios that might be expected to exceed the expected real world challenges that the controller might need to deal with, e.g. the failure of multiple rotors in an unmanned aerial vehicle or a robotic leg that encounters an unexpected object on the ground.
[0025] The dynamic system controller may be configured to receive a current system variable of thedynamic system and to control a control variable of the dynamic system towards a setpoint value by outputting a control signal. The dynamic system controller may also be configured to receive further constraints, e.g. regions of state space that the controller should either avoid or must not access.
[0026] The dynamic system may comprise an actuator and the dynamic system controller may beconfigured to output the control signal to actuate the actuator towards the set point value. The dynamic system may further control a plurality of actuators and the controller may be configured to output control signals to actuate the plurality of actuators.
[0027] In an embodiment there is provided a robot comprising a robot limb with a joint configured tobe actuated by a motor, wherein a dynamic system controller generated according to the first aspectof the present invention is configured to control the motor to actuate the joint.
[0028] According to a second aspect of the present invention there is provided a method of using adynamic system controller generated according to the first aspect of the present invention to control a dynamic system.
[0029] According to a third aspect of the present invention there is provided the use of a dynamicsystem controller generated according to the first aspect of the present invention to control a dynamic system.
[0030] According to a fourth aspect of the present invention there is provided a dynamic systemcontroller generated according to the first aspect of the present invention. The dynamic system controller may comprise an input for receiving a set point value for a system parameter and a current system value, and an output for outputting a control signal to control the dynamic system towards the set point value wherein the optimised neural network is configured to generate the control signal in dependence on the received set point value and current system value.
[0031] According to a fifth aspect of the present invention there is provided a neural network for adynamic system controller for a dynamic system, the neural network comprising: a plurality ofneurons n and synapses e configured to form a directed graph structure, wherein at least somesynapses within the neural network are configured to have synapse strengths that can change during network operation.
[0032] Each neuron n at an execution clock instance, may produces an activation value from inputactivations calculated at a last execution clock instance. Each neuron may be configured on execution such that each neuron aggregates input activation values received from input neurons to calculate an activation value, input activation values being weighted by connecting synapse strengths.
[0033] Synapse strengths of at least some of the synapses may be configured to change duringnetwork operation based on previous neuron activation values.
[0034] The synapses e may comprise a set of processor synapses, each processor synapse beingconfigured to contribute to an activation of a neuron and a set of modulator synapses, each modulator synapse being configured to contribute to an instantaneous learning rate of a neuron. This allows the network to modulate its function in a form of unsupervised learning during operation and thereby adapt to unanticipated situations.
[0035] Neurons may comprise an activation function and an instantaneous learning rate and thesynapses may comprise learning rules encoded as a function of current synapse strength, pre and post synaptic activation and the learning rate, the activation function, instantaneous learning rate and learning rules being used to calculate the strength of each synapse.
[0036] The network may encode a multidigraph structure. In other words, network connections maybe directed, may be parallel (more than one edge can connect two nodes), and may form cycles or loops (an edge from a node to itself)
[0037] Initial synapse strengths may be configured to be quantised. Each neuron may comprise anactivation function which comprises a gain parameter.
[0038] According to a sixth aspect of the present invention there is provided a dynamic systemcontroller for controlling a dynamic system, comprising a neural network according to the fifth aspect of the present invention.
[0039] The dynamic system controller may comprise an input for receiving dynamic systemoperational data and an output for outputting a control signal to control the dynamic system, wherein neural network may be configured to receive the dynamic system operational data as its input and the controller may be configured to determine the control signal from the output of the neural network.
[0040] According to a seventh aspect of the present invention there is provided a controller forcontrolling a dynamic system, the controller comprising: an input for receiving a setpoint value for a control variable of the dynamic system and a current system variable of the dynamic system; processing means for calculating a control signal for controlling the dynamic system toward the set point value; an output for outputting the control signal wherein the processing means comprises atrained neural network having neurons and synapses connected in a directed graph structure andwherein at least some of the synapses within the neural network are configured to have synapse strengths that can change during network operation.
[0041] According to an eighth aspect of the present invention there is provided method of generatinga dynamic system controller for controlling a dynamic system, the dynamic system controller comprising a neural network, the method comprising: modelling the dynamic system to generate a system model; defining a training set of control challenges for evaluating candidate neural networks, the control challenges being derived from the system model; generating an optimised neural network using a gradient-free optimiser from a population of candidate neural networks, wherein the population of candidate neural networks are evaluated using the training set of control challenges; verifying the optimised network against a verification set of control challenges derived from the system model; deploying the optimised neural network into a hardware test controller for the system; evaluating performance of the hardware test controller in a hardware test; and, in the event the hardware test controller fails the hardware test, generating a further optimised neural network and repeating the verification, deployment and evaluation steps; or, in the event the hardware test controller passes the hardware test, deploying the optimised network into the dynamic system controller for controlling the dynamic system.
[0042] Preferred features of the various aspects of the invention described above may be applied toother aspects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] One or more embodiments of the invention will now be described, by way of example only,with reference to the accompanying drawings, in which:
[0044] Figure 1 is a block diagram of a PID controller;
[0045] Figure 2 is a block diagram showing an overview of a neural controller in accordance with anembodiment of the present invention;
[0046] Figure 3 is a flowchart showing the method of developing and training the neural controllerof Figure 2;
[0047] Figure 4 is a schematic diagram illustrating clocking of a neural network, in accordance withan embodiment of the present invention;
[0048] Figure 5 is a schematic diagram showing an example of a neural network in an arbitrarydirected graph structure;
[0049] Figure 6 is a schematic diagram showing an example of a neural network with processor andmodulator synapses;
[0050] Figure 7 is a schematic diagram showing an example of a compute neuron, illustrating neuronand synaptic properties in accordance with an embodiment of the present invention;
[0051] Figure 8 is a graph showing an example of a clamp function in accordance with anembodiment of the present invention;
[0052] Figure 9 is a flowchart showing a method of developing and training a neural network forsubsequent use in creating the neural controller of Figure 2, in accordance with embodiments of thepresent invention;
[0053] Figure 10 is a flowchart showing the method of Figure 9 incorporating feedback;
[0054] Figure 11 is a block diagram showing a system for creating the neural controller of Figure 2in accordance with an embodiment of the present invention;
[0055] Figure 12 shows examples of easing functions that may be used to construct waveforms fordevelopment of a curriculum, according to an example of the present invention;
[0056] Figure 13 shows example waveforms created using the easing functions of Figure 12;
[0057] Figure 14 is an schematic diagram showing the hardware set-up for conducting a swing-upexperiment for a robot arm;
[0058] Figure 15 shows the result of using a neural controller according to embodiments of thepresent invention to control the robot arm for the experiment illustrated in Figure 14;
[0059] Figure 16 shows the result of using a PID controller to control the robot arm for the experimentillustrated in Figure 14;
[0060] Figure 17 is an schematic diagram showing the hardware set up for conducting anexperiment testing control of a multi-joint robot limb;
[0061] Figure 18 shows the result of using a neural controller according to embodiments of thepresent invention to control the robot limb for the experiment illustrated in Figure 17;
[0062] Figure 19 is a schematic diagram showing the hardware set up for conducting an experimenttesting control of the multi-joint robot limb in a free swing experiment; and
[0063] Figure 20 shows the result of using a neural controller according to embodiments of thepresent invention to control the robot-limb for the experiment illustrated in Figure 19. DETAILED DESCRIPTION
[0064] An overview of a neural controller 20 for dynamic systems (also referred to herein as a“dynamic system controller” or “controller”) in accordance with an embodiment of the presentinvention operating to control a dynamic system 26 is shown in the block diagram in Figure 2. Asillustrated, input variables 22 are input at the input 23 to the neural controller 20. The neural controller20 comprises an optimised neural network (as described herein) which then generates a controlsignal in the form of control parameters 24. Within Figure 2, the optimised neural network is withinthe processor 21. The control parameters 24 are then output from the output 25 to a control process26 (the dynamic system), where the control parameters 24 govern process control. The inputvariables 22 may be, for example, set-points at which a control variable of the control process shouldbe maintained, process limits, or any additional contextual data. Output variables 28 (also referredto herein as a “current system value” or “current system variable”), relating to the state of the system26, that are output from the dynamic system 26 are fed back into the neural controller 20 via theinput 23. This feedback is used to update the control parameters 24, as determined by the neural controller 20, in a closed loop control process.
[0065] The flowchart in Figure 3 describes the method 30 by which the neural controller 20 of Figure2 is developed and trained. Firstly, at Step 32, a gradient-free optimiser is used to optimise a neuralnetwork (the means by which this network is created and trained is discussed in detail later). The optimised neural network is then compiled, at Step 34, into a netlist which defines the network according to a microarchitecture. The microarchitecture describes the types of neural network that can be created, and so defines the permissible structures and execution behaviour of the network. The netlist then describes the neural network that has been optimised for an application. The netlist is the intermediate representation of the network architecture, and comprises lists of input neurons, compute neurons, output neurons, and synapses. For example, the netlist may be a compressed JSON encoding of the network structure and parameters. Once the netlist is created, the neural controller 20 is then generated using the netlist and a network runner at Step 36, where the runner is the component that executes the netlist on hardware.
[0066] As described above, creating the neural controller 20 comprises optimising and training aneural network using an optimiser. In general, an optimiser takes a function ^(^) = ^ and thendetermines the value of ^ that maximises or minimises ^. When this function is applied to optimisinga neural network, ^ is the specific neural network, ^ is the curriculum evaluation, which evaluateshow ‘good’ the generated network is (discussed in greater detail later), and ^ is the computed fitness,corresponding to the network’s performance (a higher fitness means a better performing neuralnetwork).
[0067] Various methods can be used to optimise a neural network, and traditionally gradient descentand back propagation methods are used. In gradient descent, often the gradient of ^ is noisy or hasmany local optima, and so gradient-descent optimisers may converge prematurely, at a local rather than global optimum. In conventional deep learning methods, a range of fixed neural networkarchitectures are estimated by data scientists, (based on literature and expected results), and thenautomated back propagation is applied to find weightings for the network. This approach often results in large inefficient networks (many layers deep) which require a huge amount of training samples and computational power, and have poor scalability. Additionally, deep learning methods struggle to adapt to the differences between their training data and real world examples, and so there is a ‘reality gap’. The ‘reality gap’ results in poor predictive performance from a deep learning network, stemmingfrom input variables the system has not seen, under- or over-fitting, or encountering unanticipatedsystem changes and situations. Thus, a different method of optimisation may be better suited. In anembodiment of the present invention, a gradient-free optimiser is used to optimise the neural network.
[0068] Many gradient-free optimisers exist, an example being an optimiser that uses neuro-evolution. Neuro-evolution is a machine learning technique that uses evolutionary algorithms to construct artificial neural networks. Neuro-evolution is particularly suited for control applications as the networks produced using neuro-evolution are extremely efficient: networks are much morecompact, and thus more computationally efficient. A neuro-evolution optimiser is also more able tofind global maxima for a network, and is less likely than other optimisation methods to prematurelyconverge to local maxima. While neuro-evolution optimisers are particularly suited for the presentapplication, any other suitable optimiser, such as a random-search optimiser, may be used.
[0069] Returning to Figure 3, the microarchitecture is a mathematical description defining the set ofpossible neural networks that can be produced, and includes the mathematical rules governing the neurons and synapses of the network. The microarchitecture may define one or more properties ofthe network that is to form the controller 20, which are discussed in detail below in relation to Figure4.
[0070] Firstly, the microarchitecture comprises neurons (nodes) 40 and synapses (edges) 42, wherethe nodes 40 have interconnecting synaptic links. Together the nodes 40 and synapses 42 form adirected graph. On execution of the network, each neuron 40 aggregates the activations from itsinput neurons, where the input activations are weighted by the connecting synapse strengths, to calculate the neuron’s own activation.
[0071] Another feature defined by the microarchitecture, is that in the resultant neural network thesynapse strengths may be plastic. Plastic synapses allow the synapse strengths to change, based on previous neuron activations, during network operation. The degree of plasticity (that is the learning rate, which relates to how much the synapse strength can change during network execution) of a synapse can be modulated by the neurons 40 the synapse 42 connects to, which enables self-tuning and mode-switching in the network. Essentially, as data flows through the neural network, thenetwork will adjust synapse weights as it executes. It should be appreciated that in the neuralnetworks produced there will be a mix of plastic and fixed synapses, and the ratio of these types ofsynapse 42 will vary dependent upon the specific network required for the control application.
[0072] The microarchitecture also specifies that the neural networks are configured so that at eachexecution (clock) of the network, each compute neuron 40 produces the activation that was calculated for the last execution, that is the activation at time t-1, as illustrated in Figure 4. Thisnetwork property is referred to herein as ‘clocking’. In traditional networks, the neuron 40 producesthe activation for the current execution, time t. Clocking allows the network to encode arbitrary(possibly cyclic) directed graphs, as shown in Figure 5, and also multidigraphs (where edges have the same end nodes), meaning that the network connections are directed, can be parallel (more than one edge can connect two nodes), and can form cycles or loops (an edge from a node to itself). The ability to encode loops and arbitrary directed graphs means that certain relationships can be encoded with less neural structure, and accordingly, the networks can learn complex functions with less computational units. Clocking offers benefits since without this feature, the neuron activations could not be calculated due to infinite recursion.
[0073] The microarchitecture may also define synapses 42 as excitatory (with a positive weighting)or inhibitory (with a negative weighting), and in some embodiments, the initial synapse strengths are quantised, but this is dependent on the optimiser used. Neural Network Properties
[0074] An example of neural network properties for a neural network that may be incorporated intoa controller 20 according to embodiments of the present invention is now described.
[0075] The network defined by the microarchitecture is a graph network, that can be represented as^ = (^, ^, ^^^^^^, ^^^^^^), where:N is the set of neurons 40;S is the set of synapses; source: ^ → ^ is a map from each synapse to its input neuron;target: ^ → ^ is a map from each synapse to its output neuron.
[0076] The neurons ^ 40 and synapses ^ 42 within the network have specific properties. Regardingfirstly the neurons 40, each neuron ^ 40 has several properties, examples of which are listed below.: gain ^^ ≥ 1 (setting the gain to be greater than 1 combats a tendency of neuron activationsto grow smaller as a network deepens); base activation (or bias) 52, relating to the activation a neuron exhibits each clock if theneuron has no non-zero processor inputs; base learning rate, also known as the maximum learning rate ^^, where ^^ is the maximummagnitude of a plastic synaptic update; activation at t, ℎ ^^ ∈ [0,1];learning rate at t, ^ ^^ ∈ [0, ^^].It should be appreciated that in other embodiments, the specific values of the neuron 40 propertiesmay change, and some neuron 40 properties may not be defined by the microarchitecture at all.
[0077] Turning to the synapses ^ 42 within the network, each synapse ^ 42 is characterised in thatthe synapse 42 has a: learning rule ^^^^^; type: excitatory or inhibitory; strength ^^^(may or may not be quantised); weight ^ ^^ , where if the synapse is excitatory, ^ ^^ = ^ ^^ , and if the synapse is inhibitory^ ^ ^^ = −^^ .
[0078] At least some of the synapses 42 within the network may be plastic and be able to modulatetheir plasticity. This modulation is a form of unsupervised learning, and allows the network to easily adapt to unanticipated situations or new data. A feature of the network architecture rules that allowsthis plasticity modulation is that the set of neuron synapses ^ are separated into a set of processorsynapses ^, and a set of modulator synapses ^ (^ = ^ ∪ ^), as illustrated in Figure 6. Theprocessor synapses ^ 42a contribute to the neuron’s activation, while the modulator synapses ^42b contribute to the instantaneous learning ratethe neuron 40, that is how much thestrength can change at time t. Each neuron 40 in the network having both processor and modulator synapses is advantageous, as it enables a signal to contribute to the output of the network (through the processor synapses 42a), or to have an effect on the internal network parameters (through the modulator synapses 42b), or both.
[0079] A compute neuron 40, showing the modulator 42b and processor 42a synapses, alongsideother neuron and synaptic properties, is illustrated in the diagram in Figure 7. It should be noted thatin Figure 7 ‘a’ is an activation function, from which the neuron activation ℎ^ is computed, and the“0.5” (labelled with reference number 54) is the base modulation. The base modulation 54 is thesame to modulator synapses 42b as the base activation “b” (labelled with reference number 52) isto processor synapses 42a, that is, base modulation 54 contributes towards modulation regardless of other modulator inputs. However, unlike base activation 52, base modulation 54 it is set to a constant, which in this example is 0.5.
[0080] The modulator synapses 42b allow the neurons 40 to modulate their plasticity. Modulation isbased on neuron activation ℎ^^^^, the instantaneous learning rate ^^^^^of the neuron 40, and the learning rules ^^^^^encoded on the neuron 40. Using these parameters the plastic strength ^^of the synapses 42 at various times after initialisation (t = 0) can be determined.
[0081] For a given neuron ^ 40 at time ^ + 1, the activation is given by:
[0082] where ^^ is the set of processor synapses 42a that are inputs to neuron ^ 40,isthe activation of the input processor synapses 42a. The activation follows a clamp function, which isdefined as:clamp^ = ^^^ (0,^^^(^, 1))[2]
[0083] An example of a clamp function 55 is illustrated in Figure 8. While the activation function isdescribed as following a clamp function, it should be appreciated that in other embodiments, othernon-linear functions ^ → [0,1] may be used.
[0084] It should be appreciated that the neuron activation (Equation 1) resembles the classicequation for neuron activations in a perceptron (a well-known rule in machine learning that computesactivations for deep learning networks), with two differences: the addition of the gain parameter, and the non-linear activation function being a clamp function.
[0085] The instantaneous learning rate ^^ of a neuron 40 at time ^ + 1 follows a similar pattern tothe neuron activation function (in that the instantaneous learning rate function also includes a multiply accumulate followed by a non-linear activation function) and is given by:where ^ is the set of modulator synaps ^^ es 42b that are inputs to neuron ^ 40, and ℎ ^^^^^^(^)is theactivation of the input modulator synapses 42b. The constant 0.5 refers to the base modulation 54 of the modulator synapse 42b.
[0086] Each synapse 42 within the network has learning rules ^^^^^ encoded, where each learningrule is encoded as a function of the current synapse strength ^, pre- and post-synaptic activation ^and ^, and the instantaneous learning rate ^. The learning rule encoded on each synapse describeshow the synapse strength ^^^ changes over time based on the previous network executions.Examples of learning rules are listed below, but many other rules may be used:
[0087] The activation function and instantaneous learning rate of a neuron 40, and the learning rulesencoded on the synapses, can be used to calculate the strength ^^at any time t for each synapse42, which governs plasticity of the neurons 40:In this way, each neuron 40 can modulate their plasticity, allowing the network to easily and efficiently adapt to new data, including self-tuning, mode-switching, and adapting to unanticipated inputs.
[0088] Prior to being executed, the network is initialised from the netlist according to themicroarchitecture (as the microarchitecture defines the allowable network structure and behaviour). Initialisation involves initialisation of the compute neurons 40 (which includes the output neurons)and synapses 42. For initialisation, at time ^ = 0 the activation of each compute neuron 40 is set to0 (ℎ^^^ = 0) and each synapse 42 in the network has a fixed strength ^^^^(the initial synapse strengthis defined in the netlist).
[0089] In addition to initialising the compute neurons 40 and synapses 42, the input and outputneurons within the network are normalised, in order to scale the inputs and outputs to be appropriate for the target control system. As a subset of compute neurons 40, output neurons behave in the same way as compute neurons 40, except the activations of the output neurons are collected to form the output of the network for each clock. Input neurons however are different to both computeneurons 40 and output neurons. Input neurons cannot have input synapses, and so their activationcannot be calculated using Equation 1 above. Input neuron activation is instead set by the user at the beginning of each clock to be specific to the target application. In practice, a network wrapper specific to the target application is used to set the activation of the input neurons, which performs the scaling, unit conversion, and any other requirement needed for normalising the input and outputof the neural network for use as a controller 20.
[0090] While the properties described above are preferable options for the network, it should benoted that there are some properties of the network that may change. For example, alternativelearning rules may be encoded. However, it is important that the learning rule is parameterised per synapse, which allows some graph structures to be fixed, and the optimiser to find the best learning rule for each synapse during network training. Similarly, alternative mathematical update rules forneuron activation may be used. For example, alternative activation functions to Equation 1 may beused, and / or gain modulation may be incorporated. The neuron activation function may comprisea gain parameter, to counter the tendency of the magnitude of neuron activations to decrease indeeper layers of the network.
[0091] Once the trained neural network is compiled into a netlist, the netlist and runner togethercreate the neural controller 20. The runner is a hardware or software implementation of the microarchitecture, which allows the neural controller 20 to operate on hardware. A software runner will typically load a netlist compiled from a trained and optimised network, while a hardware runner may be a manufacturing process which embodies the netlist in target hardware (for example IC, FPGA, or ASIC).
[0092] Various runners may be used, and the choice often depends on the implementation. Forexample, a Python / Cython runner may be chosen for consumer-grade CPUs, a C runner is suitable for embedded hardware, or a CUDA runner may be chosen for GPGPU implementation.
[0093] In some embodiments the netlist may include cryptographic signing. Signing the netlist actsas a safety feature, and prevents an experimental network being deployed to production hardware. When the netlist is signed, the network runner verifies that the release status (specifying the stages of development that have been completed), which is encoded into the netlist as part of the cryptographic signing, is appropriate for the application. Additionally, in some embodiments the netlist may be encrypted to prevent unauthorised use.
[0094] Once the netlist and runner form the neural controller 20, the neural controller 20 can bedeployed to control a target control process. Luffy Training Methodology
[0095] As described briefly in reference to Figure 3, creating the neural controller 20 firstly requiresa trained neural network, which is then compiled, according to the microarchitecture, into the netlist. The method 60 used to develop and train a neural network for subsequent use in creating the neural controller 20 in accordance with embodiments of the present invention is illustrated by the flowchart in Figure 9, and discussed in greater detail below.
[0096] As illustrated by Figure 9, creating the trained network for subsequently compiling into anetlist involves several stages of development. Firstly, the target control system is characterised and modelled at Step 62. To characterise the system, the relevant control system theory, the physical dynamics, and the sensor and actuator characteristics of the system are all taken into account. System characterisation is achieved through, for example, direct experiment and data collection, communication with stakeholders, and consultation of the existing technical and academic literature. Once the system has been characterised, this can be used to develop a model of the system, and both mathematical and software models can be generated. During model development the relevant phenomenology (relating to non-obvious physical phenomena that occur in the system) and dynamics may be captured to a sufficient fidelity to create an accurate model. It should be noted thatthe model is used only during training, validation, and verification of the network, and in use, thecontroller 20 runs independently of the model. As a result, the method is not computationallyexpensive, even though an accurate model is required. MPC (an alternate control method for APC) also requires an accurate model, however in MPC the model runs and continually updates in the control loop, which makes MPC extremely inefficient.
[0097] Once the target control system has been modelled (and thus the control problem issufficiently understood), an AI curriculum is developed at Step 64. The AI curriculum defines a series of control challenges for the AI network to perform within the system model. The curriculum comprises multiple levels which incrementally increase in complexity, in order to guide the AI network during training, and expose the network to the full space of possible behaviours. The AI curriculumalso includes one or more reward functions, which are mathematically formulated descriptions of thenetwork’s performance on a given AI curriculum.
[0098] Once the AI curriculum is developed, a candidate AI network is trained using the curriculumand an optimiser at Step 66. A gradient-free optimiser, for example, a neuro-evolution optimiser (as described previously) is used, and the optimisation process continues until either a predefined fitnessis reached, or the optimisation process is terminated by an AI developer. The result of the AI trainingis one or more candidate networks that may be effective for use as an AI controller. The candidatenetworks are then validated against a validation curriculum, which has similar statistical properties to the training curriculum.
[0099] Once training and validation is complete, AI verification is carried out at Step 68. Verificationcomprises candidate networks being tested one or more verification challenges. The verification challenges are designed to test the network’s extrapolation beyond its training set. For example, verification challenges may be performed in a higher-fidelity simulation if one is available. Training,validation, and verification ensures that the solution is stable and does not diverge rapidly onexposure to unseen conditions.
[0100] Once the candidate network has been verified, hardware integration and testing is performedat Step 70. The neural controller 20 is integrated with the target hardware, which is dependent upon the target control process. The hardware is then tested, which involves devising a series of experiments in the real world, and then comparing the results with the same experiments implemented in the system model, in order to test divergence between the model and reality. The neural controller’s performance is monitored in these controlled tests, and results then feed back into the system modelling (Step 62), curriculum development (Step 64) and AI training (Step 66) in aniterative fashion. Once the neural controller 20 passes all tests in hardware, it is considered to havesuccessfully crossed the reality gap, and be a suitable controller for the target control system.
[0101] Finally, at Step 72, the neural controller 20 is deployed to the target control system. KeyPerformance Indicators (KPIs) or operational thresholds may be specified for the target controlsystem, and are monitored during hardware deployment. For any controller, there is a trade-offbetween several KPIs. Examples of KPIs include the time to reach a setpoint, the degree of overshoot, and the reduction in energy usage. More specific examples might include ‘the network must function correctly for winds of up to 10 mph’, or ‘the network must reach the setpoint within 10s’. Ideal controller behaviour is context-dependent: some applications may need to reach the setpoint as quickly as possible regardless of overshoot, while others may need to minimise energy usage at the expense of speed, and so the specific trade-off between KPIs is application specific. Often, additional constraints are defined for the motion of the target control system, such as positional constraints or velocity ‘soft-stops’ (regions of state space which the controller should avoid) and ‘hard-stops’ (regions of state space the controller must not access). The KPIs are determined during system characterisation, and are based on customer needs. Therefore, while KPIs are monitored during hardware deployment, they feed into both model and curriculum development, and form part of AI verification and hardware testing.
[0102] Figure 10 shows a similar flowchart to that illustrated in Figure 9. However, while Figure 9shows an ordered flowchart, where each step feeds into the next, Figure 10 shows a method 80 ofdeveloping and training the neural controller 20 that incorporates feedback 82 between the differentmethod steps. As illustrated in Figure 10, every step can feed into every other stage, and in other embodiments (not illustrated), some of the method steps can be performed in parallel. An example of this was described above, where hardware testing results can feed into model and curriculum development.
[0103] A schematic block diagram showing a system for creating a neural controller 20 according toan embodiment of the present invention is shown in Figure 11. Networks (i.e. a population ofcandidate neural networks) are provided to a neuro-evolution training process 90 comprising agradient-free optimiser, for example an evolutionary algorithm 92, a curriculum 94 (containing botha training curriculum and a derived validation curriculum), and a simulation 96 of the target control system. The neuro-evolution system is initialised with one or more seed solutions, which may include pre-existing neural structure. The evolutionary algorithm 92 trains and optimises a network in thesimulation 96 using the training curriculum and then validates the network using the validationcurriculum, to generate a candidate network for the controller 20. The candidate network is evaluatedagainst the training and validation curriculum 94, and a fitness value, representing the accuracy ofthe candidate network, is assigned to the candidate network. If the candidate network is considered to have passed the training and validation (the fitness value is above some threshold), the candidatenetwork undergoes verification 97, and is then output to hardware testing 98 (in a hardware testcontroller). If the candidate network passes all the hardware tests 98, the optimised network is outputfrom the system and deployed as the neural controller 20 for the target control system. If thecandidate network does not pass hardware testing 98, a new network is provided and trained by theoptimiser.
[0104] Each of the steps involved in creating and training the neural controller 20 are discussed infurther detail below.described with reference to Figures 9 and 10, the first step in developing the neuralcontroller 20 involves creating a model, or physics simulation 96, of the target control system. Thephysics simulation 96 provides the environment for training the AI network, and the capability for testing failure scenarios, which would be impractical or prohibitively expensive to test with a real system.
[0106] In order to create a capable AI controller, a sufficiently accurate simulation 96 of the targetcontrol system is required. To assess the quality of the simulation 96 for AI training, three factors are evaluated: the physical phenomena of the system (mechanics, control latencies, sensor noise, etc.); the fidelity to which such phenomena are modelled; and the degree of domain randomisation which can be applied (for example, change in initial conditions, variation of sensor parameters, etc.). These three factors should be sufficient such that the neural network can cross effectively from the simulation 96 to real hardware.
[0107] However, as well as creating a sufficiently accurate model 96, it is desirable to accelerate thetime it takes to generate a network, and thus network training time. In order to find a trade-offbetween model accuracy and the time it takes to train a network in the simulation 96, simulationdevelopment targets the dominant physical phenomenon of the system rather than the exact physics, wherever possible.
[0108] As well as considering the physical hardware and relevant physics of the target controlsystem, the neural controller 20 should also be able to overcome any challenges associated with controlling real hardware. These challenges include the imperfections of sensory systems, manufacturing errors and tolerances, and mechanical degradation. In order take these challenges into account, the simulation 96 should include control parameters that represent these challenges. All parameters included in the simulation 96 may also be subject to domain randomisation, as determined by a developer. Curriculum Learning
[0109] An AI curriculum 94 is also developed, either after or in parallel with the system modelling,which is used to guide the evolution and training of the AI network. The AI curriculum 94 comprises a sequence of training levels ^^, and each level contains a set of scenarios. The set of scenariosdefines different challenges with varying degrees of realism, often corresponding to real-worldsituations for the target control system. As the AI network evolves, it passes through each level in turn to expand the network’s range of capabilities. The training levels typically increase in complexity, with later levels including more complex challenges and confounding factors, so that the AI network can learn all eventualities that may occur in real-world system. A portion of completed scenarios areusually preserved (that is, included in multiple levels of the curriculum 94) to prevent a previouslylearnt scenario being un-learnt as the AI network learns new scenarios (which is a challenge of training neural networks, known as catastrophic forgetting). Another method used to prevent catastrophic forgetting is to include scenarios that are of the same type, but not exactly the same, ineach level of the curriculum 94.
[0110] In the AI curriculum 94is the set of all possible scenarios, usually an infinite set, for level1 ^^, the level of the AI curriculum 94 that contains the simplest scenarios. For example, to develop a curriculum 94 for training a controller to control a robot limb, ^^may include the initial angularposition p and the set point q. ^^ would then be the set of all possible combinations from p and q.From the set of all possible combinations, four scenarios may be chosen to be the scenarios usedin ^^. The scenarios may be chosen using any appropriate method, for example, a grid search. ^^is thus a subset of n scenarios sampled from ^^, i. e.
[0111] An AI curriculum 94 successfully trains the AI network to control the target system. Oncedeveloped, the training process embodied by a working curriculum 94 for a particular problem or control domain often advantageously generalises to similar problems. Network Fitness and Reward Functions
[0112] In order to evaluate the accuracy of the trained AI network, and thus the success of the neuralcontroller 20, one or more reward functions are defined and included in the AI curriculum 94, wherea higher reward corresponds to a better solution. Reward shaping is a major component of the AI training process, as good reward functions are able to encourage target behaviours and deter unwanted ones. The behaviour of the network on each scenario is evaluated by a reward function(which may be different for each scenario). Solutions of the reward functions are assigned a fitness,where higher fitness corresponds to higher aggregate reward across all training scenarios.
[0113] Reward functions may be calculated either after each execution of the network, or at the endof a scenario based on some aggregate statistic. They may also be composed of other reward functions, for example in weighted combination.
[0114] During training, the global-optimiser tunes the parameters to improve the performancemeasured by a provided reward function. The parameters are saved periodically during optimisation.In the example described where the global optimiser is an evolutionary algorithm 92, the algorithmstores the network topology and parameters that provide the highest fitness, which are known as the champion parameters. Every time a new set of parameters emerge that have increased fitness compared to the previous champion parameters, the new set of parameters are stored by the training process.
[0115] In some embodiments a term is included in the reward calculation which penalises networkcomplexity, specifically for network size and graph complexity. Validation and Verification
[0116] A common challenge in training neural networks is that of overfitting, where networks learnbehaviours specific to the scenarios against which they are trained, but which do not generalise to the full problem domain from which those scenarios are sampled. Overfit networks may perform poorly or fail outright when exposed to new scenarios, even if those scenarios are qualitatively the same as those seen in training. In order to detect overfitting, all champion networks (networkscreated using champion parameters) are passed through a validation curriculum, which is createdby resampling the same set of scenarios used to define the training curriculum.
[0117] In the above example, the training curriculum contained a levelcomprising scenarios^^,^, … , ^^,^ ~^^. An overfit network would be one which performs well on the specificscenarios,but poorly over the whole set ^^. A validation curriculum would have a level ^′^, which wouldcomprise ^′^,^,i.e., a new sampling of the same set of scenarios ^^. It is also possiblethat validation levels may sample a larger possibility space, or be biased towards harder scenarios.
[0118] Validation exists to detect overfitting, and overfitting often demonstrates that the optimisationprocess has stagnated. It is possible to use validation fitness to perform early stopping, specifyingthat training should terminate if N successive champion parameters have a validation fitness lessthan the maximum observed validation fitness.
[0119] Once a high-performing network, or champion, has completed the training curriculum andperformed well in validation, the champion undergoes verification. Verification consists of testing the network on more extreme challenges, testing the extrapolation behaviour and edge-cases. The challenges used exceed the expected difficulty of controlling real hardware. Often, these challenges are made incrementally harder in order to determine the champion parameters performance envelope. This verification set is different to ‘test sets’ which are typically used in training of networks to detect overfitting or under-generalisation. Once verified, the champion network is tested inhardware, and if it passes all hardware tests 98, is deployed to the target application.Example Application – Controlling a BLDC Motor
[0120] A neural controller 20 generated according to embodiments of the present invention may beused to control many different systems, and one example is to control a motor. A method of creating a neural controller 20 for controlling a motor, as well as the results obtained when the neural- controlled motor is used to actuate a robot limb, are described in detail below.
[0121] For this application, the neural controller is required to actuate a generic current-controlledBrushless Direct Current (BLDC) motor, and include closed-loop control of angular position andvelocity, with additional configuration inputs to mediate behaviour (for example safety constraints or contextual information). System Characterisation and Modelling
[0122] Developing the neural controller 20 comprises firstly modelling the target control system, theBLDC motor. It should be appreciated that any appropriate modelling technique may be used, but in this example a physical model 96 based on a first-principles electromechanical derivation was implemented. While the target control system is a BLDC motor, the physical model 96 for a brushed direct current motor is described, as despite the differences in operation, the physical model 96 used to represent the brushed motor is also able to represent a BLDC motor.
[0123] In a brushed DC motor, a cylindrical rotating segment known as a rotor is wound along itsaxis of rotation with current-carrying wires, known as the armature. The rotor surrounds a stationary component, the stator, which generates a magnetic field across the rotor perpendicular to its axis of rotation (often created by opposite poles of a permanent magnet). When the armature is powered, the magnetic field generates a Lorentz force on each armature wire. As the winding carries current in opposite directions on each side of the rotor axis, opposite forces on opposite sides of the axis are generated, and torque is produced on the rotor until the armature is aligned with the field. To maintain the torque, the current direction in the armature is switched based on angle by a rotary switch known as a commutator. Typically, the commutator consists of a ring of segmented connectors; the ring rotates with the rotor, making and breaking connections with a fixed DC circuit via brushes.
[0124] In a brushed motor, for N armature windings, the commutator will have 2N segments, whichdetermine the granularity of current switching. For a two-segment commutator, the torque on therotorto achieve a constant torque ^ → ∞. The idealised torque generated in thisway is:^^ = ^^^^[5] where ^^is the magnitude of the armature current, and ^^is the motor torque constant, which can be considered constant for a particular motor.
[0125] However, when a particular armature wire rotates through the magnetic field, this alsogenerates a Lorentz force, opposing the current. This force constitutes a back-EMF ^ based on theangular velocity of the rotor ^^̇:^ = ^^^^^̇[6]which has the effect of reducing the armature current:where ^^and ^^are the armature resistance and voltage (ignoring inductance). As motors typically have a maximal operational voltage ^^^^, the back-EMF ^ produces an upper limit on armature current and hence on angular velocity of the rotor.
[0126] In BLDC motors the electromechanical relationship between the rotor and the stator isinverted: a typical BLDC motor uses permanent magnets in the rotor, and the rotor is controlled by electromagnets in the stator. Commutation is purely electrical, where rotor position is synchronised with alternating current through each phase of the stator armature. The permanent magnets of the rotor may be internal, but commonly they are surface-mounted on the outside of the cylinder in alternating pole-pairs at regular angles. The stator electromagnets are solenoids wrapped around stator teeth, with the gaps between teeth known as slots. BLDC motors additionally comprise a DC to AC converter.
[0127] As mentioned briefly above, despite the significant change in operation, the idealised physicalmodel of a brushed DC motor equally applies to BLDC motors. However, the idealised model described is not accurate in real-world systems, where motors are subject to many forces, all of which must be incorporated when characterising the BLDC motor. The forces considered for the present control system are discussed below.
[0128] Inertia, load, and gearing all impact the operation of the motor, and so are included whencharacterising the BLDC system. In general, for a motor to be useful, the motor must be loaded. Theload acts with or against the motion of the rotor in two ways: by increasing the inertia of the systemJ and by exerting a load torque ^^. In a direct drive motor (a motor with no gears), the load torquecan simply be subtracted from the rotor torque. The equation of motion for the motor is then: ^^ − ^ ^^̇^ = ^^^[8]It should also be appreciated that the rotor itself has inertia so ^ = ^^ + ^^.
[0129] Some motors have gears, which are used to increase the output torque ^′^ of the motor atthe expense of output angular velocity ^̇^: ^̇^ = ^̇[9] ^^ ^^ = ^^^
[0010] Gears also have inertia, and thus the inertia of the system is different if measured at the input or output to the gears, with the rotor inertia scaling with the gear ratio squared. The equation ofmotion must therefore be adjusted, and for a motor comprising gears becomes:
[0130] While ideal motors do not experience friction effects, real motors do, and thus friction is alsoincorporated. There are several classic models for motor friction, four of which are shown in Table 1 below. Table 1. Types of Friction in Motors. Considering only viscous and coulomb friction, the equation of motion for a direct drive motorbecomes:and for a geared motor becomes:where ^^^ and ^′^represent the additional viscous friction introduced by the gears.
[0131] Cogging torque, caused by the attraction of the rotor permanent magnets to the stator teeth,may also impact motor operation, and thus in some embodiments a model for cogging torque may also be included when characterising the BLDC system. Many different models may be used, and so are not discussed in detail here. Similarly, in some embodiments, a model representing any heating effects or backlash on the motor may be incorporated.
[0132] Electronic control, inductance and heating may also impact motor operation. Typical motorsaccept a target current ^^as input. In a BLDC motor, the input DC voltage is converted into three-phase AC voltage, with the voltage through each phase modulated by pulse-width modulation (PWM). Inductance smooths out the input voltage, allowing the phases to operate at different target currents, which are typically coordinated with the rotor position by integrated electronics to eliminatetorque ripple. The effective current ^^ with respect to the rotor torque is given by Equation 7. Theabove Equations 8 to 12 define the characteristics of the BLDC motor, and are used to develop themodel 96.
[0133] It should be noted that while the various secondary effects such as cogging and friction allowdevelopment of a more accurate simulation 96, secondary effects can be removed from the equationby choosing appropriate model parameters. The only four non-zero parameters required to specify a motor are ^^, ^^, ^^^^, and ^^. It should also be appreciated that although not incorporated in this instance, other secondary effects may be taken into account when developing a model 96 of a BLDC motor. For example, it is known that the mechanical and electrical aspects of the motor generate heat, and temperature may also be considered, as temperature affects the resistance of the electrical system, affecting the torque output. In some BLDC control systems backlash may also be consideredand incorporated into the system model 96. In a system relaying mechanical transmission of force,such as a gear train, the mechanical connections are never perfect. Gaps between gear teeth result in a “dead zone” in which no force acts between the gear teeth, causing the first gear to move freely while the second gear does not transmit any force. This is known as backlash.
[0134] Using the above defined characteristics (Equations 8 to 13, and if included, a model forcogging torque, any heating effect, or backlash), a dynamic motor simulation 96 of the BLDC motorcan be developed, where a series of equations are used to define the simulation 96.
[0135] The instantaneous angular acceleration ^̈(^) at time ^, advanced at discrete time steps ∆^ isgiven by:
[0014] where is the torque produced by the motor (that includes all contributing factors such as friction and cogging) and is the total inertia of the motor, including gearing.
[0136] The control variable is the target output current ^^ of the motor, from which the true armaturecurrent ^^ is calculated using ^^^^and the instantaneous back-EMF ^ produced by the rotor’s motion:^ 1^(^) =^ clamp^
[0015] Therefore, the state of the system at time t can be given by angular position ^, angular velocity ^̇,load torque ^^(^), and load inertia ^^(^).
[0137] It should be noted that the state observations seen by the controller 20 exhibit a lag of up toone time-step and are offset by noise, and this lag may be incorporated into the model 96.
[0138] Equations for updating angular velocity ^̇and position ^ are given by independent Verletintegrators as determined by:^(^ + ∆^) = 2^(^) − ^(^ − ∆^) + ^̈(^)∆^^
[0017] Where ^ represents the variable under consideration, and thus the above equations apply to bothangular velocity ^̇ and angular position ^. As the equations above integrate the second derivative,the value of ^̇(^ + ∆^)depends on the jerk ^^(^), which is simply calculated using the backwarddifference:
[0139] The equations above (Equations 14 to 17) define the simulation dynamics, defining theenvironment in which the neural network is trained. During model development, the first model iteration may not be sufficiently accurate, as it may omit many secondary physical effects. However,hardware testing 98, which can occur in parallel with system characterisation and modelling (asillustrated in Figure 10), shows any divergence of physical behaviour from the model 96, for example, high-frequency responses (which are consequences of unmodelled dynamics). The model 96 canthen be enhanced and refined, based on feedback 82 from the hardware testing 98, in order togenerate a more accurate model 96.
[0140] The controller 20 is configured to control a motor to actuate a robot limb. Thus, the controller20 sets an appropriate target current ^^ that drives the motor to a set point ^^. The aim is to reachthe set point in a critically damped manner, that is, to reach ^^in the least time possible while maintaining, based on the observed state of the motor, no steady-state error, no overshoot, and no oscillation. The controller 20 is designed to drive the motor to the defined set point for a variety of load torques and inertias, including those which vary over time or position (e.g. gravitational torques).
[0141] The controller 20 is also designed to constrain the achievable position of the motor betweentwo hard-stops, which act like physical barriers: when the simulation exceeds the hard-stop values, the position is clamped to the hard-stop, the velocity is set to zero, and the integrator states(Equations 16 and 17, corresponding to the value from the previous time step) are reset. Thecontroller 20 is also restricts movement between two configurable boundaries known as soft-stops,which in the present example are two angular positions ^^^^and ^^^^. When ^^exceeds a soft- stop, the controller 20 is settles at the soft-stop position and does not exceed the soft-stop. During training, the neural controller 20 hence learns a policy which maps from the state measurements to the control variables such that:Curriculum Design
[0142] An AI curriculum 94 is designed to train the neural network, and this can follow, or occur inparallel with, system modelling. The first level of the AI curriculum 94, which defines the simplestscenarios for the target motor, captures the simplest version of the problem. The simplest problem is to achieve a given set-point when there are no external forces, i.e. where there is no load on the motor, and no model parameter variation (for example, no variation in inertia, the torque is constantetc), no sensor noise, no friction, and no cogging. At the simplest level of the AI curriculum 94, thereward function, which is used to evaluate the accuracy of the network, does not consider soft-stopsor hard-stops, and the setpoint is defined as a constant value.
[0143] In subsequent levels of the curriculum, the factors described above such as load, sensornoise, friction, cogging and model parameter variation are gradually introduced, alongside randomly generated setpoint waveforms which can vary over time.
[0144] The specific curriculum 94 developed for the BLDC motor is described in detail below. Firstly,domain randomisation was used to minimise overfitting during the neuro-evolution optimisation. Thedomain variables to be randomised are the motor parameters ^^,^^ ,^^^^, ^^ , and ^, the waveforms^^(^), ^^(^), ^^(^), and the constants ^^^^, ^^^^, ^^,(sensor noise), and ^ (delay magnitudes)(which were sampled from a uniform distribution over an appropriate range of values). When randomising the soft stops ^^^^and ^^^^, the distance between soft stops, rather than their absolute position, was the property used.
[0145] The AI curriculum 94 is designed to cover a wide range of representative control tasks, andthus two classes of waveforms are used to define variables: step waveforms (in which the value is held constant between randomly chosen points and then changes to its next value); arbitrary waveforms, which comprise both smooth and discontinuous changes in value over time. Thewaveforms are constructed from piecewise combinations of easing functions 101 (shown in Figure12), a common formulation of interpolators for animation. Example waveforms 103 can be seen inFigure 13. Any or all of ^^(^), ^^(^), ^^(^) can take on one of these waveforms, or be held constantthroughout a scenario. AI Training
[0146] In the curriculum 94, every scenario evaluates the performance of a network. An evolutionaryalgorithm 92 is used as the neuro-evolution optimiser, and is a maximisation optimiser. Therefore,the performance evaluation is expressed in terms of reward. However, in this specific application related to motor control, the performance of the network is considered using loss-functions, as this is more intuitive, and each loss can then be converted into a reward. The primary objective during optimisation is for the network to achieve the setpoint waveform ^^(^). Therefore, for the loss function distance from the setpoint is penalised. An example loss function may be:where ^^is the setpoint, and ^^^^^^^^^^is the permissible setpoint error. Any appropriate loss function may be used, but it should be a composite reward function including terms for distance to setpoint, velocity, and jerk magnitude.
[0147] The total reward for a scenario is simply the average reward achieved per time step, and ascenario is considered complete when this value is greater than some threshold chosen from [0, 1] through iterative experimentation.
[0148] In this example, the curriculum 94 consists of several levels, and each level introduces a newdegree of challenge. The 1st level has shared motor parameters (^^,^^ ,^^^^ , ^^ ,∧ ^) load inertia,and soft-stops for all scenarios, with no friction, control latency, or sensor noise. Variation in these parameters is gradually introduced in subsequent levels, in order to create more challengingscenarios for training the network. Across the scenarios in each curriculum level, each parameter iseither the same, a constant or a time-varying value , and for some levels, the sampling method doesnot change from the previous level.
[0149] Each level of the curriculum 94 presents a balance of setpoint waveforms for the controller20 to match: stationary (^^(^) = ^, ^ ≈ ^(0)), fixed (^^(^) = ^, ^~^(0,2^)), and step and arbitraryfunctions (as described previously). Loss functions are manually tuned to make deviation from the setpoint most punishing on stationary waveforms, then fixed waveforms, then step and arbitrarywaveforms. The rationale for including this mix is to foreclose catastrophic forgetting, whereby a network most recently trained to follow step functions may struggle to follow smooth waveforms, ora network which can handle changing waveforms may lose the ability to stay perfectly still whenappropriate.
[0150] Later curriculum levels also contain some scenarios from earlier levels, in order to counteractcatastrophic forgetting across increasingly difficult problem domains. Based on the characteristic speeds of the motors used in this example, training scenarios last 30 s, and the simulation runs with a frequency of 500 Hz. For validation curricula, the length of the scenarios is extended to 150 s in order to ensure plastic stability.
[0151] The neural network was trained using an evolutionary algorithm 92 following the curriculumplan laid out above. The evolutionary algorithm 92 was distributed across a computer cluster utilising320 CPUs. The evolutionary algorithm 92 uses ranked parent selection, meaning that fitterindividuals are preferred for mutation, and a mix of steady state and tournament selection for survivor selection, with elitism, meaning that some of the previous generation survive, including the fittestindividual. The network was evaluated against a training curriculum and then a validation curriculum.Once candidate networks have been trained, the best-performing networks on the validation curriculum (determined by fitness) are then tested against the verification set. Verification is performed by extending the range of system latencies, load torques and inertias, to encompass challenges beyond what is expected in the real world. Once a candidate network passed verification,the network was tested in simulated experiments and on the real hardware, demonstrating sim2realtransfer, (the ability to effectively translate from simulation 96 to hardware).Hardware testing
[0152] Following training and validation, hardware integration and testing 98 is carried out. Hardwaretesting 98 is an integral step in creating a trained neural network for the neural controller 20, as theresults from hardware testing 98 feeds back into model development, in order to create an accuratemodel 96, and also into curriculum development. In this example, experiments were performed withan ODrive 270kv from ODrive Robotics, known as the test motor. The network was then retrainedusing two other motors, an RMD X8-Pro and X10, and further experiments were conducted. Thespecifications of the three motors are listed in Table 2.Table 2. Specifications of motors used for hardware testing experiments
[0153] The trained and validated neural network must be integrated into hardware for testing, andthe integration and testing means are dependent upon the specific BLDC motor. For example for the test motor, an ODrive control board connects to a consumer PC via USB, but the RMD X8-Pro and X10 motors perform input / output to and from the motor’s onboard electronics over a Controller Area Network (CAN) bus using a bespoke protocol. To integrate the network into the two types of controller a script is written which loads the netlist (generated from the trained neural network and the microarchitecture) and executes the resulting network with a Python runner. Considering firstly the test motor, to test the trained network, the ODrive may be installed on a test bench, connected to a metal rod to demonstrate angle. Weights may be added to the rod to change the mass and inertia of the system.
[0154] The validated and verified network is first tested on the test motor (the ODrive) in a test rig,with an arm of length l attached to the output side of the gear capable of holding an adjustable mass m. This allows variation of the load torque and inertia between scenarios. The motor axis wasperpendicular to the gravity direction. Hence, the load torque ^^(^) ≈ │^│ ∙ ^^^^^ where is a zero-point set at the start of the experiment; the load inertia ^ ≈ ^ ^^ ^^ ^^^ │, which is constant throughouta scenario. The controller 20 was tested against two step responses, from vertical down position to vertical up, and from vertical down to horizontal right (i.e. step responses of one half turn and onequarter turn), varying the load mass ^ ∈ {50^, 150^, 300^} This experiment set up is referred to asa swing up experiment.
[0155] Once the curriculum 94 is able to reliably produce neural networks which perform well on thetest motor, the second class of motor, the X10 motor, is tested. Switching to the X10 motor means that the curriculum 94 must be re-tuned, and the slower dynamics of the X10 motor allow the network operational frequency to be reduced, increasing the computational efficiency of the optimisation process. Once re-tuned, the AI curriculum produces a neural network which successfully controlsthe X10 motor. The new AI network (tuned using the X10 motor) is equally capable of controlling the X8 motor, which is a similar motor. However, more surprisingly, the resultant network is also capable of controlling the ODrive 270 motor (the test motor, which lies very far beyond the training domain of the X10 motor) provided that sufficient external load was added. Without load, the network createdfrom an X10 adapted AI curriculum 94 produced high frequency oscillations when controlling anODrive motor but was otherwise passable. Thus, the neural controller 20 developed was able to control a broad range of BLDC motor.
[0156] Figure 14 shows the experimental set-up for a swing-up experiment, where a BLDC motor108 is configured to actuate a robot arm 110 from a vertical down position to either a horizontal orvertical-up position, as illustrated. Weights 112 can be attached to the end of the robot arm 110.Figure 15 shows the ability of the neural controller 20 to control the BLDC motor 108 to actuate therobot arm 110 for 6 different motion scenarios (each of which have different set points and / or differentexternal forces). For each graph illustrated the letter corresponds to the desired setpoint position(V=vertical-up, H=horizontal), and the number corresponds to the number of weight plates 112attached to the robot arm 110 (1, 3, or 6). The setpoint target band is + / - 5 degrees from the set pointin each graph. Figure 15 shows that the neural controller 20 meets the setpoint as rapidly as thesystem will allow, and meets and maintains the setpoint within the defined target range, for all six tested motion scenarios. It should be noted that the small ripple and mild undershoot present in the6 plate tests is an artefact of the network training being terminated too early, and further training andtuning of the fitness function would eliminate this behaviour.
[0157] For comparison, an example of PID control is illustrated in Figure 16, which shows a PIDcontroller’s ability to control the same motor 108 for actuating the robot arm 110 for the same 6motion scenarios as illustrated in Figure 15. The PID controller is tuned separately for each of the six motion scenarios (PID tuning may be carried out by, for example, manual tuning or, as is the case in the present example, the Ziegler-Nichols method, both of which require the activeparticipation of a control engineer), resulting in 6 singularly-tuned PID controllers. In Figure 16 eachrow corresponds to a single-tuned PID controller, and each column corresponds to one test motion scenario. Firstly, the starred graphs highlight the cases where the tuning of the PID matches themotion scenario. In these situations, the PID controller can exactly control the robot arm 110, androtate it to the desired set-point. Each appropriately tuned PID exhibits fast rise time with littleovershoot or offset. However, the graphs without a star in Figure 16 show the resulting motion whenthe PID controller aims to control a motion scenario for which it was not tuned, that is, the PID and the scenario are mismatched. The results show that a PID controller tuned to one specific motion undershoots or overshoots the desired set point when used to control a different motion. Forexample, when the V6 (Vertical, 6 plates) PID is applied to the V1 (vertical, 1 plate) scenario, there is a marked drop in speed of approach to the setpoint due to the controller dynamics being tuned for a heavier weight than the PID was applied to.
[0158] Being unable to control a type of motion for which it is not tuned is a huge limitation of PIDcontrollers. Operations for which the PID controller has not been tuned for are known as being outside of the linear region of the PID. The linear region is where PID control is accurate, however, this region is very small. Thus, a PID controller is likely to only be suitable for a very narrow domain.The results presented in Figures 15 and 16 show that the neural controller 20 of the present inventionhas a wider operating range than a tuned PID.
[0159] Two further experiments were conducted, in order to assess the neural controller’s ability tocontrol a BLDC motor 108 in a system for which it has not been trained. In the two experiments discussed below, a PID controller could not be tuned for comparison due to the prohibitive time cost of tuning a gain-scheduling controller capable of matching such dynamic conditions.
[0160] In these experiments the target control system is a multi-joint robotic limb 120, modelled onthe torque and geometric characteristics of a human leg (3-dof hip joint 124 comprising three motors,1-dof knee joint 122 comprising one motor). The controller 20 was instantiated once per motor (i.e.in a multi-network configuration), and each agent was able to operate well under dynamic loads of increasing mass, external forces and torques, and was able to cancel out reflected torques from the other agents. In simple scenarios where PIDs were applicable, one agent was qualitativelyequivalent to a perfectly tuned PID, as described below.
[0161] In the first subsequent test, two motors 122, 124 were mechanically coupled in aconfiguration mimicking a human leg 120, in an arrangement not explicitly seen in network training.The hardware set up for this experiment is shown in Figure 17. The “hip” motor 124 and “knee” motor122 (both RMD X10s) are actuated at the same time. Each motor 122, 124 was controlled by aseparate instance of the neural controller 20, and so each controller 20 had to cancel out the actionof the other in order to achieve its own setpoint. The controllers 20 were given waveforms composedof piecewise step functions much like the step waveforms seen in training, with some steps occurring in tandem, while others occurred out of sequence with each other. The extent of the challenge was further compounded by an unplanned mechanical failure in the hip gear train before the experiment, which introduced extreme backlash effects into the system at the hip. Prior to this experiment, the network was retrained to target the RMD motors, with its performance envelope encompassing both motors.
[0162] The results of this experiment are shown in Figure 18. Graphs A and B show the joint position(angle) 130 measured by the motor’s encoder, which sits between the output side of the motor’sinternal gearing and the externally imposed gear train; the setpoint waveform 132 with a toleranceof 1.5° is also shown. Graphs C and D show the controller 20 output current 142, with the positionwaveform of the opposing joint 144 superimposed to help highlight changes in current to cancel outreflected torques. Graphs C and D in Figure 18 also show the current limit 140. The effect of themechanical failure on the hip was obvious in the extreme thrashing at, e.g., 14–19 s in both thecurrent output 142 and angular position 130 (Graphs C and A respectively). Nevertheless, the hipmotor manages to make an effort of following the setpoint from the perspective of its encoder (due to extreme slippage in the gears, the real angle of the joint 124 on the output side of the gear train did not vary directly with encoder angle, and could not be measured); and the downstream knee 122controller is able to adjust for both the reflected torques from the upstream joint 124 and the high-frequency oscillations in its own relative angular position. This can be seen in the distance betweenthe angle 130 and the setpoint 132, which never strays outside of the 1.5° tolerance after settling(Graph B).
[0163] There are clear signs of each controller 20 responding to the other’s motion, for example thehip’s output at ~ 32 s (Graph C), and the variation of knee output from ~ 12.5–32 s (Graph D) whilemaintaining a near perfectly stationary position. Note also that as the hip 124 angle changes, the equilibrium current output of the knee joint 122 changes, e.g. the different output current at 20 scompared to 30 s in Graph D. Thus, the neural controller 20 was able to control the motors 122, 124in this experiment, in a system for which it has not been trained.
[0164] In a third test-scenario, the upstream motor 124 (the hip motor) is deactivated, so that thedownstream motor 122 swings freely. This configuration thus resembles a motorised joint midwaydown a free-swinging pendulum as illustrated in Figure 19. The physical orientation, and hence the gravity direction as well as reflected torques, oscillates every time the joint exerted torque itself. Inthis experiment additional mass ^ was attached to the end of the robot limb, and was varied so that^{0,1^^, 2^^, 3^^} . A staircase was used as the test waveform from full extension at ^ = 65°to fullcontraction at ^ = - 65°, with each step changing the relative gravity direction.
[0165] The results of this configuration can be seen in Figure 20. It is noted that all the graphs A toD comprise trace lines relating to each of the experiment mass options above (namely corresponding to the knee joint with no weight 154, a 1kg weight 156, a 2kg weight 158, and a 3kg weight 160). InGraphs A, B and C it is further noted that the four trace lines generally overlap. Graph A shows theangular position ^ of the knee joint 122, and while four lines are included, the angular position ^ issimilar across all experiments. Graph B then shows the position error │^^ − ^│for each experimentalongside the setpoint ^^ 152 and tolerance region (shaded area), and Graph C shows the positiondeviation for each experiment. In both Graphs B and C the graph lines for each of the experiments are similar, except the 3kg experiment 160 demonstrates an overshoot beyond the tolerance region,and the controller 20 was unable to match the setpoint ^^ 152 at every angle (at t >110 s) due to theimposition of an artificial 5 A current limit 162. The current limit 162 is shown in Graph D, which alsoshows the target current for each of the four experiments 154, 156, 158, 160.
[0166] At ^ = 65°, the knee 122 is perfectly straight; at ^ = -65°, the knee 122 is perfectly bent. Asthe knee 122 bends, the centre of mass for the system moves, causing the free hip joint 124 to rotatesuch that the centre of mass remains directly below the hip joint 124. This ensures that the gravitydirection ^^varies smoothly with ^^. Additionally, when the controller 20 moves, this causes thesystem to swing, causing small variations in the gravity direction to which the controller 20 mustadjust. As ^^changes, the effect of the changing ^^is easily visible in the settled signed positionerror ^^ − ^. The trained error tolerance is │^^ − ^│ < 1⁄ 2.2 when the external gear train isaccounted for. The controller 20 always settles to within the tolerance (backlash permitting), with nodiscernible overshoot and very fast rise times.
[0167] Remarkably, the rise times, settle times and overshoot (i.e. the ^(^)waveform) are consistentacross different loads, so that the deviation from the mean position ^̅ − ^ remains less than 0.5°throughout the entire experiment. The differing target currents ^^requested by the controller 20 for each load and setpoint demonstrate that this is due to an active rejection of external torques, rather than the luck of a passive controller.
[0168] The results presented demonstrate that a neural controller 20 developed for control of aBLDC motor far outperforms PID control in a number of challenging control tasks. Firstly, the neuralcontroller 20 has a much broader operational envelope than PID, as shown in Figures 15 and 16,where the neural controller 20 competes favourably across many scenarios compared to specifically tuned PIDs. Secondly, the subsequent experiments showing the results in coupled motors (Figure18) and in a free-swinging pendulum (Figure 20) demonstrate the neural controller’s ability toaccurately actuate a robot limb in scenarios where PID would be difficult if not impossible to tune.
[0169] The example application described above, where the AI controller 20 of the present inventionis used to control motorised joints, has the potential to be deployed in several applications inengineering, robotics, mechatronics, and manufacturing, and many other industries where motorised joints require automation. Some examples include controlling the actuation of: robotic arms in factories; joints in walking robots; cranes in docks carrying inertial loads; ball mill grinders; conveyer belts; automatic doors or pistons; and grinding machines for ore processing, but many other applications exist. In particular, the AI controller 20 of the present invention is suited to applicationswhich require fast reactions to unpredictable external forces or changes in environment, due to thehigh operating frequency of the network. For example, a typical control loop for an unmanned arielvehicle is 400 Hz; a typical deep network may execute at less than 50 Hz on the available hardware, which is too slow to control the vehicle. Faster networks are thus able to target more applications.
[0170] In the presence of external forces, the neural controller 20 of the present inventionoutcompetes PID and other control techniques on energy efficiency, by giving way to the force when the force is towards the controller’s setpoint. Additionally, the AI training process can producenetworks in quite different operating domains using the same or a slightly tuned curriculum, withsimple reconfiguration (modifying the baseline model parameters to match the target motor, and alsoreducing network frequency). In general, this shows the capability of the AI controller 20 of thepresent invention to model second order (or lower) dynamic systems at high frequencies, and to out- compete any one PID controller in second order (or greater) systems. As PID control is also a generally (and very widely) applied control technique, it follows anywhere where a single tuned PIDis used, the AI controller 20 of the present invention could take its place, and be a more efficient,accurate, method of control.
Claims
CLAIMS1. A method of generating a dynamic system controller for controlling a dynamic system, thedynamic system controller comprising a neural network, the method comprising: modelling the dynamic system to generate a system model; defining a training set of control challenges for evaluating candidate neural networks, thecontrol challenges being derived from the system model;generating an optimised neural network using a neuro-evolution optimiser from a population of candidate neural networks, wherein the population of candidate neural networks are evaluated using the training set of control challenges; verifying the optimised neural network against a verification set of control challenges derived from the system model; deploying the optimised neural network into a hardware test controller for the system; evaluating performance of the hardware test controller in a hardware test; and, in the event the hardware test controller fails the hardware test, generating a furtheroptimised neural network and repeating the verification, deployment and evaluation steps;or, in the event the hardware test controller passes the hardware test, deploying the optimised neural network into the dynamic system controller for controlling the dynamic system.
2. A method of generating a dynamic system controller for controlling a dynamic system asclaimed in Claim 1, wherein the population of candidate neural networks are seed networks with a pre-defined neural structure.
3. A method of generating a dynamic system controller for controlling a dynamic system as claimed in Claim 1 or Claim 2 comprising generating a further optimised neural network comprises updating the system model and / or defining a further set of control challenges.
4. A method of generating a dynamic system controller for controlling a dynamic system as claimed in any preceding claim, wherein the neural network comprises neurons and synapses configured to form a directed graph structure and wherein at least some synapses within the neural network are configured to have strengths that can change during network operation.
5. A method of generating a dynamic system controller for controlling a dynamic system as claimed in any preceding claim, wherein the training set of control challenges comprise a series of control challenges that successively increase in complexity.
6. A method of generating a dynamic system controller for controlling a dynamic system, asclaimed in any preceding claim, wherein the verification set of control challenges comprise control challenges that exceed expected difficulty of controlling the dynamic system.
7. A method of generating a dynamic system controller for controlling a dynamic system as claimed in any preceding claim, wherein the dynamic system controller is configured to receive a current system variable of the dynamic system and to control a control variable of the dynamic system towards a setpoint value by outputting a control signal.
8. A method of generating a dynamic system controller for controlling a dynamic system as claimed in Claim 7 wherein the dynamic system comprises an actuator and the dynamic system controller is configured to output the control signal to actuate the actuator towards the set point value.
9. A method of using a dynamic system controller according to any of Claims 1 to 8 to control a dynamic system.
10. The use of a dynamic system controller according to any of Claims 1 to 8 to control adynamic system.
11. A dynamic system controller generated according to the method of any of Claims 1 to 8.
12. A dynamic system controller for controlling a dynamic system as claimed in Claim 11,comprising an input for receiving a set point value for a system parameter and a current system value, and an output for outputting a control signal to control the dynamic system towards the set point value wherein the optimised neural network is configured to generate the control signal in dependence on the received set point value and current system value.
13. A neural network for a dynamic system controller for a dynamic system, the neural networkcomprising: aplurality of neurons n and synapses e configured to form a directed graph structure,wherein at least some synapses within the neural network are configured to have synapse strengths that can change during network operation.A neural network as claimed in Claim 13, wherein, each neuron n at an execution clockinstance, produces an activation value from input activations calculated at a last execution clock instance. A neural network as claimed in Claim 13 or Claim 14, wherein the synapse strengths of at least some of the synapses are configured to change during network operation based on previous neuron activation values. Aneural network as claimed in any of Claims 13 to 15, wherein the synapses e comprise aset of processor synapses, each processor synapse being configured to contribute to an activation of a neuron and a set of modulator synapses, each modulator synapse being configured to contribute to an instantaneous learning rate of a neuron. A neural network as claimed in Claim 16, wherein the neuron comprises an activation function and an instantaneous learning rate and the synapses comprise learning rules encoded as a function of current synapse strength, pre and post synaptic activation and the learning rate, the activation function, instantaneous learning rate and learning rules being used to calculate the strength of each synapse. A neural network as claimed in any of claims 13 to 17, wherein the network encodes a multidigraph structure. A neural network as claimed in any of claims 13 to 18, wherein initial synapse strengths are configured to be quantised. A neural network as claimed in any of claims 13 to 19, wherein each neuron comprises an activation function which comprises a gain parameter. Adynamic system controller for controlling a dynamic system prising a neural networkaccording to Claims 13 to . A dynamic system controller as claimed in Claim 21, comprising an input for receiving a current system value of the dynamic system and an output for outputting a control signal to control the dynamic system, wherein the neural network is configured to receive the current system value as its input and the controller is configured to determine the control signal from the output of the neural network.
23. A controller for controlling a dynamic system comprising:an input for receiving a setpoint value for a control variable of the dynamic system and a current system variable of the dynamic system processing means for calculating a control signal for controlling the dynamic system toward the set point value an output for outputting the control signal wherein the processing means comprises a trained neural network having neurons and synapses connected in a directed graph structure and wherein at least some of the synapseswithin the neural network are configured to have synapse strengths that can change duringnetwork operation.
Citation Information
Patent Citations
Network training process for hardware definition
US11151447B1
Machine-learning-based architecture search method for a neural network
US20210073612A1
Quantization training and image processing methods and devices, and storage medium
WO2021233069A1
Efficient hardware accelerator configuration exploration
WO2023278712A1