System controller

WO2026202392A2PCT designated stage Publication Date: 2026-10-01INTUICELL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/059042
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-27
Publication Date
2026-10-01

Smart Images

  • Figure IMGF000048_0001
    Figure IMGF000048_0001
  • Figure 00000052_0000
    Figure 00000052_0000
  • Figure 00000053_0000
    Figure 00000053_0000
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

[0001] 1 1560.P002

[0002] SYSTEM CONTROLLER

[0003] Technical Field

[0004] The present invention relates to a system controller forming part of a control loop to generate a control signal for a physical system.

[0005] Background

[0006] The use of feedback control loops, or closed loop control, within a physical system is well known. A parameter of the physical system is measured and a drive signal is generated, based on the difference between the measured parameter value and a target value for the parameter, to vary the parameter such that the measured parameter value approaches the target value. Such feedback control loops generally link one measured parameter, for example temperature, with one actuator, for example a heater, that can alter the measured parameter under the influence of the drive signal.

[0007] By way of example, one type of system controller is the PID (Proportional Integral Derivative) controller, which generates a drive signal based on a difference between a measured value for the parameter and the target value (a proportional component), a cumulative difference over time between the measured value for the parameter and the target value (an integrative component), and a rate of change of the difference between the measured value for the parameter and the target value (a derivative component). The integrative component improves the rate at which the measured value approaches the target value in comparison with a purely proportional feedback control system, while the derivative component seeks to reduce any overshoot from the target value.

[0008] An aim of a feedback control loop is to minimise the extent of the departure of the measured parameter value from the target value, both in terms of the magnitude of the departure and the duration of the departure.

[0009] The operation of conventional system controllers in physical systems having multiple measured parameters and multiple actuators can be challenging, particularly when the value of a measured parameter may be affected by more than one actuator and / or when the measured parameters correspond to different modalities, for example2 1560.P002

[0010] temperature and position. This makes reducing the departure of the measured parameter value from the target value particularly challenging.

[0011] This disclosure discusses a system controller that adopts a novel approach to system control that addresses such or similar challenges.

[0012] Summary

[0013] According to a first aspect of the present invention, there is provided a method of generating a control signal for a physical system in dependence upon sensor data for operational parameters of the physical system using a control system comprising a first artificial neural network, ANN, and a second ANN, wherein each of the first ANN and the second ANN has a plurality of nodes and a plurality of edges that interconnect nodes to propagate activity between nodes. The sensor data is associated with setpoint data corresponding to a target performance for the system. The method comprises the control system receiving the sensor data and determining first input to the first ANN based on a difference between the sensor data and the setpoint data, and providing an input to the second ANN corresponding to the first input to the first ANN. The second ANN processes the input to the second ANN to generate output, and the control system provides second input to the first ANN corresponding to the output from the second ANN. The first ANN generates output in dependence on the first input and the second input to the first ANN, and the control system uses the output of the first ANN to generate the control signal for the system. A local learning rule is applied at each node of the first ANN and the second ANN to adjust weights associated with the edges of the first ANN and the second ANN such that the control signal generated using the output from the first ANN reduces a departure of the sensor data from the setpoint data. At least some of the plurality of nodes of the second ANN are stateful nodes with each stateful node being configured to output activity having a first contribution dependent on input activity from other nodes and a second contribution dependent on a latent activity for that stateful node, the latent activity being dependent upon previous input activity received by that stateful node from the other nodes. For at least some of the stateful nodes, the second contribution inhibits activity by an amount dependent on the latent activity, and the latent activity is accumulated and depreciated responsive to received activity such that activity relating3 1560.P002

[0014] to an initial portion of a departure of the sensor data from the setpoint data propagates more strongly through the second ANN than activity relating to later portions of the departure from the setpoint data. As activity propagates through the first ANN and the second ANN, the local learning rule for each node varies weights associated with incoming edges in dependence on activity input to that node so as to contribute to the control signal reducing the magnitude of the first input to the first ANN.

[0015] The introduction of stateful nodes that are able to accumulate and depreciate latent activity in dependence on input activity results in the second ANN as a whole being stateful. By having stateful nodes that inhibit activity propagating through the second ANN such that activity relating to an initial portion of a departure of the sensor data propagates more strongly through the second ANN than activity relating to later portions of the departure of the sensor data from the setpoint data, the weights corresponding to edges in the second ANN can be adjusted primarily based on the activity associated with the initial portion so that the second ANN generates output that provides additional input to the first ANN which is taken into account when adjusting the weights corresponding to the edges within the first ANN in order to generate a control signal that reduces departures of the sensor data from the setpoint data more effectively. For example, if the first ANN is a stateless ANN for which the output at any instant is determined by the current input without taking into account previous inputs and no second ANN is present, an instance of a difference between the sensor data and the setpoint data is treated the same if it occurs at the start of a departure or at the end of a departure, whereas with the stateful second ANN present, that particular instance of a difference in the initial portion of a departure will cause additional activity to be input to the first ANN from the second ANN to assist reducing the departure whereas if the particular difference occurs at the end of a departure, when the corrective action is close to completion, the activity input to the first ANN from the second ANN will typically be less.

[0016] The local learning rule for a node may be configured to increase the weight associated with an incoming edge to that node in dependence on the contribution of the activity received via that incoming edge to the output activity from the node and / or to decrease the weight associated with the incoming edge in dependence upon the magnitude of the output activity such that the weights for incoming edges that4 1560.P002

[0017] propagate activity which contributes stably to the control signal reducing any departure of the sensor data from the setpoint data are potentiated. The exact form of the local learning rule may vary between nodes.

[0018] By using ANNs, complex interrelationships between the sensor readings and the actuator operations can be taken into account when determining how to achieve a target performance. In addition, using a local learning rule to adjust the weights allows continual learning through continuous interaction with the environment, without relying on large training datasets and backpropagation. This continual learning ability allows the control system to correct and adapt to unmodelled dynamics in the real world. The continual learning ability also allows correction for changes within a physical system itself, for example sensor drift.

[0019] The ongoing training of the first ANN and the second ANN enables the weights associated with edges to be adjusted to learn to address previously encountered problems. The stateful behavior of the second ANN assists in the handling of a previously unencountered problem by allowing the second ANN to provide input activity to the first ANN that is expected to cause the first ANN to initiate generation of a control signal that reduces the departure of the sensor data from the setpoint data for the previously unencountered problem based on similarity between the initial portion of the departure for the previously unencountered problem and the initial portions of the departure for one or more previously encountered problems. In this way, the ability of the control system to deal with previously unencountered problems is enhanced.

[0020] In some examples, the target performance may relate to the controlled physical system attaining a target state while in other examples the target performance may relate to the controlled physical system exhibiting a target behavior or a target sequence of states. As such the target performance may be static or dynamic. In an example, a target behavior involves a dynamic response to the sensed environment in which the controlled physical system is operating or based on prior system behavior. At a general level, a target state may define a state that the system should attain whereas a target behavior may define how the system attains that state.

[0021] In example implementations, the second artificial neural network comprises a first sub-network of nodes and a second sub-network of nodes, the second sub-5 1560.P002

[0022] network having a greater number of nodes than the first sub-network and having the stateful nodes for which the second contribution inhibits the output activity by an amount dependent on the latent activity. The input to the second artificial neural network forms a first input to the first sub-network and a first output of the second sub-network forms a second input to the first sub-network, whereby the first subnetwork generates output to the second sub-network dependent on the input to the second artificial neural network and the first output of the second sub-network. The second sub-network generates the first output and a second output in dependence on the output from the first sub-network, the first output enabling activity to cycle through the first sub-network and the second sub-network and the second output being the output from the second artificial neural network to the first artificial neural network.

[0023] The edges associated with the nodes of the first sub-network may be configured to provide inputs to the second sub-network that correspond to different combinations of inputs to the first sub-network. In this way, the first sub-network can be configured to provide distinct inputs to the second sub-network for many different possible problem states of the physical system represented by the inputs to the first sub-network.

[0024] The stateful nodes of the second sub-network may comprise stateful excitatory nodes and the second sub-network may further comprise stateless inhibitory nodes. The second contribution to the output activity for the stateful excitatory nodes corresponds to a reduction of the output activity by a proportion of the latent activity. The latent activity for a stateful excitatory node accumulates responsive to an initial increase in received activity from the first sub-network and then depreciates responsive to a subsequent increase in inhibitory activity in the second sub-network. The latent activity for a stateful excitatory node further depreciates responsive to a subsequent increase in received activity following the initial increase.

[0025] The nodes of the first ANN and the second ANN may introduce a temporal delay to the output activity, with the magnitude of the temporal delay varying between nodes.6 1560.P002

[0026] In example implementations, the second ANN may further comprise a third sub-network of nodes comprising stateful inhibitory nodes, wherein outbound edges from excitatory nodes of the first sub-network and the second sub-network are connected to stateful inhibitory nodes of the third sub-network and outbound edges from inhibitory nodes of the third sub-network are connected to excitatory nodes of the first sub-network. The stateful inhibitory nodes of the third sub-network output activity to the stateful excitatory nodes of the first sub-network in the absence of input activity from the excitatory nodes of the first sub-network and the second subnetwork, and in response to receiving inhibitory activity from the third sub-network, the latent activity of the stateful excitatory nodes of the first sub-network is adjusted such that the second contribution to the output activity increases. In this way, activity is input into the second ANN allowing continued accumulation and discharge of latent activity so that excitatory activity is propagated to the second sub-network and the third sub-network.

[0027] The physical system may be, for example, a robotic system.

[0028] Further features and advantages of the invention will become apparent from the following description of preferred embodiments of the invention, given by way of example only, which is made with reference to the accompanying drawings.

[0029] Brief Description of the Drawings

[0030] Figure l is a block diagram schematically showing the main components of a feedback control loop for a first example system;

[0031] Figure 2 is a block diagram schematically showing in more detail the components of a pre-processing unit forming part of a system controller illustrated in Figure 1;

[0032] Figure 3 is a block diagram showing functional components of a first artificial neural network forming part of the system controller illustrated in Figure 1;

[0033] Figure 4 is a block diagram showing functional components a node forming part of the first artificial neural network illustrated in Figure 3;

[0034] Figure 5 is a block diagram showing functional sub-networks of a second artificial neural network forming part of the system controller illustrated in Figure 1;7 1560.P002

[0035] Figure 6 is a block diagram showing functional components of a stateful node forming part of the second artificial neural network illustrated in Figure 5;

[0036] Figure 7 is a block diagram schematically showing the main components of a feedback control loop for a second example system;

[0037] Figure 8 is a block diagram schematically showing functional components of a pre-processing unit and a first artificial neural network of the feedback control loop of Figure 7;

[0038] Figure 9 is a block diagram showing functional components of an artificial neural network forming part of the system controller of Figure 8;

[0039] Figure 10 is a block diagram schematically showing the main components of a feedback control loop for a third example system;

[0040] Figure 11 is a block diagram schematically showing the main components of a feedback control loop for a fourth example system;

[0041] Figure 12 is a block diagram schematically showing the main physical components of the system controller of any of the first to fourth example systems;

[0042] Figure 13 is a flow chart schematically showing operations performed by the system controller of any of the first to fourth example systems;

[0043] Detailed Description

[0044] FIRST EXAMPLE SYSTEM

[0045] System Overview

[0046] Figure 1 schematically shows the main components of a feedback control loop 1 that could be utilised in many types of system. For example, the feedback control loop 1 could be utilised within a manufacturing system, a tracking system for a camera, a robotic system, an auditory system, or an autonomous vehicle.

[0047] As shown in Figure 1, multiple sensors 3a-3c (hereafter collectively referred to as sensors 3) measure parameters of the system and the surrounding environment and respectively send sensor data signals SA, SB, SC conveying values for the measured parameters to a system controller 5. While for ease of illustration Figure 1 shows three sensors 3, generally there may be any number of sensors 3 and more specifically there may be a plurality of sensors 3. All the sensors 3 may have the same modality,8 1560.P002

[0048] or alternatively the sensors 3 may include sensors having different modalities. For example, one or more of the sensors 3 may be temperature sensors, while others of the sensors 3 may be photosensors, others of the sensors 3 may be pressure sensors, and still others of the sensors 3 may be position sensors.

[0049] The system controller 5, which can alternatively be referred to as a control system, processes the sensor data conveyed by sensor data signals SA, SB, SC to generate control signals CA, CB and Cc which are respectively applied to actuators 7a-7c (hereafter collectively referred to as actuators 7). While for ease of illustration Figure 1 shows three control signals CA, CB and Cc and three actuators 7, generally there may be any number of actuators 7. All the actuators 7 may have the same modality, or alternatively the actuators 7 may include actuators having different modalities. For example, one or more of the actuators 7 may be heaters, while others of the actuators 7 may be motors. The actuators 7 interact with actuated system components 9, and this interaction modifies the parameters sensed by the sensors 3.

[0050] As shown in Figure 1, the system controller 5 includes a pre-processor 11, which in this example compares each of the sensor data signals SA, SB, SC with respective setpoint data 13 to generate input signals for a first ANN 15 and a second ANN 19. The input signals generated by the pre-processor 11 correspond to the differences between the values conveyed by the sensor data signals SA, SB, SC for the measured parameters and corresponding values of the setpoint data. As will be described in more detail, the second ANN 19 processes the input signals from the preprocessor 11 to generate output that is also input to the first ANN 15.

[0051] The first ANN 15 processes the input signals from the pre-processor 11 and the second ANN 19 to generate output signals which, when supplied to a control signal generator 17, cause the control signal generator 17 to generate the control signals CA, CB and Cc in such a way that the interaction between the actuators 7 and the actuated system components 9 may result in the measured parameter values conveyed by the sensor data signals SA, SB, SC being modified to be closer to the corresponding setpoint data 13. As such, the first ANN 15 operates to reduce the values of the input signals generated by the pre-processor 11 to as close as possible to null signals.9 1560.P002

[0052] The second ANN 19 has a plurality of nodes that propagate activity through the second ANN 19 via a plurality of edges that interconnect nodes of the second ANN 19. The second ANN 19 applies a local learning rule at each node to adjust the weights associated with the incoming edges for that node so that the values of the input signals generated by the pre-processor 11 are reduced to be as close as possible to null signals. In this way, the local learning rule is configured to adjust the weights of the incoming edges so that the second ANN 19 generates output to the first ANN 15 that complements the input signals that are output by the pre-processor 11.

[0053] At least some of the nodes of the second ANN 19 are stateful nodes that are configured to accumulate latent activity when the input activity satisfies a threshold condition, and to propagate output activity to other nodes of the second ANN 19 in dependence on the magnitudes of the input activity and the accumulated latent activity. In this way, the stateful nodes of the second ANN 19 can cause the propagation of activity through the second ANN 19 even when the input signals generated by the pre-processor 11 are substantially null signals. As mentioned above, the local learning rule applied at each node within the second ANN 19 still aims to cause the values of the input signals generated by the pre-processor 11 to be as close as possible to null signals.

[0054] It will be appreciated that for each node of the first ANN 15 and the second ANN 19, the local learning rule applied will adjust edges based on whatever activity is received by that node. In this way, the weights associated with edges in the first ANN 15 are adjusted based on previous sensor data representing a departure from the setpoint data to initiate generation of a control signal that reduces that departure. The second ANN 19 also allows the learning developed from previous sensor data to address more effectively previously-encountered departures and also at least some new departures based on the new departures triggering a growth of activity within part of the second ANN 19 that resembles growths of activity triggered by previously encountered departures.

[0055] In this example, the second ANN 19 includes stateful nodes that provide the activity within the second ANN 19 in the absence of substantial difference between the values of the measured parameters and the setpoint data, allowing the first ANN10 1560.P002

[0056] 15 and the second ANN 19 to perform some training in the absence of departures of the sensor data from the setpoint data.

[0057] The System Controller - Function

[0058] Pre-Processor

[0059] As shown in Figure 2, in this example the pre-processor 11 includes separate processing streams for each of the sensor data signals SA, SB, SC. In particular, the sensor data signals SA, SB, SC are input to respective pre-processing functions 2 la-210, with each pre-processing function including a corresponding comparator 23a-23c that compares the input sensor data signal with corresponding setpoint data 25a-25c and outputs a signal corresponding to the difference between the parameter value conveyed by the input signal and the parameter value indicated by the corresponding setpoint data. In this way, each pre-processing function 23 outputs a signal based on the difference between the parameter value conveyed by the input sensor data signal and the parameter value indicated by the corresponding setpoint data such that the greater the difference is, the larger is the magnitude of the output signal.

[0060] The First ANN

[0061] As schematically shown in Figure 3, the first ANN 15 is configured as a random network including a plurality of artificial neurons (hereafter referred to as nodes) that are configured into three sets, in particular a set of input nodes 27, a set of basic nodes 29 and a set of output nodes 31. Hereafter, for ease of reference, the first ANN 15 will be referred to as the reflexive ANN 15 and the second ANN 19 will be referred to as the predictive ANN 19. This reflects the differing functionalities of the reflexive ANN 15 and the predictive ANN 19. The reflexive ANN 15 acts based on input signals from the pre-processor 11 and the predictive ANN 19 with the aim of reducing the magnitude of the input signals from both the pre-processor 11 and the predictive ANN to zero. The predictive ANN 19 complements the functionality of the reflexive ANN 15 by inputting signals to the reflexive ANN 15 that cause the reflexive ANN 15 to reduce departures of the values of the measured parameters from the setpoint data more effectivity.11 1560.P002

[0062] While three input nodes 27a-27c are shown in Figure 3 for ease of explanation, typically there is one input node 27 for each pre-processing function and one input node 27 for each edge interconnecting the reflexive ANN 15 and the predictive ANN 19. Similarly, while three output nodes 33a-33c are shown in Figure 3 for ease of illustration, typically there is one output node 33 for each actuator 7. While four basic nodes 29 are shown in Figure 3 for ease of illustration, typically there will be many more basic nodes 29, for example from ten to five thousand.

[0063] All the nodes of the reflexive ANN 15 are labelled either “excitatory” (Exc in Figure 3) or “inhibitory” (Inh in Figure 3). More particularly, all the input nodes 27 and all the output nodes 31 are stateless excitatory nodes while a subset of the basic nodes (represented in Figure 2 by the basic nodes 29a and 29d) are stateless excitatory nodes with the remainder being stateless inhibitory nodes (represented in Figure 3 by the basic nodes 29b and 29c).

[0064] A node outputs activity to other nodes such that activity propagates between nodes of the reflexive ANN 15. A node is a stateless node if the activity output by that node as configured at any instant is based entirely on the magnitude of the latest activity signals received by that node from other nodes, with the configuration of a node encompassing the weight co-efficients associated with incoming edges from other nodes. The difference between an excitatory node and an inhibitory node is that activity signals received by a recipient node from an excitatory node contribute to increasing the magnitude of the activity signal output by the recipient node whereas activity signals received by a recipient node from an inhibitory node contribute to reducing the magnitude of the activity signals output by the recipient node, as will explained in more detail hereafter.

[0065] In this example, nodes within the same set and having the same label are configured to have a predefined number of output edges. Accordingly, each of the input nodes 27 has a single input edge and a first predefined number of output edges interconnecting the input node with basic nodes 29 and output nodes 31. Each of the excitatory basic nodes 29 has a second predefined number of output edges interconnecting the excitatory basic node with other basic nodes 29 and output nodes 31. Each of the inhibitory basic nodes 29 has a third predefined number of output edges interconnecting the inhibitory basic node 29 with other basic nodes 29 and12 1560.P002

[0066] output nodes 31. Each of the output nodes 31 has a single output edge connected to a respective one of a set of signal generators, which form the control signal generator 17 of Figure 1, and input edges as mentioned above.

[0067] In this example, the edges within the reflexive ANN 15 are initially assigned taking into account knowledge of the relationship between sensors 3 and actuators 7. For example, it may be known that operation of a first actuator 7 will have a comparatively strong impact on the parameter value detected by a first sensor 3, whereas the operation of a second actuator 7 will have a comparatively strong impact on the parameter value detected by a second sensor 3. Accordingly, the reflexive ANN 15 is configured taking into account knowledge of the physical system so that, for example, a first population of the basic nodes 29 is assigned comparatively densely with edges to allow direct paths through the first population of basic nodes 29 from a first input node 27 corresponding to the first sensor 3 to a first output node 31 corresponding to the first actuator 7, a second population of the basic nodes 29 is assigned comparatively densely with edges to allow direct paths through the second population of basic nodes 29 from a second input node 27 corresponding to the second sensor 3 to a second output node 31 corresponding to the second actuator 7, and edges are comparatively sparsely assigned between nodes of the first population and nodes of the second population to account for the influence of the first actuator on the signal sensed by the second sensor and the influence of the second actuator on the signal sensed by the first sensor.

[0068] The topology of the sensory distribution in the physical system may influence distribution of edges within at least one of the first artificial neural network and the second artificial neural network. More generally, the insertion of the plurality of directed edges may be determined based on probabilistic function that is adapted for intra-population and inter population edge assignments based on knowledge of the physical system when the reflexive ANN 15 is constructed, such that pre-established relationships between sensors and actuators are emphasized in the resulting connectivity.

[0069] In this example, the signals output from the output nodes 31 can have a positive or negative effect on a subset of the input signals, and more generally the13 1560.P002

[0070] output signals in combination can have a positive or negative effect on the input signals in combination.

[0071] The Nodes of the First ANN

[0072] Figure 4 schematically shows the processing of signals received by a recipient node 41 from multiple excitatory nodes, represented in Figure 4 by three excitatory nodes 43a-43c and hereafter referred to as excitatory nodes 43, and multiple inhibitory nodes, represented in Figure 3 by two inhibitory nodes 45a-45b and hereafter referred to as inhibitory nodes 45. It will be appreciated that the actual number of excitatory nodes 43 and inhibitory nodes 45 in practical implementations will generally be significantly higher.

[0073] The activity signals received by the recipient node 41 from the excitatory nodes 43 and the inhibitory nodes 45 are input to respective different input functions 47a-47e. For each input function 47, the output y is determined in a periodic manner according to the function:

[0074] y(prev_y, x):=MAX(prev_y*label_specific_decay, x)

[0075] where x is the value of the input to the input function 47, prev_y is a value corresponding to previous output from the input function 47, and label specific delay is a parameter between zero and one that may have different values when the input signal being processed is from an excitatory node and when the input signal being processed is from an inhibitory node. The effect of the input function is that if there is a reduction in the value x of the input of the input function 47 that results in the value x decaying faster than the decay of the previous output of the input function 47 corresponding to the value of the label specific decay parameter, then the value of the decayed previous output y is used in preference to the value of the input x to the input function as the output y of the input function y. This reduces high-frequency signals, which assists in determining a solution for the weights of the edges throughout the reflexive ANN 15. Such high frequency signals may be generated within recurrent ANNs because there is always a risk of creating positive feedback loops that saturate the network activity, rendering it unresponsive to actual sensory14 1560.P002

[0076] input, and although such positive feedback loops can be at least partially quenched by the inhibitory nodes 45, the inhibitory quenching lags the build-up of excitatory activity thereby creating high frequency self-amplifying transients. By smoothing the activity of the individual neurons using a decay function for its activity, such selfamplifying transients can be reduced or even avoided, thereby allowing the network activity to focus on determining control signals that reduce the difference between the sensor data and the setpoint data.

[0077] The value y of the output from each input function 47 is then input to a respective weight function 49, where the value y is multiplied by a weight w corresponding to the edge via which the input signal for that input function 47 was received by the recipient node 41.

[0078] The outputs of the weight functions 49 for signals received from excitatory nodes 43 are then input to a first combiner function 51a, which sums the outputs together to generate a sum L_exc, where:

[0079] L_exc = Syi*Wi over all the outputs i corresponding to excitatory inputs.

[0080] Similarly, the outputs of the weight functions 49 for signals received from inhibitory nodes 45 are then input to a second combiner function 51b, which sums the outputs together to generate a sum L_inh, where:

[0081] L_inh = Eyi*Wi over all the outputs i corresponding to inhibitory inputs.

[0082] The values of the parameters L_exc and L_inc are output by the first combiner 51a and the second combiner 51b respectively and input to a base function 53, which determines the magnitude of the activity signal output by the recipient node 41 to other nodes. In this example, the output x of the base function is determined by the expression:

[0083] x := MAX(0, F)

[0084] where F = a*L_exc - b*L_inh - c*L_exc*L_inh.15 1560.P002

[0085] where a, b and c are constants. In some examples, the constant c may be set to zero.

[0086] This expression for the base function mitigates against the possibility of positive feedback loops being present within the reflexive ANN 15, with the L_inh parameter being a determining factor for the rate at which activity in the ANN 15 is reduced. The values of the constants a, b and c can set to many different combinations of values. In some combinations, the value of the constant c can be set to zero. So one possible combination of values is a=l, b=0.5 and c=0, whereas another possible combination of values can be set to a=l, b=l and c=l.

[0087] The output x of the base function 53 is input to an output function 55 which propagates the output x along the output edges of the node 41. In this example, the output function 55 introduces a latency to the propagation of the output x, with a latency value being specified for each output edge. In this example, the latency values are specified in a random manner.

[0088] Introducing latencies to the propagated signals introduces non-linearities into the reflexive ANN 15, which allows the activity of different nodes in the reflexive ANN 15 to be differentiated. In this way, the time-varying signal in each node is more unique, thereby increasing the number of options to find solutions in the reflexive ANN 15 by amplifying the weights of the edges from those nodes. This assists in the reflexive ANN 15 converging to a robust solution.

[0089] The base function 53 also outputs the L_exc parameter and the L_inh parameter to a learning function 57 which adjusts the weights corresponding to a subset of the input edges so as to reduce the activity propagated in the reflexive ANN 15 over time and bring the reflexive ANN 15 into a stable solution. More particularly, the learning function 57 is a local learning function which adjusts weights for input edges to the corresponding node based on parameters associated with that node.

[0090] Broadly speaking, the learning function increases the weight associated with an incoming edge to that node in dependence on the contribution of the activity received via that incoming edge to the output activity from the node and / or decreases the weight associated with the incoming edge in dependence upon the magnitude of the output activity such that the weights for incoming edges that propagate activity which contributes stably to the control signal reducing any departure of the sensor16 1560.P002

[0091] data from the setpoint data are potentiated. In some example implementations, if the application of these weight changes to an incoming edge of a node causes a positive feedback loop to form, then that incoming edge may be removed and a replacement incoming edge added. In other example implementations, which may for example involve the input activity from inhibitory nodes having a stronger effect on the output activity, no modification of the configuration of edges is required.

[0092] A specific learning function will now be described by way of example. In the example, the learning function 57 only adjusts the weights for the input edges for which the output of the input function 47 is among the highest. By focussing the learning on the received activity signals that are strongest, in effect the reflexive ANN 15 prioritises reducing the highest areas of activity. This approach assists in reaching a solution, particularly for complex systems where there is no one-to-one correspondence between an actuator and a sensed parameter.

[0093] In this example, the learning function 57 only modifies the weights for input edges for which the expression

[0094] prev_y > ALL_y* 0.75

[0095] is satisfied, where ALL_y is the maximum value of the signal y output by an input function 47 for that node, although it will be appreciated that many different expressions could be used to arrive at the result of selecting the edges providing the strongest incoming activity signals to that node. For example, the value 0.75 could be replaced by a higher or lower value in the expression given above, or alternatively a predetermined number of the strongest incoming activity signals or a predetermined proportion of the strongest incoming activity signals could be selected.

[0096] For each of the weights being modified by the learning function 57, if the corresponding input edge connects to an excitatory node and the value of the output y of the corresponding input function is greater than the value of the L_exc parameter, then the learning function 57 increases that weight w by an amount dw that may be expressed as:

[0097] dw += MAX(O, y - 1 + s - L_exc)*rate17 1560.P002

[0098] where s is a suitability parameter for the node 41 and rate is a learning rate, which is a scalar value used to control the size of dw. If the value of the L_exc parameter is greater than a threshold value T, then the learning function 57 reduces that weight w by an amount dw that may be expressed as:

[0099] dw -= MAX(0, L_exc - T)*rate.

[0100] Similarly, if the input edge corresponding to a weight being modified by the learning function 57 connects to an inhibitory node then if the output y of the corresponding input function is greater than the value of the L_inh parameter, then the learning function 57 increases that weight w by an amount dw that may be expressed as:

[0101] dw += MAX(0, y - 1 + s - L_inh)*rate

[0102] and if the value of the L_inh parameter is greater than a threshold value T, then the learning function 57 diminishes that weight w by an amount dw given by the expression:

[0103] dw -= MAX(0, L inh - T)*rate.

[0104] As the condition for increasing a weight is dependent on the output y and the condition for reducing a weight is dependent on the summation L_exc, L_inh, it is possible for both conditions to be satisfied in which case the weight is adjusted by the final value of dw after addition and subtraction.

[0105] In this example, the suitability parameter s of the node 41 is modified in dependence on changes to the length of a vector V = (y, L_exc, L_inh). In particular, if dV is zero or negative, suggesting that one or both of L_exc and L_inh is decreasing, then the suitability s is increased, for example by a fixed amount, whereas if dV is positive, suggesting an increase in one or both of L_exc and L_inh, then the suitability s is diminished, for example by a fixed amount.18 1560.P002

[0106] Increasing the suitability s has the effect that for the same difference between the output y for an edge and L_exc when the node 41 is an excitatory node, or the same difference between the output y for an edge and L_inh when the node 41 is an inhibitory node, the weight w corresponding to that edge can be potentiated by a greater amount. If, however, such an increase in weight results in worse performance of the system and the input y, V will increase over time, leading to the suitability reducing.

[0107] If a positive feedback loop develops in the reflexive ANN 15, then the output y corresponding to an edge forming part of the positive feedback loop will grow. The resultant increase in activity results in the suitability s of edges associated with that positive feedback loop diminishing, thereby reducing or eliminating any increase in the weight w for those edges. In this example, in the event that the output y for an edge exceeds the threshold T (for example y > 1) and the weight w cannot be potentiated, then that edge may be randomly assigned a different endpoint node or that edge may be removed and another edge randomly inserted elsewhere in the reflexive ANN, thereby assisting to break any positive feedback loop.

[0108] The operation of the reflexive ANN 15 described above results in the weights for edges being modified until the dw reaches zero for all nodes and the values of the sensed parameters match the values of the corresponding setpoint data. If the values of L_exc and L_inh are also stably under the threshold T, then all the pathways between the input nodes and the output nodes form part of a negative feedback control system.

[0109] The Second ANN

[0110] As shown in Figure 5, the second ANN 19, also referred to as the predictive ANN 19, is formed of three groups of nodes, a THA sub-network 61, a CTX subnetwork 63 and a RTN sub-network 65. As explained previously, the predictive ANN 19 also receives the signals from the pre-processor 11 and is configured, in cooperation with the reflexive ANN 15, with the aim of generating control signals that reduce the magnitude of the input signals from the pre-processor 11 to zero.

[0111] The signals input from the pre-processor 11, and more generally the signals input to any node, can be referred to as input activity. As will be explained in more19 1560.P002

[0112] detail hereafter, some of the nodes of the THA sub-network 61, the CTX sub-network 63 and the RTN sub-network 65 are able to generate output activity in dependence upon latent activity associated with that node in addition to the input activity. Further, the nodes of the THA sub-network 61, the CTX sub-network 63 and the RTN subnetwork 65 are configured in a recurrent manner to allow activity to propagate cyclically. Local learning rules are applied at the nodes that adjust the weights associated with the connections of the nodes in order to support activity propagating stably in the predictive ANN 19, in response to input activity, that contributes to the control signal reducing departures of the sensor data from the setpoint data.

[0113] The ability of some nodes of the predicitve ANN 19 to accumulate and depreciate latent activity allows the predicitve ANN 19 to arrive at solutions in which the output to the first ANN 15 causes the first ANN 15 to initiate control signals that more effectively reduce departures of the sensor signals from the setpoint data base on learning from previous departures. The nodes that have corresponding latent activity will hereafter be referred to as stateful nodes.

[0114] The Stateful Nodes of the Predicitve ANN

[0115] As shown in Figure 6, in which functional components that are the same as corresponding components of the stateless node 41 of the reflexive ANN 15 illustrated in Figure 4 are referenced using the same reference numerals, in this example a stateful node 71 differs from a stateless node 41 by a modified base function and the inclusion of a store 75 of latent activity Q. It will be appreciated that the store of latent activity 75 represents an additional parameter value that is taken into account in the routine that is executed to implement the base function 73.

[0116] In this example, for each stateful node, values for five parameters are specified that relate to how the value of the latent activity Q changes and how the value of the latent activity Q affects the output x of the base function 73 associated with the latent activity. In particular, the value of a parameter K determines how the base function 73 operates in dependence on the values of F, where in this example F = L_exc -(L_inh / 2).

[0117] During the time period when F < K, the value of Q increases in accordance with a rate value increase and the output x of the base function 73 is determined by20 1560.P002

[0118] sum of the value of F, which forms a first contribution to the output, and a value proportional to the value of Q with a proportionality constant srtl, which forms a second contribution to the output, provided that if that sum is less than zero than the output x is set to zero. These two expressions can be represented as:

[0119] Q: [Q_next:= Q + Q*increase]

[0120] x := MAX(0, F + Q*strl)

[0121] During the time period when F > K, the value of Q decreases according with a constant decay, which is between 0 and 1, and the output x of the base function 73 is determined by sum of the value of F and a value proportional to the value of Q with a proportionality constant srtl, whose value is greater than strl, again provided that if that sum is less than zero than the output x is set to zero.. These two expressions can be represented as:

[0122] Q: [Q_next:= Q + Q*decay]

[0123] x := MAX(0, F + Q*str2)

[0124] As will be described hereafter, the values of the five parameters (K, increase, decay, strl, strl) may vary between stateful nodes.

[0125] The CTX Sub-Network

[0126] The CTX sub-network 63 is an artificial neural network containing

[0127] ctx excitatory nodes and basic inhibitory nodes. The ctx excitatory nodes are stateful nodes for which the parameter K is set to zero. In this way, the ctx excitatory nodes only accumulate activity Q when the expression F is less than zero. The basic inhibitory nodes are stateless nodes. The number of nodes in the CTX subnetwork is significantly larger than in the THA sub-network and in the RTN subnetwork.21 1560.P002

[0128] Some of the ctx excitatory nodes have incoming edges that are connected to nodes within the THA sub-network 61. Via these incoming edges, activity is injected into the CTX sub-network 63.

[0129] Some of the ctx excitatory nodes have outgoing edges connected to nodes within the reflexive ANN 300, which form the output of the predictive ANN 19 to the reflexive ANN 300. In particular, those outgoing edges enable activity from the predictive ANN 19 to be injected into the reflexive ANN 300.

[0130] Some of the ctx excitatory nodes have outgoing edges connected to nodes within the RTN sub-network 65.

[0131] Some of the ctx excitatory nodes have outgoing edges connected to nodes within the THA sub-network 61.

[0132] For the ctx excitatory nodes, the values of strl and str2 are negative so that a build-up of latent activity results in an increasing inhibitory effect.

[0133] The THA Sub-Network

[0134] The THA sub-network 61 is a partially connected contains tha excitatory nodes, which are stateful nodes for which the parameter K is set to zero, and the increase parameter is greater than the increase parameter of the ctx excitatory nodes.

[0135] Some of the tha excitatory nodes have incoming edges via which signals from the pre-processor 11 are input to the predictive ANN 19.

[0136] As mentioned above, some of the tha excitatory nodes have outgoing edges connected to ctx excitatory nodes, and some of the tha excitatory nodes have incoming edges connected to tha excitatory nodes.

[0137] Some of the tha excitatory nodes have outgoing edges connected to nodes within the RTN sub-network 65, and some of the tha excitatory nodes have incoming edges connected to nodes within the RTN sub-network 65.

[0138] The RTN Sub-Network

[0139] The RTN sub-network 65 is a partially connected network contains rtn inhibitory nodes, which are stateful nodes for which K is greater than zero and the increase parameter is less than the increase parameter for the ctx excitatory nodes.22 1560.P002

[0140] Some of the rtn inhibitory nodes have incoming edges connected to tha excitatory nodes. Some of the rtn inhibitory nodes have incoming edges connected to tha excitatory nodes.

[0141] Some of the rtn inhibitory edges have outgoing edges connected to the tha excitatory nodes.

[0142] By setting K to be greater than zero, latent activity Q is accumulated in the RTN sub-network 65 even when there is no incoming activity.

[0143] Operation of the Predictive ANN

[0144] The general aim of the predictive ANN 19 is to input signals into the reflexive ANN 15 that cause the reflexive ANN 15 to initiate the generation of a control signal that reduces departures of the sensor data from the setpoint data more effectively. In effect, the input signals to the reflexive ANN 15 presents a problem vector and the predictive ANN 19 assists the reflexive ANN to learn solutions (i.e. configurations of weightings and / or connections) that reduce the length of the problem vector more effectively.

[0145] To achieve these solutions, the THA sub-network 61 and the CTX subnetwork 63 form an inverted autoencoder, in the sense that the number of nodes of the CTX-network 63 is significantly greater than the number of nodes of the THA subnetwork 61. The CTX sub-network 63 expands and evolves the problem state provided by the THA sub-network 61 by propagating activity using the node dynamics explained above to amplify beneficial trajectories within the CTX subnetwork 63 and dampen undesirable trajectories within the CTS sub-network 63, and then compresses the problem state into the output that is fed back into the THA subnetwork 61. The stateful ctx excitatory nodes allow the activity associated with an initial portion of a departure of the sensor data from the setpoint data to propagate comparatively strongly through the CTX sub-network 63, while inhibiting the propagation of activity associated with later portions of the departure. In this way, the weights of the CTX sub-network 63 are affected more by the initial portion than later portions.

[0146] In the absence of activity in the THA sub-network 61 and the CTX subnetwork 63, the RTN sub-network 65 accumulates latent activity, causing inhibitory23 1560.P002

[0147] activity to be injected into the THA sub-network 61. This inhibitory activity causes the parameter F to fall below zero, resulting in a build-up of latent activity in the tha excitatory nodes. This, in turn, causes activity to be injected into the CTX subnetwork 63. In this way, the CTX sub-network 63 is able to search for beneficial trajectories at times when the sensors 3 have not yet detected any departure from the target performance.

[0148] SECOND EXAMPLE SYSTEM

[0149] System Overview

[0150] By way of example, Figure 7 schematically shows the main components of a robotic system 101, including a system controller 105 and an articulated system 109. The system controller 105 and the articulated system 109 are shown separately in Figure 5 for ease of illustration, but it will be appreciated that the system controller 105 and the articulated system 109 may both form part of the same robotic apparatus, such as a biped robot or a quadruped robot.

[0151] As shown schematically in Figure 7, the articulated system 109 includes multiple sections interconnected by joints. Although only three sections and two joints are shown in Figure 7, it will be appreciated that the articulated system may involve a complex arrangement of sections and joints. In some other examples, the articulated system 109 may comprise two rigid components that are connected by a joint, each of the components not having any additional joint.

[0152] The joints enable the sections to move relative to each other in order to change a state of the articulated system 109. In this example, multiple sensors 103a-103c (hereafter collectively referred to as sensors 103) within the articulated system 109 measure parameters of the articulated system 109 and the surrounding environment and respectively generate sensor data signals SA, SB, SC conveying values for the measured parameters. While, for ease of illustration, Figure 7 shows three sensors 103, generally there may be any number of sensors 103 and more specifically there may be a plurality of sensors 103. All the sensors 103 may have the same modality, or alternatively the sensors 103 may include sensors having different modalities. For example, one or more of the sensors 103 may be position sensors, e.g. for detecting the positions of the joints, while others of the sensors 103 may be photosensors, for24 1560.P002

[0153] example for detecting the environment around the robotic system. Others of the sensors 103 may be pressure sensors, and still others of the sensors 103 may be temperature sensors.

[0154] The articulated system 109 also includes multiple actuator systems 107a-107c (hereafter collectively referred to as actuators 107). All the actuators of the actuator systems 107 may have the same modality, or alternatively the actuator systems 107 may include actuator systems with actuators having different modalities. For example, one or more of the actuators of the actuator systems 107 may be clamps, while others of the actuators may be motors. At least some of the actuators of the actuator systems 107 interact with sections and joints of the articulated system 109, and this interaction modifies the current state sensed by the sensors 103.

[0155] As shown in Figure 7, in this example the sensor data signals SA, SB, SC are respectively input to different ones of the actuator systems 107. In each actuator system 107, the measured value conveyed by the respective sensor data signal is input to a feedback controller (not shown) which compares the measured value to a corresponding target value and supplies a drive signal to the corresponding actuator in dependence upon the difference between the measured value and the target value. The actuators may involve one or both of linear and rotary actuators. The actuators may control joints or flexible elements without joints such as flexible elements configured as grippers.

[0156] The sensor data signals SA, SB, SC are also supplied to the system controller 105, which generates control signals CA, CB and Cc that are applied to respective different ones of the actuator systems 107. While for ease of illustration Figure 5 shows three control signals CA, CB and Cc and three actuator systems 107, generally there may be any number of actuator systems 107.

[0157] In this example, each actuator system 107 processes the respective control signal to determine the target value that is input to the feedback controller of that control signal for comparison with the measure value conveyed by the respective sensor signal. As will be described in more detail hereafter, the system controller 105 is able to determine a control signal for an actuator system 107 based on the measured values conveyed by multiple sensor signals. In this way, the operations performed by the actuators of the actuator systems 107 can co-operate with each other to achieve an25 1560.P002

[0158] object for the robotic system or part of the robotic system, rather than simply aiming to achieve an object for that particular actuator system 107.

[0159] While only an articulated system 109 is shown in Figure 7, the system controller 105 may also be used to control other sub-systems of the robotic system 101 comprising mechanisms such as actuators for limb systems, propulsion systems including those based on wheels, pose and positioning mechanisms.

[0160] As shown in Figure 7, the system controller 105 includes a pre-processor 111, which compares each of the sensor data signals SA, SB, SC with respective set point data 113 to generate input signals for a first ANN 115 and a second ANN 119. The set point data 113 represents a target behavior of the robotic system, and the first ANN 115 processes the input signals to generate output signals which, when supplied to a control signal generator 117, cause the control signal generator 117 to generate the control signals CA, CB and Cc in such a way that the interaction between the actuators 107 and the actuated system components 109 may result in the measured parameter values conveyed by the sensor data signals SA, SB, SC being modified to be closer to the corresponding set point data 113, thereby causing the robotic system to achieve its target behavior. The operation of the first ANN 115 and the second ANN 119 is substantially the same as for the first ANN 15 and the second ANN 19 of the first example system described above.

[0161] By using the first ANN 115 and the second ANN 119, the system controller 105 can handle feedback control for systems where there is a complex interaction between the actuators 107 and the parameters sensed by the sensors 103. For example, one of the actuators 107 may impact multiple sensed parameters or multiple actuators 107 may impact a single sensed parameter.

[0162] The System Controller - Function

[0163] In this example, as shown in Figure 8, the pre-processor 111 includes separate processing streams for each of the sensor data signals SA, SB, SC. In particular, the sensor data signals SA, SB, SC are input to respective pre-processing functions 121a-121c, with each pre-processing function including a comparator function 123 a- 123c that compares the input sensor data signal with corresponding set point data 125a-125c and outputs a signal corresponding to the difference between the parameter26 1560.P002

[0164] value conveyed by the input signal and the parameter value indicated by the corresponding set point data. The output of each comparator function 23a-23c is normalised by a respective normalisation function 127a-127c, which determines the absolute value of the output of the comparator function and normalises the resultant absolute value to a value between zero and a given maximum value, for example a maximum value of one may be used in some embodiments as an upper limit for normalisation.

[0165] In this example, the first ANN 115 includes a first ANN sub-network 131a and a second ANN sub-network 131b. The output of the normaliser function is routed to one of the first ANN sub-network 131a and the second ANN sub-network 13 lb in dependence on whether the output of the comparator function is positive or negative. In figure 8, this is schematically represented by each sensor signal being input to a comparator 123, with the output of each comparator 123 being input to a normaliser 127, which applies the normalisation function, and the output of the normaliser 127 being passed to a first ANN sub-network 13 la if the output of the comparator 123 is positive and to a second ANN sub-network 13 lb if the output of the comparator 123 is negative.

[0166] In this way, each pre-processing stream 123 outputs a signal conveying a positive value between zero and a given maximum value that is determined based on the difference between the parameter value conveyed by the input sensor data signal and the parameter value indicated by the corresponding set point data such that the greater the difference is, the larger is the magnitude of the output signal.

[0167] Although not shown in Figure 8, the second ANN 119 also includes two ANN sub-networks, with the output of the normaliser being sent to one of the ANN subnetworks if it is positive and to the other of the ANN sub-networks if it is negative. The ANN sub-network of the second ANN 119 which processes positive values provides output to the first ANN sub-network 131a and the ANN sub-network of the second ANN sub-network 119 which processes negative values provides output to the second ANN sub-network 131b.

[0168] In this example, the form and operation of the first ANN sub-network 131a and the second ANN sub-network 131b, which together form the ANN 115 of Figure 5, are substantially the same. As shown in Figure 9, in this example, each ANN sub-27 1560.P002

[0169] network 131 is a random network that is substantially the same as the ANN 15 described above with reference to Figure 2 (which has been indicated by using the same reference numerals as Figure 2), with nodes that are substantially the same as described above with respect to Figure 3.

[0170] In this example, the signals output from the output nodes 45 can have a positive or negative effect on a subset of the input signals, and more generally the output signals in combination can have a positive or negative effect on the input signals in combination.

[0171] THIRD EXAMPLE SYSTEM

[0172] System Overview

[0173] As described above, actuators may interact with actuated system components, and the sensors measure parameters associated with the actuated system components and the surrounding environment. Alternatively, as shown in Figure 10, one or more of the sensors 203 could sense parameters associated with the actuators 207 themselves. In particular, in Figure 10 each actuator system 201a-201c includes a sensor 203a-203c, an actuator 205a-205c and a feedback controller, which may comprise a PID feedback controller 207a-207c as illustrated in Figure 10. Each sensor 203 detects a parameter of the corresponding actuator 205 and sends a measurement signal S conveying the measured parameter value to the measurement input of the PID feedback controller 207 and also to the system controller 205, which functions in the same manner as the system controller of Figure 8 to provide a control signal C to the reference input of the feedback controller 207.

[0174] FOURTH EXAMPLE SYSTEM

[0175] System Overview

[0176] As shown in Figure 11, in a third example system, the second example system is modified by including a task manager 301 that outputs high level action signals to the control system 305 in response to input identifying a desired behavior. The control system 305 may be the same as the control system of any of the first to third example systems, except for the addition of a processing component that processes the28 1560.P002

[0177] action signal to identify corresponding set point data, and then updates the set point data for the artificial neural network accordingly.

[0178] The task manager 301 receives sensor signals from an actuated, for example articulated, system 309, and bases the action signals on those sensor signals. In some examples, the task manager 301 employs reinforcement learning to determine the appropriate action signal. The action signal may indicate a trajectory through a sequence of states, for example a walking motion for a bipedal robot.

[0179] Generally, the task manager 301 will operate at a rate that is slower from the rate at which the control system operates. In an example, the task manager operates at a rate in the order of 100Hz, whereas the control system operates at a rate of the order of 1kHz.

[0180] PHYSICAL DEVICE FEATURES

[0181] Figure 12 shows, by way of example, the main components for a software implementation of the system controller. As shown, the system controller 401 includes input / output devices 403, a processor 405 and memory 407.

[0182] The input / output devices 403 include one of more input devices for receiving the sensor data signals SA, SB and Sc. In some examples, there is one input device for each sensor data signal while in other examples there is a single input device having multiple ports allowing the single input device to receive the sensor data signals SA, SB and Sc. It will be appreciated that the input devices may conform to standard specifications as are well known in the art.

[0183] The input / output devices 403 also include one or more output devices for transmitting the control signals CA, CB and Sc. In some examples, there is one output device for each control signal while in other examples there is a single output device which transmits a multiplexed control signal allowing the single output device to transmit the control signals CA, CB and Cc. It will be appreciated that the output devices may conform to standard specifications as are well known in the art.

[0184] While the processor 405 is illustrated as a single component, it will be appreciated that the processor 405 may include multiple processing devices. For example, the processing operations may be distributed between multiple processing devices within the system controller 401.29 1560.P002

[0185] The memory 407 may include multiple memory devices having respective different properties, such as access times and permanence, in a manner well known in the art. The memory 407 stores data 409, program routines 411 and also provides working memory 413. The data 419 includes, for example, ANN parameters 415 providing configuration details and edge weights for the ANN and set point data 417. The routines 407 include a pre-process sensor data routine 419, a propagate activity routine 421, a learning rule routine 423 and a generate control signal routine 425.

[0186] The pre-process sensor data routine 419 determines the input signals for the ANN based on differences between the sensor data signals SA, SB and Sc received by the input / output devices 403 and the set point data 417 stored in the memory 407. The generate control signal routine 425 processes output signals from the ANN to generate the control signals CA, CB and Cc, and outputs the control signals CA, CB and Cc using the input / output devices 403.

[0187] The propagate activity routine 421 propagates activity signals through the ANN based on the ANN parameters 415 in the manner described above. The learning rule routine 423 modifies the ANN parameters 415 in the manner described above.

[0188] It will be appreciated that the system controller could alternatively be implemented in hardware, or a different combination of hardware and software, and perform the same processing operations.

[0189] MODIFICATIONS AND FURTHER EMBODIMENTS

[0190] Figure 13 is a flow chart summarising the common operations of the example embodiments set out above to generate a control signal for a physical system in dependence upon sensor data for operational parameters of the physical system using a control system comprising a first artificial neural network, ANN, and a second ANN, wherein each of the first ANN and the second ANN has a plurality of nodes and a plurality of edges that interconnect nodes to propagate activity between nodes. The sensor data is associated with setpoint data corresponding to a target performance for the system.

[0191] Initially, the first ANN and the second ANN are configured, at SI, by inserting edges and setting initial values for the weights associated with those edges. The control system then receives, at S3, the sensor data and determines, at S5, first input to30 1560.P002

[0192] the first ANN based on a difference between the sensor data and the setpoint data. The control system then provides, at S7, an input to the second ANN corresponding to the first input to the first ANN and the second ANN processes the input to the second ANN to generate output. The control system provides, at S9, second input to the first ANN corresponding to the output from the second ANN and the first ANN generates output in dependence on the first input and the second input to the first ANN. The control system then determines, at SI 1, the output of the first ANN and generates, at S13, the control signal for the system. A local learning rule is applied at each node of the first ANN and the second ANN to adjust, at SI 7, weights associated with the edges of the first ANN and the second ANN such that the control signal generated using the output from the first ANN reduces a departure of the sensor data from the setpoint data. At least some of the plurality of nodes of the second ANN are stateful nodes with each stateful node being configured to output activity having a first contribution dependent on input activity from other nodes and a second contribution dependent on a latent activity for that stateful node, the latent activity being dependent upon previous input activity received by that stateful node from the other nodes. For at least some of the stateful nodes, the second contribution inhibits activity by an amount dependent on the latent activity, and the latent activity is accumulated and depreciated responsive to received activity such that activity relating to an initial portion of a departure of the sensor data from the setpoint data propagates more strongly through the second ANN than activity relating to later portions of the departure from the setpoint data. As activity propagates through the first ANN and the second ANN, the local learning rule for each node varies weights associated with incoming edges in dependence on activity input to that node so as to contribute to the control signal reducing the magnitude of the first input to the first ANN.

[0193] While some different implementations have been described above in relation to the configuration of the system controller and the physical system, and the performance of the operations, further possible modifications will now be discussed.

[0194] In the descriptions of the reflexive ANN 15 and the predictive ANN 19, specific examples have been given of expressions for the local learning rule and the generation of output activity. It will be appreciated that these are given by way of example only, and alternative expressions may be used that achieve the same purpose.31 1560.P002

[0195] A purpose of the predictive ANN 19 is to learn to produce output to the reflexive ANN 15 based on the initial portion (that is the portion that occurs temporally first) of a departure of the sensor data from the setpoint data that assists the reflexive ANN 15 to adjust the weights of the nodes of the reflexive ANN 15 to arrive at a solution which results in a control signal being generated when a departure having the same or substantially similar initial portion occurs that reduces the departure more effectively.

[0196] In CTX sub-network 63, the rules governing the manner in which latent energy is accumulated and depreciated and the manner in which the latent energy contributes to the output activity are determined so that the initial portion of a departure propagates through the CTX sub-network while the propagation of later portions of the departure are inhibited. In this way, the output of the CTX subnetwork 63 and the adjustment of the weights within the CTX sub-network 63 are predominantly determined by the initial portion of the departure. It will be appreciated that many different mathematical expressions could be used in the rules to implement this functionality. In particular, while the mathematical expression for the contribution to the output activity set out above by way of example involve switching from a first rate strl to a second rate str2 when the stored latent activity reaches a threshold value, this is not necessary. In addition, a mathematical expression could be employed in which as input activity increase the latent activity decreases so as to move further from a reference value, and the contribution to the output activity is determined from the different between the amount of latent energy and the reference value.

[0197] As described above, the THA sub-network 61 is configured to provide inputs to the CTX sub-network 63 that are distinctive of different problem states. A purpose of the latent activity in the THA sub-network is to provide a mechanism which generates output activity to the CTX sub-network 63 responsive to inhibitory activity being received from the RTN sub-network 65. Again, there are many different mathematical expressions that could be used in the rules governing the accumulating / depreciating of latent energy and how the value of the latent energy contributes to the output activity to implement this functionality. For example, the rules governing the accumulation / depreciation of latent energy and the contribution of32 1560.P002

[0198] latent energy to the output activity may cause an initially slow build-up of output activity which sharply increases as the excitatory input activity becomes larger than the inhibitory input activity and then rapidly falls.

[0199] A purpose of the RTN sub-network 65 is to supply the inhibitory activity to the THA sub-network 61 in the event that the activity in the THA sub-network 61 and the CTX sub-network 63 is low. Again, there are many different mathematical expressions that could be used in the rules governing the accumulating / depreciating of latent energy and how the value of the latent energy contributes to the output activity to implement this functionality. For example, the latent activity may increase in the presence of input excitatory activity but otherwise decay, and the contribution to the output activity may be dependent on the difference between the value of the latent activity and a reference value such that as the latent activity decays below the reference value the output inhibitory activity increases.

[0200] With respect to the learning rule, it is not necessary for all nodes of the reflexive ANN and the predictive ANN to apply the same learning rule. For example, one learning rule may be applied in the reflexive ANN and a different learning rule may be applied in the predictive ANN. Further, different learning rules could be applied in different sub-networks and / or different populations of nodes, or even between nodes in the same population of nodes. Further, different learning rules may be applied in stateful nodes to those applied in stateless nodes.

[0201] While a learning rule is described above which, in addition to the adjusting of weights associated with edges, may remove an edge and insert a replacement edge between a different pair of nodes, by increasing the interconnectedness of the nodes within the neural network, example implementations may not require the ability to remove and re-insert edges as described above. This also removes the requirement for a suitability value in the learning rule.

[0202] An example learning rule that does not require suitability is to determine the weight w for an edge interconnecting a source node and a destination node by increasing the weight by 1 *LR (where LR is a learning rate) if the source node activity is greater than twice the next destination node activity, and to decrease the weight by 1*LR if the next destination node activity is greater than one, with both an increase and a decrease being simultaneously possible. If the determined weight is33 1560.P002

[0203] greater than one, then the weight is set at one while if the determined weight is less than zero, the weight is set at zero.

[0204] One way of representing this learning rule is the following pseudo-code.

[0205] for t = 0:n

[0206] {

[0207] if (SRC(output) > 2 * NEU(next_output)) SYN(weight) += LR;

[0208] if (NEU(next_output) > 1) SYN(weight) -= LR;

[0209] if (SYN(weight) < 0) SYN(weight) = 0;

[0210] if (SYN(weight) > 1) SYN(weight) = 1;

[0211] }

[0212] where SRC(output) is the activity output by the source node of the edge in a first activity output interval, NEU(next_output) is the activity output by the destination node of the edge in the second activity output interval, SYN is the weight of an input edge or an output edge, in other words a synapse, of a node..

[0213] In this way, the adjusting of the weight is influenced by the contribution that the activity from the source node makes on the activity output by the destination node in the next time step. While such an arrangement does not require disconnection and re-insertion elsewhere of edges, this may be performed if, for example, the activity from the source node and the activity from the destination node both exceed one.

[0214] Modifications can also be made to the configuration of the first and second artificial neural network. For example, while input nodes are included in the artificial neural networks described above (as shown in Figure 3) to act as a buffer between the output of the pre-processing unit and the other nodes of the ANN, these input nodes are not required if the outputs of the pre-processor are suitably configured that they can be treated by nodes to which they are connected simply as inputs from an excitatory node. Further, while the artificial neural network shown in Figure 3 has edges interconnecting input nodes and output nodes, these are not necessary.

[0215] Although in the example systems each sensor data signal is compared to respective different set point data to generate an input signal for the ANN, alternative configurations in which sensor data received by the system controller generates one or34 1560.P002

[0216] more input signals for the ANN based on a comparison between the sensor data and associated set point data, which may determine the magnitude of the difference between sensor data and set point data, are possible. For example, a single sensor data signal could be pre-processed to derive three parameters which are each compared with respective set point data to generate input signals for the ANN.

[0217] Further, one or more sensor signals may be compared with set point data to generate a single input signal for the ANN. Alternatively, the sensor data from two or more sensor data signals could be pre-processed to determine a sensor data value which is compared with a corresponding set point data value. The set point data may comprise a time-series of set points in some embodiments.

[0218] Similarly, although in the example systems each output signal from the ANN is input into a respective different signal generator, with each signal generator generating a control signal for a respective different actuator, alternative configurations are possible in which one or more control signals are determined based on output from the ANN. For example, multiple outputs may be input to a single signal generator to generate a control signal for one actuator.

[0219] More particularly, in example implementations of the present invention there may not be a one-to-one correspondence between the number of sensors and the number of pre-processing functions, or a one-to-one correspondence between the number of sensors and the number of actuators. Further, the sensors may detect parameters associated with the actuators, or operations thereof. More generally, the system controller is able to determine control signals for any number of actuators based on any number of measured parameter values, with there being at least one actuator that affects multiple measured parameter values and at least one sensor providing a sensor signal conveying multiple measured parameter values. In an example implementation, an IMU may be positioned in the torso of a biped robot, and sensor signals from the IMU may be used to control all the actuators within the biped robot.

[0220] Many modifications to the way in which the system controller processes the sensor signals to generate control signals are possible. For example, the normalisation function of the pre-processor may apply a super-linearization function to reach a desired state of the ANN more quickly than if a purely linear function was applied. In35 1560.P002

[0221] a desired state of the ANN, the activity of the ANN is in equilibrium and results in a desired behavior of the robotic system. An example of such a super-linearization function is:

[0222] / w =f l - (! - «)• tfx > 0

[0223]

[0224] —1 + (1 + x)aotherwise

[0225] where the exponent may, for example, be equal to three.

[0226] In some examples, the control system and / or the task manager may be implemented on the physical system. Alternatively, the control system and / or the task manager could be implemented remotely from the physical system, e.g. on one or more cloud servers, or distributed between the cloud and the physical system. The control system and a task manager could implement edge computing.

[0227] While in the illustrated example the system controller utilises software routines, it will be appreciated that at least some of these software routines may alternatively be implemented by hardware, and that there may be performance benefits in so doing.

[0228] APPLICATIONS

[0229] By way of example only, various applications of a system controller as described above will now be described.

[0230] Robotic Systems

[0231] The disclosed technology may be used to control various robotic systems, such as robots with articulated limbs (e.g. biped or quadruped robots). Other robotic systems may comprise autonomous propulsion systems for movement which do not use articulated limbs. By way of example, only, the disclosed technology may be used for autonomously controlling and operating systems such as vehicles, for example, cars and heavy-duty vehicles, trains, aircraft, surface vessels or submersibles which do not use articulated limbs for movement, including unmanned airborne vehicles (e.g. drones) or unmanned underwater vehicles or parts thereof. Another example of a system or system component which may be controlled using the36 1560.P002

[0232] disclosed technology, includes a motor. A motor may be configured, for example, to control the actuation of a system component such as a valve or the like for flow regulation or to regulate the speed of a propulsion system.

[0233] For an example such as biped and quadruped robots, the articulated limbs include motors and sensors, and there have previously been successful attempts to train such robotic systems to walk. The previous attempts have, however, required extensive training of the robotic system, particularly as movement of one limb as a result of actuation of a motor may affect multiple sensed signals. In an application of the system controller described above to robots with articulated limbs, the sensors may detect positional information for different locations on the robotic system. This positional information may, for example, be the distance of each sensed location above the ground. The control signals may be applied to respective motors causing movement of the articulated limbs. By setting the set point data to correspond to positions for the sensed location at which the robotic system is in a standing state, the robotic system can in effect learn to stand based only on data from the sensors.

[0234] The controller systems and methods for controlling a system disclosed herein may configure the actuators individually as well as collectively to generated control signals for each joint or limb to adjust the positions of the limbs about a joint, in other words, the control signals will adjust the amount of rotation and / or linear movement in a lateral or vertical direction in the one or more degrees of freedom each joint can move within with the aim of achieving a target behavior of that joint of the joint system which results in that joint or the joint system of the robot as a whole achieving a target pose, position or velocity.

[0235] It will be appreciated however by those skilled in the art of robotics that the control principles disclosed may be applied to other types of robotic joint systems, for example, to quadruped robots and robotic arms which comprise articulated joints.

[0236] These sensors may detect rotary and / or lateral movement in a number of degrees of freedom, for example, one or more or all of the six degrees of freedom associated with x,y,z co-ordinates and degrees of pitch, yaw, and roll according to the permitted movements of the joint to which they are attached. In some embodiments, one or more or all of the joints of a robot may not be capable of moving in all degrees of freedom.37 1560.P002

[0237] For example, one or more IMU sensors may be used to provide inertial measurement unit (IMU) sensor pitch information about a y-axis, IMU sensor roll data about an x-axis, and IMU sensor yaw about a z-axis (assuming an x,y,z coordinate system is used which may be centred on the centre of mass of the robot in some embodiments).

[0238] Limits of limb motion may be configured by limiting joint rotation may in some embodiments of the controller. These limitations can be modelled by preprocessing received sensory data in some embodiments prior to the pre-processed sensory data being input to the controller.

[0239] A permitted range of joint movement about a set or target point need not always be symmetrical. For example, an elbow joint may be modelled to limit rotational movement to within a certain angle. To prevent joint damage, the angle of rotational movement may be limited in the real-world in a non-symmetrical way, with movement limited more in one direction than in another.

[0240] In some embodiments, received sensor data accordingly may need to be pre-processed to be zero-centred with a permitted movement range that is off-set and / or scaled for some limb movements. For example, in some embodiments, a zero-centred autoscaling is implemented in which joint movement parameters are bound to within given range endpoints for a given joint. This may allow scaling to the maximum value of a range, while keeping the position control symmetric in both directions. This may be in some embodiments at the expense of introducing a "dead-zone" on one side if the joint endpoints are asymmetric in the real-world.

[0241] Some embodiments of the disclosed technology use in addition a system of stratified sensors which allows extreme sensory inputs to be mapped to different behaviours while avoiding large areas of synaptic movement within the artificial network and / or within specific motor neuron populations and yet still retaining sensitivity to small movements.

[0242] The disclosed control systems and methods of controlling a system comprising a robotic joint system may be used to control robot movement using one or both of a position model, where the robot joints are controlled directly using sensors which detect position, force, and velocity, and a muscle model, where muscles act in pairs and are contracted and elongated to move limbs about a joint. In some embodiments,38 1560.P002

[0243] the controller is configured to receive sensory input data from a plurality of position, force, and velocity sensors S mounted on a robot, such as the sensors S located at the joints of the robot limb system of the robot. This results in, for example, sensor data signals such as those shown as SA, SB, and Sc in Figure 1 being received by the system controller comprising at least positional data indicating sensed positions of robotic joints, and in some embodiments in addition or instead, force and velocity data from the robotic joint system sensors.

[0244] A Vision System

[0245] A vision system may include a multi-pixel camera mounted on a motor-driven platform. Each pixel of the multi-pixel camera may input a sensor signal to the system controller, and the system controller may generate one or more control signals for the motor-driven platform. By using set point data that corresponds to a still image in the border areas of the camera images and clearing the set point during motor actuation, the system controller can cause the motor driven platform to centre moving objects within the field of view of the camera

[0246] Such a vision system may have utility on a vehicle such as an automobile or an airborne vehicle, e.g. a drone, either tracking movement of other vehicles or movement of the vehicle relative to stationary objects.

[0247] An Auditory System

[0248] A cochlear implant can be used to amplify audio signals to produce output signals that can be more readily perceived by a person with hearing difficulties. In an application, one sensor detects the time-varying audio signal in the environment and inputs the sensed audio data into the system controller, while another sensor detects the time-varying output signal from the cochlear implant and inputs the sensed output data into the system controller.

[0249] A pre-processing function in the system controller decomposes the sensed audio data into individual spectral components using, for example, Fast Fourier Transform (FFT) or the Short-Term Fourier Transform (STFT). Similarly, the preprocessing function decomposes the sensed output data into the same individual spectral components using, for example, Fast Fourier Transform (FFT) or the Short-39 1560.P002

[0250] Term Fourier Transform (STFT). The pre-processing function is then able to calculates a gain value for each individual spectral component based on the corresponding decomposed sensed signal data and output signal data, and then compare the gain value for each spectral component with corresponding set point data to generate the input signals for the artificial neural network. The output signals from the artificial neural network are then used to control the gain applied within the cochlear implant.

[0251] By setting the set point data in dependence on the hearing of an individual in a frequency dependent manner, high performance can be achieved across all frequencies for that individual.

[0252] Telecommunications Network

[0253] A telecommunications network, such as a wireless communications network, has multiple network parameters which may have an impact on performance. The performance of such a telecommunications network can be characterized by multiple parameters, including parameters associated with noise, for example bit error rate, data rates and data volumes. The exploration of the network parameters to achieve desired physical performance characteristics, which may for example be determined so as to satisfy one or more service level agreement, is challenging.

[0254] In an application, the physical performance characteristics of the telecommunications network are measured, and the resultant sensor data corresponding to the measurements is input to the system controller. The set point data in the system controller is determined based on desired physical performance characteristics, which may be dynamically updated as the desired physical performance characteristics change. The control signals are then used to modify the network parameters with the aim of achieving the desired physical performance characteristics.

[0255] For example, if the aim is meet desired service levels for one or more users of a wireless communications network by optimizing signal coverage and quality, this may be achieved by providing control signals to hardware actuators for one or more of the following:

[0256] adjusting the tilt angle of one or more antennas in some embodiments;40 1560.P002

[0257] adjusting the power levels of base station transmitters or user equipment to manage interference and ensure adequate signal strength;

[0258] adjusting the direction of an antenna beam, also known as beam forming, dynamically towards specific areas or users for better coverage and capacity;

[0259] adjusting small cells and relays to enhances network capacity and remove dead zones in high-density under-served areas;

[0260] dynamically reconfiguring distributed antenna systems for better load balancing and coverage; and

[0261] optimizing spatial streams for higher network throughput by adjusting the configuration of multiple input multiple output, MIMO, antennas.

[0262] For example, to optimize signal coverage and quality in a physical network, the disclosed technology may be used to provide control signals to software actuators configured to actuate physical components, for example, one or more of the following:

[0263] a software actuator for a self-organization network which automatically optimizes network coverage, capacity, and reduces interference by dynamically reconfiguring the network responsive to receiving control signals according to the disclosed technology;

[0264] a software actuator for adjusting thresholds for handover decisions and / or adjusts other handover optimization parameters responsive to receiving control signals according to the disclosed technology to minimize cell drops and improve user experience;

[0265] a software actuator for dynamic spectrum allocation, DSA, which causes frequency bands to be adjusted responsive to receiving control signals according to the disclosed technology based on real-time traffic demand and interference conditions;

[0266] a software actuator for carrier aggregation which combines multiple frequency bands to increase user throughput and overall network capacity responsive to receiving control signals according to the disclosed technology;

[0267] a software actuator for network slicing responsive to receiving control signals according to the disclosed technology to create virtual networks optimized for specific use cases, for example, loT, video streaming, etc.41 1560.P002

[0268] a software actuator for traffic offloading responsive to receiving control signals according to the disclosed technology to redirect cellular traffic to Wi-Fi or other networks to reduce congestion on that cellular network;

[0269] a software actuator for implementing load balancing responsive to receiving control signals according to the disclosed technology to redistribute traffic across cells or frequency layers to avoid overloading; and

[0270] a software actuator for QoS (quality of service) tuning responsive to receiving control signals according to the disclosed technology which prioritizes certain types of traffic, e.g. Video streaming, VoIP, based on service agreements.

[0271] Another example of a system or system component which may be controlled using the disclosed technology comprises a network traffic management optimiser configured to maximise available uplink connectivity in a wireless. For example, the system may be used to configure multiple network resources to optimise one or more of: network capacity, network connectivity, spectrum allocation, and the like.

[0272] Other examples of use cases of the disclosed technology in a networking context include but are not limited to: predictive analytics, e.g. for faults or congestion for proactive or pre-emptive action, e.g. load-balancing, proactive or pre-emptive interference management, energy efficiency, resource scheduling and allocation, content caching, anomaly detection and / or correction management.

[0273] The disclosed technology may accordingly be used in conjunction with a variety of different sensors, including but not limited to sensors configured to sense physical properties, for example: temperature, humidity, proximity, ultrasound, light including one or more or all of ambient visible light, infra-red light, ultra-violet light, pressure, acceleration, colour, touch, level, position, hall effect, tilt, vibration, gas, chemical(s), vibration.

[0274] Such physical properties may be sensed using sensor systems comprising one or more of the following types of sensors, which is not intended to be a complete list: optical sensors, image sensors, temperature sensors, depth imaging sensors, event imaging sensors, gyroscopic sensors, position sensors, speed sensors, accelerometers, chemical sensors, pressure sensors, electromagnetic field sensors, magnetic field sensors, spectral sensors, electrical current or voltage sensors.42 1560.P002

[0275] Actuators may comprise pneumatic actuators, hydraulic actuators, electric actuators, linear actuators, rotary actuators, piezoelectric actuators, magnetic actuators, mechanical actuators, electric motors, solenoids, thermal actuators e.g. heaters or heat-sinks, valves, diaphragm actuators, stepper motors etc.

[0276] In some examples, a robotic limb actuator comprise a plurality of different types of actuators, for example, a combination of linear actuators and rotary actuators.

[0277] In the above description, the term state may represent a static state, for example a bipedal or quadrupedal robot maintaining a stationary pose such as standing up, or a dynamic state, for example a bipedal or quadrupedal robot sidestepping, swaying onto one side, jogging, running and jumping. Some general examples of target behavior accordingly may include maintaining a stable position or pose when the robotic is performing a task which requires limb movement and / or limb movement subject to external forces such occur when a robot lifts up an object.

[0278] The term configuration may be used in place of the word state to the extent that the term configuration encompasses a dynamic configuration as well as a static configuration.

[0279] The above embodiments are to be understood as illustrative examples of the invention. Further embodiments of the invention are envisaged. For example, [add possibilities]. It is to be understood that any feature described in relation to any one embodiment may be used alone, or in combination with other features described, and may also be used in combination with one or more features of any other of the embodiments, or any combination of any other of the embodiments. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the accompanying claims.

Claims

43 1560.P002CLAIMS1. A method of generating a control signal for a physical system in dependence upon sensor data for operational parameters of the physical system using a control system comprising a first artificial neural network and a second artificial neural network, wherein each of the first artificial neural network and the second artificial neural network has a plurality of nodes and a plurality of edges that interconnect nodes to propagate activity between nodes, each edge having a corresponding weight, and wherein the sensor data is associated with setpoint data corresponding to a target performance for the physical system, the method comprising:receiving, by the control system, the sensor data;determining, by the control system, first input to the first artificial neural network based on a difference between the sensor data and the setpoint data;providing, by the control system, input to the second artificial neural network corresponding to the first input to the first artificial neural network;processing, by the second artificial neural network, the input to the second artificial neural network to generate output;providing, by the control system, second input to the first artificial neural network corresponding to the output from the second artificial neural network;generating, by the first artificial neural network, output in dependence on the first input and the second input to the first artificial neural network;using, by the control system, the output of the first neural network to generate the control signal for the physical system; andadjusting weights associated with edges of the first artificial neural network and the second artificial neural network by applying a local learning rule at each of the plurality of nodes of the first artificial neural network and each of the plurality of nodes of the second neural network such that the control signal from the first artificial neural network reduces a departure of the sensor data from the setpoint data, wherein at least some of the plurality of nodes of the second artificial neural network are stateful nodes with each stateful node being configured to output activity having a first contribution dependent on input activity from other nodes and a second44 1560.P002contribution dependent on a latent activity for that stateful node, the latent activity being dependent upon previous input activity received by that stateful node from the other nodes,wherein for at least some of the stateful nodes, the second contribution inhibits activity by an amount dependent on the latent activity, and wherein the latent activity is accumulated and depreciated responsive to received input activity such that activity relating to an initial portion of a departure of the sensor data from the setpoint data propagates more strongly through the second artificial neural network than activity relating to later portions of the departure of the sensor data from the setpoint data, and wherein as activity propagates through the first artificial neural network and the second artificial neural network, the local learning rule for each node varies weights associated with incoming edges in dependence on activity input to that node so as to contribute to the control signal reducing the magnitude of the first input to the first artificial neural network.

2. The method of claim 1, wherein the local learning rule for a node is configured to increase the weight associated with an incoming edge to that node in dependence on the contribution of the activity received via that incoming edge to the output activity from the node and / or to decrease the weight associated with the incoming edge in dependence upon the magnitude of the output activity such that the weights for incoming edges that propagate activity which contributes stably to the control signal reducing any departure of the sensor data from the setpoint data are potentiated.

3. The method of claim 1 or claim 2, wherein the second artificial neural network comprises a first sub-network of nodes and a second sub-network of nodes, the second sub-network having a greater number of nodes than the first sub-network and having the stateful nodes for which the second contribution inhibits the output activity by an amount dependent on the latent activity,wherein the input to the second artificial neural network forms a first input to the first sub-network, and wherein a first output of the second sub-network forms a45 1560.P002second input to the first sub-network, whereby the first sub-network generates output to the second sub-network dependent on the input to the second artificial neural network and the first output of the second sub-network, andwherein the second sub-network generates the first output and a second output in dependence on the output from the first sub-network, the first output enabling activity to cycle through the first sub-network and the second sub-network and the second output being the output from the second artificial neural network to the first artificial neural network.

4. The method of claim 3, wherein the edges associated with the nodes of the first sub-network are configured to provide inputs to the second sub-network that correspond to different combinations of inputs to the first sub-network.

5. The method of claim 3 or claim 4, wherein the stateful nodes of the second sub-network comprise stateful excitatory nodes and the second sub-network further comprises stateless inhibitory nodes.

6. The method according to claim 5, wherein the second contribution to the output activity for the stateful excitatory nodes corresponds to a reduction by a proportion of the latent activity.

7. The method according to claim 5 or claim 6, wherein the latent activity for a stateful excitatory node accumulates responsive to an initial increase in received activity and then depreciates responsive to a subsequent increase in inhibitory activity in the second sub-network.

8. The method according to claim 7, wherein the latent activity for a stateful excitatory node further depreciates responsive to a subsequent increase in received activity following the initial increase.

9. The method of claim 7, wherein the first contribution corresponds to an expression of the form:46 1560.P002(aE-bl) - cE*I,where E is the total excitatory activity, I is the total inhibitory activity and a, b and c are constants.

10. The method of any of claims 1 to 9, wherein the activity signals output from the nodes of the first artificial neural network and the second artificial neural network incorporate a temporal delay that can vary between nodes.

11. The method of any of claims 3 to 10, wherein the first sub-network comprises stateful excitatory nodes and the artificial neural network further comprises a third sub-network of nodes comprising stateful inhibitory nodes,wherein outbound edges from excitatory nodes of the first sub-network and the second sub-network are connected to stateful inhibitory nodes of the third subnetwork,wherein outbound edges from inhibitory nodes of the third sub-network are connected to excitatory nodes of the first sub-network,wherein the stateful inhibitory nodes of the third sub-network output activity to the stateful excitatory nodes of the first sub-network in the absence of input activity from the excitatory nodes of the first sub-network and the second sub-network, and in response to receiving inhibitory activity from the third sub-network, the stateful excitatory nodes of the first sub-network accumulate and discharge latent activity so that excitatory activity is propagated to the second sub-network and the third subnetwork.

12. The method of any preceding claim, wherein the second contribution to the output activity for each stateful node in the first sub-network increases as the magnitude of the latent activity increases.

13. The method of any of the preceding claims, wherein the topology of the sensory distribution in the physical system influences distribution of edges within at least one of the first artificial neural network and the second artificial neural network.47 1560.P00214. A control system comprising at least one processor and memory storing instructions that, when implemented by the at least one processor, perform a method as claimed in any preceding claim.

15. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any of claims 1 to 13.

16. A system controller comprising:at least one input device to receive sensor data associated with a physical system;a pre-processor configured to determine one or more input signals based on the sensor data and setpoint data corresponding to a target performance for the physical system;a first artificial neural network having a plurality of nodes interconnected by a plurality of edges to propagate activity between nodes;a second artificial neural network having a plurality of nodes interconnected by a plurality of edges to propagate activity between nodes, wherein at least some of the plurality of nodes of the second artificial neural network are stateful nodes with each stateful node being configured to output activity having a first contribution dependent on input activity from other nodes and a second contribution dependent on a latent activity for that stateful node, the latent activity being dependent upon previous input activity received by that stateful node from the other nodes; and at least one output device configured to output a control signal for controlling at least one actuator of the physical system,wherein the system controller is configured to provide input signals from the pre-processor to the second artificial neural network, which processes the input signals to generate output,wherein the system controller is configured to provide a first input corresponding to input signals from the pre-processor and a second input corresponding to output from the second artificial neural network to the first artificial48 1560.P002neural network, and the first artificial neural network is configured to process the first input and the second input to generate one or more output signals,wherein the system controller is configured to determine the control signals in accordance with the one of more output signals from the first artificial network and to provide the control signal to the at least one output device,wherein for at least some of the stateful nodes of the second artificial neural network, the second contribution inhibits activity by an amount dependent on the latent activity, and wherein the latent activity is accumulated and depreciated responsive to received input activity such that activity relating to an initial portion of a departure of the sensor data from the setpoint data propagates more strongly through the second artificial neural network than activity relating to later portions of the departure of the sensor data from the setpoint data,and wherein the nodes of the first artificial neural network and the second artificial neural network are configured to adjust the weights corresponding to incoming edges to the node using a local learning rule so that as activity propagates through the first artificial neural network and the second artificial neural network, the local learning rule for each node varies weights associated with incoming edges in dependence on activity input to that node so as to contribute to the control signal reducing the magnitude of the first input to the first artificial neural network.

17. The system controller of claim 16, wherein the local learning rule for a node is configured to increase the weight associated with an incoming edge to that node in dependence on the contribution of the activity received via that incoming edge to the output activity from the node and / or to decrease the weight associated with the incoming edge in dependence upon the magnitude of the output activity such that the weights for incoming edges that propagate activity which contributes stably to the control signal reducing any departure of the sensor data from the setpoint data are potentiated.

18. The system controller of claim 16 or claim 17, wherein the second artificial neural network comprises a first sub-network of nodes and a second subnetwork of nodes, the second sub-network having a greater number of nodes than the49 1560.P002first sub-network and having the stateful nodes for which the second contribution inhibits the output activity by an amount dependent on the latent activity, wherein the input to the second artificial neural network forms a first input to the first sub-network, and wherein a first output of the second sub-network forms a second input to the first sub-network, whereby the first sub-network generates output to the second sub-network dependent on the input to the second artificial neural network and the first output of the second sub-network, andwherein the second sub-network generates the first output and a second output in dependence on the output from the first sub-network, the first output enabling activity to cycle through the first sub-network and the second sub-network and the second output being the output from the second artificial neural network to the first artificial neural network.

19. The system controller of claim 18, wherein the edges associated with the nodes of the first sub-network are configured to provide inputs to the second subnetwork that correspond to different combinations of inputs to the first sub-network.

20. The system controller of claim 18 or claim 19, wherein the stateful nodes of the second sub-network comprise stateful excitatory nodes and the second sub-network further comprises stateless inhibitory nodes.

21. The system controller according to claim 20, wherein the second contribution to the output activity for the stateful excitatory nodes corresponds to a reduction by a proportion of the latent activity.

22. The system controller according to claim 19 or claim 20, wherein the latent activity for a stateful excitatory node accumulates responsive to an initial increase in received activity and then depreciates responsive to a subsequent increase in inhibitory activity in the second sub-network.50 1560.P00223. The system controller according to claim 22, wherein the latent activity for a stateful excitatory node further depreciates responsive to a subsequent increase in received activity following the initial increase.

24. The system controller of claim 22, wherein the first contribution corresponds to an expression of the form:(aE-bl) - cE*I,where E is the total excitatory activity, I is the total inhibitory activity and a, b and c are constants.

25. The system controller of any of claims 16 to 24, wherein the activity signals output from the nodes of the first artificial neural network and the second artificial neural network incorporate a temporal delay that can vary between nodes.

26. The system controller of any of claims 18 to 25, wherein the first subnetwork comprises stateful excitatory n odes and the artificial neural network further comprises a third sub-network of nodes comprising stateful inhibitory nodes, wherein outbound edges from excitatory nodes of the first sub-network and the second sub-network are connected to stateful inhibitory nodes of the third subnetwork,wherein outbound edges from inhibitory nodes of the third sub-network are connected to excitatory nodes of the first sub-network,wherein the stateful inhibitory nodes of the third sub-network output activity to the stateful excitatory nodes of the first sub-network in the absence of input activity from the excitatory nodes of the first sub-network and the second sub-network, and in response to receiving inhibitory activity from the third sub-network, the stateful excitatory nodes of the first sub-network accumulate and discharge latent activity so that excitatory activity is propagated to the second sub-network and the third subnetwork.51 1560.P00227. The system controller of any of claims 16 to 26, wherein the second contribution to the output activity for each stateful node in the first sub-network increases as the magnitude of the latent activity increases.

28. The system controller of any of claims 16 to 27, wherein the topology of the sensory distribution in the physical system influences distribution of edges within at least one of the first artificial neural network and the second artificial neural network.

29. A robotic system comprising a system controller according to any of claims 15 to 27.