System controller for a robotic system

WO2026162594A1PCT designated stage Publication Date: 2026-08-06INTUICELL AB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
INTUICELL AB
Filing Date
2026-01-28
Publication Date
2026-08-06

Smart Images

  • Figure EP2026052221_06082026_PF_FP_ABST
    Figure EP2026052221_06082026_PF_FP_ABST
Patent Text Reader

Abstract

There is provided a method of generating a control signal for a robotic system in dependence on sensor data associated with a current state of the robotic system. The method comprises receiving, by a control system including an artificial neural network having a plurality of nodes interconnected by a plurality of edges, the sensor data associated with the current state of the robotic system. The control signal determines one or more input signals for the artificial neural network dependent on the sensor data and set point data corresponding to a target state or behavior of the robotic system, and inputs the one or more input signals to the artificial neural network, which generates one or more output signals dependent upon the input signals. The control signals then determines a control signal based on the one or more output signals from the artificial neural network, and inputs the control signal to an actuator system to cause a modification of the current state of the robotic system. The artificial neural network adjusts weights associated with edges of the artificial neural network in accordance with a local learning rule to reduce a difference between the sensor data and the corresponding set point data. By using an artificial neural network, complex interrelationships between the sensor readings and the actuator operations can be taken into account when determining how to achieve a target state or behavior.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SYSTEM CONTROLLER FOR A ROBOTIC SYSTEM

[0002] Technical Field

[0003] This disclosure relates to a system controller for a robotic system, a robotic system including the system controller, and a method of operation of the system controller. In particular, but not exclusively, the system controller may control one or more joints of a robot, for example a biped or quadruped robot, to maintain a stable pose and / or a position, such as a balanced pose in a standing or walking configuration.

[0004] Background

[0005] Robotic systems such as robotic arms, biped robots and quadruped robots are known. A robotic system generally has actuators that change the configuration of the robotic system in accordance with a control signal from a system controller to achieve a desired behavior. Feedback control loops, or closed loop control, are commonly present within the robotic system. In a simple example of a feedback control loop, a parameter of the robotic system is measured, and a control signal for an actuator is determined based on the difference between the measured value for the parameter and a target value so as to vary the parameter such that the measured value approaches the target value.

[0006] One type of system controller is the PID (Proportional Integral Derivative) controller, which generates a control signal based on a difference between the measured value for the parameter and the target value (a proportional component), a cumulative difference over time between the measured value for the parameter and the target value (an integrative component), and a rate of change of the difference between the measured value for the parameter and the target value (a derivative component). The integrative component improves the rate at which the measured value approaches the target value in comparison with a purely proportional feedback control system, while the derivative component seeks to reduce any overshoot from the target value.

[0007] A problem with incorporating simple feedback control loops, where the control signal for a single actuator is determined based on the measured value received from a single sensor, in robotic systems is that the value of a measured parameter may be affected by more than one actuator and / or operation of one parameter may affect themeasured values from multiple sensors. Such a scenario can be encountered, for example, in articulated systems, for example in a biped or quadruped robot.

[0008] The use of machine learning techniques in robotic systems is also known. These machine learning techniques typically require the use of large training datasets and backpropagation. Even after extensive training, the robotic system is typically only able to function correctly in the same contexts or environments as the training took place, and is not able to adapt to new contexts.

[0009] This disclosure discusses a novel machine learning technique that can be implemented in a system controller to control robotic systems. Some of the disclosed technology relates to, for example, controlling one or more articulated limbs of a robotic limb system comprising a robot, such as a biped or quadruped robot, to maintain one or more of a stable pose, position or velocity. The disclosed controller may be used, for example, to attain, maintain, and regain if perturbed, a target behavior such as a target state comprising a balanced static pose when standing or a balanced dynamic pose when moving. Movement may be movement of the limb itself or the robot body it is attached to. For example, a biped robot may need to maintain a pose to carry an item when walking or other types of locomotion such as side-stepping, swaying onto one side, jogging, running and jumping. Some general examples of target behavior accordingly may include maintaining a stable position or pose when the robotic is performing a task which requires limb movement and / or limb movement subject to external forces such occur when a robot lifts up an object.

[0010] Summary

[0011] According to a first aspect of the invention, there is provided a method of generating a control signal for a robotic system in dependence on sensor data associated with a current state of the robotic system. The method comprises receiving, by a control system including an artificial neural network having a plurality of nodes interconnected by a plurality of edges, the sensor data associated with the current state of the robotic system. The control signal determines one or more input signals for the artificial neural network dependent on the sensor data and set point data corresponding to a target state or behavior of the robotic system, and inputs the one or more input signals to the artificial neural network, which generates one or more output signals dependent uponthe input signals. The control signals then determines a control signal based on the one or more output signals from the artificial neural network, and inputs the control signal to an actuator system to cause a modification of the current state of the robotic system. The artificial neural network adjusts weights associated with edges of the artificial neural network in accordance with a local learning rule to reduce a difference between the sensor data and the corresponding set point data.

[0012] By using an artificial neural network, complex interrelationships between the sensor readings and the actuator operations can be taken into account when determining how to achieve a target state or behavior. In addition, by using a local learning rule to adjust the weights allows continual learning without relying on large training datasets and backpropagation, but rather through continuous interaction with the environment. This continual learning ability allows the artificial neural network to correct and adapt to unmodelled dynamics in the real world. The continual learning ability also allows correction for changes within the robotic system itself, for example sensor drift, which allows for increased deployment times of the robotic system.

[0013] According to a second aspect, there is provided a method of generating a control signal for an articulated system having two or more sections interconnected by one or more joints, the articulated system further comprising an actuator system having an actuator for causing relative movement of at least two sections of the articulated system and a feedback controller that provides an actuation signal to the actuator in dependence on a measurement signal associated with the current state of the articulated system and a target signal. The method includes receiving, by a control system comprising an artificial neural network having a plurality of nodes interconnected by a plurality of edges, sensor data associated with a current state of the articulated system, and determining, by the control system, input signals for the artificial neural network dependent on the sensor data and set point data corresponding to a target state of the articulated system. The determined input signals are input to the artificial neural network, which generates one or more output signals. The control system then determines the control signal based on the one or more output signals. The control signal is input to the actuator system, where the actuator system determines the target signal based on the control signal. Weights for a plurality of input edges for which the node receives activity output by other nodes of the artificial neural network are adjustedto reduce deviation between the sensor data and the corresponding set point data. The weights may be adjusted using a local learning rule.

[0014] In this way, the weights can be adjusted without relying on large training datasets and backpropagation, but rather may be adjusted based on continuous interaction with the environment. This continual learning ability allows the artificial neural network to correct and adapt to unmodelled dynamics in the real world.

[0015] In some example implementations, by adjusting the weights for only a subset of the input edges to a node, the artificial network can more efficiently arrive at a set of weights for which the sensor data signals stably match their target behavior. In an implementation, the subset of the plurality of edges is selected in dependence on the magnitude of the activity received from each of the other nodes. For example, the subset of edges may correspond to the edges via which the node receives the highest activity output by other nodes. The selection of the subset of the plurality of edges in dependence on the magnitude of the activity received from each of the other nodes may provide a form of prioritisation for resolving one or more latent, in other words hidden, problems represented in the sensor data.

[0016] The method may further comprise a preliminary step of determining, by a task manager, the set point data and inputting the set point data to the control system. The task manager may determine the set point data using reinforcement learning. The task manager may update the set point data (representative of a new target state or behaviour), for example in response to the input of a new task. By incorporating the control system as middleware between the task manager and the feedback controllers of the actuating systems, the task manager need only address high level signalling, while the control system can handle low level signalling which may be dependent on the specifics of the robotic system. In this way, the implementation of complex robotic systems, which may involve many sensors and many actuators with complex interrelationships between the operations of the actuators and the resultant changes in sensor readings, may be facilitated.

[0017] Further, if the specifics of the robotic system vary, the continual learning functionality of the control system enables the control system to adapt to resultant changes to the specifics of the robotic system or the environment. The operation of the task manager is not necessarily affected. There are many possible causes of varianceof the specifics of a robotic system, for example mechanical wear and actuator drift (response, compliance joint friction), gradual sensor drift and calibration changes, slow shifts in system dynamics over many cycles, environmental variability, for example variable terrain conditions, external perturbations arising from object contact, human-in-the-loop operator interactions, environmental factors such as wind and temperature (for example icy conditions), and growing mismatch between models and physical hardware in unstructured settings. In effect, the control system provides a mapping between desired actions and the required actuator signals.

[0018] In this way, the control system addresses a behavioral, for example a performance, issue with current robotic systems. In particular, many current robotic systems may perform well in controlled demonstrations or short-time tests, but performance deteriorates over longer time deployments and outside of tightly controlled conditions. The control system allows for a prolonged usage without requiring some form of servicing, for example the control system may reduce operational down-time, may avoid needing to retrain a task manager, and if retraining is needed, reduce the time duration required to retrain the task manager using any state of the art Al system, for example, compared to the retraining that a reinforcement Al model may need.

[0019] According to another aspect of the invention, there is provided a recurrent Al system (for example as described herein) configured as a middleware between a task manager and one or more actuator systems in an articulated system. The recurrent Al system is configured to adjust in real time to high level signals from the task manager indicating a target state for the articulated system into control signals for the one or more actuator systems.

[0020] In some embodiments of the disclosed aspects, the control system may receive a task from the task manager defining a trajectory, causing the control system to send control signals to cause the articulated system to sequentially move through a plurality of target states. For example, such a task may correspond to a walking action for a bipedal robot.

[0021] It will be appreciated that the aspects may relate to a negative feedback controller. For a negative feedback controller, it is detrimental to introduce positive feedback into the environment. This disclosure provides techniques for removing positive feedback loops inside a recurrently connected ANN, and positive feedbackloops interacting with the environment through the sensors and actuators of a recurrent ANN, by altering that recurrent ANN, allowing the recurrent ANN to effectively control an environment through sensors and actuators with negative feedback control.

[0022] A node may process a first set of inputs from excitatory nodes, which increase activity in the node, and a second set of inputs from inhibitory nodes, which reduce activity in the node. For each input, the node selects a larger of the value of the input and a previous value of that input modified by a decay function, and multiplies the selected value by a weight for the corresponding edge to generate a weighted input. The node then generates a first summation of the weighted inputs corresponding to the first set of inputs and a second summation of the weighted inputs corresponding to the second set of inputs, and calculates an output in dependence upon the first summation and the second summation. By selecting a decayed value for a previous value of the input when that decayed value is greater than the current value of the input, performance deterioration caused by high frequency artefacts which may arise in the artificial neural network is ameliorated.

[0023] In an implementation, each output edge for a node is associated with a latency which determines the timing at which the output of the node is propagated along that edge to a different node. The latencies may be randomly assigned to edges. Introducing latencies in this way has the effect of introducing non-linearity into the artificial neural network, which assists in reaching a robust solution set of weights.

[0024] According to a third aspect, there is provided a computer-implemented method of generating a control signal for a robotic system in dependence upon sensor data associated with the robotic system. The sensor data is associated with set point data corresponding to a target behavior or target state for the robotic system. A control system, comprising an artificial neural network having a plurality of nodes interconnected by a plurality of directed edges, receives the sensor data, determines one or more input signals for the artificial neural network and inputs the one or more input signals to the artificial neural network. The artificial neural network generates one or more output signals and the control system determines the control signal based on the one or more output signals. The method further comprises removing an input edge to a recipient node and inserting a directed edge elsewhere in the artificial neural network in dependence upon determining that the activity received via the input edge is high andthe suitability of the recipient node to reduce a difference between the sensor data and the corresponding set point data is low.

[0025] The above aspects above provide feedback loop apparatus that has wide applicability.

[0026] According to another aspect, there is provided a computer-implemented method for generating a control signal for a robotic joint system of a robot in dependence upon sensor data associated with one or more robotic joints forming the robotic joint system, wherein the sensor data is associated with set point data corresponding to a target behavior comprising one or more or all of a target pose, a target position, or a target velocity of for the robotic joint system. The method comprises receiving, by a control system comprising an artificial neural network comprising a plurality of neuron populations, at least some of the neuron populations being associated with a degree of freedom of a joint of the robotic joint system, each population having a plurality of nodes interconnected by a plurality of edges, the sensor data. The control system determines one or more input signals for the artificial neural network and inputs the one or more input signals to the artificial neural network, which generates one or more output signals. The control system then determines a control signal based on the one or more output signals from the artificial neural network. The artificial neural network adjusts weights associated with edges of the artificial neural network to reduce a difference between the sensor data and the corresponding set point data, wherein the adjusting of weights comprises a node using a local learning rule to adjust the weights for a subset of the plurality of edges for which the node receives activity output by other nodes.

[0027] In some embodiments, at least one joint of the articulated system is actuatable with one or more degrees of freedom by one or more actuators responsive to said one or more actuators receiving the control signal.

[0028] In some embodiments, at least one joint of the articulated system is not actuatable by an actuator using a motor component in at least one degree of freedom.

[0029] In some embodiments, local learning rules are applied by nodes to edges within a neuron population differ from the local learning rules applied by nodes to edges between neuron populations.In some embodiments, the control system configures the actuators to control movement of the joints individually to attain the target behavior.

[0030] In some embodiments, the control system configures the actuators to control the movement of the joints collectively to maintain the target behavior for a minimum duration of time.

[0031] In some embodiments, the control system configures the actuators to control the movement of the joints of the articulated system collectively to regain the target behavior responsive to one or more perturbations in the target behavior caused by a perturbing force acting on the articulated system.

[0032] In some embodiments, regaining the target behavior results in a stable or balanced pose of a robot.

[0033] In some embodiments, at least one joint of the articulated system comprises a one or more control points and wherein each force applied by an actuator to the joint to regain the target pose is correlated to the perturbing force or a component of the perturbing force.

[0034] In some embodiments, the correlation is linear.

[0035] In some embodiments, the number of control points enables a superlinear response to the perturbing force.

[0036] In some embodiments, the articulated system is controlled using a joint position model and the controller provides differential sensor signal prioritisation for position control of individual joints.

[0037] In some embodiments, the articulated system is controlled using a muscle model and the controller provides differential sensor signal prioritisation for muscle control of individual joints.

[0038] In some embodiments, the robot is one of a robotic arm; a biped robot; and a quadruped robot.

[0039] Various implementations, given by way of example only, will now be described with reference to the accompanying drawings.

[0040] Brief Description of the Drawings

[0041] Figure 1 is a block diagram schematically showing the main components of a feedback control loop for a system;Figure 2 is a block diagram schematically showing functional components of a system controller forming part of the feedback control loop of Figure 1;

[0042] Figure 3 is a block diagram showing functional components of an artificial neural network forming part of the system controller of Figure 2;

[0043] Figure 4 is a block diagram showing functional components a node forming part of the artificial neural network of Figure 3;

[0044] Figure 5 is a flow chart schematically showing operations performed by the system controller forming part of the control loop of Figure 1;

[0045] Figure 6 is a block diagram schematically showing the main physical components of the system controller of Figure 1;

[0046] Figure 7A is a schematic diagram of a biped robot in a standing pose;

[0047] Figure 7B is a schematic diagram of a biped robot in a different pose to that of Figure 7A;

[0048] Figure 7C is a schematic diagram illustrating limited degrees of freedom of movement of the torso of the biped robot of Figures 7A and 7CB

[0049] Figure 8 is an example control system for implementing an actuated muscle sensor system according to some embodiments of the disclosed technology;

[0050] Figure 9A is an example of a control system model for controlling a joint system of a biped robot according to some embodiments of the disclosed technology; and Figure 9B is a schematic diagram showing an example of a robotic joint system model in the control system model of Figure 9 A;

[0051] Figure 10 is a block diagram schematically showing the main components of a further feedback control loop for a system; and

[0052] Figure 11 is a block diagram schematically showing the main components of an alternative feedback control loop to the feedback control loop illustrated in Figure 1.

[0053] Detailed Description

[0054] System Overview

[0055] Figure 1 schematically shows the main components of a robotic system 1, including a system controller 5 and an articulated system 9. The system controller 5 and the articulated system 9 are shown separately in Figure 1 for ease of illustration,but it will be appreciated that the system controller 5 and the articulated system 9 may both form part of the same robotic apparatus, such as a biped robot or a quadruped robot.

[0056] As shown schematically in Figure 1, the articulated system 9 includes multiple sections interconnected by joints. Although only three sections and two joints are shown in Figure 1, it will be appreciated that the articulated system may involve a complex arrangement of sections and joints. In some other examples, the articulated system 9 may comprise two rigid components that are connected by a joint, each of the components not having any additional joint.

[0057] The joints enable the sections to move relative to each other in order to change a state of the articulated system 9. In this example, multiple sensors 3a-3c (hereafter collectively referred to as sensors 3) within the articulated system 9 measure parameters of the articulated system 9 and the surrounding environment and respectively generate sensor data signals SA, SB, SC conveying values for the measured parameters. While, for ease of illustration, Figure 1 shows three sensors 3, generally there may be any number of sensors 3 and more specifically there may be a plurality of sensors 3. All the sensors 3 may have the same modality, or alternatively the sensors 3 may include sensors having different modalities. For example, one or more of the sensors 3 may be position sensors, e.g. for detecting the positions of the joints, while others of the sensors 3 may be photosensors, for example for detecting the environment around the robotic system. Others of the sensors 3 may be pressure sensors, and still others of the sensors 3 may be temperature sensors.

[0058] The articulated system 9 also includes multiple actuator systems 7a-7c (hereafter collectively referred to as actuators 7). All the actuators of the actuator systems 7 may have the same modality, or alternatively the actuator systems 7 may include actuator systems with actuators having different modalities. For example, one or more of the actuators of the actuator systems 7 may be clamps, while others of the actuators may be motors. At least some of the actuators of the actuator systems 7 interact with sections and joints of the articulated system 9, and this interaction modifies the current state sensed by the sensors 3.

[0059] As shown in Figure 1, in this example the sensor data signals SA, SB, SC are respectively input to different ones of the actuator systems 7. In each actuator system7, the measured value conveyed by the respective sensor data signal is input to a feedback controller (not shown) which compares the measured value to a corresponding target value and supplies a drive signal to the corresponding actuator in dependence upon the difference between the measured value and the target value. The actuators may involve one or both of linear and rotary actuators. The actuators may control joints or flexible elements without joints such as flexible elements configured as grippers.

[0060] The sensor data signals SA, SB, SC are also supplied to the system controller 5, which generates control signals CA, CB and Cc that are applied to respective different ones of the actuator systems 7. While for ease of illustration Figure 1 shows three control signals CA, CB and Cc and three actuator systems 7, generally there may be any number of actuator systems 7.

[0061] In this example, each actuator system 7 processes the respective control signal to determine the target value that is input to the feedback controller of that control signal for comparison with the measure value conveyed by the respective sensor signal. As will be described in more detail hereafter, the system controller 5 is able to determine a control signal for an actuator system 7 based on the measured values conveyed by multiple sensor signals. In this way, the operations performed by the actuators of the actuator systems 7 can co-operate with each other to achieve an object for the robotic system or part of the robotic system, rather than simply aiming to achieve an object for that particular actuator system 7.

[0062] While only an articulated system 9 is shown in Figure 1, the system controller 5 may also be used to control other sub-systems of the robotic system 1 comprising mechanisms such as actuators for limb systems, propulsion systems including those based on wheels, pose and positioning mechanisms.

[0063] As shown in Figure 1, the system controller 5 includes a pre-processor 11, which compares each of the sensor data signals SA, SB, SC with respective set point data 13 to generate input signals for an artificial neural network (ANN) 15. The set point data 13 represents a target state or behavior of the robotic system, and the ANN 15 processes the input signals to generate output signals which, when supplied to a control signal generator 17, cause the control signal generator 17 to generate the control signals CA, CB and Cc in such a way that the interaction between the actuators 7 and the actuated system components 9 may result in the measured parameter values conveyed by thesensor data signals SA, SB, SC being modified to be closer to the corresponding set point data 13, thereby causing the robotic system to achieve its target state or behavior. In particular, as will be described in more detail hereafter, each node within the ANN 15 utilises a local learning rule to modify weights associated with input edges to that node to reduce the activity for that node until a stable solution is reached where the control signals CA, CB and Cc have caused the measured values conveyed by the sensor data signals SA, SB, SC to approach their respective data values of the set point data 13. Such a stable solution may not involve all the measured values conveyed by the sensor data signals SA, SB, SC actually matching their respective set point data 13 as the interrelationships between the sensor data signals SA, SB, S and the control signals CA, CB and Cc may not permit this as a stable solution.

[0064] By using the ANN 15, the system controller 5 can handle feedback control for systems where there is a complex interaction between the actuators 7 and the parameters sensed by the sensors 3. For example, one of the actuators 7 may impact multiple sensed parameters or multiple actuators 7 may impact a single sensed parameter.

[0065] The System Controller - Function

[0066] In this example, as shown in Figure 2, the pre-processor 11 includes separate processing streams for each of the sensor data signals SA, SB, SC. In particular, the sensor data signals SA, SB, SC are input to respective pre-processing functions 21a-21c, with each pre-processing function including a comparator function 23a-23c that compares the input sensor data signal with corresponding set point data 25a-25c and outputs a signal corresponding to the difference between the parameter value conveyed by the input signal and the parameter value indicated by the corresponding set point data. The output of each comparator function 23a-23c is normalised by a respective normalisation function 27a-27c, which determines the absolute value of the output of the comparator function and normalises the resultant absolute value to a value between zero and a given maximum value, for example a maximum value of one may be used in some embodiments as an upper limit for normalisation. The output of the normaliser function is routed to one of a first ANN sub-network 31a and a second ANN subnetwork 3 lb in dependence on whether the output of the comparator function is positive or negative. In figure 2, this is schematically represented by each sensor signal beinginput to a comparator, with the output of each comparator 23 being input to a normaliser 27, which applies the normalisation function, and the output of the normaliser 27 being passed to a first ANN sub-network 31a if the output of the comparator 23 is positive and to a second ANN sub-network 3 lb if the output of the comparator 23 is negative.

[0067] In this way, each pre-processing stream 23 outputs a signal conveying a positive value between zero and a given maximum value that is determined based on the difference between the parameter value conveyed by the input sensor data signal and the parameter value indicated by the corresponding set point data such that the greater the difference is, the larger is the magnitude of the output signal.

[0068] In this example, the form and operation of the first ANN sub-network 31a and the second ANN sub-network 31b, which together form the ANN 15 of Figure 1, are substantially the same. As shown in Figure 3, in this example, each ANN sub-network 31 is a random network including a plurality of artificial neurons (hereafter referred to as nodes) that are configured into three sets, in particular a set of input nodes 41, a set of basic nodes 43 and a set of output nodes 45.

[0069] While three input nodes 41a-41c are shown in Figure 3 for ease of explanation, typically there is one input node for each pre-processing function and accordingly there is one input node corresponding to each sensor data signal. Similarly, while three output nodes 45a-45c are shown in Figure 3 for ease of illustration, typically there is one output node 45 for each control signal generator and accordingly there is one output node 45 for each actuator 7. While four basic nodes 43 are shown in Figure 3 for ease of illustration, typically there will be many more basic nodes, for example from ten to five thousand.

[0070] All the nodes of the ANN 15 are labelled either “excitatory” (Exc in Figure 3) or “inhibitory” (Inh in Figure 3). More particularly, all the input nodes 41 and all the output nodes 45 are excitatory nodes while a subset of the basic nodes (represented in Figure 2 by the basic nodes 43a and 43d) are excitatory nodes with the remainder being inhibitory nodes (represented in Figure 3 by the basic nodes 43b and 43c). The difference between an excitatory node and an inhibitory node is that signals received by a recipient node from an excitatory node generally contribute to increasing the activity output by the recipient node whereas signals received by a recipient node froman inhibitory node generally contribute to reducing the activity output by the recipient node, as will explained in more detail hereafter.

[0071] In this example, nodes within the same set and having the same label are configured to have a predefined number of output edges. Accordingly, each of the input nodes 41 has a single input edge connected to the output of a corresponding one of the pre-processing streams 23 and a first predefined number of output edges interconnecting the input node with basic nodes 43 and output nodes 45. Each of the excitatory basic nodes 43 has a second predefined number of output edges interconnecting the excitatory basic node with other basic nodes 43 and output nodes 45. Each of the inhibitory basic nodes 43 has a third predefined number of output edges interconnecting the inhibitory basic node 43 with other basic nodes 43 and output nodes 45. Each of the output nodes 45 has a single output edge connected to respective one of a set of signal generators 33a-33c, which form the control signal generator of Figure 1, and input edges as mentioned above.

[0072] In this example, the edges within an ANN sub-network 31 are assigned taking into account knowledge of the relationship between sensors 3 and actuators 7. For example, it may be known that operation of a first actuator 7 will have a comparatively strong impact on the parameter value detected by a first sensor 3, whereas the operation of a second actuator 7 will have a comparatively strong impact on the parameter value detected by a second sensor 3. Accordingly, a first population of the basic nodes 43 will be assigned to allow direct paths through the first population of basic nodes 43 from a first input node 41 corresponding to the first sensor 3 to a first output node 45 corresponding to the first actuator 7, and the insertion of directed edges into the ANN sub-network 29 will be biassed, using a probabilistic function governing the addition of edges, to add directed edges between the first input node 41, the first population of basic nodes 43 and the first output node 45. Similarly, a second population of the basic nodes 43 will be assigned to allow direct paths through the second population of basic nodes 43 from a second input node 41 corresponding to the second sensor 3 to a second output node 45 corresponding to the second actuator 7, and the insertion of directed edges into the ANN sub-network 29 will be biassed, using the probabilistic function, to add directed edges between the second input node 41, the second population of basic nodes 43 and the second output node 45. The probabilistic function will allow the additionof edges connecting the basic nodes 43 of the first population and the basic nodes 43 of the second population, either directly or via other basic nodes, permitting activity in the first population of basic nodes 43 to interact with the second population of basic nodes 43, and vice versa. More generally, the insertion of each of the plurality of directed edges is specified by probabilities to a subset of the plurality of nodes for each source node when the ANN is constructed, such that pre-established relationships between sensors and actuators are emphasized in the resulting connectivity.

[0073] In this example, the signals output from the output nodes 45 can have a positive or negative effect on a subset of the input signals, and more generally the output signals in combination can have a positive or negative effect on the input signals in combination.

[0074] The Nodes

[0075] Figure 4 schematically shows the processing of activity signals received by a recipient node 41 from multiple excitatory nodes, represented in Figure 3 by three excitatory nodes 43a-43c and hereafter referred to as excitatory nodes 43, and multiple inhibitory nodes, represented in Figure 3 by two inhibitory nodes 45a-45b and hereafter referred to as inhibitory nodes 45. It will be appreciated that the actual number of excitatory nodes 43 and inhibitory nodes 45 in practical implementations will generally be significantly higher.

[0076] The activity signals received by the recipient node 41 from the excitatory nodes 43 and the inhibitory nodes 45 are input to respective different input functions 47a-47e. For each input function 47, the output y is determined in a periodic manner according to the function

[0077] y(prev_y, x):=MAX(prev_y*label_specific_decay, x)

[0078] where x is the value of the input to the input function 47, prev_y is a value corresponding to previous output from the input function 47, and label specific delay is a parameter between zero and one that may have different values when the input signal being processed is from an excitatory node and when the input signal being processed is from an inhibitory node. The effect of the input function is that if there isa reduction in the value x of the input of the input function 47 that results in the value x decaying faster than the decay of the previous output of the input function 47 corresponding to the value of the label specific decay parameter, then the value of the decayed previous output y is used in preference to the value of the input x to the input function as the output y of the input function y. This has the effect of reducing high-frequency signals which assists in determining a solution. Such high frequency signals may be generated within recurrent artificial neural networks because there is always a risk of creating positive feedback loops that saturate the network activity, rendering it unresponsive to actual sensory input, and although such positive feedback loops can be at least partially quenched by the inhibitory nodes, the inhibitory quenching lags the build-up of excitatory activity thereby creating high frequency self-amplifying transients. By smoothing the activity of the individual neurons using a decay function for its activity, such self-amplifying transients can be reduced or even avoided, thereby allowing the network activity to focus on determining control signals that reduce the difference between the sensor data and the set point data.

[0079] The value y of the output from each input function 47 is then input to a respective weight function 49, where the value y is multiplied by a weight w corresponding to the edge via which the input signal for that input function 47 was received by the recipient node 41.

[0080] The outputs of the weight functions 49 for signals received from excitatory nodes 43 are then input to a first combiner function 51a, which sums the outputs together to generate a sum L_exc, where:

[0081] L_exc = Xy*w over all the outputs corresponding to excitatory inputs.

[0082] Similarly, the outputs of the weight functions 49 for signals received from inhibitory nodes 45 are then input to a second combiner function 5 lb, which sums the outputs together to generate a sum L_inh, where:

[0083] L_inh = Xy*w over all the outputs corresponding to inhibitory inputs.The values of the parameters L_exc and L_inc are output by the first combiner 51a and the second combiner 51b respectively and input to a base function 53, which determines the magnitude of the activity signal output by the recipient node 41 to other nodes. In this example, the output x of the base function is determined by the expression:

[0084] x := MAX(0, L_exc - (L_inh / 2)).

[0085] This expression mitigates against the possibility of positive feedback loops being present within the ANN 15, with the L_inh parameter being a determining factor for the rate at which activity in the ANN 15 is reduced. It will be appreciated that variations to this expression can be made while achieving the same effect.

[0086] The output x of the base function 53 is input to an output function 55 which propagates the output x along the output edges of the node 41. In this example, the output function 55 introduces a latency to the propagation of the output x, with a latency value being specified for each output edge. In this example, the latency values are specified in a random manner.

[0087] Introducing latencies to the propagated signals introduces non-linearities into the artificial neural network, which allows the activity of different nodes in the artificial neural network to be differentiated. In this way, the time-varying signal in each node is more unique, thereby increasing the number of options to find solutions in the network by amplifying the weights of the edges from those nodes. This assists in the artificial neural network converging to a robust solution.

[0088] The base function 53 also outputs the L_exc parameter and the L_inh parameter to a learning function 57 which adjusts the weights corresponding to a subset of the input edges so as to reduce activity in the ANN 15 over time and bring the ANN into a stable solution. More particularly, the learning function 57 is a local learning function which adjusts weights for input edges to the corresponding node based on parameters associated with that node.

[0089] The learning function 57 only adjusts the weights for the input edges for which the output of the input function 47 is among the highest. By focussing the learning on the received activity signals that are strongest, in effect the ANN 15 acts first to reduce the highest areas of activity. This approach assists in reaching a solution, particularlyfor complex systems where there is no one-to-one correspondence between an actuator and a sensed parameter.

[0090] In this example, the learning function 57 only modifies the weights for input edges for which the expression

[0091] prev_y > ALL_y* 0.75

[0092] is satisfied, where ALL_y is the maximum value of the signal y output by a input function 47 for that node, although it will be appreciated that many different expressions could be used to arrive at the result of selecting the edges providing the strongest incoming activity signals to that node. For example, the value 0.75 could be replaced by a higher or lower value in the expression given above, or alternatively a predetermined number of the strongest incoming activity signals or a predetermined proportion of the strongest incoming activity signals could be selected.

[0093] For each of the weights being modified by the learning function 57, if the corresponding input edge connects to an excitatory node and the value of the output y of the corresponding input function is greater than the value of the L_exc parameter, then the learning function 57 increases that weight w by an amount dw that may be expressed as:

[0094] dw += MAX(0, y - 1 + s - L_exc)*rate

[0095] where s is a suitability parameter for the node 41 and rate is a learning rate, which is a scalar value used to control the size of dw, and if the value of the L_exc parameter is greater than a threshold value T, then the learning function 57 reduces that weight w by an amount dw that may be expressed as:

[0096] dw -= MAX(0, L_exc - T)*rate.

[0097] Similarly, if the input edge corresponding to a weight being modified by the learning function 57 connects to an inhibitory node then if the output y of the corresponding input function is greater than the value of the L_inh parameter, then thelearning function 57 increases that weight w by an amount dw that may be expressed as:

[0098] dw += MAX(0, y - 1 + s - L_inh)*rate

[0099] and if the value of the L_inh parameter is greater than a threshold value T, then the learning function 57 diminishes that weight w by an amount dw given by the expression:

[0100] dw -= MAX(0, L inh - T)*rate.

[0101] As the condition for increasing a weight is dependent on the output y and the condition for reducing a weight is dependent on the summation L_exc, L_inh, it is possible for both conditions to be satisfied in which case the weight is adjusted by the final value of dw after addition and subtraction.

[0102] The suitability parameter s of the node 41 is modified in dependence on changes to the length of a vector V = (y, L_exc, L_inh). In particular, if dV is zero or negative, suggesting that one or both of L_exc and L_inh is decreasing, then the suitability s is increased, for example by a fixed amount, whereas if dV is positive, suggesting an increase in one or both of L_exc and L_inh, then the suitability s is diminished, for example by a fixed amount.

[0103] Increasing the suitability s has the effect that for the same difference between the output y for an edge and L_exc when the node 41 is an excitatory node, or the same difference between the output y for an edge and L_inh when the node 41 is an inhibitory node, the weight w corresponding to that edge can be potentiated by a greater amount. If, however, such an increase in weight results in worse performance of the system and the input y, V will increase over time, leading to the suitability reducing.

[0104] If a positive feedback loop develops in the ANN 15, then the output y corresponding to an edge forming part of the positive feedback loop will grow. The resultant increase in activity results in the suitability s of edges associated with that positive feedback loop diminishing, thereby reducing or eliminating any increase in the weight w for those edges. In this example, in the event that the output y for an edgeexceeds the threshold T (for example y > 1) and the weight w cannot be potentiated, then that edge may be randomly assigned a different endpoint node or that edge may be removed and another edge randomly inserted elsewhere in the ANN 15, thereby assisting to break any positive feedback loop.

[0105] The operation of the ANN 15 described above results in the weights for edges being modified until the dw reaches zero for all nodes and the values of the sensed parameters match the values of the corresponding set point data. If the values of L_exc and L_inh are also stably under the threshold T, then all the pathways between the input nodes and the output nodes form part of a negative feedback control system.

[0106] Operation of the System Controller

[0107] As shown in Figure 5, the operation of the system controller 5 starts by the system controller configuring, at SI, an initial configuration for each of the first ANN sub-network 31a and the second ANN sub-network 31b, which includes determining which edges of the random network are initially present and which edges of the random network are initially not present and also initial weights for the present edges. In this example, all the initial weights are set to zero but other initial configurations can be used.

[0108] Following configuring the ANN 15, the system controller 5 receives, at S3, sensor data in the form of the sensor data signals SA, SB and Sc. The system controller 5 then determines, at S5, input signals for the first ANN sub-network 3 la and the second ANN sub-network 31b using the pre-processor 11 in the manner described above. In particular, input signals are determined in dependence on the magnitude of the difference between parameter values conveyed by the sensor data signals SA, SB and Sc and set point data associated with the parameters.

[0109] The input signals are then input to the first ANN sub-network 31a and the second ANN sub-network 3 lb, which process, at S7, the input signals to generate output signals. While processing the input signals, the ANN sub-networks 31 adjust, at S9, weights for a subset of the input edges to a node using a local learning rule for that node. The system controller then determines, at SI 1, a control signal based on the output signals. This control signal can affect the parameter values conveyed by the sensor data signals SA, SB and Sc.The adjustment of the weightings has the aim of reducing the magnitude of the difference between the parameter values conveyed by the sensor data signals SA, SB and Sc and the set point data associated with the parameters. In general terms, the local learning rule aims to increase the weights of input edges into a node to result in an output signal that causes the parameter values for subsequent sensor data signals SA, SB and Sc, such that the activity in the ANN 15 reduces, while decreasing the weights for input edges to nodes for which the activity within the node is too large. The learning rule also maintains a suitability value for each node which is increased if the adjustment of weights tends to reduce activity within the node but is decreased if the adjustment of weights tends to increase activity within the node. In the event that activity signals indicative of a positive feedback loop are detected, the ANN 15 removes the input connection of the edge conveying the largest activity signal into the node with the lowest suitability and randomly connects the removed edge endpoint as an input to a different node within the ANN 15 such that the ANN 15 has a new configuration. Over time, the changes of weights in, and the configuration of, the ANN 15 results in the ANN 15 entering a low activity state in which the output signals result in a control signal for which the resultant parameter values in the sensor data signals SA, SB and Sc generally match the set point data, and hence the input signals are low.

[0110] The System Controller - Physical Device Features

[0111] Figure 6 shows, by way of example, the main components for a software implementation of the system controller. As shown, the system controller 61 includes input / output devices 63, a processor 65 and memory 67.

[0112] The input / output devices 63 include one of more input devices for receiving the sensor data signals SA, SB and Sc. In some examples, there is one input device for each sensor data signal while in other examples there is a single input device having multiple ports allowing the single input device to receive the sensor data signals SA, SB and Sc. It will be appreciated that the input devices may conform to standard specifications as are well known in the art.

[0113] The input / output devices 63 also include one or more output devices for transmitting the control signals CA, CB and Sc. In some examples, there is one output device for each control signal while in other examples there is a single output devicewhich transmits a multiplexed control signal allowing the single output device to transmit the control signals CA, CB and Cc. It will be appreciated that the output devices may conform to standard specifications as are well known in the art.

[0114] While the processor 65 is illustrated as a single component, it will be appreciated that the processor 65 may include multiple processing devices. For example, the processing operations may be distributed between multiple processing devices within the system controller 61.

[0115] The memory 67 may include multiple memory devices having respective different properties, such as access times and permanence, in a manner well known in the art. The memory 67 stores data 69, program routines 71 and also provides working memory 73. The data 69 includes, for example, ANN parameters 75 providing configuration details and edge weights for the ANN 15 and set point data 77. The routines 67 include a pre-process sensor data routine 79, a propagate activity routine 81, a learning rule routine 83 and a generate control signal routine 85.

[0116] The pre-process sensor data routine 79 determines the input signals for the ANN 15 based on differences between the sensor data signals SA, SB and Sc received by the input / output devices 63 and the set point data 77 stored in the memory 67. The generate control signal routine 85 processes output signals from the ANN 15 to generate the control signals CA, CB and Cc, and outputs the control signals CA, CB and Cc using the input / output devices 63.

[0117] The propagate activity routine 81 propagates activity signals through the ANN 15 based on the ANN parameters 75 in the manner described above. The learning rule routine 83 modifies the ANN parameters 75 in the manner described above.

[0118] It will be appreciated that the system controller could alternatively be implemented in hardware, or a different combination of hardware and software, and perform the same processing operations.

[0119] Modifications and further examples

[0120] By way of example only, further details of various applications of a system controller as described above in a robotic system will now be described.

[0121] As discussed above, various robotic systems, such as robots with articulated limbs (e.g. biped or quadruped robots), are known. Other robotic systems maycomprise autonomous propulsion systems for movement which do not use articulated limbs. By way of example, only, the disclosed technology may be used for autonomously controlling and operating systems such as vehicles, for example, cars and heavy duty vehicles, trains, aircraft, surface vessels or submersibles which do not use articulated limbs for movement, including unmanned airborne vehicles (e.g. drones) or unmanned underwater vehicles or parts thereof. Another example of a system or system component which may be controlled using the disclosed technology, includes a motor. A motor may be configured, for example, to control the actuation of a system component such as a valve or the like for flow regulation or to regulate the speed of a propulsion system.

[0122] For an example such as biped and quadruped robots, the articulated limbs include motors and sensors, and there have previously been successful attempts to train such robotic systems to walk. The previous attempts have, however, required extensive training of the robotic system, particularly as movement of one limb as a result of actuation of a motor may affect multiple sensed signals. In an application of the system controller described above to robots with articulated limbs, the sensors may detect positional information for different locations on the robotic system. This positional information may, for example, be the distance of each sensed location above the ground. The control signals may be applied to respective motors causing movement of the articulated limbs. By setting the set point data to correspond to positions for the sensed location at which the robotic system is in a standing state, the robotic system can in effect learn to stand based only on data from the sensors.

[0123] Figure 7A shows schematically by way of example a robotic joint system of a robot which may be controlled using a method of controlling a system by implementing a controller as disclosed herein.

[0124] In Figure 7A, a biped robot has a plurality of joints and the movements of limbs around the joints of the robotic joint system are controlled using motor components activated by actuators under the control of a controller system.

[0125] The controller systems and methods for controlling a system disclosed herein may configure the actuators individually as well as collectively to generated control signals for each joint or limb to adjust the positions of the limbs about a joint, in other words, the control signals will adjust the amount of rotation and / or linear movement ina lateral or vertical direction in the one or more degrees of freedom each joint can move within with the aim of achieving a target state or behavior of that joint of the joint system which results in that joint or the joint system of the robot as a whole achieving a target pose, position or velocity.

[0126] It will be appreciated however by those skilled in the art of robotics that the control principles disclosed may be applied to other types of robotic joint systems, for example, to quadruped robots and robotic arms which comprise articulated joints.

[0127] In Figure 7A, the biped robot is illustrated in an example balanced standing pose whereas Figure 7B shows an example of a biped robot in a different pose. The pose shown in Figure 7B may be a balanced or stable pose if the robot is moving in a stable manner but could also be an unbalanced or unstable pose if the robot has moved responsive to a perturbing force and has not yet regained its balance.

[0128] The term balance as used herein with reference to the robot as a whole takes its conventional meaning in the art, that is to say, the centre of mass of the robot is maintained within a set range associated with stability, meaning that the centre of mass does not deviate from that set range for a set duration of time, for example, for 10 seconds, 1 minute or another suitable time interval, which may be context dependent on a task the robot is to perform.

[0129] In Figure 7A the robot is shows schematically having a plurality of sensors, S, located at various locations of its joints. In addition, but optionally in some embodiments, and as illustrated, sensors may be mounted on other components of the robot such as on the robot's limbs and / or torso as well as points of inflexion such the robot's waist.

[0130] These sensors may detect rotary and / or lateral movement in a number of degrees of freedom, for example, one or more or all of the six degrees of freedom associated with x,y,z co-ordinates and degrees of pitch, yaw, and roll according to the permitted movements of the joint to which they are attached. In some embodiments, one or more or all of the joints of a robot may not be capable of moving in all degrees of freedom.

[0131] For example, as illustrated in Figure 7C, one or more IMU sensors may be used to provide inertial measurement unit (IMU) sensor pitch information about a y-axis, IMU sensor roll data about an x-axis, and IMU sensor yaw about a z-axis (assuming an x,y,z co-ordinate system is used which may be centred on the centre of mass of therobot in some embodiments). Figure 7C does not show any degree of yaw movement about the z-axis, which is in the vertical direction. For example, an elbow joint may be modelled with only the capability of rotating a robot upper arm limb relative to a robot lower arm limb in a 2-D plane, in other words for example, within a particular x-y plane, whereas a robot wrist joint may be configured to be capable of moving in all three dimensions.

[0132] Limits of limb motion may be configured by limiting joint rotation may in some embodiments of the controller. These limitations can be modelled by pre-processing received sensory data in some embodiments prior to the pre-processed sensory data being input to the controller.

[0133] A permitted range of joint movement about a set or target point need not always be symmetrical. For example, as shown in Figure 7A, an elbow joint may be modelled to limit rotational movement to within a certain angle a3. To prevent joint damage, the angle of rotational movement may be limited in the real-world in a non-symmetrical way, with movement limited more in one direction than in another, shown by the different angles al and a2 in Figure 7A.

[0134] In some embodiments, received sensor data accordingly may need to be pre-processed to be zero-centred with a permitted movement range that is off-set and / or scaled for some limb movements. For example, in some embodiments, a zero-centred autoscaling is implemented in which joint movement parameters are bound to within given range endpoints for a given joint. This may allow scaling to the maximum value of a range, while keeping the position control symmetric in both directions. This may be in some embodiments at the expense of introducing a "dead-zone" on one side if the joint endpoints are asymmetric in the real -world.

[0135] Some embodiments of the disclosed technology use in addition a system of stratified sensors which allows extreme sensory inputs to be mapped to different behaviours while avoiding large areas of synaptic movement within the artificial network and / or within specific motor neuron populations and yet still retaining sensitivity to small movements.

[0136] The disclosed control systems and methods of controlling a system comprising a robotic joint system may be used to control robot movement using one or both of a position model, where the robot joints are controlled directly using sensors which detectposition, force, and velocity, and a muscle model, where muscles act in pairs and are contracted and elongated to move limbs about a joint. In some embodiments, the controller is configured to receive sensory input data from a plurality of position, force, and velocity sensors S mounted on a robot, such as the sensors S located at the joints of the robot limb system of the robot. This results in, for example, sensor data signals such as those shown as SA, SB, and Sc in Figure 2 being received by the system controller 5 shown in Figure 2 comprising at least positional data indicating sensed positions of robotic joints, and in some embodiments in addition or instead, force and velocity data from the robotic joint system sensors.

[0137] The controller system 5 shown in Figure 2 may be configured to control a robotic limb system comprising a plurality of robotic joints in some example implementations, for example, joints associated with limbs which are attached to the robotic body of the robot shown in Figures 7A, 7B and 7C. The limbs are attached to the robotic body or torso via joints which allow the robot to be capable of changing its pose and / or its position, in other words, the joints rotate and thus allow for movement of the robotic limbs with one or more degrees of freedom.

[0138] The resulting limb movement arising from rotation of one or more joints of the robot may result in the robot adopting a new pose which may be stable or unstable depending on the location of the centre of mass of the robot after the pose has been attained. Changes in position of a robot's centre of mass can be detected using a centre of mass IMU sensor in some embodiments which may also be configured to provide sensory data as input to the controller 5.

[0139] The new pose may be determined or categorised as a static pose if the pose is maintained over a minimum period of time. An example of a static pose is a balanced standing pose held for an interval of time such as 30 seconds. The new pose may be maintained whilst there is other movement of the robot, for example, a stable pose may be maintained by a torso and arm state even when the robot is moving. This can occur, for example, when the pose results from movement of an upright torso which is then held stable despite the legs of a biped robot still moving.

[0140] In some embodiments of the disclosed technology, the controller seeks to maintain a stable balanced standing pose of a robot by firing actuators which result in limb movement(s) which individually or collectively result in one or more limb or torsopositions being maintained within certain threshold deviation(s) from set point position(s) for a duration of time.

[0141] In some embodiments, the pose is determined to be maintained if the limbs forming the pose remain below one or both of a maximum rotational deviation from a set angle and a maximum linear deviation such that the robotic body to which they are attached also remains stable, that is to say it does not move beyond threshold amounts.

[0142] The number of degrees of freedom which each articulated section of each robotic limb has will determine the dimensions of the parameter set for establishing stability. For example, taking an extreme, if the robotic limb has no ankle flex or knee articulation, then only hip articulation can be controlled, and this may be limited to rotation about a hip axis in a forwards and backwards direction. For stability, maintaining the degree of rotation in the forwards and backwards direction below a rotation threshold of, for example, 5 degrees, for longer than a time threshold of, for example, 1 minute, may result in the robot pose being determined as stable and the robot may be considered to be balanced in that stable pose. If the limb positions and robotic pose correspond to the criteria for the pose to be considered a static standing pose of the robot, then maintaining control over the robotic limb such that it remains within the maximum allowed rotational degree of movement for a time period that lasts at least as long as the minimum length of time.

[0143] The direction of movements which result in a stable pose may be limited to allow the model to be more likely to converge on a stable solution.

[0144] Figure 8 shows an example actuated muscle sensor system according to some embodiments of the disclosed technology in which joint actuators are controlled using a controller according to an embodiment of the disclosed technology. The overall controller structure for the Actuated Muscle Sensor, AMS, shown in Figure 8 is configured to support the generation of movement of robotic joints forming a robotic joint system in a controlled manner, for example, so as to achieve a target state or behavior state of a robot comprising the robotic joint system. The target or goal state may be associated with a stable or balanced position or pose of the robot whilst the robot itself is either static (for example, a balanced standing pose) or a dynamic pose (for example, a stable walking pose).The primary function of the AMS accordingly is to generate joint movement which in analogous biological systems is achieved by controlling the movement of a set of muscles.

[0145] As shown in the example embodiments of an AMS in Figure 8, limb movement is achieved by converting the activity of the (1) Alpha-Motoneuron (A) to movement of some kind through the Muscle Model (4) all the way to an external actuator (9). That is, activity in the Alpha-Motoneuron is converted into movement (by force, position or velocity). The activity of the Muscle (4) thereby affects the output of the three sensors. The three sensor dimensions outlined are Force (6, Ib-sensor), Velocity (7, la-sensor) and Position / Muscle length (8, Il-sensor). These three dimensions are extracted from the relevant muscle model implementation such that they correspond to equivalent metrics in the external world. However, a bias can be added to the Velocity and Position sensor through the activity of either the Gamma-Dynamic Motoneuron (2, adds bias to the Velocity sensor) or the Gamma-Static Motoneuron (3, adds bias to the Position sensor).

[0146] In biology, the Gamma-Static adds some bias to the Velocity sensor (la-sensor), but in some of the embodiments of the disclosed technology, these two sensors are completely disjoint.

[0147] Figure 9A shows an example configuration of an artificial neural network of a control system for a biped robot. In Figure 9A, sensors provide input to an inhibitory tract neuron population and to an excitatory tract neuron population of the controller ANN as well as providing input to a left intersegment inhibitory population, a left intersegment excitatory population, a right intersegment inhibitory population and a right intersegment excitatory population. As shown in Figure 9A, the intersegment populations of inhibitory and excitatory neurons link the inhibitory and excitatory tract neurons to the joint model motor neurons, which in turn also provide output to the main neuron population. The main neuron population then provides output to the actuators.

[0148] An example of the joint model as shown in Figure 9A is illustrated in more detail in Figure 9B. The joint model of Figure 9B is configured for a biped robot such as is shown in Figures 7A, 7B, and 7C. It will be appreciated however, that the example illustrated may be readily extended to other types of robotic joint systems as would be apparent to anyone of ordinary skill in the art.In Figure 9B, the joint model comprises two parts, one for the left-hand side joints of a robot and one of the right-hand side joints of the robot. A series of motor neuron populations for each permitted degree of freedom of joints located along the lefthand side of the robot form the lefthand side joint system model. It will be appreciated that the motor neuron populations may differ from that shown in Figure 9B depending on the number of degrees of freedom for joint of the robot is permitted to move in.

[0149] As shown, each of the nodes in a motor neuron population for movement in a degree of freedom (pitch, roll, and yaw) of a joint on the left hand side of the robot's joint system model receives input from at least the left hand side excitatory and inhibitory intersegmental neuron populations.

[0150] A series of motor neuron populations for each possible degree of freedom of joints located along the right-hand side of the robot form the right-hand side joint system model. Each of the nodes in the righthand side joint system model receives input from at least the righthand side excitatory and inhibitory intersegmental neuron populations.

[0151] In some embodiments, the sensor data associated with the joint system may comprise sensors mounted on, embedded in, or attached to a plurality of joints and / or limbs which are configured to collaboratively act on the robot to balance the robot in a standing or walking state.

[0152] A permitted range of joint movement about a set or target point need not always be symmetrical. For example, knee and elbow joints having limited movement in one direction, although in some situations it is possible to hyperextend a knee (and some situations may require so-called double-jointedness). Received sensor data accordingly may need to be zero-centred with a permitted movement range that is off-set and / or scaled for some limb movements. For example, in some embodiments, an updated zero-centred autoscaling may be implemented in which movement parameters r land r_2 are bound to withing given range endpoints for a given joint. This allows scaling to the maximum value of a range, while keeping the position control symmetric in both directions although possibly at the expense of introducing a "dead-zone" on one side if the joint endpoints are asymmetric.Some embodiments of the disclosed technology use a system of stratified sensors which allows extreme sensory inputs to be mapped different behaviours while avoiding large areas of synaptic movement within the artificial network and / or within specific motor neuron populations and yet still retaining sensitivity to small movements.

[0153] Using an actuated muscle sensor in combination with additional control points for sensor sensitivity and in parallel with a force / motor control model for controlling joints may allow the controller to implement a differential priority or importance scheme for sensor signals from different muscles while enabling maintained actual force generated by them.

[0154] This also may advantageously allow for pre-emptive adjustment of sensor importance that can change faster than the actual muscle model decreasing the latency of informational flow due to inertia.

[0155] In some embodiments, joint position and velocity sensors are directly connected to populations of motor neurons associated with the respective joint in contrast to the embodiment shown in Figures 9A and 9B which provide activity via a tract and / or intersegment population of neurons.

[0156] In some embodiments, position control uses joint angles when the pivot locations of each joint are known, however, in some embodiments, position control may be implemented directly in Cartesian space. In some embodiments, sensors provide sensory data comprising joint positions and joint velocities and the distance in Cartesian space and / or provide sensory data indicating a difference in the joint space between the current and target position, and the controller neural network as a whole will drive movement to minimise distance / difference.

[0157] In some embodiments of the disclosed controller system which implements force control, a muscle model a spring-damper model, for example, a Hill muscle model is used. In this model, different sensors provide sensory data indicating the position, velocity, and force of each individual muscle. The output of the neuron network then drives muscle contraction and, in some embodiments, opposing muscle pairs may be implemented for elongation.

[0158] In some embodiments, the joint system target state or behavior is sought by the controller responsive to sensory data input comprising of one or more of: the actual poses / positions of one or more joints or limb parts of the robot, the joint or limbvelocities and the forces acting on or generated by one or more or all of the joints and limb parts of the robot.

[0159] In some embodiments, for example, where the robot comprises a biped, populations of motor neurons are defined for all lower-body motors, but the ranges of the motors themselves are locked, except those that we allow movement for (e.g. hip pitch). The intra-limb movement abstraction may be captured by providing "intersegment" populations for each leg as shown schematically in Figures 8A and 8B.

[0160] In some embodiments of the controller system shown schematically in Figure 9A and 9B, sensory data received from the sensors comprises a pitch of an IMU where the IMU is located in the centre of the waist, and a difference in height of IMU from original position.

[0161] Training the model may comprise initially locking all motors apart from the hip pitch. This may train the model so that it learns to balance quickly and is fairly robust to perturbations already. Later, complexity may be increased and movement of the knees and roll of the IMU can be introduced.

[0162] Some embodiments of the disclosed technology comprise a computer-implemented method of generating a control signal for a robotic limb or joint system of a robot in dependence upon sensor data associated with one or more robotic limbs or joints respectively forming the robotic limb or joint system. The sensor data is associated with set point data corresponding to a target performance. In some examples, the target performance may relate to the robotic system attaining a target state while in other examples the target performance may relate to the controlled system exhibiting a target behavior or sequence of states. As such the targets may be static or dynamic. In an example, a target behavior involves a dynamic response to the sensed environment in which the controlled system is operating or based on prior system behavior. At a general level, a target state may define what a system should attain whereas a target behavior may define how the system attains that state. The target state or behavior may comprise one or more or all of a target pose, a target position, or a target velocity of for the robotic limb system.

[0163] The method comprises receiving, by a control system comprising an artificial neural network comprising a plurality of neuron populations, at least some of the neuron populations; being associated with a degree of freedom of a joint of the robotic limbsystem, each population having a plurality of nodes interconnected by a plurality of edges, the sensor data; determining, by the control system, one or more input signals for the artificial neural network and inputting the one or more input signals to the artificial neural network; generating, by the artificial neural network, one or more output signals; determining, by the control system, the control signal based on the one or more output signals; and adjusting weights associated with edges of the artificial neural network to reduce a difference between the sensor data and the corresponding set point data, wherein the adjusting of weights comprises a node using a local learning rule to adjust the weights for a subset of the plurality of edges for which the node receives activity output by other nodes.

[0164] The local learning rules applied by nodes to edges within a neuron population may differ from the local learning rules applied by nodes to edges between neuron populations. The controller may configure the actuators to control movement of the robotic joints or limbs individually to attain the target state or behavior. The actuators may comprise actuators for motor components that the robot is configured to use to control movement of its joints, resulting in limb movement. The controller may configure the actuators to control the movement of the robotic joints collectively to maintain the target state or behavior for a minimum duration of time. The controller may also in addition or instead configure the actuators to control the movement of the robotic joints of the robotic joint system collectively to regain the target state or behavior responsive to one or more perturbations in the target state or behavior caused by a perturbing force acting on the robot, for example, where regaining the target state or behavior results in a stable or balanced pose of the robot.

[0165] At least one robotic limb or joint of the robotic limb system may comprise one or more control points. Each force applied by an actuator to a robotic joint or limb to regain the target pose is correlated to a perturbing force or a component of the perturbing force in some embodiments. The correlation may be linear in some embodiments, or, if there are one or more so as to have a sufficient number of control points, a superlinear response to the perturbing force may be output by the controller.

[0166] In some embodiments, the robotic system is controlled using a joint position model and the controller provides differential sensor signal prioritisation for position control of individual joints.In some embodiments, wherein the robotic system is controlled using a muscle model and the controller provides differential sensor signal prioritisation for muscle control of individual joints.

[0167] The disclosed controller may be used to control joints of a robot where the robot comprises a robotic arm, for example, on a manufacturing line. Other examples of robots that can be controlled by the controller include, a biped robot or quadruped, for example, biped or quadruped robots configured to operate in a warehousing or similar logistics environment.

[0168] In some embodiments of an implementation of the method, the control signal is generated for a joint system of a robot in dependence on sensor data associated with the joint system. The joint system may comprise a plurality of joints which are associated with one or more limbs forming a robot or part of a robot. In some embodiments, the joints are part of an articulated limb system forming the robot. In some embodiments, one or more joints may attach a limb to another body part of the robot, where the other body part may be articulated and comprise one or more joints or be non-articulated and be jointless.

[0169] In some embodiments, the joint system is part of a limb system comprising one or more articulated limbs which are attached at one end to the same body part, for example, in a biped robot, robot arms and legs may be attached via joints to a torso of the biped robot. The torso may also be articulated in some embodiments.

[0170] In some embodiments at least one limb of the limb system is configured to be capable of generating one or more forces acting on the robot to balance the robot in a standing or walking state in dependence upon sensor data associated with the limb system, with the sensor data being associated with set point data corresponding to a target state or behavior for the limb system.

[0171] In some embodiments at least two limbs of the limb system are configured to collaboratively generate forces which act on the robot to achieve a target state or behavior, for example, to balance the robot in a target standing pose or in a stable pose associated with a walking state or behavior. The forces may be generated in dependence upon received sensor data associated with the joint system, with the sensor data being associated with set point data corresponding to a target behavior for the joint system.The target behavior may be determined based on parameter values for movement of the joint system to achieve a particular limb state.

[0172] In some embodiments, the sensor data associated with the limb system may comprise data derived directly or indirectly from sensors mounted on, embedded in, or attached to each of the one or more limbs which individually or collaboratively are acting on the robot to balance the robot in a standing or walking state.

[0173] Advantageously, by adding additional control points for sensor sensitivity in parallel to the already existing force / motor control, differential importance of sensor signal from different muscles while enabling maintained actual force generated by them is enabled, as is pre-emptive adjustment of sensor importance that can change faster than the actual muscle model decreasing the latency of informational flow due to inertia. In some embodiments, a computer-implemented method of generating a control signal for a robotic joint system in dependence upon sensor data associated with the robotic joint system is provided where the sensor data is associated with set point data corresponding to a target state or behavior such as a target pose, velocity or force applied by the robotic joint system and where the method comprises receiving, by a control system comprising an artificial neural network having a plurality of nodes interconnected by a plurality of directed edges, the sensor data; determining, by the control system, one or more input signals for the artificial neural network and inputting the one or more input signals to the artificial neural network; generating, by the artificial neural network, one or more output signals; and determining, by the control system, the control signal based on the one or more output signals, wherein the method further comprises removing an input edge connection to a recipient node and inserting a directed edge elsewhere in the artificial neural network in dependence upon determining that the activity received via the input edge is high and the suitability of the recipient node to reduce a difference between the sensor data and the corresponding set point data is low.

[0174] By way of another example, for unmanned airborne systems and unmanned underwater systems, navigation can be an issue as environmental factors such as wind or water currents can affect navigation in an unpredictable manner. In an application of the system controller described above to unmanned airborne or underwater vehicles, the sensor signals could be produced by positional sensors, for example utilising asuitable global positioning system. This sensor data is compared with set point data based on a planned journey path to generate input signals for the artificial neural network. The output signals from the artificial neural network are then used to generate control signals for steering devices to maintain the unmanned airborne or underwater vehicle along a desired journey path.

[0175] As shown in Figure 10, a task manager 61 is provided that outputs high level action signals to the control system 5 in response to input identifying a desired action. The control system 5 is the same as the control system 5 of Figure 1 except for the addition of a processing component that processes the action signal to identify corresponding set point data, and then updates the set point data for the artificial neural network accordingly.

[0176] The task manager 61 receives sensor signals from the articulation system 9, and bases the action signals on those sensor signals. In some examples, the task manager employs reinforcement learning to determine the appropriate action signal. The action signal may indicate a trajectory through a sequence of states, for example a walking motion for a bipedal robot.

[0177] Generally, the task manager will operate at a rate that is slower from the rate at which the control system operates. In an example, the task manager operates at a rate in the order of 100Hz, whereas the control system operates at a rate of the order of I kHz.

[0178] In the example implementation of Figure 1, the system controller determines the target values for feedback controllers within actuators in dependence upon received sensor signals. Such an arrangement is useful for implementations where the system controller is applied within a robotic system where the feedback controller in each actuator system is already present. It will be appreciated, however, that a feedback loop can be formed by the sensors sending sensor signals to the system controller, which in turn generates, in dependence on the difference between the measured values provided by the sensor signals and the set point data, control signals that are directly applied to drive circuitry for the actuators within the actuator systems. As described above with reference to Figure 10, a task manager may be configured to provide the set point data to the system controller.The disclosed technology may be used in conjunction with a variety of different sensors, including but not limited to sensors configured to sense physical properties, for example: temperature, humidity, proximity, ultrasound, light including one or more or all of ambient visible light, infra-red light, ultra-violet light, pressure, acceleration, colour, touch, level, position, hall effect, tilt, vibration, gas, chemical(s), vibration.

[0179] Such physical properties may be sensed using sensor systems comprising one or more of the following types of sensors, which is not intended to be a complete list: optical sensors, image sensors, temperature sensors, depth imaging sensors, event imaging sensors, gyroscopic sensors, position sensors, speed sensors, accelerometers, chemical sensors, pressure sensors, electromagnetic field sensors, magnetic field sensors, spectral sensors, electrical current or voltage sensors.

[0180] Actuators may comprise pneumatic actuators, hydraulic actuators, electric actuators, linear actuators, rotary actuators, piezoelectric actuators, magnetic actuators, mechanical actuators, electric motors, solenoids, thermal actuators e.g. heaters or heatsinks, valves, diaphragm actuators, stepper motors etc.

[0181] In some examples, a robotic limb actuator comprise a plurality of different types of actuators, for example, a combination of linear actuators and rotary actuators.

[0182] In some examples, the control system and / or the task manager may be implemented on the articulation system. Alternatively, the control system and / or the task manager could be implemented remotely, e.g. on one or more cloud servers, or distributed between the cloud and the articulation system. The control system and the task manager could implement edge computing.

[0183] As described above, actuators 7 may interact with actuated system components 9, and the sensors 3 measure parameters associated with the actuated system components 9 and the surrounding environment. Alternatively, as shown in Figure 11, one or more of the sensors 3 could sense parameters associated with the actuators 7. In particular, in Figure 11 each actuator system lOla-lOlc includes a sensor 103a-103c, an actuator 105a-105c and a feedback controller, which may comprise a PID feedback controller 107a-107c as illustrated in Figure 11. Each sensor 103 detects a parameter of the corresponding actuator 105 and sends a measurement signal S conveying the measured parameter value to the measurement input of the PID feedback controller 107 and also to the system controller 5, which functions in the same manner as the systemcontroller of Figure 1 to provide a control signal C to the reference input of the feedback controller 107.

[0184] Although the illustrated example shows each sensor data signal being compared to respective different set point data to generate an input signal for the ANN 15, alternative configurations in which sensor data received by the system controller generates one or more input signals 15 for the ANN 15 based on a comparison between the sensor data and associated set point data, which may determine the magnitude of the difference between sensor data and set point data, are possible. For example, a single sensor data signal could be pre-processed to derive three parameters which are each compared with respective set point data to generate input signals for the ANN 15. Further, one or more sensor signals may be compared with set point data to generate a single input signal for the ANN 15. Alternatively, the sensor data from two or more sensor data signals could be pre-processed to determine a sensor data value which is compared with a corresponding set point data value. The set point data may comprise a time-series of set points in some embodiments.

[0185] Similarly, although in the illustrated example each output signal from the ANN 15 is input into a respective different signal generator 33, with each signal generator 33 generating a control signal for a respective different actuator 7, alternative configurations are possible in which one or more control signals are determined based on output from the ANN 15. For example, multiple outputs may be input to a single signal generator to generate a control signal for one actuator.

[0186] More particularly, in example implementations of the present invention there may not be a one-to-one correspondence between the number of sensors and the number of pre-processing functions, or a one-to-one correspondence between the number of sensors and the number of actuators. Further, the sensors may detect parameters associated with the actuators, or operations thereof. More generally, the system controller is able to determine control signals for any number of actuators based on any number of measured parameter values, with there being at least one actuator that affects multiple measured parameter values and at least one sensor providing a sensor signal conveying multiple measured parameter values. In an example implementation, an IMU may be positioned in the torso of a biped robot, and sensor signals from the IMU may be used to control all the actuators within the biped robot.Many modifications to the way in which the system controller processes the sensor signals to generate control signals are possible. For example, the normalisation function of the pre-processor 11 may apply a super-linearization function to reach a desired state of the ANN more quickly than if a purely linear function was applied. In a desired state of the ANN, the activity of the ANN is in equilibrium and results in a desired behavior of the robotic system. An example of such a super-linearization function is:

[0187] if ® > 0

[0188]

[0189] otherwise

[0190] where the exponent may, for example, be equal to three.

[0191] In addition, modifications may be made to the configuration and operation of the artificial neural network. For example, by increasing the interconnectedness of the nodes within the neural network, example implementations may not require the ability to remove and re-insert edges as described above. This also removes the requirement for a suitability value in the learning rule.

[0192] An example learning rule that does not require suitability is to determine the weight w for an edge interconnecting a source node and a destination node by increasing the weight by 1*LR (where LR is a learning rate) if the source node activity is greater than twice the next destination node activity, and to decrease the weight by 1 *LR if the next destination node activity is greater than one, with both an increase and a decrease being simultaneously possible. If the determined weight is greater than one, then the weight is set at one while if the determined weight is less than zero, the weight is set at zero.

[0193] One way of representing this learning rule is the following pseudo-code.

[0194] for t = 0:n

[0195] {

[0196] if (SRC(output) > 2 * NEU(next_output)) SYN(weight) += LR;

[0197] if (NEU(next_output) > 1) SYN(weight) -= LR;if (SYN(weight) < 0) SYN(weight) = 0;

[0198] if (SYN(weight) > 1) SYN(weight) = 1;

[0199] }

[0200] where SRC(output) is the activity output by the source node of the edge in a first activity output interval, NEU(next output) is the activity output by the destination node of the edge in the second activity output interval, SYN is the weight of an input edge or an output edge, in other words a synapse, of a node..

[0201] In this way, the adjusting of the weight is influenced by the effect that the activity from the source node has on the activity output by the destination node in the next time step. While such an arrangement does not require disconnection and reinsertion elsewhere of edges, this may be performed if, for example, the activity from the source node and the activity from the destination node both exceed one.

[0202] While the artificial neural network shown in Figure 3 includes input nodes acting as a buffer between the output of the pre-processing unit and the other nodes of the ANN, these input nodes are not required if the outputs of the pre-processor are suitably configured that they can be treated by nodes to which they are connected simply as inputs from an excitatory node. Further, while the artificial neural network shown in Figure 3 has edges interconnecting input nodes and output nodes, these are not necessary.

[0203] In the above description, the term state may represent a static state, for example a bipedal or quadrupedal robot maintaining a stationary pose such as standing up, or a dynamic state, for example a bipedal or quadrupedal robot side-stepping, swaying onto one side, jogging, running and jumping. Some general examples of target behavior accordingly may include maintaining a stable position or pose when the robotic is performing a task which requires limb movement and / or limb movement subject to external forces such occur when a robot lifts up an object.

[0204] The term configuration may be used in place of the word state to the extent that the term configuration encompasses a dynamic configuration as well as a static configuration.While in the illustrated example the system controller utilises software routines, it will be appreciated that at least some of these software routines may alternatively be implemented by hardware, and that there may be performance benefits in so doing.

[0205] The above examples are to be understood as illustrative examples only. Further examples are envisaged. Furthermore, equivalents and modifications not described above may also be employed without departing from the scope of the invention, which is defined in the accompanying claims.

Claims

CLAIMS1. A method of generating a control signal for a robotic system in dependence on sensor data associated with a current state of the robotic system, the method comprising:receiving, by a control system comprising an artificial neural network having a plurality of nodes interconnected by a plurality of edges, the sensor data associated with the current state of the robotic system;determining, by the control system, one or more input signals for the artificial neural network dependent on the sensor data and set point data corresponding to a target state of the robotic system, and inputting the one or more input signals to the artificial neural network;generating, by the artificial neural network, one or more output signals dependent upon the input signals;determining, by the control system, a control signal based on the one or more output signals;inputting the control signal to an actuator system to cause a modification of the current state of the robotic system; andadjusting weights associated with edges of the artificial neural network in accordance with a local learning rule to reduce a difference between the sensor data and the corresponding set point data.

2. The method of claim 1, wherein the articulated system comprises a plurality of actuators for adjusting the orientation between two sections of the articulated system via at least one joint of the articulated system, each actuator receiving a respective measurement signal from the articulated system and a respective control signal from the control system, each control signal being determined in dependence on the one or more output signals of the artificial neural network.

3. The method of claim 1 or claim 2, wherein the articulated system comprises a plurality of sensors to provide the sensor data.

4. A method according to any preceding claim, further comprising a preliminary step of determining, by a task manager, the set point data and inputting the set point data to the control system.

5. A method according to claim 4, wherein the task manager determines the set point data using reinforcement learning.

6. A method according to claim 4 or claim 5, further comprising updating, by the task manager, the set point data in dependence on an input task, and inputting the updated set point data to the control system.

7. A method according to any preceding claim, wherein the adjusting of weights comprises a node using a local learning rule to adjust the weights for a subset of the plurality of edges for which the node receives activity output by other nodes.

8. A method according to claim 7, wherein no learning rule is applied to edges not included in the subset of the plurality of edges.

9. A method according to any preceding claim, further comprising selecting the subset of the plurality of edges in dependence on the magnitude of the activity received from each of the other nodes.

10. A method according to claim 9, wherein the subset of the plurality of edges comprises the edges via which the node receives the highest activity output by other nodes.

11. A method according to any preceding claim, wherein the input signals are determined based on a difference between the sensor data and the set point data.

12. A method according to any preceding claim, wherein the plurality of nodes comprises a set of excitatory nodes and a set of inhibitory nodes.

13. A method according to claim 12, wherein a node:receives a first set of inputs from excitatory nodes and a second set of inputs from inhibitory nodes;for each input, selects a larger of the value of the input and a value representative of previous input modified by a decay function, and multiplies the selected value by a weight for the corresponding edge to generate a weighted input;generates a first summation of the weighted inputs corresponding to the first set of inputs;generates a second summation of the weighted inputs corresponding to the second set of inputs;calculates an output in dependence upon the first summation and the second summation; andpropagates the output along the output edges.

14. A method according to claim 13, wherein each output edge is associated with a latency which determines the timing at which the output is propagated along that output edge.

15. A method according to claim 14, wherein the latency for each edge is randomly assigned.

16. A method according to any of claims 13 to 15, wherein the learning rule increases the weights associated with edges within the subset of edges for which the selected value is greater than the corresponding one of the first summation and the second summation.

17. A method according to claim 16, wherein the increase in weight is dependent on a suitability value for that node, wherein the suitability is dependent on the values of the first summation and the second summation.

18. A method according to claim 17, further comprising removing an input edge and inserting a directed edge elsewhere in the artificial neural network independence upon an expression indicating the selected value for the input edge is high and the suitability value for the node is low.

19. A method according to any preceding claim, wherein the plurality of nodes comprises:one or more input nodes respectively configured to receive the one or more input signals and to propagate the one or more input signals to the other nodes of the artificial neural network;one or more output nodes respectively configured to output the one or more output signals; anda plurality of basic nodes interconnected as a network wherein any basic node is connectable to any other basic node by a directed edge, thereby defining a plurality of possible directed edges, and a proportion of the plurality of possible directed edges are assigned.

20. A method according to claim 19, wherein the number of directed edges for each node is determined when the artificial neural network is constructed.

21. A method according to claim 19 or claim 20, wherein the insertion of each of the plurality of directed edges is specified by probabilities to a subset of the plurality of nodes for each source node when the artificial neural network is constructed, such that pre-established relationships between sensors and actuators are emphasized in the resulting connectivity.

22. A method of generating a control signal for an articulated system having two or more sections interconnected by one or more joints, the articulated system further comprising an actuator system having an actuator for causing relative movement of at least two sections of the articulated system and a feedback controller that provides an actuation signal to the actuator in dependence on a measurement signal associated with the current configuration of the articulated system and a target signal, the method comprising:receiving, by a control system comprising an artificial neural network having a plurality of nodes interconnected by a plurality of edges, sensor data associated with a current configuration of the articulated system;determining, by the control system, one or more input signals for the artificial neural network dependent on the sensor data and set point data corresponding to a target configuration of the articulated system, and inputting the one or more input signals to the artificial neural network;generating, by the artificial neural network, one or more output signals; determining, by the control system, a control signal based on the one or more output signals;inputting the control signal to the actuator system, the actuator system determining the target signal based on the control signal; andadjusting weights associated with edges of the artificial neural network to reduce a difference between the sensor data and the corresponding set point data.

23. The method of claim 22, wherein the articulated system comprises a plurality of actuators for adjusting the orientation between two sections of the articulated system via at least one joint of the articulated system, each actuator receiving a respective measurement signal from the articulated system and a respective control signal from the control system, each control signal being determined in dependence on the one or more output signals of the artificial neural network.

24. The method of claim 22 or claim 23, wherein the articulated system comprises a plurality of sensors to provide the sensor data.

25. A method according to any of claims 22 to 24, further comprising a preliminary step of determining, by a task manager, the set point data and inputting the set point data to the control system.

26. A method according to claim 25, wherein the task manager determines the set point data using reinforcement learning.

27. A method according to claim 25 or claim 26, further comprising updating, by the task manager, the set point data in dependence on an input task, and inputting the updated set point data to the control system.

28. A method according to any of claims 22 to 24, wherein the adjusting of weights comprises a node using a local learning rule to adjust the weights for a subset of the plurality of edges for which the node receives activity output by other nodes.

29. A method according to claim 28, wherein no learning rule is applied to edges not included in the subset of the plurality of edges.

30. A method according to any of claims 22 to 29, further comprising selecting the subset of the plurality of edges in dependence on the magnitude of the activity received from each of the other nodes.

31. A method according to claim 30, wherein the subset of the plurality of edges comprises the edges via which the node receives the highest activity output by other nodes.

32. A method according to any of claims 22 to 31, wherein the input signals are determined based on a difference between the sensor data and the set point data.

33. A method according to any of claims 22 to 32, wherein the plurality of nodes comprises a set of excitatory nodes and a set of inhibitory nodes.

34. A method according to claim 33, wherein a node:receives a first set of inputs from excitatory nodes and a second set of inputs from inhibitory nodes;for each input, selects a larger of the value of the input and a value representative of previous input modified by a decay function, and multiplies the selected value by a weight for the corresponding edge to generate a weighted input;generates a first summation of the weighted inputs corresponding to the first set of inputs;generates a second summation of the weighted inputs corresponding to the second set of inputs;calculates an output in dependence upon the first summation and the second summation; andpropagates the output along the output edges.

35. A method according to claim 34, wherein each output edge is associated with a latency which determines the timing at which the output is propagated along that output edge.

36. A method according to claim 35, wherein the latency for each edge is randomly assigned.

37. A method according to any of claims 34 to 36, wherein the learning rule increases the weights associated with edges within the subset of edges for which the selected value is greater than the corresponding one of the first summation and the second summation.

38. A method according to claim 37, wherein the increase in weight is dependent on a suitability value for that node, wherein the suitability is dependent on the values of the first summation and the second summation.

39. A method according to claim 38, further comprising removing an input edge and inserting a directed edge elsewhere in the artificial neural network in dependence upon an expression indicating the selected value for the input edge is high and the suitability value for the node is low.

40. A method according to any of claims 22 to 39, wherein the plurality of nodes comprises:one or more input nodes respectively configured to receive the one or more input signals and to propagate the one or more input signals to the other nodes of the artificial neural network;one or more output nodes respectively configured to output the one or more output signals; anda plurality of basic nodes interconnected as a network wherein any basic node is connectable to any other basic node by a directed edge, thereby defining a plurality of possible directed edges, and a proportion of the plurality of possible directed edges are assigned.

41. A method according to claim 40, wherein the number of directed edges for each node is determined when the artificial neural network is constructed.A method according to claim 40 or claim 41, wherein the insertion of each of the plurality of directed edges is specified by probabilities to a subset of the plurality of nodes for each source node when the artificial neural network is constructed, such that pre-established relationships between sensors and actuators are emphasized in the resulting connectivity.

42. A control system comprising at least one processor and memory storing instructions that, when implemented by the at least one processor, perform a method as claimed in any preceding claim.

43. A robotic system comprising a system controller according to claim 42.

44. The robotic system of claim 43, wherein the robot is one ofa robotic arm; a biped robot; and a quadruped robot.