Physical environment interaction with equivariant strategies

By introducing a basic weight matrix combination that respects the symmetry of the physical environment in the final layer of the neural network, the problem of large data requirements during training is solved, and the interaction efficiency of the computer control system in the real environment is improved.

CN114467094BActive Publication Date: 2025-09-05ROBERT BOSCH GMBH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080063639.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-11
Filing Date
2020-09-08
Publication Date
2025-09-05
Estimated Expiration
2040-09-08

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively utilizing the symmetry of the environment when training neural networks that control computer systems and interact with the physical environment, resulting in the need for large amounts of data and interaction time, and difficulty in obtaining sufficient observation data in real environments.

Method used

By introducing a linear combination of basic weight matrices in the final layer of the neural network, we ensure that it complies with the symmetry of the physical environment, utilize the symmetry to constrain the determination of action probabilities, reduce the number of parameters and improve training efficiency.

Benefits of technology

Effectively utilizing the symmetry of the environment reduces the amount of observation data and interaction time required for training, and improves the performance efficiency of neural networks in real environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114467094B_ABST
    Figure CN114467094B_ABST
Patent Text Reader

Abstract

The present invention relates to a computer-implemented method (800) for interacting with a physical environment according to a policy. The policy determines a plurality of action probabilities for corresponding actions based on an observable state of the physical environment. The policy comprises a neural network parameterized by a set of parameters. The neural network determines the action probabilities by determining a final layer input from the observable state and applying a final layer of the neural network to the final layer input. The final layer is applied by applying a linear combination of a set of equivariant basis weight matrices to the final layer input. The basis weight matrices are equivariant in the sense that for a set of a plurality of predefined transformations of the final layer input, each transformation results in a corresponding predefined action permutation of the basis weight matrix output for the final layer input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a computer-controlled system for interacting with a physical environment according to a policy and a corresponding computer-implemented method. The present invention also relates to a training system for configuring such a system and a corresponding computer-implemented method. The present invention also relates to a computer-readable medium comprising instructions for performing one of the above methods, parameters of such a policy, and / or weight matrix data underlying such a policy. Background Art

[0002] Applying computer-implemented methods to interact with a physical environment is well known. Typically, sensor data is acquired from one or more sensors (such as cameras, temperature sensors, pressure sensors, etc.); a computer-implemented method is applied to determine an action based on the sensor data; and actuators are used to implement the determined action in the physical environment, for example, moving a robotic arm, activating the steering or braking system of an autonomous vehicle, or controlling the movement of an invasive medical (robotic) tool within a patient's body. The process used to determine the action is often referred to as a strategy for computer-controlled interaction.

[0003] Computer-controlled systems include robotic systems in which a robot (e.g., under the control of an external device or embedded controller) can automatically perform one or more tasks. Other examples of systems that can be computer-controlled are vehicles and their components, household appliances, power tools, manufacturing machines, personal assistants, access control systems, drones, nanorobots, and heating control systems. Various computer-controlled systems can operate autonomously in an environment, for example, autonomous robots, autonomous agents, or intelligent agents.

[0004] Examples of healthcare robotics, particularly in image-guided therapy, include controlling the movement of imaging systems (e.g., X-ray, MRI, ultrasound systems) around a patient, taking into account the patient's anatomy, obstructions, and operating room equipment; robotically guiding intracavitary or extracavitary diagnostic imaging equipment, such as, but not limited to, a bronchoscope in a lung bronchi or an intravascular ultrasound device in a blood vessel; and steering deployable or non-deployable medical tools (e.g., flexible or non-flexible needles, catheters, guidewires, balloons, stents, etc.) toward a target based on an X-ray or ultrasound image or other image to treat and / or measure biophysical parameters. A common example of autonomous computer control in healthcare is the dynamic adjustment of multiple imaging and display parameters (e.g., integration time, contrast) and filters for X-ray or ultrasound based on the current image content.

[0005] While policies can be manually formulated in some cases, it is interesting to note that machine learning models (e.g., neural networks) can also be used as policies. Such machine learning models are typically parameterized by a set of parameters that can be trained for a specific task. In "Proximal Policy Optimization Algorithms" by John Schulman et al. (incorporated herein by reference and available at https: / / arxiv.org / abs / 1707.06347), a method for training such a policy is disclosed. This method comes from the field of reinforcement learning, in which a set of parameters is optimized based on a given reward function. The method alternates between sampling data through interaction with the environment and optimizing parameters (in this case, of a neural network) based on the sampled data. Summary of the Invention

[0006] According to a first aspect of the present invention, a computer-implemented method for interacting with a physical environment according to a policy is provided, as defined in claim 1. According to another aspect of the present invention, a computer-implemented method for configuring a system for interacting with a physical environment according to a policy is provided, as defined in claim 8. According to further aspects, a computer control system for interacting with a physical environment and a training system for configuring such a system are provided, as defined in claims 12 and 13, respectively. According to other aspects of the present invention, a computer-readable medium is provided, as defined in claims 14 and 15.

[0007] As is known per se, in various embodiments a neural network (also referred to as an artificial neural network) is used as a policy for interacting with a physical environment. Such a neural network policy can be parameterized by a set of parameters. The set of parameters can be trained based on interaction data of real and / or simulated interactions. Having been trained to perform a specific task, the policy can be deployed in a system to actually interact with the physical environment according to the set of parameters. During such interaction, an observable state of the environment (such as a camera image) can be input to the policy. Based on this, an action can be selected to be implemented by an actuator (such as a robotic arm) in the environment. The set of actions is typically finite, for example, a robotic arm can be driven left, right, up or down. As is common in reinforcement learning, the policy may be stochastic, in the sense that the policy output may include multiple action probabilities for performing the corresponding action, rather than directly returning a single action to be performed.

[0008] As the inventors have appreciated, in many cases, the physical environment in which an interaction occurs can exhibit a variety of symmetries, with respect to which actions are expected to be favorable under which observation states. For example, in a control system for an autonomous vehicle whose task is to maintain its lane by steering left or right based on a camera image, the desirability of steering left or right can be similarly reversed if the image is flipped horizontally. Various other kinds of symmetries in the observation states are also possible, for example, involving rotations, vertical or diagonal flips, and so on. Multiple sensor measurements of the observable state can be affected differently by symmetry. For example, if the autonomous vehicle uses a camera image and a left / right tilt sensor, horizontal mirror symmetry can correspond to a horizontal flip of the camera image and the negation of the angle measured by the tilt sensor. Possible actions can be similarly affected by the symmetries of the environment in a variety of ways. For example, some actions may be swapped or otherwise permuted, while other actions may not be affected at all by a particular symmetry. Generally, symmetries can be represented by transformations (in many cases linear transformations) on a set of observable states; and by permutations on a set of possible actions. Such symmetries can also be found in healthcare settings by taking into account certain symmetries within the patient's body (e.g., sagittal, frontal, and / or transverse planes and axes, movement of bones around central axes, symmetries between organs (e.g., right and left lungs), etc.) or in the environment of the operating room in an operating room. Such symmetries can be detected directly from measurements or can be found after prior processing of the results of these measurements to retrieve certain symmetries.

[0009] The inventors realized that by incorporating this symmetry into the neural network that calculates the policy, more effective policies can be obtained. For example, fewer parameters may be required to obtain policies of the same quality (e.g., with the same expected cumulative reward). Alternatively, fixing the number of parameters and introducing symmetry can allow policies with higher expected cumulative rewards to be obtained. In the process of training the neural network, data efficiency can be improved, for example, the number of observations required to achieve a policy of a certain quality can be reduced. The latter is particularly important when applying the policy to a physical (e.g., non-simulated) environment. In fact, the amount of data required to learn the policy can be very large. On the other hand, it is generally difficult to obtain a large amount of observation data in a non-simulated environment, for example, interaction time is limited and failure can be accompanied by real-world costs.

[0010] In the field of image classification, the use of symmetries in neural networks is known per se. For example, in “Group Equivariant Convolutional Networks” by T.S. Cohen and M. Welling (available at https: / / arxiv.org / abs / 1602.07576 and incorporated herein by reference), image classification of rotated handwritten digits is demonstrated. By incorporating translation, reflection, and rotation, a neural network is obtained that effectively learns to recognize digits regardless of how they are rotated. However, due to important differences between image classification and reinforcement learning, such known group equivariant convolutional networks are not suitable for determining action probabilities. For example, the final layer of such a neural network outputs the same classification regardless of how the image is rotated. This is undesirable when determining action probabilities because, as discussed above, the desirability of performing an action may change with the transformation of the output. Furthermore, although standard group equivariant convolutional networks typically rely on the invariance of images under translation, this is often not a useful type of symmetry to consider when determining action probabilities because these symmetries are generally not expected to change in a predictable manner in response to translation of the input.

[0011] Interestingly, however, the inventors have proposed a better type of neural network in which the state / action symmetry of the combination of the physical environment at hand can be effectively incorporated. In general, the neural network can include multiple layers. The action probabilities can be determined in the final layer of the neural network based on the final layer input, and thus based on the observable state. Interestingly, in various embodiments, the final layer can be applied to the final layer input by applying a linear combination of a set of carefully defined base weight matrices. The coefficients of this linear combination can be contained in the parameter set of the neural network. The output of applying the linear combination can provide a pre-nonlinear activation for some or all of the action probabilities, from which the action probabilities can be calculated (for example, using softmax). Interestingly, these base weight matrices can be defined to incorporate the state / action symmetry of the combination by being defined as equivariant: for a set of multiple predefined transformations of the final layer input, each transformation results in a corresponding predefined action permutation of the base weight matrix output for the final layer input.

[0012] For example, the transformation of the final layer input can be a matrix R θ The linear transformation represented by and the permutation of the basis weight matrix output can be similarly represented as the matrix P θ In this case, for each environmental symmetry θ to be incorporated, the basis weight matrix W can satisfy the equivariance relation P θ W=WR θ For any final layer input, this in turn means that Pθ Wz=WR θ For example, in this case, the action permutation P is outputted by the original basis weight matrix Wz of the untransformed final layer input Z. θ , transform the final layer input Z to obtain the transformed final layer input R θ z may result in the corresponding basic weight matrix output WR θ z becomes the permutation P θ Wz.

[0013] If each base weight matrix is ​​defined as equivariant, then the linear combination of such base weight matrices can also be equivariant and therefore obey the symmetries in the physical environment. Therefore, the set of base weight matrices can be effectively constrained to provide base weight matrix outputs, from which action probabilities can be determined while obeying the symmetries of the physical environment. Because the symmetry is taken into account, experience from one observed state can be effectively reused to derive actions in other transformed states. Notably, not only can the visual part of the network be reused as in existing convolutional neural networks, but exploration can also be reused, which is particularly important in reinforcement learning due to the sparsity of rewards. Therefore, by obeying the symmetries in the physical environment, fewer parameters may be required to obtain a neural network with a certain expressiveness, providing efficiency improvements in use and training.

[0014] Interestingly, the techniques described herein can be applied without the need to learn symmetries (e.g., from an environment model or shaped reward symmetries): rather, the symmetries can be specified a priori and used as disclosed herein, without the need to specify an overall policy, environment model, or reward shaping. For example, there is no need to infer a model as in many known model-based methods, avoiding the need for an associated complex architecture with many moving parts.

[0015] In more detail, with respect to the action permutations of the base weight matrix, these generally correspond to symmetries of the physical environment. For example, it can be expected that a certain transformation of the observable state of the physical environment (e.g., a horizontal swap of images of the environment) results in a certain permutation of the actions to be performed, e.g., performing a first action in the physical environment represented by the image corresponds to performing a second action in the physical environment represented by the swapped image. The action permutations can be predefined, e.g., manually defined. The action permutations can be obtained as input to training a neural network (e.g., for generating the base weight matrix), but can also be implicit as a restriction on the base weight matrix obtained from an external source.

[0016] Transformations on the final layer inputs: These can reflect symmetries of the physical environment in a variety of ways. Preferably, the final layer inputs are determined from the observable states in an equivariant manner (e.g., given a multi-state transformation of the observable state, each state transformation results in a corresponding transformation on the final layer inputs). These transformations on the final layer can then correspond to action permutations as discussed above. The transformations are typically predefined, e.g., manually defined as part of the neural network design. The transformations can be given explicitly as input to training the neural network (e.g., for generating a set of basis weight matrices), or can be implicit as restrictions on the basis weight matrices obtained from an external source.

[0017] For example, one can expect the physical environment to satisfy a symmetry (e.g., mirror symmetry θ) where observable states are equivariant to action probabilities: every transformation Q of the expected observable state x is θ x corresponds to a permutation P of the action probability y θ y. The final layer input can then be defined to be equivariant to the observable states and action probabilities: the transformation R of the final layer input z can be defined in this way θ , for the observable state x, the corresponding final layer input transformation R θ z is equal to the observable state Q for the transformed θ The final layer input of x. In addition, as discussed above, for the basic weight matrix W, the transformation R θ can correspond to the permutation P as specified above θ , for example, P θ W=WR θ .

[0018] Interestingly, by transforming the observable state, final layer inputs, and actions with corresponding transformations, the overall neural network can also satisfy equivariance, i.e., providing action probabilities that respect the symmetries of the environment corresponding to these transformations and permutations. This can be true regardless of how precisely the previous layers guarantee equivariance, and several possibilities exist for this. Preserving the symmetries of the entire network can lead to particularly efficient learning and accurate results.

[0019] For example, the final layer inputs can be determined from the observable state by applying one or more layers of a known group equivariant neural network (e.g., as disclosed in "Group Equivariant Convolutional Networks"). For example, a known neural network designed for image classification that is invariant to translation and swapping can be used. Transforming the inputs of such a network (e.g., translation or swapping) may result in corresponding transformations of the layer outputs. Such layer outputs can be used as the final layer inputs for a strategy as described herein, where transformations of the internal layer outputs corresponding to the environmental symmetries are used to define the basis weight matrices of the final layer. These environmental symmetries typically do not include translations.

[0020] However, it is not necessary to use a group equivariant neural network as disclosed in "Group Equivariant Convolutional Networks", and in particular, it is not necessary to use translation-preserving internal layers as described in "Group Equivariant Convolutional Networks". Throughout, embodiments are provided. It is not even necessary to use a neural network that explicitly takes into account the symmetries of the environment for determining the final layer inputs, for example, the neural network can be trained using observable states and their transformations as inputs, where the neural network can be stimulated via its loss function to determine final layer inputs that transform in an equivariant manner when applied to the transformed observable states. In fact, using only the final layer as described in this article may already be sufficient to stimulate the neural network to provide final layer inputs that transform according to the observable state transformation.

[0021] Regardless of the exact transformation of the final layer input and the corresponding action permutation, the base weight matrix can be defined in a variety of ways. For example, the base weight matrix can be represented as a matrix, a set of vectors applied as an inner product to the final layer input, etc. Multiple base weight matrices can be derived from a single submatrix (e.g., for one input and / or output channel). As discussed below, the base weight matrices can be predefined, for example, by computer or manual pre-calculation, or calculated when needed. In any case, typically, at least some of the base weight matrices affect multiple outputs of the final layer, thereby reflecting the fact that the symmetry of the environment limits the possible outputs of the neural network. Therefore, the influence of the base weight matrix on multiple outputs can be regarded as a kind of weight sharing between different outputs, through which a reduction in the number of parameters of the neural network can be achieved.

[0022] Although the base weight matrices can cover the entire space of weight matrices that obey transformations and action permutations, interestingly, this is not required: the set of base weight matrices can only cover a subspace of allowed base weight matrices, for example, a randomly sampled subspace. This can allow for a greater reduction in the number of parameters, especially when the number of base weight matrices would otherwise become too large. Thus, equivariance can be maintained and a balance can be achieved between performance and expressiveness of the neural network layer.

[0023] The linear combination of the base weight matrices discussed above can be applied when interacting with the physical environment according to a policy, and when configuring (e.g., training) a system for performing such environmental interaction. In both cases, the final layer of the neural network may require fewer parameters, thereby increasing the efficiency of training and using the neural network.

[0024] Alternatively, the environmental symmetries that define transformations and permutations can form a mathematical group. For example, the set of environmental symmetries can include identity symmetry, closure when composing transformations, associativity, and closure when inverting. These properties are inherent to the symmetry; for example, if a policy should obey a single 90-degree rotation of an image, then the policy should also obey repeated 90-degree rotations and -90-degree rotations. By considering the full group of symmetries, the model can better utilize the available symmetries.

[0025] Alternatively, the sensor data may include an image of the physical environment. In various applications, images provide useful information about the state of the environment, such as traffic conditions for autonomous vehicles, the medical environment for medical tools, intermediate products for manufacturing robots, and the like. Images often exhibit various symmetries corresponding to the action permutations of the actions to be performed by the actuators, such as rotations or mirror images. In such cases, it may be particularly effective to incorporate such state / action-state symmetries using techniques such as those provided herein.

[0026] Optionally, the feature transformation corresponds to a rotation of the image of the physical environment and / or the feature transformation corresponds to a reflection of the image. For example, the reflection can be a mirror image, for example on a central axis of the image or 3D scene. In one embodiment, the feature transformation corresponds to a 180 degree rotation of the image. In one embodiment, the set of feature transformations includes a 90 degree rotation, a 180 degree rotation, and a 270 degree rotation. In one embodiment, the set of transformations includes a horizontal mirror image and a vertical mirror image. Such environmental symmetries occur frequently in practice and are therefore particularly useful in incorporating them into neural networks for policy.

[0027] Optionally, the sensor data may include one or more additional sensor measurements along with the image. Symmetries of the physical environment may affect such measurements in different ways. For example, one or more additional sensor measurements may be invariant to environmental symmetries that do affect the image (e.g., images are swapped while temperature measurements are unaffected). Interestingly, however, one or more additional sensor measurements may also transform along with the input image, e.g., a measurement of an angle to the horizontal may be inverted when the input image is swapped horizontally. Interestingly, such transformations of the additional sensor measurements may also be taken into account by the neural network, allowing such sensor measurements to be effectively used to determine action probabilities.

[0028] Alternatively, in addition to the original linear combination discussed above, applying the final layer may include applying another linear combination of a set of base weight matrices to the final layer input. In addition to the coefficients of the original linear combination, the coefficients of the other linear combination may be contained in a parameter set. For example, the set of possible actions to be performed may include multiple subsets that are each affected by action permutation. For example, actions a1 and a2 may be swapped by action permutation and actions a3 and a4 may be swapped independently. In this case, instead of obtaining a set of base weight matrices for simultaneously calculating a1 to a4, a set of base weight matrices may be obtained, which is first applied to calculate the output of a first subset of actions a1 and a2 and then applied to calculate the output of a second subset of actions a3 and a4 using a different set of parameters. Thus, the set of base weight matrices may be reused, reducing the storage required to store them and, when applicable, also reducing the computational resources required to calculate them.

[0029] Optionally, applying the final layer may further include applying another linear combination of another set of base weight matrices to the final layer input. This other set of base weight matrices may be obtained similarly to the original set of base weight matrices. Interestingly, however, this other set of base weight matrices may be equivariant with respect to the other set of transformations: for another set of multiple predefined transformations of the final layer input, each transformation results in another corresponding predefined action permutation output by the other base weight matrix for the final layer input. Thus, the set of possible actions may include a first subset, the action probabilities of the first subset being determined using the original set of base weight matrices, and a second subset being determined using the other set of base weight matrices. While a single overall set of base weight matrices may also be used to determine the action probabilities of both sets of possible actions, it is interesting to note that using different base weight matrices may be more efficient, as a smaller number of base weight matrices may be sufficient, and the base weight matrices themselves may be smaller. Preferably, the final layer input is equivariant with respect to the original set of transformations and the other set of transformations, as discussed above.

[0030] Optionally, determining the action probability may further include applying a softmax to at least the output of applying a linear combination of the base weight matrices to the final layer input. The linear combination of the base weight matrices may provide a numerical value indicating the relative desirability of performing the corresponding action. By applying a softmax to such relative desirability values, a probability distribution of the actions may be obtained. For example, a softmax may be applied to the output of the linear combination of the base weight matrices and, optionally, to the output of other linear combinations that provide desirability values ​​for other possible actions.

[0031] Optionally, the final layer input may include multiple feature vectors corresponding to corresponding transformations of the observable state. For example, transforming the observable state according to state-action symmetry may result in a permutation of multiple feature vectors. In other words, the transformation of the final layer input may include a permutation of multiple feature vectors and may therefore be estimated relatively efficiently. For example, known group convolutional neural networks typically provide feature vectors of this type. For example, a feature vector may include one feature, at most or at least two features, or at most or at least five features. Similarly, some or all other layers of the neural network may include multiple feature vectors corresponding to corresponding transformations of the observable state.

[0032] Optionally, the features of the final layer input or the input of the previous layer can be determined by average pooling the feature vectors corresponding to the translation of the observable state. For example, the previous layers of the neural network can provide feature vectors corresponding to both the translation of the observable state and another transformation, as disclosed in, for example, "Group Equivariant Convolutional Networks". The feature vectors corresponding to multiple translations and other special transformations can be average pooled to obtain the feature vector for the other transformation, thereby allowing the object recognition capabilities of the translation-equivariant neural network to be used in the previous layers while providing more compressed input to the later layers.

[0033] Optionally, the plurality of action probabilities includes at least one action probability that is invariant under every action permutation, and at least one action probability that varies under some action permutations. For example, two actions (e.g., "move left" and "move right") may be swapped under an action permutation (e.g., corresponding to a mirrored input image), while another action (e.g., "do nothing") may be unaffected by this transformation of the input. Interestingly, the techniques provided herein are powerful enough to express this type of action permutation, and more generally, other types of action permutations that do not have a one-to-one correspondence with the symmetries of the physical environment.

[0034] Alternatively, another base weight matrix for another layer of the neural network may be obtained. Specifically, a set of other base weight matrices for another layer of the neural network may be obtained, wherein transforming the input of the other layer according to a transformation from a set of multiple predefined transformations results in a corresponding predefined transformation of the output of the other base weight matrix for the input of the other layer. To estimate this other layer, a linear combination of the set of other base weight matrices may be applied to the input of the other layer. For example, as discussed above, the transformations of the input of the other layer and the output of the other layer may correspond to a symmetry of a physical environment similar to that of the input of the final layer. By using linear combinations of base weight matrices in other layers of the neural network, and preferably in all layers of the neural network, a better reduction in the set of parameters of the neural network may be achieved.

[0035] Optionally, when training the policy, in other words when optimizing its set of parameters, a set of basis weight matrices for the final layer can be automatically determined from a plurality of predefined transformations and corresponding predefined action permutations. Other sets of basis weight matrices for other layers can also be automatically determined. Although in some cases the set of basis weight matrices can be determined manually, particularly for larger layer sizes and / or a larger number of symmetries, such manual calculations can become infeasible and can be cumbersome to perform multiple times. For larger layer sizes, an approximate set of basis weight matrices can be determined, for example a set of basis weight matrices that provides equivariance but does not necessarily cover the entire set of possible equivariances, thereby providing a further reduction in the number of parameters without affecting equivariance.

[0036] In various embodiments, determining the set of basis weight matrices for the final layer can be formulated as determining for each transformation R of the final layer input θ Each permutation P of the base weight matrix output θ And similarly for other layers, equation P is satisfied θ W=WR θ The weight matrix set W. Equation P θ W=WR θ This results in a linear system of entries in W that can be solved using general techniques. Similar equations can be defined for other layers.

[0037] Alternatively, the base weight matrix can be obtained by obtaining an initial weight matrix W, applying the transformation and inversion of the corresponding action permutation to the initial weight matrix, and adding the transformed and permuted initial weight matrices together. In particular, for the linear transformation P θ 、R θ , the inventors realized that calculating W′=∑ θ P θ -1 WR θ Can provide satisfaction about Pθ 、R θ The equivariance relation P for each θ W=WR θ weight matrix. Thus, a candidate basic weight matrix can be obtained. Optionally, the set of candidate basic weight matrices obtained in this manner can be further improved by orthogonalizing the basic weight matrix (e.g., vectorizing the basic weight matrix, orthogonalizing the vectorized basic weight matrix, and devectorizing the orthogonalized vectorized basic weight matrix). Thus, in a randomized manner, a set of basic weight matrices that provides a good representation of the overall set of basic weight matrices can be obtained.

[0038] Alternatively, a policy gradient algorithm can be used to optimize the set of parameters. A variety of policy gradient techniques (such as the PPO method disclosed in "Proximal Policy Optimization Algorithms" by John Schulman et al.) can be combined with the neural networks provided herein. Due to the incorporation of state / action symmetry in the neural network and the resulting reduction in the number of parameters, the techniques provided herein allow for significant improvements in the data efficiency of the policy gradient algorithm. It should be noted that although PPO is a so-called model-free reinforcement learning technique, the techniques described herein are also applicable to model-free reinforcement learning.

[0039] It will be appreciated by those skilled in the art that two or more of the above-mentioned embodiments, implementations, and / or alternative aspects of the invention may be combined in any way deemed useful.

[0040] Those skilled in the art can perform modifications and variations of any system and / or any computer-readable medium based on this description, which correspond to the described modifications and variations of the corresponding computer-implemented method. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] These and other aspects of the invention will be apparent from and further elucidated with reference to the embodiments described by way of example in the following description and with reference to the accompanying drawings, in which:

[0042] Figure 1 A computer control system for interacting with a physical environment according to a policy is shown;

[0043] Figure 2 A training system for configuring a computer-controlled system that interacts with a physical environment according to a policy is shown;

[0044] Figure 3 A computer control system for interacting with a physical environment (in this case, an autonomous vehicle) is shown;

[0045] Figure 4 A detailed embodiment of a neural network for a strategy for interacting with a physical environment is shown;

[0046] Figure 5a An embodiment of the transformation of observable states is shown;

[0047] Figure 5b One embodiment of the transformation of the final layer input is shown;

[0048] Figure 5c One embodiment of an action permutation of actions to be performed is shown;

[0049] Figure 6 A computer-implemented method for interacting with a physical environment according to a policy is shown;

[0050] Figure 7 A computer-implemented method of configuring a system to interact with a physical environment according to a policy is shown;

[0051] Figure 8 A computer readable medium containing data is shown.

[0052] It should be noted that these figures are merely schematic and not drawn to scale.In the figures, elements corresponding to elements already described may have the same reference numerals. DETAILED DESCRIPTION

[0053] Figure 1 A computer control system 100 is shown for interacting with a physical environment 081 according to a policy. The policy can determine a plurality of action probabilities for corresponding actions based on an observable state of the physical environment 081. The policy can include a neural network parameterized by a parameter set 040. The neural network can determine the action probabilities by determining a final layer input from the observable state and applying a final layer of the neural network to the final layer input. The system 100 can include a data interface 120 and a processor subsystem 140 that can communicate internally via data communication 121. The data interface 120 can be used to access the parameter set of the policy 040. The data interface 120 can also be used to access the base weight matrix data 030 as discussed below. It can be, for example, Figure 2 The system 200 determines the parameter set 042 and / or the base weight matrix data 030 according to the methods described herein.

[0054] The processor subsystem 140 may be configured to access the data 030, 040 during operation of the system 100 and use of the data interface 120. For example, Figure 1As shown in FIG, the data interface 120 may provide access 122 to an external data store 021, which may contain the data 030, 040. Alternatively, the data 030, 040 may be accessed from an internal data store that is part of the system 100. Alternatively, the data 030, 040 may be received from another entity via a network. For example, when configuring the system 100, the data 030, 040 may be received from an internal data store. Figure 2 The system 200 obtains data 030, 040, for example, by performing corresponding environmental interactions multiple times. Generally, the data interface 120 can take various forms, such as a network interface with a local area network or a wide area network (e.g., the Internet), a storage interface with internal or external data storage, etc. The data storage 021 can take any known and suitable form.

[0055] The system 100 may include an image input interface 160 or any other type of input interface for acquiring sensor data 124 from one or more sensors, such as a camera 071, the sensor data 124 indicating an observable state of the physical environment. For example, the camera may be configured to capture image data 124, and the processor subsystem 140 may be configured to determine the observable state based on the image data 124 obtained from the input interface 160 via data communication 123. The input interface may be configured for various types of sensor signals indicating physical quantities of the environment and / or the device 100 itself, and combinations thereof, such as video signals, radar / LiDAR signals, ultrasonic signals, and the like.

[0056] In some embodiments, the sensor may be arranged in the environment 081. In other embodiments, the sensor may be arranged away from the environment 081, for example if one or more quantities can be measured remotely. For example, a camera-based sensor may be arranged outside the environment 081, but may measure quantities associated with the environment, such as the position and / or orientation of a physical entity in the environment. The sensor interface 180 may also access sensor data from elsewhere (e.g., from a data storage device or a network location). The sensor interface 180 may have any suitable form, including but not limited to, for example, a low-level communication interface based on I2C or SPI data communication, and including but not limited to a data storage interface (such as a memory interface or a persistent storage interface), or a personal network interface, a local area network interface, or a wide area network interface (such as a Bluetooth interface, a Zigbee interface, a Wi-Fi interface, an Ethernet interface, or a fiber optic interface). The sensor may be part of the system 100.

[0057] The system 100 may include an actuator interface 180 for providing actuator data to the actuator that causes the actuator to perform an action in the physical environment 081 of the system 100. For example, the processor subsystem 140 may be configured to determine the actuator data based at least in part on the probability of action determined by the strategy as described herein. For example, the strategy may detect an exception (e.g., the risk of a collision) and, based on this, activate a safety system (e.g., a brake). There may also be multiple actuators that perform corresponding actions. The actuator may be an electric actuator, a hydraulic actuator, a pneumatic actuator, a thermal actuator, a magnetic actuator, and / or a mechanical actuator. Specific but non-limiting examples include electric motors, electroactive polymers, hydraulic cylinders, piezoelectric actuators, pneumatic actuators, solenoids, stepper motors, servo mechanisms, etc. The actuator may be part of the system 200.

[0058] The processor subsystem 140 can be configured to, during operation of the system 100 and use of the data interface 120, obtain base weight matrix data 030 representing a set of base weight matrices for a final layer of the neural network, wherein for a set of multiple predefined transformations of the final layer input, each transformation results in a corresponding predefined action permutation of the base weight matrix output for the final layer input. The processor subsystem 140 can be further configured to control interaction with the physical environment by repeatedly: obtaining sensor data indicating an observable state of the physical environment from one or more sensors via the sensor interface 160; determining an action probability; and providing actuator data to an actuator via the actuator interface 180 that causes the actuator to implement an action in the physical environment based on the determined action probability. The processor subsystem 140 can be configured to determine the action probability based on the observable state by applying the final layer of the neural network by applying a linear combination of the set of base weight matrices to the final layer input, the coefficients of the linear combination being contained in the parameter set.

[0059] refer to Figure 3-4 Various details and aspects of the operation of system 100, including optional aspects thereof, will be further elucidated.

[0060] Typically, the system 100 can be embodied as a single device or apparatus (such as a workstation or server (e.g., based on a laptop or desktop computer)) or in a single device or apparatus. The device or apparatus may include one or more microprocessors that execute appropriate software. For example, the processor subsystem can be embodied not only by a single central processing unit (CPU), but also by a combination or system of such a CPU and / or other types of processing units. The software may have been downloaded and / or stored in a corresponding memory, for example, a volatile memory (such as RAM) or a non-volatile memory (such as Flash). Alternatively, the functional units of the system (e.g., the data interface and the processor subsystem) can be implemented in a device or apparatus in the form of programmable logic (e.g., as a field programmable gate array (FPGA) and / or a graphics processing unit (GPU)). Typically, each functional unit of the system can be implemented in the form of a circuit. It should be noted that the system 100 can also be implemented in a distributed manner, for example, involving different devices and apparatuses (such as distributed servers, for example, in the form of cloud computing).

[0061] Figure 2 A training system 200 is shown for configuring a computer-controlled system that interacts with a physical environment according to strategies as described herein. For example, the training system 200 can be used to configure the system 100. The training system 200 and the system 100 can be combined into a single system.

[0062] The training system 200 may include a data interface 220 and a processor subsystem 240 that may communicate internally via data communications 221. The data interface 220 may be used to access a set of parameters for a strategy 040. The data interface 220 may also be used to access base weight matrix data 030 representing a set of base weight matrices for a final layer of a neural network, where for a set of multiple predefined transformations of the final layer input, each transformation results in a corresponding predefined action permutation of the base weight matrix output for the final layer input.

[0063] The processor subsystem 240 may be configured to access the data 030, 040 during operation of the system 200 and use of the data interface 220. For example, Figure 2As shown in FIG, data interface 220 can provide access 222 to external data storage 022, which can contain the data 030, 040. Alternatively, data 030, 040 can be accessed from an internal data storage that is part of system 200. Alternatively, data 030, 040 can be received from another entity via a network. In general, data interface 220 can take a variety of forms, such as a network interface to a local area network or a wide area network (e.g., the Internet), a storage interface to an internal or external data storage, etc. Data storage 022 can take any known and suitable form.

[0064] The processor subsystem 140 can be configured to optimize the policy's parameter set 040 during operation of the system 100 and use of the data interface 120 to maximize the expected reward of interacting with the environment. To optimize the parameter set, the processor subsystem 140 can be configured to repeatedly obtain interaction data indicating a sequence of observed environmental states and corresponding actions performed by the computer control system to be configured, determine a reward for the interaction, determine an action probability that the policy selects the corresponding action in one of the observed states in the sequence of observed environmental states, and adjust the parameter set 040 based on the determined reward and action probability to increase the expected reward. To apply the final layer of the neural network, the processor subsystem 140 can apply a linear combination of the set 030 of base weight matrices to the final layer input, and the coefficients of the linear combination can be included in the parameter set 040.

[0065] System 200 may also include a communication interface (not shown) configured to communicate with the system to be configured (e.g., system 100). For example, system 200 may obtain interaction data of one or more environmental interactions of other systems via the communication interface. The interaction data may be obtained before and / or during the optimization of the parameter set. In the latter case, in order for the other system to interact with the environment according to the current policy, system 200 may also provide the current parameter set 040 of the policy to the other system. Various known types of communication interfaces may be used, for example, arranged for direct communication with other systems 200 (e.g., using USB, IEEE 1394, or similar interfaces); or via a computer network (e.g., a wireless personal area network, the Internet, an intranet, a LAN, a WLAN, etc.). The communication interface may also be an internal communication interface, such as a bus, an API, a storage interface, etc.

[0066] refer to Figure 3-4 Various details and aspects of the operation of system 200, including optional aspects thereof, will be further elucidated.

[0067] Typically, the system 200 can be embodied as a single device or apparatus (such as a workstation or server (e.g., based on a laptop or desktop computer)) or in a single device or apparatus. The device or apparatus may include one or more microprocessors that execute appropriate software. For example, the processor subsystem can be embodied not only by a single central processing unit (CPU), but also by a combination or system of such a CPU and / or other types of processing units. The software may have been downloaded and / or stored in a corresponding memory, for example, a volatile memory (such as RAM) or a non-volatile memory (such as Flash). Alternatively, the functional units of the system (e.g., the data interface and the processor subsystem) can be implemented in a device or apparatus in the form of programmable logic (e.g., as a field programmable gate array (FPGA) and / or a graphics processing unit (GPU)). Typically, each functional unit of the system can be implemented in the form of a circuit. It should be noted that the system 200 can also be implemented in a distributed manner, for example, involving different devices and apparatuses (such as distributed servers, for example, in the form of cloud computing).

[0068] Figure 3 One embodiment is shown above, in which a vehicle control system 300 for controlling a vehicle 62 is shown, the vehicle control system including a system for interacting with a physical environment according to a strategy of one embodiment, e.g. Figure 1 The vehicle 62 may be an autonomous or semi-autonomous vehicle, but this is not required. For example, the system 300 may also be a driver assistance system for a non-autonomous vehicle 62. For example, the vehicle 62 may include an interactive system for controlling the vehicle based on images obtained from the camera 071. For example, the vehicle control system 300 may include a camera interface (not shown separately) for obtaining images of the vehicle's environment 081 from the camera 071.

[0069] The control system 300 may also include an actuator interface (not shown separately) for providing actuator data to the actuator that causes the actuator to perform actions to control the vehicle 62 in the physical environment 081. The vehicle control system 300 can be configured to determine the actuator data for controlling the vehicle based on the action probabilities determined by the policy, and to provide the actuator data to the actuator via the actuator interface. For example, the actuator can be caused to control the steering and / or braking of the vehicle. For example, the control system can control or assist the steering of the vehicle 62 by rotating the wheels 42. For example, to keep the vehicle 62 within the lane, the wheels 42 can be rotated left or right. In this case, the policy can be configured to be equivariant under horizontal swaps, for example, if the camera images 071 are swapped horizontally, the action of rotating the wheels left or right can be replaced.

[0070] Figure 4 A detailed but non-limiting embodiment of a neural network NN, 400 for a strategy for interacting with a physical environment is shown. For example, the neural network NN may be applied to Figure 1 system 100 and / or Figure 2 In the system 200.

[0071] As shown in the figure, the neural network NN can be configured to determine a plurality of action probabilities AP1, 491 to APn, 492 for corresponding actions. A variety of numbers of action probabilities are possible, for example, 2, 3, up to or at least 5, or up to or at least 10. Each action can correspond to a signal to be provided to one or more actuators to perform a specific action (e.g., "move left," "move right," "move up," "move down," etc.). One of the actions can be a no-op, for example, an action that does not affect the physical environment. The neural network NN can be configured to determine the action probabilities in such a way that their sum is 1 (e.g., via a softmax function).

[0072] , the neural network NN can determine the action probability APi based on the observable state OS of the physical environment, 410. The observable state typically includes one or more sensor measurements, for example, an image of the physical environment obtained from a camera and / or measurements of one or more additional sensors. The observable state OS can be represented by a feature vector (for example, of up to or at least 100 features, or up to or at least 1000 features). The observable state OS can include or be based on multiple previous sensor measurements, for example, the observable state OS can include a rolling history of a fixed number of recent sensor measurements, a rolling average of recent sensor measurements, etc. Depending on the application, various types of processing (for example, image processing such as scaling) can also be performed to obtain the observable state OS from the sensor measurements. Typically, the observable state OS can include multiple types of sensor data, audio data, video data, radar data, LiDAR data, ultrasonic data, or multiple individual sensor readings or their historical readings.

[0073] The neural network NN may be parameterized by a parameter set PAR, 440. For example, the parameter set PAR may include coefficients of a base weight matrix for a final layer or for other layers as described herein. For example, the number of layers of the neural network NN may be at least 5 layers or at least 10 layers, and the number of parameters PAR may be at least 1,000 or at least 10,000. From the perspective of training efficiency, it is advantageous to use a neural network NN that is suitable for gradient-based optimization, for example, the parameter set of the neural network is continuous and / or differentiable. Neural networks are also called artificial neural networks.

[0074] In various embodiments, the parameters PAR of the neural network NN can be optimized to maximize the expected reward of interacting with the environment according to the corresponding policy. For example, the expected reward can be the expected cumulative reward defined by modeling the interaction with the environment as a Markov decision process (MDP). Mathematically speaking, an MDP is a tuple in is the space of possible environment states OS, It is the space of possible actions, is the immediate reward function, is the transfer function and γ∈[0, 1] is the discount factor. The policy evaluated by the neural network NN can be defined as in is a probability simplex on the action space, i.e., the set of action probabilities AP1, ..., APn sums to 1. Here, ω represents the set of policy parameters PAR.

[0075] In various embodiments, a neural network NN can be trained to incorporate state / action symmetries of the physical environment in which the interaction occurs. The set of symmetries can be denoted as Θ. The set Θ is often assumed to have a mathematical group structure, for example, Θ can include identity symmetries, combinations of symmetries that can be and are closed when inverted and can be associated, which means that the symmetry and For example, the group of symmetries can be the set of horizontal mirror images {I, H}, where I represents identity and H represents horizontal mirror images; or it can be the set of horizontal and / or vertical mirror images Where I stands for identity, H stands for horizontal mirroring, V stands for vertical mirroring, and Indicates horizontal and vertical mirroring, etc.

[0076] Symmetries usually affect the observable state OS and the set of possible actions that determine the action probabilities AP1, ..., APn. For example, for each symmetry θ, a transformation Q of the observable state can be defined θ For example, transform Q θ The input image contained in the observable state can be rotated or reflected. In general, Q θ is a linear transformation, for example, represented by a matrix. In addition, for each symmetry θ, a permutation P of the action probability APi can be defined θ, for example, also represented by a matrix. The techniques provided herein are powerful enough to support various types of permutations, for example, in some embodiments, the action probabilities include at least one action probability that is invariant under each action permutation, and at least one action probability that varies under some action permutations. Interestingly, in various embodiments, the neural network NN can be configured to be equivariant with respect to these observable state transformations and action permutations, for example, it can be enforced or at least encouraged that transforming the state OS and then computing the action probabilities APi results in the same output as computing the action probabilities APi and then permuting them, for example,

[0077] P θ [π ω ](·|s)=π ω (·|Q θ [s]).

[0078] where Q θ is a transformation of the observable state s and P θ is the corresponding action permutation. Symmetries and corresponding transformations of intermediate layer feature vectors, action probabilities, and observable states are usually defined manually.

[0079] As shown in the figure, the neural network NN can determine the action probability APi by determining the final layer input FLI, 450 from the observable state OS in an operation Ls, 420; and then applying the final layer of the neural network to the final layer input FLI. Preferably, the operation Ls is configured to be equivariant to the symmetry Θ, in the sense that each transformation Q of the observable state OS θ , θ∈Θ leads to the corresponding transformation R of the final layer input FLi θ As noted, however, this need not be strictly enforced, for example, the equivariance may be approximate. Various ways of determining the final layer input FLI are discussed in more detail below.

[0080] Interestingly, the final layer of the neural network NN can be configured to preserve equivariance by using a set of basis weight matrices BWM, 430 that are equivariant with respect to an expected set of symmetries of the physical environment. For example, each transformation R of the final layer input z corresponding to such a symmetry θ z can result in a corresponding action permutation P of the base weight matrix output Wz applied to the final layer input z θ , for example, for the final layer input FLI,z and symmetry θ∈Θ,P θ WZz=WR θz. The parameter set PAR may include coefficients corresponding to each base weight matrix, and the final layer of the neural network NN may be applied by applying a linear combination LC, 460 of the set of base weight matrices BWM with the coefficients given by the parameter set PAR to the final layer input FLI. Interestingly, if the base weight matrices are equivariant, then the linear combination is also equivariant, and accordingly, an equivariant linear combination output LCO, 470 may be obtained. Then, a softmax SMX, 480 may be applied to at least the linear combination output LCO to obtain the action probability APi.

[0081] As an example, the figure shows the basic weight matrix W1, ..., Wk and the corresponding linear combination coefficients C1, ..., Ck. When applying a neural network NN, the basic weight matrix and coefficients are usually fixed. When training a neural network, at least the coefficients are usually trained, for example, to maximize the expected reward of interacting with the environment in terms of the coefficients.

[0082] The set of basis weight matrices BWM can be defined in many ways. For example, mathematically, the equivariance of the weight matrix W applied to the final layer of a neural network NN can be expressed as:

[0083]

[0084] P θ W=WR θ

[0085]

[0086] Therefore, the equivariance of W can be expressed as or in In particular, if P θ is a permutation and R θ is a linear transformation, and we can observe the constraint is linear, and therefore is a linear subspace of the total space of weight matrices. Thus, in some embodiments, the set BWM can be defined as the space In this case, for example, the set BWM can be obtained manually or from the transformation R θ and permutation P θ The determination is made computationally, for example, using known linear algebra techniques.

[0087] The above collection It can also be defined for a single input and / or output channel, in which case the matrix can be obtained from the matrix by applying the matrix to the corresponding input channel and / or applying the matrix to the corresponding output channel. Get the basic weight matrix BWi.

[0088] Interestingly, the set of basis weight matrices (BWMs) need not cover the entire set of equivariant weight matrices. For example, the BWMs can be sampled as a randomly sampled subspace of the space of equivariant weight matrices. In this way, the number of parameters C1,…,Ck can be reduced while still maintaining equivariance, which is particularly important when the space of equivariant weight matrices is relatively large.

[0089] In particular, a method for automatically determining the basic weight matrix BWM is as follows. First, obtain one or more initial weight matrices W i , whose coefficients are randomly sampled, for example, from a univariate Gaussian distribution or similar. The initial weight matrix W can then be obtained by i Determine the basic weight matrix Transform T θ and permutation P θ applied to the initial weight matrix and the results are added together, e.g., to obtain Therefore, effectively, the initial weight matrix W i can be symmetrized to obtain the basic weight matrix. The resulting weight matrix may actually be equivariant, for example, because:

[0090]

[0091] The base weight matrices obtained in this way can be further refined by orthogonalizing or even orthonormalizing the set of base weight matrices obtained. As a result, the weight matrices may become more independent of each other and thus facilitate training. For example, the determined Vectorize to form a vector with the value corresponding to The matrix of rows And calculate the singular value decomposition To perform orthogonalization / normalization. In this case, the set of basis weight matrices BWM can be obtained by devectorizing the columns of V corresponding to the non-zero singular values ​​in ∑. It can be noted that if enough initial weight matrices are used, this process can be used to find a complete basis or a random subspace. For example, the number of initial weight matrices can be at most or at least 100, or at most or at least 250.

[0092] Thus, we have discussed how, by computing linear combinations of basis weight matrices BWM, the linear combination output LCO of the final layer of a neural network can be determined, said linear combination output corresponding to one or more possible actions to be performed.

[0093] Another linear combination of the same set of base weight matrices BWM can also be applied to the final layer input FLI to obtain another linear combination output. In addition to the original linear combination coefficients Ci, the coefficients of the other linear combination can be included in the parameter set PAR. This is particularly attractive if the actions corresponding to the other linear combination outputs should be permuted in the same way as when the environmental symmetry is applied as the original linear combination output LCO. By applying the same base weight matrix twice, it is possible to avoid using a larger base weight matrix for simultaneously calculating two sets of linear combination outputs (which may require the use of more and / or larger base weight matrices).

[0094] Furthermore, alternatively or additionally, another linear combination of another set of base weight matrices may be applied to the final layer input FLI. For another set of multiple predefined transformations of the final layer input, each transformation may result in a corresponding another predefined action permutation of another base weight matrix output for the final layer input. In other words, the action whose probability is determined using this another linear combination may be permuted differently from the action of the original final layer output according to the environmental symmetry. The another set of translations of the final layer input may be equal to the original set of transformations of the final layer input, e.g., the action corresponding to the another linear combination may be equivariant to the same symmetry but according to a different permutation. The another set of translations of the final layer may also be different, e.g., the action corresponding to the another linear combination may be equivariant to a different set of environmental symmetries. In the latter case, the final layer input FLI is preferably equivariant under two sets of symmetries.

[0095] Now proceeding to the calculation of the final layer input FLI based on the definition of the transformation of the observable state OS and the final layer input, various implementations can be envisaged.

[0096] In general, the final layer input FLI can be obtained by applying one or more layers of a group-equivariant convolutional network to the observable state OS. The filters of the group-equivariant convolutional network can be defined as the basis The elements of the linear vector space covered by , where each transformation θ∈Θ has its own basis. The set of filters can be composed of each filter in the basis and the coefficients of each input and output channel Therefore, as described in this paper, the space can be considered as having a base {e i} linear vector space. Any can be described as a linear combination of basis vectors. The filters ω can exist in the span of this basis, in other words, they can be based on {e i}. Therefore, W and ω can be considered as corresponding. Coefficients can be learned and shared between group transformations. Therefore, the filter can be transformed with respect to the basis of the group θ Defined as

[0097]

[0098] Therefore, the filter coefficients can be effectively shared between the basis e(θ) of the transformations θ∈Θ, and thus a transformed version of ω(·) can be obtained instead of a completely new filter for each θ∈Θ. Neural networks can be applied by using these filters as convolutional network filters (e.g., by applying the transformed filter corresponding to each transformation θ∈Θ).

[0099] Specifically, in some embodiments, the final layer input FLI may include a plurality of feature vectors FV1, 451 through ... FVm, 452, each feature vector corresponding to a corresponding transformation of the observable state. For example, the transformation of the final layer input may permute the plurality of feature vectors according to the group structure of the environment symmetry, e.g., if Then the transformation of the final layer input The eigenvector corresponding to symmetry θ2 can be mapped to the eigenvector corresponding to symmetry θ3, etc. This structure of having an eigenvector for each environmental symmetry can be replicated in some or all previous layers of the neural network NN, so that equivariance can be maintained throughout the neural network.

[0100] However, regardless of how exactly the transformations of intermediate layer inputs and outputs are defined, it is also possible to use linear combinations of the basis weight matrices to compute layer outputs at some or all other layers of the neural network. For example, further basis weight matrix data representing a set of further basis weight matrices for another layer of the neural network may be obtained, wherein transforming the further layer inputs according to one transformation from a set of a plurality of predefined transformations results in a corresponding predefined transformation of the further basis weight matrix outputs for the further layer inputs. To evaluate the further layer, linear combinations of the set of further basis weight matrices may be applied to the further layer inputs, again parameterized by the set of parameters PAR. In this case, it is also possible, for example, to find The basis or subspace as described above determines the P given the layer output corresponding to the symmetry of the environment θ and the transformation R of the layer input θ A collection of other basic weight matrices .

[0101] In some embodiments, one or more preceding layers of the neural network NN may be designed to be equivariant not only to state-action symmetry Θ, but also to translations. For example, the observable state may comprise an image of a physical environment, where one or more initial layers of the neural network are equivariant to translations of the image as well as to state-action symmetry. For example, group-equivariant neural network layers of "Group Equivariant Convolutional Networks" may be used. It should be noted that translations of the observable state do not typically induce state-action symmetry, since in many applications they do not result in a permutation of the intended action. Nevertheless, by including translations in the preceding layers of the neural network, they can be used in these preceding layers for object recognition tasks, which are common in, for example, convolutional neural networks. In later layers, the translation symmetry can then be effectively factored by average pooling the translations of the observable state.

[0102] A neural network NN can be used to interact with a physical environment by repeatedly: acquiring sensor data indicating an observable state OS of the physical environment; determining action probabilities APi based on the observable state OS; and providing actuator data to actuators that cause the actuators to implement actions in the physical environment based on the determined action probabilities APi. For example, implemented actions can be sampled based on the action probabilities, or the action with the greatest likelihood can be selected, etc.

[0103] A neural network NN can be trained by repeatedly: acquiring interaction data indicating a sequence of observed environment states and corresponding actions performed by the system; determining a reward for the interaction; determining an action probability APi that the policy selects the corresponding action in one of the observed states in the sequence of observed environment states; and adjusting a set of parameters based on the determined reward and action probability to increase the expected reward. A variety of reinforcement learning techniques known per se can be applied; for example, a policy gradient algorithm, such as that disclosed in “Proximal Policy Optimization Algorithms,” can be used. As is well known, such optimization methods can be heuristic and / or reach local optima. Both on-policy methods, in which interaction data is obtained from interactions with the current set of policy parameters, and off-policy methods, in which this is not the case, can be used. In any case, while optimizing a standard neural network policy traditionally involves updating all filter weights, interestingly, using the techniques presented herein, this can instead involve updating the coefficients Ci of the underlying weight matrix, resulting in faster and more efficient learning.

[0104] Figures 5a-5cA non-limiting example of an action permutation of a set of transformations and action probabilities of observable states and final layer inputs is shown. 11 、z 12 、z 21 、z 22 and the observable state 510 of the 2 by 2 input image with the additional sensor measurement x, e.g., Figure 4 The observable state OS is also shown. The action probabilities π1, π2, π3, π4, π5 (for example, Figure 4 A vector 550 of action probabilities AP1, ..., APn).

[0105] As an example, the physical environment in which the observable state 510 is obtained and the physical environment in which the action 550 is performed can be expected to be equivariant to horizontal and vertical mirroring. In this embodiment, the physical environment includes identity I, horizontal mirroring H, vertical mirroring V, and horizontal plus vertical mirroring. The transformation set can be viewed as a mathematical group, where the operations Taking I as the identity, wait.

[0106] For example, horizontal mirroring of the observable state, represented by arrow 520, can result in transformed state 511. In this transformed state, the image is horizontally mirrored while the additional sensor measurements are negated; for example, x can represent an angle in the vertical plane. Similarly, vertical mirroring of the observable state, represented by arrow 521, can result in transformed state 521, for example, in which the image is vertically mirrored but the sensor measurements remain the same. Observable state 510 can also be mirrored both horizontally and vertically, for example, resulting in transformed state 513.

[0107] In this embodiment, a transformation Θ of the observable state can be expected to result in an action permutation in the action permutation set of the neural network. For example, based on domain knowledge, it can be expected that the set of action probabilities (π1, π2, π3, π4, π5), 550 should be permuted to the action probabilities (π2, π1, π3, π4, π5), 551 under horizontal mirror symmetry 560; in other words, the first action is expected to be equally likely in the original observable state as the second action in the transformed observable state, and vice versa. The other three actions in this embodiment are not affected by horizontal symmetry. Similarly, under vertical symmetry 561, the action probabilities should be permuted to (π1, π2, π4, π3, π5), 552; and under horizontal and vertical symmetry 553, the action probabilities should be permuted to (π2, π1, π4, π3, π5), 553.

[0108] Using the techniques presented herein, this equivariance between transformations of the observable state 510 and the corresponding permutations of the action probabilities 550 output by the neural network can be achieved by computing the final layer input 530 from the observable state 510 in an equivariant manner and then computing the action probabilities 550 from the final layer input 530 in an equivariant manner. To this end, the final layer input 530 can include multiple feature vectors corresponding to the respective transformations Θ. Shown in the figure are the feature vectors y corresponding to the transformations I, H, V, and HV, respectively. I 、y H 、y V and y HV The transformation Θ in this embodiment permutes the final layer input 530 according to the group action of Θ. For example, due to the application of H to I, H, V, Get H, I, V, so the final layer input 530 is transformed by horizontal symmetry 540 to (y I ,y H ,y V ,y HV ) is replaced by (y H ,y I ,y HV ,y V ). Similarly, the final layer input 530 is transformed by vertical symmetry 541 to obtain (y V ,y HV ,y I ,y H ), 532, and transform the final layer input 542 by horizontal and vertical symmetry 542 to obtain (y HV ,y V ,y H ,y I ), 533.

[0109] Thus, as illustrated in this embodiment, given the transformation of the observable state 510 and the action permutation 550, which can be manually defined based on existing domain knowledge of the physical environment, the transformation of the final layer input 530 can be automatically determined. As disclosed herein, given the transformation of the final layer input and the action permutation, a set of basis weight matrices can be automatically determined in which the final layer of the neural network maintains equivariance. Similarly, by computing the final layer input 530 in an equivariant manner (e.g., using additional layers formed in a similar manner), the symmetries of the physical environment can be effectively incorporated into the neural network.

[0110] Figure 6A block diagram of a computer-implemented method 800 for interacting with a physical environment according to a policy is shown. The policy may determine a plurality of action probabilities for corresponding actions based on an observable state of the physical environment. The policy may include a neural network parameterized by a set of parameters. The neural network may determine the action probabilities by determining a final layer input from the observable state and applying a final layer of the neural network to the final layer input. The method 800 may correspond to Figure 1 However, this is not a limitation, as another system, apparatus, or device may also be used to perform the method 800.

[0111] Method 800 may include accessing 810 a set of parameters for a policy in an operation titled "Accessing Policy."

[0112] Method 800 may include, in an operation titled “Obtain Base Weight Matrix Data,” obtaining 820 base weight matrix data representing a set of base weight matrices for a final layer of the neural network, wherein for a set of multiple predefined transformations of the final layer input, each transformation results in a corresponding predefined action permutation of the base weight matrix output for the final layer input.

[0113] Method 800 may include controlling 830 an interaction with a physical environment in an operation titled "Control Interaction." To control the interaction, operation 830 may include repeatedly:

[0114] - In an operation entitled "Obtain Sensor Data", sensor data indicative of an observable state of a physical environment is obtained 832 from one or more sensors;

[0115] - in an operation entitled “DETERMINE ACTION PROBABILITY,” determining 834 the action probability based on the observable state comprises applying a final layer of the neural network by applying a linear combination of a set of basis weight matrices to the final layer input, the coefficients of the linear combination being contained in the set of parameters;

[0116] - In an operation entitled "Provide Actuator Data", actuator data is provided 836 to the actuator that causes the actuator to implement an action in the physical environment based on the determined action probability.

[0117] Figure 7 A block diagram of a computer-implemented method 900 for configuring a system to interact with a physical environment according to a policy is shown. For example, the system may use Figure 8 Method 800. Method 900 may include optimizing a set of parameters of a policy to maximize an expected reward from interacting with an environment according to the policy by repeatedly:

[0118] - In an operation entitled "Obtain Interaction Data", obtaining 910 interaction data indicative of a sequence of observed environmental states and corresponding actions performed by the system;

[0119] - in an operation entitled "Determine Reward", determining 920 a reward for the interaction;

[0120] - in an operation entitled “Determine Action Probability,” determining 930 an action probability that the policy selects a corresponding action in one observed state in the sequence of observed environment states, comprising applying a final layer of a neural network by applying a linear combination of a set of basis weight matrices to the final layer input, the coefficients of the linear combination being contained in the set of parameters;

[0121] - In an operation entitled "Adjust Parameters", a set of parameters is adjusted 940 based on the determined reward and action probabilities to increase the expected reward.

[0122] It should be understood that, generally, Figure 6 Method 800 and Figure 7 The operations of method 900 may be performed in any suitable order (eg, serially, concurrently, or a combination thereof), respecting a particular order dictated by, for example, input / output relationships, where applicable.

[0123] The method can be implemented on a computer as a computer-implemented method, as dedicated hardware, or as a combination of the two. Figure 8 As illustrated in FIG, instructions for a computer (e.g., executable code) can be stored on a computer-readable medium 1000, for example, in the form of a series of machine-readable physical marks 1010 and / or as a series of elements having different electrical (e.g., magnetic or optical) properties or values. The executable code can be stored in a transient or non-transitory manner. Examples of computer-readable media include memory devices, optical storage devices, integrated circuits, servers, online software, etc. FIG11 shows an optical disc 1000.

[0124] Alternatively or in addition, the computer-readable medium 1000 may include transient or non-transitory data 1010 representing a set of parameters for a strategy for interacting with a physical environment as described herein, the strategy determining a plurality of action probabilities for corresponding actions based on an observable state of the physical environment, the strategy comprising a neural network that determines the action probabilities by determining a final layer input from the observable state and applying a final layer of the neural network to the final layer input, the final layer of the neural network being applied by applying a linear combination of a set of basis weight matrices to the final layer input, the coefficients of the linear combination being contained in the set of parameters.

[0125] Alternatively or in addition, the computer-readable medium 1000 may include transient or non-transitory data 1010 representing basis weight matrix data, wherein the basis weight matrix data represents a set of basis weight matrices for a policy for interacting with a physical environment as described herein, the policy determining a plurality of action probabilities for corresponding actions based on an observable state of the physical environment, the policy comprising a neural network that determines the action probabilities by determining a final layer input from the observable state and applying a final layer of the neural network to the final layer input, the final layer of the neural network being applied by applying a linear combination of the set of basis weight matrices to the final layer input.

[0126] Examples, embodiments or optional features, whether or not indicated as non-limiting, should not be construed as limiting the invention as claimed.

[0127] It should be noted that the above-described embodiments illustrate rather than limit the present invention, and that those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between brackets should not be construed as limiting the claims. The use of the verb "comprise" and its conjugations does not exclude the presence of elements or stages other than those described in the claims. The article "a" or "an" before an element does not exclude the presence of a plurality of such elements. Expressions such as "at least one" before a list of elements or a group of elements represent the selection of all or any subset of elements from the list or group. For example, the expression "at least one of A, B, and C" should be understood to include: only A; only B; only C; both A and B; both A and C; both B and C; or all A, B, and C. The present invention can be implemented by hardware comprising several different elements and by a suitably programmed computer. In a device claim enumerating several means, several of these means may be embodied by the same item of hardware. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

Claims

1. A computer-implemented method (800) for interacting with a physical environment according to a policy, the policy determining a plurality of action probabilities for corresponding actions based on an observable state of the physical environment, wherein the policy comprises a neural network parameterized by a set of parameters, the neural network determining the action probabilities by determining a final layer input from the observable state and applying a final layer of the neural network to the final layer input, the method comprising: - accessing (810) a set of parameters of the policy; - obtaining (820) base weight matrix data representing a set of base weight matrices for said final layer of said neural network, wherein for a set of a plurality of predefined transformations of said final layer input, each transformation results in a corresponding predefined action permutation of a base weight matrix output for said final layer input; - controlling (830) interaction with the physical environment by repeatedly: - obtaining (832) sensor data indicative of an observable state of the physical environment from one or more sensors; - determining (834) the action probability based on the observable state, comprising applying a final layer of the neural network by applying a linear combination of the set of basis weight matrices to the final layer input, the coefficients of the linear combination being contained in the set of parameters; - providing (836) to an actuator actuator data that causes the actuator to effect an action in the physical environment based on the determined action probability.

2. The method (800) of claim 1, wherein the sensor data comprises an image of the physical environment.

3. The method (800) according to claim 2, wherein the feature transformation corresponds to a rotation and / or a characteristic transformation corresponds to a reflection.

4. The method (800) of claim 2, wherein the sensor data additionally comprises one or more additional sensor measurements.

5. The method (800) according to any one of claims 1-4, wherein applying the final layer further comprises applying another linear combination of the set of basis weight matrices to the final layer input, the coefficients of the another linear combination being contained in the set of parameters.

6. A method (800) according to any one of claims 1-4, wherein applying the final layer further comprises applying another linear combination of another set of basis weight matrices to the final layer input, wherein for another set of multiple predefined transformations of the final layer input, each transformation results in a corresponding another predefined action permutation of another basis weight matrix output for the final layer input.

7. A method (800) according to any one of claims 1-4, wherein a layer input of a layer of the neural network includes a plurality of feature vectors corresponding to respective transformations of the observable state, and features of the layer input are determined by average pooling the feature vectors corresponding to the translations of the observable state.

8. A computer-implemented method (900) for configuring a system for interacting with a physical environment according to a policy using the method of any one of claims 1 to 4, comprising optimizing a set of parameters of the policy to maximize an expected reward for interacting with the environment according to the policy, the optimizing being performed by repeatedly: - obtaining (910) interaction data indicating a sequence of observed environmental states and corresponding actions performed by the system; - determining (920) a reward for said interaction; - determining (930) an action probability that the policy selects the corresponding action in one of the observed states of the environment, comprising applying a final layer of the neural network by applying a linear combination of the set of basis weight matrices to a final layer input, the coefficients of the linear combination being contained in the set of parameters; - adjusting (940) the set of parameters based on the determined reward and action probabilities to increase the expected reward.

9. The method (900) of claim 8, comprising obtaining the set of base weight matrices by determining the set of base weight matrices from the plurality of predefined transformations and corresponding predefined action permutations.

10. The method (900) of claim 9, comprising determining a base weight matrix by obtaining an initial weight matrix, applying a transformation and inversion of a corresponding action permutation to the initial weight matrix, and adding together the transformed and permuted initial weight matrices.

11. The method (900) of claim 10, further comprising orthogonalizing the determined set of basis weight matrices.

12. A computer control system (100) for interacting with a physical environment (081) according to a policy, the policy determining a plurality of action probabilities for corresponding actions based on an observable state of the physical environment, wherein the policy comprises a neural network parameterized by a set of parameters, the neural network determining the action probabilities by determining a final layer input from the observable state and applying a final layer of the neural network to the final layer input, the system comprising: - a data interface (120) for accessing the set of parameters of said strategy (040); - a sensor interface (160) for obtaining sensor data indicative of an observable state of the physical environment from one or more sensors; - an actuator interface (180) for providing actuator data to an actuator causing said actuator to perform an action in said physical environment; a processor subsystem (140) configured to obtain base weight matrix data representing a set of base weight matrices for a final layer of the neural network, wherein for a set of a plurality of predefined transformations of the final layer input, each transformation results in a corresponding predefined action permutation of the base weight matrix output for the final layer input; and the processor subsystem is configured to control interaction with the physical environment by repeatedly: - obtaining sensor data (124) indicative of an observable state of the physical environment from the one or more sensors via the sensor interface; - determining the action probability based on the observable state, comprising applying a final layer of the neural network by applying a linear combination of the set of basis weight matrices to the final layer input, the coefficients of the linear combination being contained in the set of parameters; - providing actuator data (126) to the actuator via the actuator interface that causes the actuator to effect an action in the physical environment based on the determined action probability.

13. A training system (200) for configuring a computer control system that interacts with a physical environment according to a strategy using the method of any one of claims 1 to 4, the training system comprising: - a data interface (220) for accessing a set of parameters of the strategy (040) and base weight matrix data (030) representing a set of base weight matrices for a final layer of the neural network, wherein for a set of a plurality of predefined transformations of the final layer input, each transformation results in a corresponding predefined action permutation of the base weight matrix for the final layer input; a processor subsystem (240) configured to optimize a set of parameters of the policy to maximize an expected reward from interacting with the environment according to the policy, the optimization being performed by repeatedly: - obtaining interaction data indicative of a sequence of observed environmental states and corresponding action permutations executed by said computer control system; - determining a reward for said interaction; - determining an action probability that the policy selects a corresponding action in one of the observed states of the environment, comprising applying a final layer of the neural network by applying a linear combination of the set of basis weight matrices to the final layer input, the coefficients of the linear combination being contained in the set of parameters; - adjusting the set of parameters based on the determined reward and action probabilities to increase the expected reward.

14. A computer-readable medium (1000) comprising transient or non-transitory data (1010), the data representing one or more of: - instructions which, when executed by a processor system, cause the processor system to perform the computer-implemented method according to claim 1; - a parameter set of a policy for interacting with a physical environment as claimed in claim 1, said policy determining a plurality of action probabilities for corresponding actions based on an observable state of the physical environment, said policy comprising a neural network, said neural network determining said action probabilities by determining a final layer input from the observable state and applying a final layer of said neural network to said final layer input, said final layer of said neural network being applied by applying a linear combination of a set of basis weight matrices to said final layer input, said coefficients of said linear combination being contained in said parameter set; - Basic weight matrix data representing a set of basic weight matrices of a strategy for interacting with a physical environment as mentioned in claim 1, the strategy determining a plurality of action probabilities of corresponding actions based on an observable state of the physical environment, the strategy comprising a neural network that determines the action probabilities by determining a final layer input from the observable state and applying a final layer of the neural network to the final layer input, the final layer of the neural network being applied by applying a linear combination of the set of basic weight matrices to the final layer input.

15. A computer-readable medium (1000) comprising transitory or non-transitory data (1010), the data representing instructions which, when executed by a processor system, cause the processor system to perform the computer-implemented method of claim 8.

Citation Information

Patent Citations

  • Reinforcement learning with auxiliary tasks

    CN110114783A

  • Reinforcement learning with auxiliary tasks

    US20190258938A1