Generation of rotationally invariant or -covariant descriptors of point configurations
The system generates rotationally invariant or covariant descriptors by transforming feature vectors into moment matrices, addressing the insufficient geometric information in existing models and enhancing the predictive capabilities of machine learning models for three-dimensional configurations.
Patent Information
- Application Number
- DE202025104173
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-12-11
- Estimated Expiration
- 2035-07-31
AI Technical Summary
Existing machine learning models struggle to accurately predict properties of three-dimensional point configurations due to insufficient inclusion of high-level geometric information in rotationally invariant or covariant descriptors, leading to reduced performance in tasks such as atomistic simulations.
A system and method for generating rotationally invariant or covariant descriptors by transforming feature vectors into moment matrices, which are then combined to determine descriptors, avoiding the need for Clebsch-Gordan operations and enabling efficient computation using hardware accelerators like GPUs or TPUs.
This approach allows for the efficient generation of descriptors that capture high-level geometric information, improving the performance of machine learning models in predicting properties of three-dimensional configurations, particularly in molecular dynamics simulations.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
GENERAL STATE OF THE ART
[0001] This specification relates to the generation of rotationally invariant or -covariant descriptors of point configurations for processing by machine learning models.
[0002] Machine learning models receive an input and generate an output, such as a predicted output, based on the received input. Some machine learning models are parametric and generate the output based on the received input and the values of the model's parameters.
[0003] Some machine learning models are deep models that use multiple layers of models to produce an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers, each of which applies a nonlinear transformation to a received input to produce an output. SUMMARY
[0004] This specification generally describes a system and a method, implemented as computer programs on one or more computers at one or more locations, for generating rotationally invariant or covariant descriptors of a three-dimensional configuration of points. The points can, for example, correspond to atoms, in which case the invariant or covariant descriptors can be used by machine learning models to predict physical properties of configurations of the atoms, such as energies or forces in atomic or molecular systems.
[0005] Invariant or covariant descriptors can be used to uniquely characterize configurations of points, independent of the orientation of the coordinate system used to define the points' locations. In atomic / molecular systems, examples of rotationally invariant descriptors include bond lengths, which contain two-body information, and bond angles, which contain three-body and sometimes higher-level information. Generally, to uniquely characterize different configurations of points, the descriptors for each configuration must contain sufficiently "high-level" geometric information about the configuration; that is, the descriptors must be functions of the (relative) positions of many points. For example, n-body information is needed, where n > 2, 3, 4, etc.Descriptors can be characterized in terms of a "field order," where polynomials of order d are denoted as having field order d + 1 (this convention includes a point (e.g., an atom) at the origin of the coordinate system). Failure to include such information in the descriptors can reduce the performance of machine learning models that use the descriptors to predict properties of the point configuration, as in atomistic simulations.
[0006] The invention is set out in claim 1; further aspects are set out in the dependent claims.
[0007] In one aspect, a system is provided herein comprising one or more computers and one or more storage devices that store instructions ready to execute, induce the one or more computers to generate rotationally invariant or covariant descriptors of a three-dimensional configuration of points. The instructions are configured to control the one or more computers to perform a procedure that involves using the coordinates of the points to determine a plurality of feature vectors. Each feature vector has a corresponding degree (l) and comprises a respective plurality of features (which may be called "fundamental features"). Each feature is determined using a spherical harmonic. (Ylm) the degree (l) of the feature vector and a respective order (m) is determined by linearly combining (e.g. summing) values of the spherical harmonic, which are evaluated at the respective coordinates of the points.
[0008] The procedure performed by the system further includes transforming each of the plurality of feature vectors into a corresponding moment matrix (M). a,b,l Each moment matrix 112 corresponds to a respective irreducible representation of the 3D rotation group in a direct summation representation of a tensor product of irreducible representations of the 3D rotation group. The procedure performed by the system further includes the use of the moment matrices to determine one or more invariant or covariant descriptors of the three-dimensional configuration of the points.
[0009] Conventionally, descriptors expressed using irreducible representations of the 3D rotation group can be determined using a Clebsch-Gordan operation. This involves determining the tensor (outer) product of two feature vectors and then forming linear combinations defined by Clebsch-Gordan coefficients that include elements of the tensor product to determine the elements of the irreducible representation. Such a process can be referred to as "coupling" the feature vectors to encode higher-order geometric information in the irreducible representations.
[0010] For example, the tensor product can be derived from a first feature vector (a) of degree a with a respective second feature vector (b) of degree b (where a and b are integers), in which case the tensor product a⊗b a matrix of the form (2a + 1) × (2b + 1).
[0011] A direct sum representation of the tensor product can then denote a direct sum of irreducible representations (which can be referred to as "irreps") that are considered a⊗b=(|a−b|)⊕⋯⊕(a+b) can be designated, where each of the irreducible representations (|a−b|),⋯,(a+b) The form (2a + 1) × (2b + 1) and the corresponding degree l = |a - b|, |a - b| + 1, ... , a + b. Every irreducible representation can be determined from linear combinations of the elements of the tensor product using Clebsch-Gordan coefficients. Thus, irreducible representations encoding higher-order geometric information can be obtained from feature vectors that include only 2-body information (i.e., dependent on particle coordinates relative to an origin).
[0012] In contrast, the present system can avoid the need to form the tensor product of the feature vectors by determining irreducible representations as moment matrices from a single feature vector and then combining (e.g., multiplying) different moment matrices to obtain the feature vectors that encode higher-order geometric information about the 3D configuration of points.
[0013] Each moment matrix (M a,b,l ) can be generated using a corresponding feature vector of the same degree l as the moment matrix, e.g., using a feature vector with respective features of degree m, given by ∑riYlm(ri), where r i the coordinates (e.g. x-, y-, z-coordinates) of a point designated with an index i and the sum goes over all points.
[0014] In some implementations, transforming each of the multiple feature vectors into a corresponding moment matrix involves determining elements of the moment matrix using respective linear combinations of the features of the feature vector. For example, each of the linear combinations might include a sum of each of the features weighted by a respective coefficient that depends on the degree (l) of the corresponding feature vector and the respective order (m) of the feature. Each coefficient could, for example, be a Clebsch-Gordan coefficient.
[0015] As an example, an element at position (m1, m2) in the moment matrix (M) a,b,l ) are obtained from a sum over the features of the corresponding feature vector, in which each of the features of the respective order m is assigned a respective Clebsch-Gordan coefficient. Ca,m1,b,m2l,m Clebsch-Gordan coefficients are weighted. They can be calculated using standard formulas and / or programming libraries. See, for example, "E3x:E(3)-Equivariant Deep Learning Made Easy" arXiv:2401.07595. As used herein, the term Clebsch-Gordan coefficient also includes other analogous coefficients, such as Wigner-3-j symbols and so on, e.g., those that use a different phase convention or scaling factor.
[0016] In general, a descriptor can be a numerical value or a collection of numerical values (e.g., an ordered collection of numerical values) that characterizes a configuration of points, such as the configuration of atoms / molecules in a chemical system. A rotationally invariant descriptor can refer to a scalar value determined from the coordinates of the points, which remains unchanged when the coordinates are expressed in different coordinate systems related by a rotation, e.g., an element of the 3D rotation group, SO(3). A rotationally covariant descriptor can be a vector or tensor determined from the coordinates of the points, which rotates "in the same way" as the coordinate system when the coordinate system is rotated. Invariant descriptors can, for example, be determined by the scalar product (dot product) of two covariant vectors.
[0017] In some implementations, using the moment matrices to determine one or more invariant or covariant descriptors involves multiplying two or more of the moment matrices and a vector that comprises a linear combination of one or more of the feature vectors to obtain a corresponding invariant or covariant descriptor.
[0018] In some implementations, using moment matrices to determine one or more invariant or covariant descriptors involves determining one or more combined moment matrices. Each combined moment matrix is determined using a specific linear combination of moment matrices that have the same form as the combined moment matrix. For example, the moment matrices that have the same form as the combined moment matrix can include moment matrices of different degrees (l). For example, the moment matrices that have the same form as the combined moment matrix can include at least one moment matrix of each degree (l) in a range from |a - bl to (a + b), where the specific form of the combined moment matrix is (2a + 1) × (2b + 1). The linear combination can be defined by learnable parameters, e.g.,Parameters that are adjusted during the training of a machine learning model.
[0019] The system can then use the combined moment matrices to determine the one or more invariant or covariant descriptors.
[0020] As an example, the one or more combined moment matrices may comprise one or more square matrices. Using the moment matrices to determine the one or more invariant or covariant descriptors may then involve determining one or more invariant descriptors using respective traces of the one or more square matrices, for example, using linear combinations of the respective traces of the one or more square matrices.
[0021] In some implementations, using moment matrices to determine one or more invariant or covariant descriptors of the three-dimensional configuration of the points involves determining a plurality of moment block matrices. Each moment block matrix can comprise a plurality of blocks (submatrices), each containing one of the moment matrices or combinations thereof. A moment block matrix can alternatively be referred to as a "matrix of matrices."
[0022] The system can then multiply the multitude of moment block matrices and a feature block matrix comprising a linear combination of one or more of the feature vectors to obtain a corresponding invariant or covariant descriptor. Using moment block matrices can enable the efficient determination of invariant or covariant descriptors using one or more hardware accelerators optimized for matrix multiplication (i.e., of large matrices), such as graphics processing units (GPUs) or tensor processing units (TPUs). For example, by multiplying fewer, larger moment block matrices, the one or more hardware accelerators can determine invariant or covariant descriptors more quickly (e.g., by multiplying the invariant or covariant descriptors).measured in real time), compared to separately multiplying products of the moment matrices or combined moment matrices to determine the invariant or covariant descriptors.
[0023] In some implementations, each moment block matrix can be updated ("shifted") by adding the identity matrix (e.g., a scaled identity matrix) to the moment block matrix. This approach can enable more efficient training (e.g., determining weighting parameters for linear combinations of feature vectors or moment matrices), even for large numbers of moment block matrices, similar to skip connections in ResNets.
[0024] Corresponding blocks of each of the moment block matrices can have the same respective shapes. In other words, each of the moment block matrices can have the same block structure. In some implementations, each of the moment block matrices is a square matrix. Thus, products of the moment block matrices remain square and fit into the allocated memory.
[0025] In some implementations, the feature block matrix comprises a linear combination of a plurality of feature vectors of the same degree (l).
[0026] In some implementations, the feature block matrix comprises a concatenation of feature blocks. Each feature block can contain one or more rows or columns, each containing a linear combination of feature vectors of the same degree (l). For example, weighting parameters defining each linear combination can be learnable parameters of a machine learning model, such as a neural network.
[0027] In some implementations, multiplying the plurality of moment block matrices and the feature block matrix comprising one or more of the feature vectors to obtain a corresponding invariant or covariant descriptor may involve: determining another feature block matrix comprising a linear combination of a plurality of feature vectors of the same degree (l); and multiplying (i) one or more covariant descriptors obtained by multiplying the plurality of moment block matrices and the feature block matrix, and (ii) the further feature block matrix (or its transpose) to obtain a corresponding plurality of invariant descriptors.
[0028] In some implementations, the use of moment matrices to determine one or more invariant or covariant descriptors involves: determining (i) one or more linear combinations of the feature vectors; and / or (ii) one or more linear combinations of the moment matrices, where each linear combination is determined using an appropriate set of feature or moment matrix weighting parameters (e.g., learnable parameters).
[0029] Optionally, the system can further adjust the feature or moment matrix weighting parameters to optimize an objective function that depends on one or more invariant or covariant descriptors.
[0030] In some implementations, the procedure performed by the system further includes: determining one or more linear combinations of the invariant or covariant descriptors, with each linear combination being determined using an appropriate set of descriptor weighting parameters (e.g., learnable parameters), and adjusting the descriptor weighting parameters to optimize an objective function that depends on the one or more linear combinations of the invariant or covariant descriptors.
[0031] In general, the procedure performed by the system may further involve processing one or more invariant or covariant descriptors using a machine learning model to generate a corresponding model output. For example, the model output may be indicative of one or more physical properties of the point configuration.
[0032] In some implementations, optimizing the objective function may involve: obtaining a variety of training data items. Each training data item may include (a) a training input comprising the coordinates of a particular configuration of points, and (b) a target output comprising one or more physical properties of the configuration of points. The procedure performed by the system may, for each of the training data items, include: determining one or more invariant or covariant descriptors for the configuration of points in the training input; and processing the one or more invariant or covariant descriptors using the machine learning model to produce a corresponding model output for the configuration of points in the training input. The objective function may consist of a comparison (e.g.,Differences between the model outputs and the corresponding target outputs depend on this. Therefore, the generation of the descriptors can be optimized to improve the machine learning model's performance. In general, any suitable target (or "loss") function can be used, depending on the model and target outputs being compared, e.g., a least-squares target function, a cross-entropy loss function, etc.
[0033] In some implementations, the descriptors can be generated by a different machine learning model (e.g., a neural network) that is trained end-to-end with the machine learning model. For example, the end-to-end training of the two models might involve backpropagating gradients of the objective function with respect to the learnable parameters using the respective neural networks of the models.
[0034] In some implementations, the configuration of points can correspond to a configuration of atoms or molecules, with each point corresponding to a respective atom or molecule in the configuration of atoms or molecules.
[0035] In implementations, descriptors can include radial functions (e.g., radial basis functions) to contain both radial and angular information. For example, Chebyshev polynomials can be used as radial functions, such as Chebyshev polynomials of the logarithm of the radius.
[0036] In some examples, the procedure performed by the system can be carried out on a multitude of real subsets of the atoms to generate, for each subset, one or more rotationally invariant or covariant descriptors of the respective configuration of atoms in the subset. For example, each subset can consist of atoms of the same respective type (e.g., a first subset consisting of carbon atoms, a second subset consisting of hydrogen atoms, and so on). In some implementations, the procedure performed by the system can involve multiplying moment matrices determined using feature vectors of different subsets of the atoms.
[0037] In some implementations, the descriptors can include 116 radial functions (e.g., radial basis functions) to contain both radial and angular information. For example, Chebyshev polynomials can be used as radial functions, such as the Chebyshev polynomial of the logarithm of the radius.
[0038] Special embodiments of the item described in this specification may be implemented in such a way as to realize one or more of the following advantages.
[0039] The systems and procedures described in this specification can be used to obtain rotationally invariant or covariant descriptors of spatial configurations (3D arrangements) of points in a more computationally efficient way than existing approaches that use Clebsch-Gordan operations (“coupling”) to generate descriptors from feature vectors (higher degree) obtained from tensor products of feature vectors (lower degree).
[0040] For example, the calculation of descriptors by performing Clebsch-Gordan operations with feature vectors whose respective degrees cover a wide range, e.g. from 0 to a maximum value (l), scales poorly, e.g. with a computational complexity of O(l). 6 In contrast, the methods described in this specification (“matrix multiplication”) can scale more favorably, e.g., with a computational complexity of O(l). 3In particular, the present disclosure makes it possible to obtain effective descriptors without having to perform all the operations that would be required to couple feature vectors using the (complete) Clebsch-Gordan approach.
[0041] The improved scaling is particularly advantageous for generating descriptors of chemical systems (e.g., comprising atoms, ions, and / or molecules) for use in molecular dynamics calculations, which typically require the generation of large numbers of descriptors over a large number of time steps, such as molecular dynamics simulations of condensed-phase systems. In this way, the presented systems and methods can enable a more efficient use of computing resources.
[0042] As mentioned above, the procedures performed by the system, as described in this specification, can also be efficiently implemented using block matrices and hardware accelerators such as GPUs or TPUs.
[0043] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 shows an example of a system for predicting molecular properties. Fig. Figure 2 shows body orders that are required to distinguish between exemplary molecular configurations. Fig. Figure 3 shows an example set of moment matrices. Fig. Figure 4 shows an example of how the multiplication of moment matrices can be used to obtain invariant or covariant descriptors. Fig. Figure 5 shows an example moment block matrix. Fig. Figure 6 is a flowchart of an example process for generating rotationally invariant or -covariant descriptors of a three-dimensional configuration of points.
[0044] Identical reference numbers and designations in the different drawings indicate identical elements. DETAILED DESCRIPTION
[0045] Fig. Figure 1 shows an example of a molecular property prediction system 100. The molecular property prediction system 100 is an example of a system implemented as computer programs on one or more computers at one or more locations, in which the systems, components, and techniques described below are implemented.
[0046] System 100 is configured to process a chemical structure 102, which defines a three-dimensional configuration of atoms or molecules, to generate values of one or more predicted properties 104 of the chemical structure. As used herein, a chemical structure 102 can refer to any data defining any three-dimensional configuration of atoms or molecules. In general, the chemical structure 102 is received as, or converted into, a three-dimensional set of points defining the locations of each atom or molecule in the chemical structure. Each point in the set of points can be assigned one of a variety of “colors” γ, each color corresponding to a different point type, such as an atom or molecule type. Mathematically expressed, each point r i ∈ ℝ 3 . is of color γ i ∈ ℝ 3. C (the set of colors) and i = 1, ..., n (the number of points), where the points of the same color are a respective subset S γ the set of points S. It is assumed that references to color serve only for easier explanation and that the term "color" is merely intended to represent designations for different sets of particles (e.g., atoms or molecules).
[0047] System 100 includes a feature generator 106 configured to process the coordinates of points in the chemical structure 102 to determine a variety of feature vectors 108. Each feature vector has a corresponding degree (l) and comprises one or more features. Each feature is determined using a spherical harmonic function. (Ylm) the degree (l) of the feature vector and a respective order (m) are determined by linearly combining values of the spherical harmonic, which are evaluated at the respective coordinates of the points.
[0048] The system 100 further comprises a moment matrix generator 110, configured to transform each of the plurality of feature vectors 108 into a corresponding moment matrix 112. Each moment matrix 112 corresponds to a respective irreducible representation of the 3D rotation group, SO(3), in a direct sum representation of a tensor product of irreducible representations of the 3D rotation group. Each moment matrix 112 can be determined by the Clebsch-Gordon relation H(a)⊗H(b)≅H(|a−b|)⊕H(|a−b|+1)⊕⋯⊕H(a+b) "backwards" to identify features in H(|a−b|)⊕H(|a−b|+1)⊕⋯⊕H(a+b) to be encoded as a (2a + 1) × (2b + 1) matrix, where each H(l) an irreducible (2l + 1)-dimensional (real) representation of the 3D rotation group.
[0049] For example, every moment matrix (for the atoms / molecules with the color γ) can be defined by the formula: Ma,b,l(γ)=la,b,l∑r∈SγYl(r) where M a,b,l a (2a + 1) × (2b + 1) moment matrix 112 is and l a,b,l a transformation that is applied to the feature vector (i.e., an embedding), i.e. la,b,l:H(l)→Mat2a+1,2b+1 is a mapping from the feature vector to the moment matrix 112.
[0050] Fig. Figure 3 illustrates the moment matrices 312A-C for the case a = b = 1, i.e., 3 × 3 matrices. The moment matrices 312A-CM 1,1,l correspond to the respective irreducible representations H(0),H(1),H(2) the 3D rotation group in a direct summation representation of a tensor product H(a=1)⊗H(b=1) of irreducible representations of the 3D rotation group. The moment matrix 312A M 1,1,0 , the H(0) This corresponds to a diagonal matrix formed from a single value derived from the feature values, the moment matrix 312B M 1,1,1 , the H(1) This corresponds to an antisymmetric matrix formed from three values derived from the feature values, and the moment matrix 312C M 1,1,2 , the H(2) This corresponds to a traceless symmetric matrix formed from five values derived from the feature values.
[0051] Explicit formulas for the moment matrices 312A-CM 1,1,l can be expressed as follows (where, for the sake of simplicity, the formulas are shown for a single point with a corresponding (x, y, z) coordinate): M1,1,0=(100010001) M1,1,1=(0−zyz0−x−yx0) M1,1,2=(2x2−y2−z23xyxzxy2y2−x2−z23yzxzyz2z2−x2−y23)
[0052] For a = b = 2, i.e., 5x5 matrices, the corresponding moment matrices M are 2,2,l : M2,2,0=(1000001000001000001000001) M2,2,1=(02xz−y0−2x0yz0−z−y0x−y3y−z−x03z003y−z30) M2,2,2=(−2x2+y2+z203xy3xz−23yz0−2x2+y2+z2−3xz3xy3(z2−y2)3xy−3xzx 2−2y2+z23yz3xz3xz3xy3yzx2+y2−2z23xy−23yz3(z2−y2)3xz3xy2x2−y2−z2)
[0053] The moment matrices M 2,2,i are antisymmetric for odd i and can be considered Di−DiT are written, while for even i the moment matrices M 2,2,i are symmetrical and considered Di+diag(di)+DiT can be written, with upper triangular matrices D i and diagonal matrices with entries d i Using r 2 = x 2 + y 2 +z 2 , the remaining a = b = 2 moment matrices are 112: D2,2,3=(03r2x−5x310z3−6r2z6x2y−4y3+6yz253(xz2−xy2)00−6x2y−y3+9yz 2−6x2z+9y2z+z3103xyz00010x3−6r2x3(r2y−5x2y)0000−3(r2z−5x2z)00000) D2,2,4=(070yz(y2−z2)−20xy(r2−7z2)−20xz(r2−7y2)−103(r2−7x2)0010xz(2x2+9y2−5z2)−10xy(2 x2−5y2+9z2)53(6x2(y2−z2)−y4+z4)000−20yz(r2−7x2)103xz(7x2−3r2)0000103xy(7x2−3r2)00000) d4=(4x4−12x2y2−12x2z2−16y4+108y2z2−16z44x4−12x2y2−12x2z2−19y4−102y2z2+19z4−16x4−12x2y2−108 x2z2+4y4−12y2z2−16z4−16x4+108x2y2−12x2z2−16y4−12y2z2+4z424x4−72x2y2−72x2z2+9y4+18y2z2+9z4)
[0054] In general, an element can be located at position (m1, m2) in the moment matrix M a,b,l from a sum over the features of the corresponding feature vector, in which each of the features of the respective order m is assigned a respective Clebsch-Gordan coefficient Ca,m1,b,m2l,m The Clebsch-Gordan coefficients are weighted. They can be calculated using standard formulas and / or programming libraries. See, for example, "E3x:E(3)-Equivariant Deep Learning Made Easy" arXiv:2401.07595.
[0055] The system 100 further includes a descriptor generator 114 which is configured to process the moment matrices 112 in order to generate one or more invariant or covariant descriptors 116 of the three-dimensional configuration of the points.
[0056] The descriptor generator 114 can generate the descriptors 116 by multiplying the moment matrices 112. For example, two or more of the moment matrices can be used to multiply a vector that comprises a linear combination of one or more of the feature vectors to obtain invariant or covariant descriptors.
[0057] Fig. Figure 2 shows examples of 3D configurations of chemical structures 1a, 1b, 2a, 2b, 3a, 3b that cannot be distinguished by descriptors constructed from 2-body information, 3-body information, etc. From the perspective of the central black atom in each chemical structure, the surrounding chemical environments in examples 1a and 1b appear identical when only 2-body information (e.g., distances) is considered, but are easily distinguished by 3-body information (e.g., angles), whereas examples 2a and 2b require 4-body information (e.g., dihedral angles). The chemical structures in examples 3a and 3b even require higher-order information to distinguish them. System 100 can be used to generate higher-order descriptors 116 for distinguishing between different chemical structures without the computational cost associated with existing methods.For example, structures 1a and 1b can be distinguished using a respective invariant descriptor, which is calculated by. Tr(M12) is obtained, i.e. from the trace of a moment matrix M1 (correspondingly H(1) ) squared. Examples of invariant descriptors that similarly apply to the different chemical structures in Fig. The two results that were determined are listed in the following table. invariant 1a 1b 2a 2b 3a 3b Tr(M0) 4.47 4.47 8.94 8.94 15.7 15.7 Tr(M12) 0.0 -2.0 -1.067 -1.067 -1.859 -1.859 Tr(M1 M2 M3) 0.0 -0.321 -0.675 0.626 1.251 1.251 Tr(M1 M3 M42) 0.0 -0.15 -0.685 0.033 1.686 1.556
[0058] Ideal descriptors are unique, computationally efficient, and covariant. System 100 can enable the practical construction of demonstrably complete systems of features with these desired properties, valid for all 3D point configurations.
[0059] Fig. Figure 4 illustrates an example of how the multiplication of moment matrices 112 can be used to obtain invariant or covariant descriptors. In this example, the descriptors 416 are generated by a product 400 of moment matrices 112 that are in Fig. 4 each by the corresponding irreducible representation H(l) are characterized by the 3D rotation group. In general, the matrix product 400 can be expressed as follows: Mam−1,am,lm(γm)⋅…⋅Ma1,a2,l2(γ2)⋅M0,a1,l1(γ1) where l1 = a1 and |a1 - a2| ≤ l2 ≤ a1 + a2, ..., |a m-1 - a m | ≤ l m ≤ a m-1 + a m , and which leads to descriptors 416, the covariant a m × 1 matrices are, i.e., vectors in H(am), given by polynomials of degree l1 + ... + l m The calculation of descriptors 416 requires O(m⋅a3) Steps for an upper bound a ≥ a i .
[0060] Back to Fig. 1, the property prediction system 100 further comprises a property prediction machine learning model 118 configured to process a model input comprising the descriptors 116 to produce a model output comprising respective predicted values 104 of one or more properties of the chemical structure 102.
[0061] Any suitable machine learning model for trait prediction 118 can be used. For example, the machine learning model for trait prediction 118 can include a neural network trained to process descriptors to generate values of one or more predicted traits. The neural network can include any suitable types of neural network layers (e.g., fully connected layers, attention layers, convolutional layers, recurrent layers, etc.) in any suitable number (e.g., 3 layers, or 10 layers, or 50 layers) and connected in any suitable configuration (e.g., as a directed graph of layers). In a particular example, the neural network can have a transformer architecture, such as that described in Ashish Vaswani et al., “Attention is all you need,” Advances in Neural Information Processing Systems 30 (NIPS 2017).
[0062] The one or more physical properties may include, for example, one or more of the following: an energy of the configuration of atoms; a force acting on one or more of the atoms; an optical or spectroscopic property of the configuration of atoms; an electrical, magnetic, or electromagnetic property of the configuration of atoms; an acoustic or mechanical property of the configuration of atoms. Some examples of such physical properties include an energy of formation (such as free energy or enthalpy), a band gap energy, a conductivity or other electrical property (including, for example, superconducting properties such as the superconducting transition temperature); a magnetic property (e.g., permeability, Curie temperature); a mechanical property (e.g., bulk modulus, elastic modulus, density, ductility, strength, hardness); and a phase transition property (e.g.,a phase transition temperature, such as a melting or boiling point) and so on.
[0063] The machine learning model for property prediction 118 can, for example, be trained to predict the values 104 of one or more properties using a supervised learning procedure. For example, the machine learning model for property prediction 118 can be trained using a variety of training data elements, each comprising a training input containing a respective set of descriptors for a 3D configuration of atoms or molecules, determined using moment matrices 112 as described above, and a respective target output comprising target values of the one or more properties for the 3D configuration of atoms or molecules. The target values can, for example, be determined for the 3D configuration of atoms or molecules by performing a respective (e.g., ab initio) quantum chemical calculation for the configuration of atoms, e.g.,A density functional theory (DFT) or a coupled-cluster (e.g., CCSD) calculation can be used. Alternatively or additionally, the target values can be obtained from experimental data, such as spectroscopic or thermochemical experimental data.
[0064] Fig. Figure 5 shows an example moment block matrix 500 (or “matrix of matrices”), which comprises a multitude of blocks, each containing one of the moment matrices 112 or a combination of moment matrices. This matrix can be applied to column vectors from H(l1)⊕H(l2)⊕⋯⊕H(lr) can be used to generate covariant descriptors.
[0065] For many practical applications, it is important to organize matrix multiplications efficiently. Especially when using hardware accelerators such as GPUs / TPUs with hardware support for matrix multiplication, it can be more cost-effective to perform calculations with a few large matrices, such as the moment block matrix 500, rather than with many small matrices.
[0066] The moment block matrix 500 can be constructed by linear combinations of M a,b,l (γ) for l = |a - b|, ..., a + b are used to fill respective (2a + 1) × (2b + 1) matrices, and then by replacing r × r of these matrices for a, b in {l1, l2, ..., l r} into a larger square matrix of side length (2l1 + 1) + ... + (2l r + 1) are packed to form the moment block matrix 500. Traces of the square matrices correspond to components in H(0).
[0067] Multiplying k - 1 such block matrices 500 produces a matrix composed of covariant descriptors of field order k. This matrix can then be multiplied by n1 column vectors from H(l1)⊕H(l2)⊕⋯⊕H(lr) These methods can be used to generate covariant vectors of field order k + 1. Scalar descriptors (invariants) can be obtained by taking scalar products of the n1 column vectors in H(l1) with n2 new covariants of the field order k.
[0068] This approach can be integrated into machine learning architectures to achieve significant efficiency gains compared to existing techniques based on Clebsch-Gordon operations. H(l1)⊗H(l2)→H(l3). Some implementations may use a deep neural network instead of a linear combination of invariants, or employ nonlinear activation functions to modify the matrices obtained in intermediate steps. For existing architectures that use multiple layers of Clebsch-Gordon operations, activation functions are restricted to functions of the scalar channel (since a transcendental function cannot be applied to a vector). However, any analytic function applied to (2l + 1) × (2l + 1) matrices generated by the present procedure (not element-wise, but defined, for example, by a Taylor series for matrices) is also a covariant operation. This property means that analytic functions such as matrix exponentiation can be integrated into neural network architectures that process the descriptors.
[0069] More generally, since the present methods can use the descriptors to generate a complete representation of 3D configurations of points, the descriptors can make it possible to distinguish all types of (e.g. molecular) configurations in a computationally efficient manner.
[0070] An exemplary algorithm for determining invariant descriptors follows. The inputs to the algorithm include respective coordinates (points), such as for each of a multitude of atoms in a chemical system. Each point can be assigned one of a multitude of "colors" γ (e.g., corresponding to the atom type). In mathematical notation, each point r i ∈ ℝ 3 is of color γ i ∈ ℝ 3 . C (the set of colors) and i = 1, ..., n (the number of points).
[0071] Hyperparameters for the algorithm include: the number of vectors for scalar products (n vec); the number of matrix products (n Mat ); integers 0 ≤ l1 ≤ ... ≤ l r , the matrix pages 2l i + 1 of submatrices correspond to; for i = 1...n Mat : integers b i ≥ 0 (corresponding to the body orders b) i +2).
[0072] The algorithm includes: Step 1: Spherical harmonics: For l = 0, ... ,2l: for each color y, determine Y l (γ) = Σ r∈Sγ Y l (r), where S γ the set of points for the color γ. Step 2: Vectors: For i = 1,...,r, calculate 2 · n Mat · n vec linear combinations of Y li (γ). Step 3: Matrices: For (a, b) ∈ ℝ 3 . {l1, ...,l r} 2 Calculate b1, ... , b nMat Matrices of the form a × b by linear combinations of l a,b,l (Y l (γ)) for l = |a - b|, ..., a + b. They to b1, ..., b nMatassemble square matrices. Step 4: Products: For i = 1, ..., n Mat mount n vec Column vectors from step 2 into a matrix V, use matrices from step 3 to calculate products W = M1 · M bi · to calculate V. Take all scalar products of irreducible parts of columns of W with vectors from step 2, which nMat⋅r⋅nvec2 This results in invariants that can be output as a linear combination.
[0073] The algorithm can be implemented, for example, in the Python programming language using the E3x library (“E3x:E(3)-Equivariant Deep Learning Made Easy” arXiv:2401.07595), as shown in the following code listing, which for simplicity is for a matrix product and a color:
[0074] Fig. Figure 6 is a flowchart of an example process 600 for generating rotationally invariant or rotationally covariant descriptors of a three-dimensional configuration of points. For simplicity, process 600 is described as being performed by a system of one or more computers at one or more locations. For example, a property prediction system, e.g., the property prediction system 100, may be used. Fig. 1, which is programmed according to this specification, perform process 600.
[0075] The system uses (step 602) the coordinates of the points to determine a multitude of feature vectors. Each feature vector has a corresponding degree (l) and comprises one or more features. Each feature is defined using a spherical harmonic. (Ylm) the degree (l) of the feature vector and a respective order (m) are determined by linearly combining values of the spherical harmonic, which are evaluated at the respective coordinates of the points.
[0076] The system then transforms (step 604) each of a plurality of feature vectors into a corresponding moment matrix, M. a,b,l Each moment matrix corresponds to a respective irreducible representation H(l) of the 3D rotation group in a direct summation representation H(|a−b|)⊕H(|a−b|+1)⊕⋯⊕H(|a+b|) a tensor product of irreducible representations (H(a)⊗H(b)) the 3D rotation group.
[0077] The system then uses (step 606) the moment matrices to determine one or more invariant or covariant descriptors of the three-dimensional configuration of the points.
[0078] The systems and procedures described in this specification can be used in a wide variety of different areas.
[0079] For example, the predicted properties 104 of the chemical structure 102, determined by the system 100, can be used for a variety of chemical or biochemical applications. For instance, some applications might involve selecting a chemical structure from a multitude of candidate chemical structures based on the respective predicted properties 104 generated by the system 100 for each of the candidate chemical structures. For example, the selection of the chemical structure might involve performing respective molecular dynamics calculations for each of the candidate chemical structures. At each of a multitude of time steps following an initial time step, the coordinates of the atoms in the candidate chemical structure are updated based on one or more model outputs generated by the machine learning model at one or more previous time steps.For example, the machine learning model can be trained to predict a particular force on each of the atoms and to use the forces to update the coordinates of the atoms.
[0080] As a specific example, the selected chemical structure could be a drug or a ligand of an industrial enzyme. The selection of a chemical structure from the multitude of candidate chemical structures may involve evaluating the interaction of each candidate chemical structure with a target molecule and / or a target molecule's binding site (e.g., a protein or a nucleic acid such as DNA or RNA). For example, the target molecule and / or the target molecule's binding site could be a receptor or an enzyme, and the selected chemical structure could be an agonist or antagonist of the receptor or enzyme.
[0081] A "ligand" can refer to a molecule or other compound capable of binding to a target molecule, such as a protein, and forming a complex. Ligands can include, for example, small organic compounds, macromolecules, and so on. A ligand can associate or interact (e.g., through chemical bonds, hydrogen bonds, van der Waals forces, hydrophobic interactions, electrostatic interactions, etc.) to form a common structure with the target molecule. In some implementations, the ligands may include low-molecular-weight complex ligands, such as organic compounds with a molecular weight of <900 Daltons. In some other implementations, the candidate ligands may include polypeptide ligands, i.e., defined by an amino acid sequence. In some implementations, the ligand is a polypeptide ligand, a polynucleoside ligand, or a polynucleotide ligand.
[0082] In some implementations, the interaction evaluation may include assessing the binding of the candidate ligand to the structure of the target molecule (such as a biological molecule). For example, the interaction evaluation may involve identifying a ligand that binds with sufficient affinity to exert a biological effect. In other implementations, the interaction evaluation may include assessing an association between the candidate ligand and the target molecule that has an effect on a function of the target molecule, such as an enzyme. The evaluation may include assessing the affinity between the candidate ligand and the target molecule or complex, or assessing the selectivity of the interaction. Candidate ligands may be selected based on those with the highest affinity.The evaluation of the interaction can additionally include the simulation of dynamic behavior of the ligand and the target molecule, for example through molecular dynamics simulations, which can allow the consideration of kinetic aspects of the interaction.
[0083] The evaluation of a candidate ligand's interaction with the target molecule can be performed using a computer-aided approach, where graphical models of the candidate ligand and target molecule structures are displayed for user manipulation, and / or the evaluation can be performed partially or fully automatically, for example, using standard molecular (e.g., protein-ligand) docking software. In some implementations, the evaluation may include determining an interaction score for the candidate ligand, where the interaction score provides a measure of the interaction between the candidate ligand and the target molecule. The interaction score may depend on the strength and / or specificity of the interaction, e.g., a score that depends on the free energy of the binding. A candidate ligand can then be selected based on its score.
[0084] In some implementations, the target molecule comprises a receptor or an enzyme, and the ligand is an agonist or antagonist of the receptor or enzyme. In some implementations, the method can be used to identify the structure of a cell surface marker. This can then be used to identify a ligand, such as an antibody, an aptamer, or a label like a fluorescent label, that binds to the cell surface marker. This can be used for the identification and / or treatment of cancer cells.
[0085] In some implementations, the ligand is a drug, and the interaction of each of a variety of target molecules (such as a biological molecule) with each of the candidate ligands is evaluated. Then, one or more of the candidate ligands can be selected, either to obtain a ligand that interacts (functionally) with each of the target molecules or to obtain a ligand that interacts (functionally) with only one of the target molecules. For example, in some implementations, it may be desirable to obtain a drug that is effective against multiple drug targets. Alternatively, it may also be desirable to screen a drug for off-target effects. For example, in agriculture, it may be useful to determine that a drug developed for one plant species does not interact with another, different plant species and / or animal species.
[0086] As another example, the selected chemical structure could be a drug, and the selection of the chemical structure from the multitude of candidate chemical structures could involve: evaluating the interaction of each candidate chemical structure with each of a multitude of target molecules and / or binding sites of target molecules in order to either (i) obtain a chemical structure that interacts with each of the target molecules and / or binding sites, or (ii) obtain a chemical structure that interacts with only one of the target molecules and / or binding sites of the target molecules.
[0087] As another example, the selected chemical structure may be a catalyst for an industrial chemical process, and the one or more physical properties may include one or more measures or predictors of the catalytic activity of the selected chemical structure for the industrial chemical process (e.g. rate coefficients, bond energies, bond lifetimes, diffusion constants, etc.).
[0088] In some examples, the selected chemical structure (and / or any of the candidate chemical structures) may be an inorganic compound (e.g., a ceramic, a superconductor, an organometallic compound, etc.) or an alloy.
[0089] In some implementations, the ligand can be synthesized. The biological activity of the ligand can then be tested in vitro and / or in vivo. For example, the ligand can be tested for ADME (absorption, distribution, metabolism, excretion) and / or toxicological properties to eliminate unsuitable ligands. Testing may involve, for example, contacting the candidate small molecule, polypeptide, or polynucleotide ligand with the target molecule (e.g., protein) and measuring any change in the expression or activity of the target molecule.
[0090] Although the systems and methods in this specification are generally described in terms of predicting the properties of chemical structures, it is recognized that analogous systems can be used for many other applications. For example, in other implementations, each of the points can correspond to a specific location in an environment (e.g., a real-world environment), and the method further includes determining the coordinates of each of the points from one or more images of the environment, e.g., obtained using one or more image sensors. For example, each image can comprise a plurality of pixels, with each pixel having an associated set of one or more pixel values. Determining the coordinates of each of the points from one or more images of the environment can involve processing the pixel values.For example, each set of one or more pixel values can include a corresponding depth value for that pixel. Thus, in some examples, the one or more images can comprise one or more point clouds of the environment.
[0091] Descriptors determined using the system and procedure can be used for a wide variety of applications, including scene classification, object detection or localization, action or gesture recognition, semantic separation, image generation, control of a mechanical agent or robot, and so on. For example, the descriptors can be used to predict or simulate future properties of the environment, such as future configurations of the environment.
[0092] For example, some implementations use the machine learning model to perform one or more of the following tasks: (i) a classification task in which the model output classifies the environment into one or more categories; (ii) an object detection or localization task in which the model output is indicative of whether one or more objects are present in the environment, or the model output defines coordinates of a respective region for one or more objects in the environment; (iii) an action or gesture recognition task in which the model output is indicative of whether one or more actions or gestures are being performed in the environment; (iv) a semantic segmentation task in which the model output assigns spatial locations in the environment to respective segmentation categories; (v) a keypoint detection task in which the model output includes coordinates of one or more keypoints in the environment and (vi) an environment similarity task in which the model output is indicative of a similarity of the environment to one or more predefined environments.
[0093] In some implementations, the environment can be a real-world environment, and the agent comprises (i) a mechanical agent or robot that interacts with the real-world environment to perform the specified task, or (ii) an electronic agent that controls equipment in the real-world environment to perform the specified task. The coordinates of the points can be obtained, for example, by processing sensor data acquired by one or more sensors in the real-world environment, such as one or more sensors on the mechanical agent or robot.
[0094] The process may then involve providing instructions or control signals to the mechanical or electronic agent. In some examples, the process may involve processing sensor data acquired by one or more sensors in the real-world environment (e.g., sensors of the mechanical agent or robot) to determine descriptors that characterize the environment based on coordinates of points extracted from the sensor data. In some implementations, the descriptors may be provided as input to a machine learning policy model (e.g., a policy neural network), which processes the updated representation according to the machine learning policy model's parameters to determine an output that includes an action to be performed by the mechanical or electronic agent.
[0095] In some examples, the descriptors can be processed to determine an expected return for one or more actions, e.g., using an action-value model for machine learning (e.g., a Q-function neural network).
[0096] As a more specific example, descriptors can be used to provide input for the control system of a mechanical agent, such as a robot or vehicle, operating in a real-world environment. The control system can provide output that directs the robot's or vehicle's operation to perform a task, such as manipulating an object in the environment or moving within it. Descriptors can be used, for example, to identify objects for the robot to manipulate, or obstacles or paths for the mechanical agent to move along. The control system can then use these descriptors, for example, to make decisions about how to perform a task assigned to the robot, or to control the direction or speed of the agent's movement.
[0097] The agent can be a mechanical agent, such as a robot or vehicle, controlled to perform actions in the real environment in response to observations, in order to accomplish the task, for example, manipulating an object or navigating the environment. Thus, the agent can be, for example, a real or simulated robot; as some other examples, the agent can be a control system to operate one or more machines or pieces of equipment in an industrial plant.
[0098] In cases where the environment is virtual, the descriptors can be used to train the agent to perform the task in the virtual environment. The trained agent can then be used to perform the task in a real-world environment. This process can be called "sim-to-real" training.
[0099] The real (physical) environment can be any suitable environment, such as a real physical environment, e.g., a manufacturing environment, a warehouse environment, a road environment, and so on. The physical environment can include one or more agents that interact with the environment, e.g., robot agents, vehicles, and so on. The physical environment can include a variety of objects, e.g., tools, packages, mechanical parts, electrical parts, walls, floors, road surfaces, conveyor belt surfaces, and so on. In some implementations, the systems and procedures described above can be used for real-world control, especially for optimal control tasks, e.g., to assist a robot in manipulating a deformable or rigid object. Thus, the physical environment can include a real-world environment containing a physical object, e.g., an object to be picked up or manipulated.The descriptors can be used to generate a representation of the physical environment at the current time step or at one or more other (e.g., previous) time steps, and can, for example, define a representation of the shape or configuration of the physical object at that time step. The representation of the physical environment at the new time step can define a predicted representation of the shape or configuration of the physical object, for example, when it is subjected to a force or deformation, such as from a robot actuator. The method can further involve controlling the robot using the predicted representation to manipulate the physical object, for example, by...Using the actuator, the robot moves toward a target location, shape, or configuration of the physical object by being controlled to optimize an objective function that depends on a difference between the predicted representation and the target location, shape, or configuration of the physical object. Controlling the robot can involve providing control signals to the robot based on the predicted representation to cause the robot to perform actions, such as using an actuator of the robot to manipulate the physical object to perform a task. For example, this can involve controlling the robot, such as its actuator, using a reinforcement learning process with a reward that is at least partially based on a value of the objective function, to learn to perform a task that involves manipulating the physical object.
[0100] In some implementations of agent control, the environment is a real-world environment, and the agent is a mechanical agent that interacts with the real-world environment, such as a robot or an autonomous or semi-autonomous land, air, or sea vehicle that operates in or navigates through the environment. The actions are actions performed by the mechanical agent in the real-world environment to accomplish the task. For example, the agent might be a robot or other mechanical agent that interacts with the environment to perform a specific task, such as locating or manipulating an object of interest in the environment, moving an object of interest to a specific location in the environment, or navigating to a specific destination in the environment. In these implementations, the observations might include, for example,One or more of the following are included: images, object position data, and sensor data to capture observations as the agent interacts with the environment. The actions can define control signals for controlling the robot or other mechanical agent, such as positions, torques, or other control signals for the parts of the mechanical agent, or higher-level control commands.
[0101] In some examples, the points can be keypoints identified in one or more images, which might, for example, define landmarks of an object depicted in the image. Thus, the descriptors can be used to generate a representation that characterizes an object, at least partially, based on its keypoints. This representation can then be processed by a machine learning model to perform a machine learning task, such as classifying the object according to one or more predefined classes, estimating its 3D pose, and so on.
[0102] This specification uses the term "configured" in the context of systems and computer program components. For a system consisting of one or more computers to be configured to perform certain operations or actions, this means that software, firmware, hardware, or a combination thereof is installed on the system which, when operating, causes the system to perform the operations or actions. For one or more computer programs to be configured to perform certain operations or actions, this means that the one or more programs contain instructions which, when executed by a data processing apparatus, cause the apparatus to perform the operations or actions.
[0103] Embodiments of the subject matter and the functional operations described in this specification may be implemented in digital electronic circuitry, in physically embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of these. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a physical non-volatile storage medium for execution by or control of the operation of a data processing apparatus.The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access storage device, or a combination of one or more of these. Alternatively or additionally, the program instructions can be encoded on an artificially generated, propagating signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is produced to encode information for transmission to a suitable receiving apparatus for execution by a data processing apparatus.
[0104] The term "data processing apparatus" refers to data processing hardware and encompasses all types of apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. The apparatus may also include, or additionally comprise, specialized logic circuits, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit). Optionally, the apparatus may include, in addition to hardware, code that creates an execution environment for computer programs, such as code representing processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these. Thus, a system, an artificial neural network, or a trained artificial neural network, as described herein, can be implemented in hardware using electronic circuits, for example,in a physical box. Similarly, computer code described herein may be code to emulate such hardware, or code for a hardware description language.
[0105] A computer program, which may also be called or described as a program, software, software application, app, module, software module, script, or code, may be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it may be provided in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computer environment. A program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file containing other programs or data, such as one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in several coordinated files, such as...Files that store one or more modules, subroutines, or parts of code. A computer program can be deployed to run on one computer or on multiple computers located at one site or distributed across multiple sites and connected by a data communication network.
[0106] In this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine is implemented as one or more software modules or components installed on one or more computers at one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines may be installed and run on the same computer or computers.
[0107] The processes and logic flows described in this specification can be executed by one or more programmable computers running one or more computer programs to perform functions by operating on input data and producing an output. The processes and logic flows can also be executed by dedicated logic circuitry, such as an FPGA or an ASIC, or by a combination of dedicated logic circuitry and one or more programmable computers.
[0108] Computers capable of running a computer program may be based on general-purpose or specialized microprocessors, or both, or on any other type of central processing unit. Generally, a central processing unit receives instructions and data from read-only memory, random-access memory, or both. The elements of a computer are a central processing unit for carrying out or executing instructions and one or more storage devices for storing instructions and data. The central processing unit and memory may be augmented by, or integrated into, specialized logic circuitry. Generally, a computer will also include, or be operationally coupled to, one or more mass storage devices for storing data, or both, such as magnetic disks, magneto-optical disks, or optical disks.However, a computer does not necessarily have to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a portable audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name just a few.
[0109] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM and flash memory devices; magnetic disks such as internal hard disks or removable media; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0110] To enable interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer that has a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for showing information to the user, and a keyboard and pointing device, such as a mouse or trackball, with which the user can input information into the computer. Other types of devices can also be used for interaction with a user; for example, the feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including acoustic, verbal, or tactile input.Furthermore, a computer can interact with a user by sending and receiving documents to and from a device used by the user; for example, by sending web pages to a web browser on a user's device in response to requests received by the web browser. A computer can also interact with a user by sending text messages or other forms of messages to a personal device, such as a smartphone running a messaging application, and receiving reply messages from the user.
[0111] Data processing equipment for implementing machine learning models may, for example, include special hardware accelerator units for processing common and computationally intensive parts of machine learning training or production (i.e., inference) workloads.
[0112] Machine learning models can be implemented and deployed using a machine learning framework, such as a TensorFlow framework, a Microsoft Cognitive Toolkit framework, an Apache Singa framework, or an Apache MXNet framework.
[0113] Implementations of the subject matter described in this specification may be implemented in a computer system comprising a back-end component, such as a data server; a middleware component, such as an application server; a front-end component, such as a client computer with a graphical user interface, a web browser, or an application through which a user can interact with an implementation of the subject matter described in this specification; or any combination of one or more such back-end, middleware, or front-end components. The system components may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), such as the internet.
[0114] A computer system can comprise clients and servers. A client and a server are generally located remotely and typically interact via a communication network. The client-server relationship arises from computer programs running on the respective computers, which have a client-server relationship with each other. In some embodiments, a server transmits data, such as an HTML page, to a user device, for example, for the purpose of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, for example, as a result of user interaction, can be received by the device at the server.
[0115] Although this specification contains many specific implementation details, these should not be interpreted as limitations on the scope of an invention or the scope of what can be claimed, but rather as descriptions of features that may be specific to certain embodiments of certain inventions. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in several embodiments or in any suitable subcombination.Furthermore, although features described above may be effective in certain combinations and may even be initially claimed as such, in some cases one or more features from a claimed combination may be extracted from the combination, and the claimed combination may be directed to a subcombination or a variation of a subcombination.
[0116] Similarly, while operations are illustrated in the drawings and listed in the claims in a specific order, this should not be interpreted as meaning that such operations must be performed in the specific order shown or in sequential order, or that all illustrated operations must be performed to achieve desirable results. Multitasking and parallel processing may be advantageous under certain circumstances. Furthermore, the separation of different system modules and components in the embodiments described above should not be interpreted as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0117] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions mentioned in the claims can be performed in a different order and still achieve desirable results. As an example, the processes shown in the accompanying figures do not necessarily require the specific order shown or a sequential order to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Zitierte Nicht-Patentliteratur
[0000] E3x:E(3)-Equivariant Deep Learning Made Easy“ arXiv:2401.07595 [0015, 0054, 0073] Ashish Vaswani et al., „Attention is all you need“, Advances in Neural Information Processing Systems 30 (NIPS 2017
[0061]
Claims
[1] System comprising one or more computers and one or more storage devices storing instructions capable of causing, when executed by the one or more computers, the one or more computers to generate rotationally invariant or covariant descriptors of a three-dimensional configuration of points, wherein the instructions are configured to control the one or more computers to perform a procedure comprising: Using the coordinates of the points to determine a plurality of feature vectors, each feature vector having a corresponding degree (l) and comprising one or more respective features, each feature being determined using a spherical harmonic. (Ylm) the degree (l) of the feature vector and a respective order (m) is determined by linearly combining values of the spherical harmonic function, which are evaluated at the respective coordinates of the points; Transforming each of the multitude of feature vectors into a corresponding moment matrix (M) a,b,l ), where each moment matrix represents a respective irreducible representation (H(l)) of the 3D rotation group in a direct summation representation (H(|a−b|)⊕H(|a−b|+1)⊕⋯⊕H(a+b)) a tensor product of irreducible representations (H(a)⊗H(b)) the 3D rotation group); and Using the moment matrices to determine one or more invariant or covariant descriptors of the three-dimensional configuration of the points. [2] System according to claim 1, wherein transforming each of the plurality of feature vectors into a corresponding moment matrix comprises determining elements of the moment matrix using respective linear combinations of the features of the feature vector. [3] System according to claim 2, wherein each of the linear combinations comprises a sum of each of the features weighted with a respective Clebsch-Gordan coefficient which depends on the degree (l) of the corresponding feature vector and the respective order (m) of the feature. [4] System according to one of claims 1-3, wherein each of the moment matrices is the irreducible representation of the 3D rotation group with the degree (l) of the corresponding feature vector and has a respective form (2a + 1) × (2b + 1), wherein a and b are chosen such that the degree (l) of the corresponding feature vector lies in a range from |a - b| to (a + b). [5] System according to any of the preceding claims, wherein the use of the moment matrices to determine one or more invariant or covariant descriptors comprises: Multiplying two or more of the moment matrices and a vector comprising a linear combination of one or more of the feature vectors to obtain a corresponding invariant or covariant descriptor. [6] System according to any of the preceding claims, wherein the use of the moment matrices to determine one or more invariant or covariant descriptors comprises: Determining one or more combined moment matrices, wherein each combined moment matrix is determined using a respective linear combination of moment matrices that have the same form as the combined moment matrix; and Using the combined moment matrices to determine the one or more invariant or covariant descriptors. [7] System according to claim 6, wherein the moment matrices having the same form as the combined moment matrix comprise moment matrices of different degrees (l). [8] System according to claim 7, wherein the moment matrices having the same form as the combined moment matrix comprise at least one moment matrix of each degree (l) in a range from |a - b| to (a + b), wherein the respective form of the combined moment matrix is (2a + 1) × (2b + 1). [9] System according to one of claims 6-8, wherein the one or more combined moment matrices comprise one or more square matrices and the use of the moment matrices to determine the one or more invariant or covariant descriptors comprises: Determining one or more invariant descriptors using respective traces of one or more square matrices. [10] System according to claim 9, wherein the one or more invariant descriptors are determined using linear combinations of the respective traces of the one or more square matrices. [11] System according to one of the preceding claims, wherein the use of the moment matrices to determine one or more invariant or covariant descriptors of the three-dimensional configuration of the points comprises the following: Determining a plurality of moment block matrices, wherein each moment block matrix comprises a plurality of blocks, each comprising one of the respective moment matrices or combined moment matrices; and Multiplying the multitude of moment block matrices and a feature block matrix comprising a linear combination of one or more of the feature vectors to obtain a corresponding invariant or covariant descriptor. [12] System according to claim 11, wherein corresponding blocks of each of the moment block matrices have the same respective shapes. [13] System according to claim 11 or 12, wherein each of the moment block matrices is a square matrix. [14] System according to one of claims 11-13, wherein the feature block matrix comprises a linear combination of a plurality of feature vectors of the same degree (l). [15] System according to one of claims 11-14, wherein the feature block matrix comprises a concatenation of feature blocks, each feature block comprising one or more rows or columns, each comprising a respective linear combination of feature vectors of the same degree (l). [16] System according to one of claims 11-15, wherein multiplying the plurality of moment block matrices and a feature block matrix comprising one or more of the feature vectors to obtain a corresponding invariant or covariant descriptor comprises: Determine another feature block matrix comprising a linear combination of a plurality of feature vectors of the same degree (l); and Multiplying (i) one or more covariant descriptors obtained by multiplying the plurality of moment block matrices and the feature block matrix, and (ii) the further feature block matrix to obtain a corresponding plurality of invariant descriptors. [17] System according to any of the preceding claims, wherein the use of the moment matrices to determine one or more invariant or covariant descriptors comprises: Determining (i) one or more linear combinations of the feature vectors; and / or (ii) one or more linear combinations of the moment matrices, each linear combination being determined using an appropriate set of feature or moment matrix weighting parameters. [18] System according to claim 17, wherein the method further comprises: Adjusting the feature or moment matrix weighting parameters to optimize an objective function that depends on one or more invariant or covariant descriptors. [19] System according to any of the preceding claims, wherein the method further comprises: Determining one or more linear combinations of the invariant or covariant descriptors, wherein each linear combination is determined using an appropriate set of descriptor weighting parameters, and Adjusting the descriptor weighting parameters to optimize an objective function that depends on one or more linear combinations of the invariant or covariant descriptors. [20] System according to one of the preceding claims, further comprising processing the one or more invariant or covariant descriptors using a machine learning model to generate a corresponding model output. [21] System according to claim 20, wherein the model output is indicative of one or more physical properties of the point configuration. [22] System according to claim 20 or 21, if dependent on claim 18 or 19, wherein optimizing the objective function comprises: Receiving a plurality of training data elements, wherein each training data element comprises (a) a training input comprising coordinates of a respective configuration of points, and (b) a target output comprising one or more physical properties of the configuration of points; for each of the training data elements: Determining one or more invariant or covariant descriptors for the configuration of training input points; and Processing one or more invariant or covariant descriptors using the machine learning model to generate a corresponding model output for the configuration of points of the to generate training input; and wherein the objective function depends on a comparison between the model outputs and the corresponding target outputs. [23] System according to claim 21 or 22, wherein the configuration of points corresponds to a configuration of atoms, each point corresponding to a respective atom in the configuration of atoms. [24] System according to claim 23, wherein the instructions are configured to control the one or more computers to carry out the method for a plurality of real subsets of the atoms in order to generate for each subset one or more rotationally invariant or -covariant descriptors of the respective configuration of atoms in the subset. [25] System according to claim 21, wherein the one or more physical properties comprise one or more of the following: an energy of the configuration of atoms; a respective force on one or more of the atoms; an optical or spectroscopic property of the configuration of atoms; an electrical, magnetic or electromagnetic property of the configuration of atoms; an acoustic or mechanical property of the configuration of atoms. [26] System according to one of claims 23-25, when dependent on claim 22, wherein obtaining a plurality of training data elements comprises determining one or more physical properties of each configuration of atoms by performing a respective quantum chemical calculation for the configuration of atoms. [27] System according to one of claims 21-26, further comprising the selection of a chemical structure from a plurality of candidate chemical structures based on respective model outputs for each of the candidate chemical structures. [28] System according to claim 27, wherein the selection of the chemical structure from the plurality of candidate chemical structures comprises the following: Performing respective molecular dynamics calculations for each of the candidate chemical structures, where at each of a plurality of time steps after an initial time step, the coordinates of the atoms in the candidate chemical structure are updated based on one or more model outputs generated by the machine learning model at one or more previous time steps. [29] System according to claim 27 or 28, wherein the selected chemical structure is a drug or a ligand of an industrial enzyme, and the selection of the chemical structure from the plurality of candidate chemical structures comprises the following: Evaluating the interaction of each candidate chemical structure with a target molecule and / or a binding site of a target molecule. [30] System according to claim 29, wherein the target molecule and / or the binding site of the target molecule comprises a receptor or an enzyme and wherein the selected chemical structure is an agonist or antagonist of the receptor or enzyme. [31] System according to claim 27 or 28, wherein the selected chemical structure is a drug, and the selection of the chemical structure from the plurality of candidate chemical structures comprises the following: Evaluating the interaction of each candidate chemical structure with each of a multitude of target molecules and / or binding sites of target molecules in order to either (i) obtain a chemical structure that interacts with each of the target molecules and / or binding sites, or (ii) obtain a chemical structure that interacts with only one of the target molecules and / or binding sites of the target molecules. [32] System according to claim 27 or 28, wherein the selected chemical structure is a catalyst for an industrial chemical process and the one or more physical properties comprise one or more measures or predictors of the catalytic activity of the selected chemical structure for the industrial chemical process. [33] System according to one of claims 27-32, wherein the selected chemical structure is an inorganic compound or alloy. [34] System according to one of claims 1-22, wherein each of the points corresponds to a respective location in an environment and the method further comprises determining the coordinates of each of the points from one or more images of the environment. [35] System according to claim 34, wherein each image comprises a plurality of pixels, each pixel being associated with a set of one or more pixel values, and determining the coordinates of each of the points from one or more images of the environment comprises processing the pixel values. [36] System according to claim 35, wherein each set of one or more pixel values comprises a respective depth value for the corresponding pixel. [37] System according to one of claims 34-36, wherein the environment is a real environment. [38] System according to one of claims 34-37, depending on one of claims 21-23, wherein the machine learning model is used to perform one or more of the following tasks: (i) a classification task in which the model output classifies the environment into one or more categories; (ii) an object detection or localization task in which the model output is indicative of whether one or more objects are present in the environment, or the model output defines coordinates of a respective region for one or more objects in the environment; (iii) an action or gesture recognition task in which the model output is indicative of whether one or more actions or gestures are being performed in the environment; (iv) a semantic segmentation task in which the model output assigns spatial locations in the environment to respective segmentation categories; (v) a keypoint detection task in which the model output includes coordinates of one or more keypoints in the environment and (vi) an environment similarity task in which the model output is indicative of a similarity of the environment to one or more predefined environments. [39] System according to one of claims 34-38, further comprising the use of the model output to determine one or more actions to be performed by an agent interacting with the environment to perform a specified task. [40] System according to claim 39, wherein the environment is a real environment and the agent comprises (i) a mechanical agent or robot that interacts with the real environment to perform the specified task, or (ii) an electronic agent that controls equipment in the real environment to perform the specified task. [41] System according to claim 40, wherein the coordinates of the points are obtained by processing sensor data acquired by one or more sensors in the real environment. [42] Computer-readable storage medium encoded with instructions which, when executed by one or more computers, cause the one or more computers to perform operations to implement a computer-implemented method for generating rotationally invariant or covariant descriptors of a three-dimensional configuration of points, the method comprising: Using the coordinates of the points to determine a plurality of feature vectors, each feature vector having a corresponding degree (l) and comprising one or more respective features, each feature being determined using a spherical harmonic. (Ylm) the degree (l) of the feature vector and a respective order (m) is determined by linearly combining values of the spherical harmonic function, which are evaluated at the respective coordinates of the points; Transforming each of the multitude of feature vectors into a corresponding moment matrix (M) a,b,l ), where each moment matrix represents a respective irreducible representation (H(l)) of the 3D rotation group in a direct summation representation (H(|a−b|)⊕H(|a−b|+1)⊕⋯⊕H(a+b)) a tensor product of irreducible representations (H(a)⊗H(b)) corresponds to the 3D rotation group; and Using the moment matrices to determine one or more invariant or covariant descriptors of the three-dimensional configuration of the points. [43] Computer-readable storage medium according to claim 42, wherein transforming each of the plurality of feature vectors into a corresponding moment matrix comprises determining elements of the moment matrix using respective linear combinations of the features of the feature vector. [44] Computer-readable storage medium according to claim 43, wherein each of the linear combinations comprises a sum of each of the features weighted with a respective Clebsch-Gordan coefficient which depends on the degree (l) of the corresponding feature vector and the respective order (m) of the feature. [45] Computer-readable storage medium according to one of claims 42-44, wherein each of the moment matrices is the irreducible representation of the 3D rotation group with the degree (l) of the corresponding feature vector and has a respective form (2a + 1) × (2b + 1), wherein a and b are chosen such that the degree (l) of the corresponding feature vector is in a range from |a - b| to (a + b). [46] Computer-readable storage medium according to one of claims 42-45, wherein the use of the moment matrices to determine one or more invariant or covariant descriptors comprises: Multiplying two or more of the moment matrices and a vector comprising a linear combination of one or more of the feature vectors to obtain a corresponding invariant or covariant descriptor. [47] Computer-readable storage medium according to one of claims 42-46, wherein the use of the moment matrices to determine one or more invariant or covariant descriptors comprises: Determining one or more combined moment matrices, wherein each combined moment matrix is determined using a respective linear combination of moment matrices that have the same form as the combined moment matrix; and Using the combined moment matrices to determine the one or more invariant or covariant descriptors. [48] Computer-readable storage medium according to claim 47, wherein the moment matrices having the same shape as the combined moment matrix comprise moment matrices of different degrees (l). [49] Computer-readable storage medium according to claim 48, wherein the moment matrices having the same form as the combined moment matrix comprise at least one moment matrix of each degree (l) in a range from |a - b| to (a + b), wherein the respective form of the combined moment matrix is (2a + 1) × (2b + 1). [50] Computer-readable storage medium according to one of claims 47-49, wherein the one or more combined moment matrices comprise one or more square matrices and the use of the moment matrices to determine the one or more invariant or covariant descriptors comprises the following: Determining one or more invariant descriptors using respective traces of one or more square matrices. [51] Computer-readable storage medium according to claim 50, wherein the one or more invariant descriptors are determined using linear combinations of the respective tracks of the one or more square matrices. [52] Computer-readable storage medium according to one of claims 42-51, wherein the use of the moment matrices to determine one or more invariant or covariant descriptors of the three-dimensional configuration of the points comprises the following: Determining a plurality of moment block matrices, wherein each moment block matrix comprises a plurality of blocks, each comprising one of the respective moment matrices or combined moment matrices; and Multiplying the multitude of moment block matrices and a feature block matrix comprising a linear combination of one or more of the feature vectors to obtain a corresponding invariant or covariant descriptor. [53] Computer-readable storage medium according to claim 52, wherein corresponding blocks of each of the moment block matrices have the same respective shapes. [54] Computer-readable storage medium according to claim 52 or 53, wherein each of the moment block matrices is a square matrix. [55] Computer-readable storage medium according to one of claims 52-54, wherein the feature block matrix comprises a linear combination of a plurality of feature vectors of the same degree (l). [56] Computer-readable storage medium according to one of claims 52-54, wherein the feature block matrix comprises a concatenation of feature blocks, each feature block comprising one or more rows or columns, each comprising a respective linear combination of feature vectors of the same degree (l). [57] Computer-readable storage medium according to one of claims 52-56, wherein multiplying the plurality of moment block matrices and a feature block matrix comprising one or more of the feature vectors to obtain a corresponding invariant or covariant descriptor comprises: Determine another feature block matrix comprising a linear combination of a plurality of feature vectors of the same degree (l); and Multiplying (i) one or more covariant descriptors obtained by multiplying the plurality of moment block matrices and the feature block matrix, and (ii) the further feature block matrix to obtain a corresponding plurality of invariant descriptors. [58] Computer-readable storage medium according to one of claims 42-57, wherein the use of the moment matrices to determine one or more invariant or covariant descriptors comprises: Determining (i) one or more linear combinations of the feature vectors; and / or (ii) one or more linear combinations of the moment matrices, each linear combination being determined using an appropriate set of feature or moment matrix weighting parameters. [59] Computer-readable storage medium according to claim 58, wherein the method further comprises: Adjusting the feature or moment matrix weighting parameters to optimize an objective function that depends on one or more invariant or covariant descriptors. [60] Computer-readable storage medium according to any one of claims 42-59, wherein the method further comprises: Determining one or more linear combinations of the invariant or covariant descriptors, wherein each linear combination is determined using an appropriate set of descriptor weighting parameters, and Adjusting the descriptor weighting parameters to optimize an objective function that depends on one or more linear combinations of the invariant or covariant descriptors. [61] Computer-readable storage medium according to one of claims 42-60, further comprising processing one or more invariant or covariant descriptors using a machine learning model to generate a corresponding model output. [62] Computer-readable storage medium according to claim 61, wherein the model output is indicative of one or more physical properties of the point configuration. [63] Computer-readable storage medium according to claim 61 or 62, when dependent on claim 59 or 60, wherein optimizing the objective function comprises: Receiving a plurality of training data elements, wherein each training data element comprises (a) a training input comprising coordinates of a respective configuration of points, and (b) a target output comprising one or more physical properties of the configuration of points; for each of the training data elements: Determining one or more invariant or covariant descriptors for the configuration of training input points; and Processing one or more invariant or covariant descriptors using the machine learning model to generate a corresponding model output for the configuration of points of the training input; and wherein the objective function depends on a comparison between the model outputs and the corresponding target outputs. [64] Computer-readable storage medium according to claim 62 or 63, wherein the configuration of points corresponds to a configuration of atoms, each point corresponding to a respective atom in the configuration of atoms. [65] Computer-readable storage medium according to claim 64, wherein the instructions are configured to control the one or more computers to carry out the method for a plurality of real subsets of the atoms in order to generate for each subset one or more rotationally invariant or -covariant descriptors of the respective configuration of atoms in the subset. [66] Computer-readable storage medium according to claim 65, wherein the one or more physical properties comprise one or more of the following: an energy of the configuration of atoms; a respective force on one or more of the atoms; an optical or spectroscopic property of the configuration of atoms; an electrical, magnetic or electromagnetic property of the configuration of atoms; an acoustic or mechanical property of the configuration of atoms. [67] Computer-readable storage medium according to one of claims 64-66, when dependent on claim 63, wherein obtaining a plurality of training data elements comprises determining one or more physical properties of each configuration of atoms by performing a respective quantum chemical calculation for the configuration of atoms. [68] Computer-readable storage medium according to one of claims 64-67, further comprising the selection of a chemical structure from a plurality of candidate chemical structures based on respective model outputs for each of the candidate chemical structures. [69] Computer-readable storage medium according to claim 68, wherein the selection of the chemical structure from the plurality of candidate chemical structures comprises the following: Performing respective molecular dynamics calculations for each of the candidate chemical structures, where at each of a plurality of time steps after an initial time step, the coordinates of the atoms in the candidate chemical structure are updated based on one or more model outputs generated by the machine learning model at one or more previous time steps. [70] Computer-readable storage medium according to claim 68 or 69, wherein the selected chemical structure is a drug or a ligand of an industrial enzyme, and the selection of the chemical structure from the plurality of candidate chemical structures comprises the following: Evaluating the interaction of each candidate chemical structure with a target molecule and / or a binding site of a target molecule. [71] Computer-readable storage medium according to claim 70, wherein the target molecule and / or the binding site of the target molecule comprises a receptor or an enzyme and wherein the selected chemical structure is an agonist or antagonist of the receptor or enzyme. [72] Computer-readable storage medium according to claim 68 or 69, wherein the selected chemical structure is a drug, and the selection of the chemical structure from the plurality of candidate chemical structures comprises: Evaluating the interaction of each candidate chemical structure with each of a multitude of target molecules and / or binding sites of target molecules in order to either (i) obtain a chemical structure that interacts with each of the target molecules and / or binding sites, or (ii) obtain a chemical structure that interacts with only one of the target molecules and / or binding sites of the target molecules. [73] Computer-readable storage medium according to claim 68 or 69, wherein the selected chemical structure is a catalyst of an industrial chemical process and the one or more physical properties comprise one or more measures or predictors of the catalytic activity of the selected chemical structure for the industrial chemical process. [74] Computer-readable storage medium according to one of claims 68-73, wherein the selected chemical structure is an inorganic compound or alloy. [75] Computer-readable storage medium according to one of claims 42-63, wherein each of the points corresponds to a respective location in an environment and the method further comprises determining the coordinates of each of the points from one or more images of the environment. [76] Computer-readable storage medium according to claim 75, wherein each image comprises a plurality of pixels, each pixel being associated with a set of one or more pixel values, and determining the coordinates of each of the points from one or more images of the environment comprises processing the pixel values. [77] Computer-readable storage medium according to claim 76, wherein each set of one or more pixel values comprises a respective depth value for the corresponding pixel. [78] Computer-readable storage medium according to one of claims 75-77, wherein the environment is a real environment. [79] Computer-readable storage medium according to any one of claims 75-78, depending on any one of claims 61-63, wherein the machine learning model is used to perform one or more of the following tasks: (i) a classification task in which the model output classifies the environment into one or more categories; (ii) an object detection or localization task in which the model output is indicative of whether one or more objects are present in the environment, or the model output defines coordinates of a respective region for one or more objects in the environment; (iii) an action or gesture recognition task in which the model output is indicative of whether one or more actions or gestures are being performed in the environment; (iv) a semantic segmentation task in which the model output assigns spatial locations in the environment to respective segmentation categories; (v) a keypoint detection task in which the model output includes coordinates of one or more keypoints in the environment and (vi) an environment similarity task in which the model output is indicative of a similarity of the environment to one or more predefined environments. [80] Computer-readable storage medium according to one of claims 75-79, further comprising the use of the model output to determine one or more actions to be performed by an agent that interacts with the environment to perform a specified task. [81] Computer-readable storage medium according to claim 80, wherein the environment is a real environment and the agent comprises (i) a mechanical agent or robot that interacts with the real environment to perform the specified task, or (ii) an electronic agent that controls equipment in the real environment to perform the specified task.